IP Library Granted Patent US 9,646,605
Granted Patent B2
US 9,646,605 · App. 13/746,687 · Granted May 9, 2017

False alarm reduction in speech recognition systems using contextual information

Inventors: Konstantin Biatov (Sankt Augustin, DE); Aravind Ganapathiraju (Hyderabad, IN); Felix Immanuel Wyss (Zionsville, IN)
Assignee: Interactive Intelligence Group, Inc.
G10L15/063G10L15/183G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,646,605
App. No.
13/746,687
Granted
May 9, 2017
Kind
B2
Abstract

A system and method are presented for using spoken word verification to reduce false alarms by exploiting global and local contexts on a lexical level, a phoneme level, and on an acoustical level. The reduction of false alarms may occur through a process that determines whether a word has been detected or if it is a false alarm. Training examples are used to generate models of internal and external contexts which are compared to test word examples. The word may be accepted or rejected based on comparison results. Comparison may be performed either at the end of the process or at multiple steps of the process to determine whether the word is rejected.

Claims (35)

1. A computerized method for reducing false alarms in a speech recognition system, the method comprising:

receiving a plurality of training examples;

generating a model of a left internal context based at least in part on the plurality of training examples, wherein the generation of the model includes compact representation of the left internal context in the form of spectral, cepstral or sinusoidal descriptions;

generating a model of a right internal context based at least in part on the plurality of training examples, wherein the generation of the model includes compact representation of the right internal context in the form of spectral, cepstral or sinusoidal descriptions;

generating a model of a left external context based at least in part on the plurality of training examples, wherein the generation of the model includes compact representation of the left external context in the form of spectral, cepstral or sinusoidal descriptions;

generating a model of a right external context based at least in part on the plurality of training examples, wherein the generation of the model includes compact representation of the right external context in the form of spectral, cepstral or sinusoidal descriptions;

receiving at least one test word, the at least one test word comprising an external context;

comparing the external context of the at least one test word against a threshold associated with each of the model of the left internal context, the model of the right internal context, the model of the left external context, and the model of the right external context; and

rejecting the at least one test word if it is not within the thresholds.

2. The method of claim 1 , wherein the test word is an analog context.

3. The method of claim 2 , further comprising converting the test word from an analog context to a digital format.

4. The method of claim 1 , further comprising:

learning an acceptable threshold for each of the model of the left internal context, the model of the right internal context, the model of the left external context, and the model of the right external context based at least in part on cross-validating sets; and

wherein the comparing step is performed using each acceptable threshold.

5. The method of claim 1 , wherein each training example in the plurality of training examples comprises a representation of a test word and a local context; and

wherein each local context is based on average phoneme and syllable duration from similar word types.

6. The method of claim 1 , wherein the comparing step comprises the additional step of evaluating the at least one word with a perplexity test.

7. The method of claim 1 , wherein each of the model of the left internal context, the model of the right internal context, the model of the left external context, and the model of the right external context include compact representations.

8. A computerized method for reducing false alarms in a speech recognition system, the method comprising:

receiving a plurality of training examples, each training example comprising a representation of a spoken word and a local context;

generating at least one model of an acoustic context based on the plurality of training examples, wherein the generation of the model includes compact representation of the acoustic context in the form of spectral, cepstral or sinusoidal descriptions;

generating at least one model of a phonetic context based on the plurality of training examples, wherein the generation of the model includes compact representation of the phonetic context in the form of spectral, cepstral or sinusoidal descriptions;

generating at least one model of a linguistic context based on the plurality of training examples, wherein the generation of the model includes compact representation of the linguistic context in the form of spectral, cepstral or sinusoidal descriptions;

receiving at least one test word, the at least one test word comprising an external context;

comparing the at least one test word against a threshold associated with each of the model of the acoustic context, the model of the phonetic context, and the model of the linguistic context; and

rejecting the at least one test word if it is not within the thresholds.

9. The method of claim 8 , wherein the spoken word is an analog context.

10. The method of claim 9 , further comprising converting the spoken word from an analog context to a digital format.

11. The method of claim 8 , further comprising:

learning an acceptable threshold for each of the model of the acoustic context, the model of the phonetic context, and the model of the linguistic context based at least in part on cross-validating sets; and

wherein the comparing step is performed using each acceptable threshold.

12. The method of claim 8 , wherein each training example in the plurality of training examples comprises a representation of a spoken word and a local context; and

wherein each local context is based on average phoneme and syllable duration from similar word types.

13. The method of claim 8 , wherein the comparing step comprises the additional step of evaluating the at least one word with a perplexity test.

14. The method of claim 8 , wherein the generating at least one model of an acoustic context step includes generating a left internal model, a right internal model, a left external model, and a right external model for each spoken word in the plurality of training examples.

Assignments (8)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 04814/0387 Recorded Feb 5, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070115/0445 →
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 040815/0001 Recorded Feb 3, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070498/0001 →
CHANGE OF NAME Recorded Jun 6, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067644/0877 →
SECURITY AGREEMENT Recorded Feb 22, 2019
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.; ECHOPASS CORPORATION; GREENEDEN U.S. HOLDINGS II, LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 048414/0387 →
MERGER Recorded Jul 1, 2018
From: INTERACTIVE INTELLIGENCE GROUP, INC.
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 046463/0839 →
SECURITY AGREEMENT Recorded Dec 5, 2016
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC., AS GRANTOR; ECHOPASS CORPORATION; INTERACTIVE INTELLIGENCE GROUP, INC.; BAY BRIDGE DECISION TECHNOLOGIES, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 040815/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2016
From: INTERACTIVE INTELLIGENCE, INC.
To: INTERACTIVE INTELLIGENCE GROUP, INC.
Reel/Frame 040647/0285 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2013
From: BIATOV, KONSTANTIN; GANAPATHIRAJU, ARAVIND; WYSS, FELIX IMMANUEL
To: INTERACTIVE INTELLIGENCE, INC.
Reel/Frame 029670/0001 →
Continuity (1)
Related Publication 20140207457A1 · Jul 24, 2014