IP Library Granted Patent US 10,019,983
Granted Patent B2
US 10,019,983 · App. 13/599,297 · Granted Jul 10, 2018

Method and system for predicting speech recognition performance using accuracy scores

Inventors: Aravind Ganapathiraju (Hyderabad, IN); Yingyi Tan (Carmel, IN); Felix Immanuel Wyss (Bloomington, IN); Scott Allen Randal (Redmond, WA)
G10L15/01G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,019,983
App. No.
13/599,297
Granted
Jul 10, 2018
Kind
B2
Abstract

A system and method are presented for predicting speech recognition performance using accuracy scores in speech recognition systems within the speech analytics field. A keyword set is selected. Figure of Merit (FOM) is computed for the keyword set. Relevant features that describe the word individually and in relation to other words in the language are computed. A mapping from these features to FOM is learned. This mapping can be generalized via a suitable machine learning algorithm and be used to predict FOM for a new keyword. In at least one embodiment, the predicted FOM may be used to adjust internals of speech recognition engine to achieve a consistent behavior for all inputs for various settings of confidence values.

Claims (81)

1. A method for predicting speech recognition performance in a speech recognition system, the system comprising a recognition engine, a database, a model learning module, and a performance prediction module, the method comprising the steps of:

a. determining, by the performance prediction module, at least one feature vector for an input into the speech recognition system, wherein the at least one feature vector includes features that comprise at least two features selected from the group comprising: the number of phonemes, the number of syllables, and the number of stressed vowels;

b. creating a prediction model by:

i. selecting a set of keywords;

ii. computing an other feature vector of desired features for each of the keywords;

iii. inputting the other feature vector into the model learning module, wherein the model learning module adjusts parameters to minimize a cost function; and

iv. saving the results from the model learning module as the prediction model for prediction of a figure of merit of the input;

c. passing the at least one feature vector into the prediction model;

d. applying, by the performance prediction module, the prediction model to predict a figure of merit for the speech recognition system, wherein the figure of merit is indicative of the accuracy of performance of the speech recognition system, wherein the figure of merit (fom) is predicted using a mathematical expression

fom

=

i

=

1

N

a

i

(

x

i

-

b

i

)

2

N represents an upper limit on a number of features based on the determined feature vector used to learn the prediction, i represents the index of features, x i represents the i-th feature in the determined feature vector, and the equation parameters a and b are learned values;

e. reporting, by the performance prediction module, the predicted figure of merit for the speech recognition system performance; and

f. adjusting the recognition engine based on the predicted figure of merit.

2. The method of claim 1 , wherein the figure of merit prediction has a detection rate averaging 5 FA/Kw/Hr.

3. The method of claim 1 , wherein the input comprises at least one word.

4. The method of claim 1 , wherein the input comprises a phonetic pronunciation.

5. The method of claim 1 , wherein the method is performed in real-time as additional input is provided.

6. The method of claim 1 , wherein the at least one feature vector is determined comprising the steps of:

converting the input into a sequence of phonemes; and

performing morphological analysis of words in a language.

7. The method of claim 6 , wherein the converting is performed using statistics for phonemes and phoneme confusion matrix.

8. The method of claim 7 , further comprising the step of computing the phoneme confusion matrix using a phoneme recognizer.

9. The method of claim 1 , further comprising the step of automatically adjusting internal scores of the recognition engine based on the prediction reported by the performance prediction module.

10. A computer system with a digital microprocessor and associated memory configured for executing software programs, the system being configured for predicting speech recognition performance, comprising:

using a performance prediction module to determine at least one feature vector for an input into the speech recognition system, wherein the at least one feature vector includes features that comprise at least two features selected from the group comprising: the number of phonemes, the number of syllables, and the number of stressed vowels;

creating a prediction model by:

selecting a set of keywords;

computing an other feature vector of desired features for each of the keywords;

inputting the other feature vector into the model learning module, wherein the model learning module adjusts parameters to minimize a cost function; and

saving the results from the model learning module as the prediction model for prediction of a figure of merit of the input;

passing the at least one feature vector into the prediction model;

using the performance prediction module to apply the prediction model to predict a figure of merit for the speech recognition system, wherein the figure of merit is indicative of the accuracy of performance of the speech recognition system, wherein the figure of merit Om) is predicted using a mathematical expression

fom

=

i

=

1

N

a

i

(

x

i

-

b

i

)

2

N represents an upper limit on a number of features based on the determined feature vector used to learn the prediction, i represents the index of features, x i represents the i-th feature in the determined feature vector, and the equation parameters a and b are learned values;

using the performance prediction module to report the predicted figure of merit for the speech recognition system performance; and

adjusting the recognition engine based on the predicted figure of merit.

11. The system of claim 10 , wherein the figure of merit has a detection rate averaging 5 FA/Kw/Hr.

12. The system of claim 10 , wherein the input comprises at least one word.

13. The system of claim 10 , wherein the input comprises a phonetic pronunciation.

14. The system of claim 10 , wherein the at least one feature vector is determined comprising:

a. converting the input into a sequence of phonemes; and

b. performing morphological analysis of words in a language.

15. The system of claim 14 , wherein the converting is performed using statistics for phonemes and phoneme confusion matrix.

16. The system of claim 15 , further comprising computing the phoneme confusion matrix using a phoneme recognizer.

17. The system of claim 10 , further comprising automatically adjusting internal scores of the recognition engine based on the prediction reported by the performance prediction module.

Assignments (8)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 04814/0387 Recorded Feb 5, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070115/0445 →
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 040815/0001 Recorded Feb 3, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070498/0001 →
CHANGE OF NAME Recorded Jun 7, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067651/0783 →
SECURITY AGREEMENT Recorded Feb 22, 2019
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.; ECHOPASS CORPORATION; GREENEDEN U.S. HOLDINGS II, LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 048414/0387 →
MERGER Recorded Jun 5, 2018
From: INTERACTIVE INTELLIGENCE GROUP, INC.
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 045994/0126 →
SECURITY AGREEMENT Recorded Dec 5, 2016
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC., AS GRANTOR; ECHOPASS CORPORATION; INTERACTIVE INTELLIGENCE GROUP, INC.; BAY BRIDGE DECISION TECHNOLOGIES, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 040815/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2016
From: INTERACTIVE INTELLIGENCE, INC.
To: INTERACTIVE INTELLIGENCE GROUP, INC.
Reel/Frame 040647/0285 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2012
From: GANAPATHIRAJU, ARAVIND; TAN, YINGYI; WYSS, FELIX IMMANUEL; RANDAL, SCOTT ALLEN
To: INTERACTIVE INTELLIGENCE, INC.
Reel/Frame 028876/0368 →
Continuity (1)
Related Publication 20140067391A1 · Mar 6, 2014