IP Library Granted Patent US 10,319,366
Granted Patent B2
US 10,319,366 · App. 15/478,108 · Granted Jun 11, 2019

Predicting recognition quality of a phrase in automatic speech recognition systems

Inventors: Amir Lev-Tov (Bat-Yam, IL); Avraham Faizakof (Kfar-Warburg, IL); Yochai Konig (San Francisco, CA)
G10L15/01G10L15/02G10L15/04G10L15/063G10L15/26G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,319,366
App. No.
15/478,108
Granted
Jun 11, 2019
Kind
B2
Abstract

A method for predicting a speech recognition quality of a phrase comprising at least one word includes: receiving, on a computer system including a processor and memory storing instructions, the phrase; computing, on the computer system, a set of features comprising one or more features corresponding to the phrase; providing the phrase to a prediction model on the computer system and receiving a predicted recognition quality value based on the set of features; and returning the predicted recognition quality value.

Claims (78)

1. A method for configuring a speech analytics system, the method comprising:

computing, on a computer system comprising a processor and memory storing instructions, a plurality of text features corresponding to a text phrase, the text phrase comprising text including at least one word;

computing, by supplying the plurality of text features corresponding to the text phrase as input to a prediction model on the computer system, a predicted recognition quality value representing a likelihood of the text phrase being correctly recognized by an automatic speech recognition system of the speech analytics system when appearing in spoken form in user speech, the prediction model being trained using:

a plurality of true transcriptions of a collection of recorded speech;

a plurality of training text features of the true transcriptions; and

a recognizer output generated by supplying the recorded speech to the automatic speech recognition system; and

displaying, on a graphical user interface, the predicted recognition quality value for the text phrase.

2. The method of claim 1 , further comprising receiving the text of the phrase via the graphical user interface.

3. The method of claim 1 , further comprising displaying, on the graphical user interface, a label corresponding to the phrase, the label being based on the predicted recognition quality value.

4. The method of claim 1 , wherein the prediction model is a neural network.

5. The method of claim 4 , wherein the neural network is a multilayer perceptron neural network and wherein the neural network is trained by applying a backpropagation algorithm.

6. The method of claim 1 , wherein the prediction model is generated by:

generating, on the computer system, a plurality of training phrases from the collection of recorded speech;

calculating, on the computer system, a target value for each of the training phrases;

calculating the plurality of training text features of each of the training phrases;

training, on the computer system, the prediction model based on the training text features; and

setting, on the computer system, a filtering threshold.

7. The method of claim 6 , wherein the generating the training phrases comprises:

segmenting the plurality of true transcriptions into a plurality of true phrases;

processing the collection of recorded speech using the automatic speech recognition system to generate the recognizer output;

comparing the recognizer output to the true phrases to identify matches between the recognizer output and the true phrases;

based on the comparison, tagging matching phrases among the true phrases that match the recognizer output as hits;

selecting the matching phrases with a number of hits greater than a threshold value as training phrases; and

returning the plurality of training phrases.

8. The method of claim 6 , wherein the filtering threshold is set by optimizing precision and recall values on a test set of phrases of the plurality of training phrases.

9. The method of claim 1 , wherein the features of the phrase comprise at least one of:

a precision of a word in the phrase;

a recall of a word in the phrase;

a phrase error rate;

a sum of the precision of the phrase and the recall of the phrase;

a number of long words in the phrase;

a number of vowels in the phrase;

a length of the phrase;

a confusion matrix of the phrase; and

a feature of a language model.

10. The method of claim 1 , further comprising:

comparing the predicted recognition quality value to a threshold value; and

computing the likelihood of the phrase being correctly recognized by the automatic speech recognition system when appearing in audio based on the comparison between the predicted recognition quality value and the threshold value.

11. A system for configuring a speech analytics system, the system comprising:

a processor; and

memory storing instructions that, when executed by the processor, cause the processor to:

compute a plurality of text features corresponding to a text phrase, the text phrase comprising text including at least one word;

compute, by supplying the plurality of text features corresponding to the text phrase as input to a prediction model, a predicted recognition quality value representing a likelihood of the text phrase being correctly recognized by an automatic speech recognition system of the speech analytics system when appearing in spoken from in user speech, the prediction model being trained using:

a plurality of true transcriptions of a collection of recorded speech;

a plurality of training text features of the true transcriptions; and

a recognizer output generated by supplying the recorded speech to the automatic speech recognition system; and

display, on a graphical user interface, the predicted recognition quality value for the text phrase.

12. The system of claim 11 , wherein the memory further stores instructions that, when executed by the processor, cause the processor to receive the text of the phrase via the graphical user interface.

13. The system of claim 11 , wherein the memory further stores instructions that, when executed by the processor, cause the processor to display, on the graphical user interface, a label corresponding to the phrase, the label being based on the predicted recognition quality value.

14. The system of claim 11 , wherein the prediction model is a neural network.

15. The system of claim 14 , wherein the neural network is a multilayer perceptron neural network and wherein the neural network is trained by applying a backpropagation algorithm.

16. The system of claim 11 , wherein the prediction model is generated by:

generating a plurality of training phrases from the collection of recorded speech;

calculating a target value for each of the training phrases;

calculating the plurality of training text features of each of the training phrases;

training the prediction model based on the training text features; and

setting a filtering threshold.

17. The system of claim 16 , wherein the generating the training phrases comprises:

segmenting the plurality of true transcriptions into a plurality of true phrases;

processing the collection of recorded speech using the automatic speech recognition system to generate the recognizer output;

comparing the recognizer output to the true phrases to identify matches between the recognizer output and the true phrases;

based on the comparison, tagging matching phrases among the true phrases that match the recognizer output as hits;

selecting the matching phrases with a number of hits greater than a threshold value as training phrases; and

returning the plurality of training phrases.

18. The system of claim 16 , wherein the filtering threshold is set by optimizing precision and recall values on a test set of phrases of the plurality of training phrases.

19. The system of claim 11 , wherein the features of the phrase comprise at least one of:

a precision of a word in the phrase;

a recall of a word in the phrase;

a phrase error rate;

a sum of the precision of the phrase and the recall of the phrase;

a number of long words in the phrase;

a number of vowels in the phrase;

a length of the phrase;

a confusion matrix of the phrase; and

a feature of a language model.

20. The system of claim 11 , wherein the memory further stores instructions that, when executed by the processor, cause the processor to:

compare the predicted recognition quality value to a threshold value; and

compute the likelihood of the phrase being correctly recognized by the automatic speech recognition system when appearing in audio based on the comparison between the predicted recognition quality value and the threshold value.

Assignments (6)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 04814/0387 Recorded Feb 5, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070115/0445 →
CHANGE OF NAME Recorded May 13, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067391/0097 →
SECURITY AGREEMENT Recorded Feb 22, 2019
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.; ECHOPASS CORPORATION; GREENEDEN U.S. HOLDINGS II, LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 048414/0387 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2017
From: KONIG, YOCHAI
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 041987/0253 →
MERGER Recorded Apr 12, 2017
From: UTOPY, INC.,
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 041987/0306 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2017
From: LEV-TOV, AMIR; FAIZAKOF, AVRAHAM
To: UTOPY, INC.,
Reel/Frame 042241/0421 →
Continuity (2)
Continuation 14067732 · Oct 30, 2013
Related Publication 20170206889A1 · Jul 20, 2017