IP Library Granted Patent US 11,195,514
Granted Patent B2
US 11,195,514 · App. 16/414,885 · Granted Dec 7, 2021

System and method for a multiclass approach for confidence modeling in automatic speech recognition systems

Inventors: Ramasubramanian Sundaram (Hyderabad, IN); Aravind Ganapathiraju (Hyderabad, IN); Yingyi Tan (Carmel, IN)
G10L15/063G06N20/00G10L15/14G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,195,514
App. No.
16/414,885
Granted
Dec 7, 2021
Kind
B2
Abstract

A system and method are presented for a multiclass approach for confidence modeling in automatic speech recognition systems. A confidence model may be trained offline using supervised learning. A decoding module is utilized within the system that generates features for audio files in audio data. The features are used to generate a hypothesized segment of speech which is compared to a known segment of speech using edit distances. Comparisons are labeled from one of a plurality of output classes. The labels correspond to the degree to which speech is converted to text correctly or not. The trained confidence models can be applied in a variety of systems, including interactive voice response systems, keyword spotters, and open-ended dialog systems.

Claims (25)

1. A method for training a confidence model in an automatic speech recognition system to obtain probability of an output hypothesis being correct, comprising the steps of:

providing training examples of audio data, wherein the training examples comprise features and a plurality of multiclass labels that are associated with the features, wherein the plurality of multiclass labels correspond to a degree that the training examples are correctly converted to text;

generating features for training by a decoding module for each audio file in the audio data, wherein the decoding module comprises the confidence model in the automatic speech recognition system;

applying a Deep Neural Network to evaluate the features generated by comparing a hypothesized segment of speech to a known segment of speech, wherein the confidence model provides a probability comprising floating-point numbers of the hypothesized segment of speech being correct to the decoding module; and

labeling comparisons of hypothesized segments to reference segments from one of a plurality of output classes.

2. The method of claim 1 , wherein the comparing comprises examining edit distance and normalized edit distance as a metric to determine class label.

3. The method of claim 2 , wherein the normalized edit distance is obtained by dividing the edit distance value by the length of the string.

4. The method of claim 1 , wherein the plurality of output classes is four.

5. The method of claim 1 , wherein the labels comprise one of a plurality of labels corresponding to the degree to which speech is converted to text correctly.

6. The method of claim 1 , wherein the training of the confidence model is performed offline using supervised learning.

7. A method for converting input speech to text using confidence modelling with a multiclass approach in an automatic speech recognition system, the method comprising the steps of:

accepting input speech into the automatic speech recognition system;

converting the input speech into a set of features by a frontend module using a speech feature extraction method;

accepting the features by a decoding module and determining a best hypotheses of output text using an acoustic model; and

applying a confidence model trained with a Deep Neural Network to the decoding module to obtain a probability using a multiclass classifier to predict class output text being correct, wherein the confidence model is trained by:

providing training examples of audio data, wherein the training examples comprise features and a plurality of multiclass labels that are associated with the features, wherein the plurality of multiclass labels correspond to a degree that the training examples are correctly converted to text;

generating features for training by a decoding module for each audio file in the audio data, wherein the decoding module comprises the confidence model in the automatic speech recognition system;

evaluating the features generated by comparing a hypothesized segment of speech to a known segment of speech, wherein the confidence model provides a probability comprising floating-point numbers of the hypothesized segment of speech being correct to the decoding module; and

labeling comparisons of hypothesized segments to reference segments from one of a plurality of output classes.

8. The method of claim 7 , wherein the speech feature extraction method comprises Mel-frequency Cepstrum Coefficients.

9. The method of claim 7 , wherein the comparing comprises examining edit distance and normalized edit distance as a metric to determine class label.

10. The method of claim 9 , wherein the normalized edit distance is obtained by dividing the edit distance value by the length of the string.

11. The method of claim 7 , wherein the plurality of output classes is four.

12. The method of claim 7 , wherein the labels comprise one of a plurality of labels corresponding to the degree to which speech is converted to text correctly.

13. The method of claim 7 , wherein the training of the confidence model is performed offline using supervised learning.

Assignments (5)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 050860/0227 Recorded Feb 3, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070096/0452 →
CHANGE OF NAME Recorded Jan 30, 2025
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 070056/0995 →
CORRECTIVE ASSIGNMENT TO CORRECT THE TO ADD PAGE 2 OF THE SECURITY AGREEMENT WHICH WAS INADVERTENTLY OMITTED PREVIOUSLY RECORDED ON REEL 049916 FRAME 0454. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY AGREEMENT. Recorded Oct 29, 2019
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 050860/0227 →
SECURITY AGREEMENT Recorded Jul 31, 2019
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 049916/0454 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2019
From: SUNDARAM, RAMASUBRAMANIAN; GANAPATHIRAJU, ARAVIND; TAN, YINGYI
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 049206/0105 →
Continuity (2)
Provisional Application 62673505 · May 18, 2018
Related Publication 20190355348A1 · Nov 21, 2019