IP Library Granted Patent US 7,660,715
Granted Patent B1
US 7,660,715 · App. 10/756,669 · Granted Feb 9, 2010

Transparent monitoring and intervention to improve automatic adaptation of speech models

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,660,715
App. No.
10/756,669
Granted
Feb 9, 2010
Kind
B1
Abstract

A system and method to improve the automatic adaptation of one or more speech models in automatic speech recognition systems. After a dialog begins, for example, the dialog asks the customer to provide spoken input and it is recorded. If the speech recognizer determines it may not have correctly transcribed the verbal response, i.e., voice input, the invention uses monitoring and if necessary, intervention to guarantee that the next transcription of the verbal response is correct. The dialog asks the customer to repeat his verbal response, which is recorded and a transcription of the input is sent to a human monitor, i.e., agent or operator. If the transcription of the spoken input is correct, the human does not intervene and the transcription remains unmodified. If the transcription of the verbal response is incorrect, the human intervenes and the transcription of the misrecognized word is corrected. In both cases, the dialog asks the customer to confirm the unmodified and corrected transcription. If the customer confirms the unmodified or newly corrected transcription, the dialog continues and the customer does not hang up in frustration because most times only one misrecognition occurred. Finally, the invention uses the first and second customer recording of the misrecognized word or utterance along with the corrected or unmodified transcription to automatically adapt one or more speech models, which improves the performance of the speech recognition system.

Claims (66)

1. A method to retrain an automatic speech recognition system, in which automatic speech recognition system a plurality of speech models is stored, the method comprising:

(a) extracting, by the automatic speech recognition system, a first user utterance from a sampled first input voice stream received from a user in response to a query;

(b) selecting, by the automatic speech recognition system and based on the first user utterance, a first speech model from among the plurality of speech models, the first speech model producing a first tentative recognition result corresponding to the first user utterance;

(c) informing the user of the first tentative recognition result;

(d) determining, by the automatic speech recognition system, from the user's response whether the first tentative recognition result was correct;

(e) performing the following steps when the first tentative recognition result is not correct:

(i) requesting the user to repeat the response to the query;

(ii) extracting, by the automatic speech recognition system, a second user utterance from a sampled second input voice stream received from the user in response to the requesting step;

(iii) selecting, by the automatic speech recognition system and based on the second user utterance, a second speech model, different than the first speech model, the second speech model producing a second tentative recognition result corresponding to the second user utterance; and

(iv) determining, by a human operator, when the second speech model correctly corresponds to at least one of the first and second user utterances,

wherein the first and second speech models are selected from a plurality of speech models.

2. The method of claim 1 , wherein the plurality of speech models are developed from a large vocabulary stored in a speech recognition adaptation database; and wherein step (e) further comprises:

(v) retraining the first speech model using the second speech model.

3. The method of claim 1 , wherein the first tentative recognition result is at least one word.

4. The method of claim 1 , wherein, when the first tentative recognition result is correct, steps (i)-(v) are not performed.

5. The method of claim 1 , wherein the automatic speech recognition system provides a transcription of at least one of the first and second user utterances and wherein the informing step comprises:

converting, by a text-to-speech resource, the transcription into speech; and

communicating, by an interactive voice response unit, the speech to the user.

6. The method of claim 5 , wherein the determining, by a human operator, step comprises:

displaying the transcription to the human operator;

playing a recording of the at least one of the first and second user utterances to the human operator; and

selecting, by the human operator, a third speech model as correctly corresponding to the recording, based on the transcription and recording.

7. The method of claim 1 , further comprising:

selecting, by the human operator, a third speech model that correctly corresponds to the second user utterance, when the second speech model does not correctly correspond to the second user utterance.

8. The method of claim 7 , further comprising:

retraining at least one speech model using said third speech model.

9. A computer readable medium comprising processor executable instructions that, when executed, perform the steps of claim 1 .

10. A method to retrain an automatic speech recognition system, in which automatic speech recognition system a plurality of speech models is stored, the method comprising:

(a) extracting, by the automatic speech recognition system, a first user utterance from a first input voice stream from a user, the first user utterance being a response to a query;

(b) selecting, by the automatic speech recognition system, a first speech model, the first speech model producing a first tentative recognition result based on the first user utterance;

(c) determining, by the automatic speech recognition system, that the first tentative recognition result does not correctly characterize the first user utterance;

(d) selecting, by a human operator and based on at least one of the first user utterance and a second user utterance received from the user, a second speech model as correctly characterizing the first user utterance, the second speech model producing a second tentative recognition result; and

(e) retraining the first speech model using at least one of the first and second user utterances and the second tentative recognition result.

11. The method of claim 10 , wherein, when the first tentative recognition result correctly characterizes the first user utterance, not performing the selecting step (d).

12. The method of claim 10 , wherein the plurality of speech models are developed from a large vocabulary stored in a speech recognition adaptation database and wherein the determining step comprises:

(C1) informing the user of the first tentative recognition result; and

(C2) determining from the first user's response whether the first tentative recognition result correctly characterizes the first user utterance.

13. The method of claim 12 , further comprising before the human operator selecting step (d):

(f) requesting the first user to repeat the response to the query;

(g) extracting the second user utterance from a sampled second input voice stream received from the first user in response to the requesting step (e); and

(h) selecting the second speech model producing the second tentative recognition result corresponding to the second user utterance.

14. The method of claim 10 , wherein the automatic speech recognition system generates a transcription of the first user utterance and wherein the determining step (c) comprises:

(C1) converting, by a text-to-speech resource, the transcription into speech; and

(C2) communicating, by an interactive voice response unit, the speech to the user.

15. The method of claim 14 , wherein the selecting step (d) comprises:

(D1) displaying the transcription to the human operator;

(D2) playing a recording of the first user utterance to the human operator; and

(D3) selecting, by the human operator and based on the transcription and recording, a third speech model as correctly corresponding to the recording.

16. The method of claim 14 , wherein an adaptation agent is operable to provide an adaptation engine improved data to retrain at least one speech model, the improved data comprising said first user utterance and at least one of (i) a corrected transcription of said first user utterance when said human operator corrects said transcription; (ii) an unmodified transcription of said first user utterance when said human operator does not correct said transcription.

17. A computer readable medium comprising instructions that, when executed, perform the steps of claim 10 .

18. The method of claim 10 , wherein the first tentative recognition result is at least one word.

19. A speech recognition system comprising:

a speech recognition resource operable to extract a first user utterance from a first input voice stream from a user, the first user utterance being a response to a query; select a first speech model producing a first tentative recognition result characterizing the first user utterance; and

determine that the first tentative recognition result does not correctly characterize the first user utterance;

a model adaptation agent operable, when the first tentative recognition result does not correctly characterize the first user utterance, to alert a human operator, based on the first user utterance, to select a second speech model, different than the first speech model, to produce a second tentative recognition result correctly characterizing the first user utterance,

wherein the first and second speech models are selected from a plurality of speech models.

20. The system of claim 19 , further comprising:

an interactive voice response unit operable to inform the user of the first tentative recognition result and wherein the speech recognition resource is operable to determine from the first user's response whether the first tentative recognition result correctly characterizes the first user utterance; and further comprising:

an adaptation engine operable to retrain at least one speech model using at least the second tentative recognition result.

21. The system of claim 19 , further comprising:

an interactive voice response unit operable to request the first user to repeat the response to the query; and wherein the automatic speech recognition system is operable to extract a second user utterance from a sampled second input voice stream received from the first user in response to the request and select a third speech model to produce a third tentative recognition result corresponding to the second user utterance.

22. The system of claim 19 wherein, when the first tentative recognition result correctly characterizes the first user utterance, the adaptation engine does not alert the human operator.

23. The system of claim 19 , wherein the speech recognition resource generates a transcription of at the first user utterance and further comprising:

a text-to-speech resource operable to convert the transcription into speech; and

an interactive voice response unit operable to communicate the speech to the user.

24. The system of claim 23 , wherein the adaptation agent is operable to display the transcription to the human operator and play a recording of the first user utterance to the human operator and wherein the human operator, based on the transcription and recording, selects a third speech model as correctly corresponding to the recording.

Assignments (27)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2024
From: AVAYA LLC
To: ARLINGTON TECHNOLOGIES, LLC
Reel/Frame 067022/0780 →
INTELLECTUAL PROPERTY RELEASE AND REASSIGNMENT Recorded Mar 25, 2024
From: CITIBANK, N.A.
To: AVAYA LLC; AVAYA MANAGEMENT L.P.
Reel/Frame 066894/0117 →
INTELLECTUAL PROPERTY RELEASE AND REASSIGNMENT Recorded Mar 25, 2024
From: WILMINGTON SAVINGS FUND SOCIETY, FSB
To: AVAYA LLC; AVAYA MANAGEMENT L.P.
Reel/Frame 066894/0227 →
(SECURITY INTEREST) GRANTOR'S NAME CHANGE Recorded Sep 21, 2023
From: AVAYA INC.
To: AVAYA LLC
Reel/Frame 065019/0231 →
RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 53955/0436) Recorded May 18, 2023
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
To: AVAYA MANAGEMENT L.P.; AVAYA INC.; INTELLISIST, INC.; AVAYA INTEGRATED CABINET SOLUTIONS LLC
Reel/Frame 063705/0023 →
RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 61087/0386) Recorded May 18, 2023
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
To: AVAYA MANAGEMENT L.P.; AVAYA INC.; INTELLISIST, INC.; AVAYA INTEGRATED CABINET SOLUTIONS LLC
Reel/Frame 063690/0359 →
RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 045034/0001) Recorded May 18, 2023
From: GOLDMAN SACHS BANK USA., AS COLLATERAL AGENT
To: ZANG, INC. (FORMER NAME OF AVAYA CLOUD INC.); AVAYA INC.; INTELLISIST, INC.; AVAYA INTEGRATED CABINET SOLUTIONS LLC; OCTEL COMMUNICATIONS LLC; VPNET TECHNOLOGIES, INC.; HYPERQUALITY, INC.; HYPERQUALITY II, LLC; CAAS TECHNOLOGIES, LLC; AVAYA MANAGEMENT L.P.
Reel/Frame 063779/0622 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded May 4, 2023
From: AVAYA INC.; AVAYA MANAGEMENT L.P.; INTELLISIST, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 063542/0662 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded May 3, 2023
From: AVAYA MANAGEMENT L.P.; AVAYA INC.; INTELLISIST, INC.; KNOAHSOFT INC.
To: WILMINGTON SAVINGS FUND SOCIETY, FSB [COLLATERAL AGENT]
Reel/Frame 063742/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS AT REEL 45124/FRAME 0026 Recorded Apr 26, 2023
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: AVAYA HOLDINGS CORP.; AVAYA INC.; AVAYA MANAGEMENT L.P.; AVAYA INTEGRATED CABINET SOLUTIONS LLC
Reel/Frame 063457/0001 →
RELEASE OF SECURITY INTEREST ON REEL/FRAME 020166/0705 Recorded Aug 26, 2022
From: CITICORP USA, INC.
To: AVAYA, INC.; SIERRA HOLDINGS CORP.; AVAYA TECHNOLOGY, LLC; OCTEL COMMUNICATIONS LLC; VPNET TECHNOLOGIES, INC.
Reel/Frame 061328/0074 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Aug 5, 2022
From: AVAYA INC.; INTELLISIST, INC.; AVAYA MANAGEMENT L.P.; AVAYA CABINET SOLUTIONS LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 061087/0386 →
BANKRUPTCY COURT ORDER RELEASING THE SECURITY INTEREST RECORDED AT REEL/FRAME 020156/0149 Recorded Jul 25, 2022
From: CITIBANK, N.A., AS ADMINISTRATIVE AGENT
To: AVAYA, INC.; AVAYA TECHNOLOGY LLC; OCTEL COMMUNICATIONS LLC; VPNET TECHNOLOGIES
Reel/Frame 060953/0412 →
SECURITY INTEREST Recorded Sep 25, 2020
From: AVAYA INC.; AVAYA MANAGEMENT L.P.; INTELLISIST, INC.; AVAYA INTEGRATED CABINET SOLUTIONS LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 053955/0436 →
SECURITY INTEREST Recorded Jan 23, 2018
From: AVAYA INC.; AVAYA INTEGRATED CABINET SOLUTIONS LLC; OCTEL COMMUNICATIONS LLC; VPNET TECHNOLOGIES, INC.; ZANG, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 045124/0026 →
SECURITY INTEREST Recorded Jan 10, 2018
From: AVAYA INC.; AVAYA INTEGRATED CABINET SOLUTIONS LLC; OCTEL COMMUNICATIONS LLC; VPNET TECHNOLOGIES, INC.; ZANG, INC.
To: GOLDMAN SACHS BANK USA, AS COLLATERAL AGENT
Reel/Frame 045034/0001 →
BANKRUPTCY COURT ORDER RELEASING ALL LIENS INCLUDING THE SECURITY INTEREST RECORDED AT REEL/FRAME 025863/0535 Recorded Dec 15, 2017
From: THE BANK OF NEW YORK MELLON TRUST, NA
To: AVAYA INC.
Reel/Frame 044892/0001 →
BANKRUPTCY COURT ORDER RELEASING ALL LIENS INCLUDING THE SECURITY INTEREST RECORDED AT REEL/FRAME 041576/0001 Recorded Dec 15, 2017
From: CITIBANK, N.A.
To: AVAYA INC.; AVAYA INTEGRATED CABINET SOLUTIONS INC.; OCTEL COMMUNICATIONS LLC (FORMERLY KNOWN AS OCTEL COMMUNICATIONS CORPORATION); VPNET TECHNOLOGIES, INC.
Reel/Frame 044893/0531 →
BANKRUPTCY COURT ORDER RELEASING ALL LIENS INCLUDING THE SECURITY INTEREST RECORDED AT REEL/FRAME 030083/0639 Recorded Dec 15, 2017
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
To: AVAYA INC.
Reel/Frame 045012/0666 →
SECURITY INTEREST Recorded Jan 27, 2017
From: AVAYA INC.; AVAYA INTEGRATED CABINET SOLUTIONS INC.; OCTEL COMMUNICATIONS CORPORATION; VPNET TECHNOLOGIES, INC.
To: CITIBANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 041576/0001 →
SECURITY AGREEMENT Recorded Mar 13, 2013
From: AVAYA, INC.
To: BANK OF NEW YORK MELLON TRUST COMPANY, N.A., THE
Reel/Frame 030083/0639 →
SECURITY AGREEMENT Recorded Feb 22, 2011
From: AVAYA INC., A DELAWARE CORPORATION
To: BANK OF NEW YORK MELLON TRUST, NA, AS NOTES COLLATERAL AGENT, THE
Reel/Frame 025863/0535 →
CONVERSION FROM CORP TO LLC Recorded May 12, 2009
From: AVAYA TECHNOLOGY CORP.
To: AVAYA TECHNOLOGY LLC
Reel/Frame 022677/0550 →
REASSIGNMENT Recorded Jun 26, 2008
From: AVAYA TECHNOLOGY LLC; AVAYA LICENSING LLC
To: AVAYA INC
Reel/Frame 021156/0082 →
SECURITY AGREEMENT Recorded Nov 28, 2007
From: AVAYA, INC.; AVAYA TECHNOLOGY LLC; OCTEL COMMUNICATIONS LLC; VPNET TECHNOLOGIES, INC.
To: CITICORP USA, INC., AS ADMINISTRATIVE AGENT
Reel/Frame 020166/0705 →
SECURITY AGREEMENT Recorded Nov 27, 2007
From: AVAYA, INC.; AVAYA TECHNOLOGY LLC; OCTEL COMMUNICATIONS LLC; VPNET TECHNOLOGIES, INC.
To: CITIBANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 020156/0149 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2004
From: THAMBIRATNAM, DAVID PRESHAN
To: AVAYA TECHNOLOGY
Reel/Frame 014931/0565 →