IP Library Granted Patent US 7,292,976
Granted Patent B1
US 7,292,976 · App. 10/447,888 · Granted Nov 6, 2007

Active learning process for spoken dialog systems

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,292,976
App. No.
10/447,888
Granted
Nov 6, 2007
Kind
B1
Abstract

A large amount of human labor is required to transcribe and annotate a training corpus that is needed to create and update models for automatic speech recognition (ASR) and spoken language understanding (SLU). Active learning enables a reduction in the amount of transcribed and annotated data required to train ASR and SLU models. In one aspect of the present invention, an active learning ASR process and active learning SLU process are coupled, thereby enabling further efficiencies to be gained relative to a process that maintains an isolation of data in both the ASR and SLU domains.

Claims (51)

1. A method in a spoken dialog system, comprising:

(1) storing transcription data that is generated by a transcription of a first set of utterances in a data store, the transcription data being used for generation of a first model that is used in an automatic speech recognition process;

(2) selecting a second set of utterances for annotation, wherein annotated utterances are used for generation of a second model that is used in a spoken language understanding process;

(3) determining whether transcription data for a chosen utterance in said selected second set of utterances is included in the data store; and

(4) retrieving transcription data from the data store if it is determined that transcription data for the chosen utterance is included in the data store.

2. The method of claim 1 , further comprising selecting the first set of utterances from an available set of utterances.

3. The method of claim 2 , wherein selecting the first set of utterances is performed using an active learning process.

4. The method of claim 1 , wherein selecting the second set of utterances is performed using an active learning process.

5. The method of claim 1 , further comprising storing annotation data in a second data store.

6. A method in a spoken dialog system, comprising:

(1) ranking a first set of audio files for transcription using an active learning automatic speech recognition process;

(2) ranking a second set of audio files for annotation using an active learning spoken language understanding process, wherein the second set of audio files includes at least one audio file in common with the first set of audio files;

(3) providing a first list of top ranked audio files to a lab for transcription, the first list of top ranked audio files corresponding to the first set ranking; and

(4) providing a second list of top ranked audio files to the lab for annotation, the second list of top ranked audio files corresponding to the second set ranking,

wherein an audio file is deleted from the first list of top ranked audio files if it is also included in the second list of top ranked audio files.

7. The method of claim 6 , wherein the first set of audio files is identical to the second set of audio files.

8. The method of claim 6 , wherein the first list of top ranked audio files includes the top N ranked audio files of the first set.

9. The method of claim 6 , wherein the second list of top ranked audio files includes the top N ranked audio files of the second set.

10. The method of claim 6 , wherein the first list and the second list are mutually exclusive.

11. The method of claim 6 , further comprising training a model for the automatic speech recognition process using transcriptions of the audio files in the first list.

12. The method of claim 6 , further comprising training a model for the spoken language understanding process using annotations of the audio files in the second list.

13. A method in a spoken dialog system, comprising:

(1) training a model for an automatic speech recognition process using transcriptions of audio files included in a first list;

(2) training a model for a spoken language understanding process using transcriptions of audio files included in a second list;

(3) ranking a first set of audio files for transcription using the active learning automatic speech recognition process;

(4) ranking a second set of audio files for annotation using the active learning spoken language understanding process;

(5) providing a third list of top ranked audio files to a lab for transcription; and

(6) providing a fourth list of top ranked audio files to a lab for annotation,

wherein the size of the third list is adjusted relative to the size of the first list and the size of the fourth list is adjusted relative to the size of the second list upon a determination that one of the models for automatic speech recognition and spoken language understanding requires a greater level of training.

14. The method of claim 13 , wherein if the automatic speech recognition model requires a greater level of training, the size of the third list is increased relative to the size of the first list, and the size of the fourth list is decreased relative to the size of the second list.

15. The method of claim 13 , wherein if the spoken language understanding model requires a greater level of training, the size of the third list is decreased relative to the size of the first list, and the size of the fourth list is increased relative to the size of the second list.

16. The method of claim 13 , wherein the third list and the fourth list are mutually exclusive.

17. A computer-readable medium that stores a program for controlling a computer device to perform a method in a spoken dialog system, the method comprising:

(1) storing transcription data that is generated by a transcription of a first set of utterances in a data store, the transcription data being used for generation of a first model that is used in an automatic speech recognition process;

(2) selecting a second set of utterances for annotation, wherein annotated utterances are used for generation of a second model that is used in a spoken language understanding process;

(3) determining whether transcription data for a chosen utterance in said selected second set of utterances is included in the data store; and

(4) retrieving transcription data from the data store if it is determined that transcription data for the chosen utterance is included in the data store.

18. A computer-readable medium that stores a program for controlling a computer device to perform a method in a spoken dialog system, the method comprising:

(1) ranking a first set of audio files for transcription using an active learning automatic speech recognition process;

(2) ranking a second set of audio files for annotation using an active learning spoken language understanding process, wherein the second set of audio files includes at least one audio file in common with the first set of audio files;

(3) providing a first list of top ranked audio files to a lab for transcription, the first list of top ranked audio files corresponding to the first set ranking; and

(4) providing a second list of top ranked audio files to a lab for annotation, the second list of top ranked audio files corresponding to the second set ranking,

wherein an audio file is deleted from the first list of top ranked audio files if it is also included in the second list of top ranked audio files.

19. A computer-readable medium that stores a program for controlling a computer device to perform a method in a spoken dialog system, the method comprising:

(1) training a model for an automatic speech recognition process using transcriptions of audio files included in a first list;

(2) training a model for a spoken language understanding process using transcriptions of audio files included in a second list;

(3) ranking a first set of audio files for transcription using the active learning automatic speech recognition process;

(4) ranking a second set of audio files for annotation using the active learning spoken language understanding process;

(5) providing a third list of top ranked audio files to a lab for transcription; and

(6) providing a fourth list of top ranked audio files to the lab for annotation,

wherein the size of the third list is adjusted relative to the size of the first list and the size of the fourth list is adjusted relative to the size of the second list upon a determination that one of the models for automatic speech recognition and spoken language understanding requires a greater level of training.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038275/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038275/0130 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2003
From: HAKKANI-TUR, DILEK Z.; RAHIM, MAZIN G.; RICCARDI, GIUSEPPE; TUR, GOKHAN
To: AT&T CORP.
Reel/Frame 014124/0509 →