IP Library Granted Patent US 7,742,918
Granted Patent B1
US 7,742,918 · App. 11/773,681 · Granted Jun 22, 2010

Active learning for spoken language understanding

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,742,918
App. No.
11/773,681
Granted
Jun 22, 2010
Kind
B1
Abstract

Disclosed is a system and method of training a spoken language understanding module. Such a module may be utilized in a spoken dialog system. The method of training a spoken language understanding module comprises training acoustic and language models using a small set of transcribed data S t , recognizing utterances in a set S u that are candidates for transcription using the acoustic and language models, computing confidence scores of the utterances, selecting k utterances that have the smallest confidence scores from S u and transcribing them into a new set S i , redefining S t as the union of S t and S i , redefining S u as S u minus S i , and returning to the step of training acoustic and language models if word accuracy has not converged.

Claims (37)

1. A non-transitory computer-readable storage medium storing instructions for controlling a computing device to generate a classifier, the instructions comprising:

(1) training a classifier using current training data S t , the training data S t generated by sampling a plurality of utterances;

(2) classifying utterances in a pool S u using the trained classifier;

(3) computing a call type confidence score for each utterance;

(4) sorting candidate utterances with respect to the confidence score of the maximum scoring call type;

(5) selecting the lowest scored k utterances from S u using the confidence scores and labeling them to define a labeled set S i ;

(6) redefining S t =S t ∪S i ; and

(7) redefining S u =S u −S i .

2. The non-transitory computer-readable storage medium of claim 1 , wherein steps 1 through 7 are practiced until labelers and utterances are no longer available.

3. The non-transitory computer-readable storage medium of claim 1 , wherein k is more than one.

4. The non-transitory computer-readable storage medium of claim 1 , wherein selecting k utterances from S u further comprises leaving out utterances with confidence scores indicating that the utterances were correctly recognized.

5. The non-transitory computer-readable storage medium of claim 1 , wherein selecting k utterances from S u further comprises selecting the lowest scoring k utterances from S u .

6. The non-transitory computer-readable storage medium of claim 1 , wherein selecting k utterances from S u further comprises selecting utterances according to a confidence score distribution that is closest to a prior distribution.

7. A non-transitory computer-readable storage medium storing instructions for controlling a computing device to generate a spoken language understanding module, the instructions comprising, from a small amount of training data S t and a larger amount of unlabeled data S u :

(1) training a plurality of classifiers independently using a training data set S t , the training data S t generated by sampling a plurality of utterances;

(2) classifying utterances in a set S u using the plurality of classifiers and computing a call type confidence score for all utterances;

(3) sorting candidate utterances with respect to a score of the maximum scoring call type according to one of the classifiers if the classifiers disagree;

(4) selecting and labeling the lowest scored k utterances from S u to define a labeled set S i and redefining S t and S u as follows:

(5) S t =S t ∪S i ; and

(6) S u =S u −S i , wherein the labeled utterances are used to generate the spoken language understanding module.

8. The non-transitory computer-readable storage medium of claim 7 , wherein the steps occur only while labelers and utterances are available.

9. The non-transitory computer-readable storage medium of claim 7 , wherein selecting k utterances from S u further comprises selecting utterances according to a confidence score distribution that is closest to a prior distribution.

10. The non-transitory computer-readable storage medium of claim 7 , wherein selecting k utterances from S u further comprises selecting the lowest scoring k utterances from S u .

11. A method of generating a spoken dialog understanding module, the method causing a processor of a computing device to perform steps comprising, from a small amount of training data S t and a larger amount of unlabeled data S u :

classifying via the processor of the computing device utterances in an unlabelled data set S u using a plurality of classifiers;

computing via the processor of the computing device a call type confidence score for all utterances;

selecting utterances for labeling from the unlabeled data S u based on whether the classification from the plurality of classifiers disagree;

redefining S t =S t ∪a labeled set S i ;

redefining S u =S u −S i

labeling the selected utterances; and

generating a spoken language understanding module using the labeled utterances.

12. The method of claim 11 , wherein the selected utterances are the lowest scored k utterances from S u to the final label set S i , wherein the method further causes the processor of the computing device to perform steps comprising redefining S t and S u as follows:

S t =S t ÅS i ; and

S u =S u −S i , wherein the labeled utterances are used to generate the spoken language understanding module.

13. The method of claim 12 , wherein the steps occur only while labelers and utterances are available.

14. The method of claim 12 , wherein selecting k utterances from S u further comprises selecting utterances according to a confidence score distribution that is closest to a prior distribution.

15. The method of claim 12 , wherein selecting k utterances from S u further comprises selecting the lowest scoring k utterances from S u .

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2016
From: HAKKANI-TUR, DILEK Z.; SCHAPIRE, ROBERT ELIAS; TUR, GOKHAN
To: AT&T CORP.
Reel/Frame 038980/0949 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038983/0256 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038983/0386 →