IP Library Granted Patent US 7,835,910
Granted Patent B1
US 7,835,910 · App. 10/448,415 · Granted Nov 16, 2010

Exploiting unlabeled utterances for spoken language understanding

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,835,910
App. No.
10/448,415
Granted
Nov 16, 2010
Kind
B1
Abstract

A system and method for exploiting unlabeled utterances in the augmentation of a classifier model is disclosed. In one embodiment, a classifier is initially trained using a labeled set of utterances. Another set of utterances is then selected from an available set of unlabeled utterances. In one embodiment, this selection process can be based on a confidence score threshold. The trained classifier is then augmented using the selected set of unlabeled utterances.

Claims (146)

1. A method for training a speech classifier in a spoken language understanding system using an available set of utterances associated with a single kind of information, the available set of utterances including a first labeled part and a second unlabeled part, the method comprising:

(1) training via a processor a speech classifier for a spoken language understanding system with the first labeled part of the available set of utterances;

(2) selecting via the processor a portion of the second unlabeled part of the available set of utterances; and

(3) augmenting via the processor the trained speech classifier with the selected portion of the second unlabeled part of the available set of utterances, wherein the augmenting further comprises minimizing a loss function.

2. The method of claim 1 , wherein said training comprises training the speech classifier with the first labeled part of the available set of utterances, wherein said first labeled part includes utterances that have been labeled manually.

3. The method of claim 1 , wherein said selecting comprises selecting a portion of the second unlabeled part of the available set of utterances based on confidence scores.

4. The method of claim 3 , wherein said selecting comprises selecting a portion of the second unlabeled part of the available set of utterances based on confidence scores that are greater than a threshold value.

5. The method of claim 3 , wherein the first labeled part of the available set of utterances are selected based on confidence scores that are less than a threshold value, and wherein said selecting comprises selecting the remaining portion of the available set of utterances after the first labeled part of the available set of utterances is identified.

6. The method of claim 1 , wherein said augmenting comprises augmenting the trained speech classifier with the selected portion of the second unlabeled part of the available set of utterances in a weighted manner.

7. The method of claim 1 , wherein the loss function further comprises:

i

(

ln

(

1

+

-

y

i

f

(

x

i

)

)

+

η

KL

(

P

(

.

|

x

i

)

ρ

(

f

(

x

i

)

)

)

)

where

KL( p∥q )= p ln( p/q )+(1− p )ln((1− p )/(1− q ))

is the Kullback-Leibler divergence between two probability distributions p and q, which correspond to the distribution from a prior model, P(.|x i ), and the distribution from a constructed model, ρ(f(x i )), respectively.

8. An unsupervised learning method for training a speech classifier in a spoken language understanding system from a single kind of information, the method comprising:

(1) receiving a collection of unlabeled utterance data that has been collected by a natural language dialog system;

(2) generating confidence values for the received collection of unlabeled utterance data;

(3) assembling a group of utterances from the received collection of unlabeled utterance data that has a confidence value indicative of usefulness in augmenting the speech classifier; and

(4) augmenting the speech classifier with the assembled group of utterances, wherein the augmenting further comprises minimizing a loss function.

9. The method of claim 8 , wherein said receiving comprises receiving raw utterance data.

10. The method of claim 8 , wherein said assembling comprises assembling a group of utterances from the received collection of unlabeled utterance data that has a confidence value greater than a threshold value.

11. The method of claim 8 , further comprising identifying a set of utterances that are candidates for manual labeling, wherein said assembling comprises assembling the remainder of the received collection of unlabeled utterance data that has not been identified for manual labeling.

12. The method of claim 8 , wherein said augmenting comprises augmenting the speech classifier with the assembled group of utterances in a weighted manner.

13. The method of claim 8 , wherein the loss function further comprises:

i

(

ln

(

1

+

-

y

i

f

(

x

i

)

)

+

η

KL

(

P

(

.

x

i

)

ρ

(

f

(

x

i

)

)

)

)

where

KL( p∥q )= p ln( p/q )+(1− p )ln((1− p )/(1− q ))

is the Kullback-Leibler divergence between two probability distributions p and q, which correspond to the distribution from a prior model, P(.|x i ), and the distribution from a constructed model, ρ(f(x i )), respectively.

14. A tangible computer-readable medium that stores instructions for controlling a computer device to train a speech classifier in a spoken language understanding system using an available set of utterances associated with a single kind of information, the available set of utterances including a first labeled part and a second unlabeled part, the instructions comprising:

(1) training a speech classifier for a spoken language understanding system with the first labeled part of the available set of utterances;

(2) selecting a portion of the second unlabeled part of the available set of utterances; and

(3) augmenting the trained speech classifier with the selected portion of the second unlabeled part of the available set of utterances, wherein the augmenting further comprises minimizing a loss function.

15. A tangible computer-readable medium that stores instructions for controlling a computer device to train a speech classifier for a natural language dialog system using utterances associated with a single kind of information, the instructions comprising:

(1) receiving a collection of unlabeled utterance data that has been collected by a natural language dialog system;

(2) generating confidence values for the received collection of unlabeled utterance data;

(3) assembling a group of utterances from the received collection of unlabeled utterance data that has a confidence value indicative of usefulness in augmenting the speech classifier; and

(4) augmenting the speech classifier with the assembled group of utterances, wherein the augmenting further comprises minimizing a loss function.

16. A spoken language understanding module that uses a method for training a speech classifier using an available set of utterances associated with a single kind of information, the available set of utterances including a first labeled part and a second unlabeled part, the method comprising:

(1) training a speech classifier for a spoken language understanding module with the first labeled part of the available set of utterances;

(2) selecting a portion of the second unlabeled part of the available set of utterances; and

(3) augmenting the trained speech classifier with the selected portion of the second unlabeled part of the available set of utterances, wherein the augmenting further comprises minimizing a loss function.

17. A spoken language understanding module that uses an unsupervised learning method for training a speech classifier from utterances associated with a single kind of information, the method comprising:

(1) receiving a collection of unlabeled utterance data that has been collected by a natural language dialog system;

(2) generating confidence values for the received collection of unlabeled utterance data;

(3) assembling a group of utterances from the received collection of unlabeled utterance data that has a confidence value indicative of usefulness in augmenting the speech classifier; and

(4) augmenting the speech classifier with the assembled group of utterances, wherein the augmenting further comprises minimizing a loss function.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038275/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038275/0130 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2003
From: HAKKANI-TUR, DILEK Z.; TUR, GOKHAN
To: AT&T CORP.
Reel/Frame 014131/0150 →