IP Library Granted Patent US 8,024,190
Granted Patent B2
US 8,024,190 · App. 12/414,587 · Granted Sep 20, 2011

System and method for unsupervised and active learning for automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,024,190
App. No.
12/414,587
Granted
Sep 20, 2011
Kind
B2
Abstract

A system and method is provided for combining active and unsupervised learning for automatic speech recognition. This process enables a reduction in the amount of human supervision required for training acoustic and language models and an increase in the performance given the transcribed and un-transcribed data.

Claims (51)

1. A method of training an automatic speech recognition module, the method comprising:

training, via a processor of a computing device, acoustic and language models using a set of transcribed data, speech recognition scores, and word confidence scores for a retrieved set of un-transcribed data, wherein the set of transcribed data, speech recognition scores, and word confidence scores is generated by steps comprising:

recognizing utterances in a set of candidates for transcription using acoustic and language models trained using an initial set of transcribed data;

computing confidence scores of the utterances;

selecting a subset of utterances that have the smallest confidence scores from the set of candidates and transcribing them into a transcribed set;

adding the transcribed set to the initial set of transcribed data to produce an updated set of transcribed data, speech recognition scores, and word confidence scores;

selecting a pre-determined amount of un-transcribed data, the un-transcribed data being associated with utterances not selected in the selecting of the subset of utterances; and

applying the pre-determined amount of un-transcribed data to train the acoustic and language models; and

iteratively performing the training if word accuracy has not converged.

2. The method of claim 1 , the method further comprising:

prior to the training, training the acoustic and language models using another set of transcribed data.

3. The method of claim 2 , further comprising:

recognizing utterances in the set of candidates for transcription using the acoustic and language models.

4. The method of claim 3 , further comprising:

computing by a processor confidence scores of the utterances.

5. The method of claim 4 , further comprising:

selecting k utterances that have smallest confidence scores from the set of candidates and transcribing the k utterances into a first additional transcribed set.

6. The method of claim 5 , wherein k is more than one.

7. The method of claim 5 , wherein selecting k utterances further comprises leaving out utterances with confidence scores indicating that the utterances were correctly recognized.

8. The method of claim 5 , further comprising:

adding the first additional transcribed set to the another set of transcribed data to produce the set of transcribed data.

9. The method of claim 8 , further comprising:

removing the first additional transcribed set from the set of candidates.

10. The method of claim 9 , wherein the retrieved set of un-transcribed data is retrieved from the set of candidates.

11. The method of claim 10 , further comprising selecting a sample of un-transcribed data.

12. The method of claim 1 , wherein word posterior probability estimates are used for word confidence scores associated with the utterances.

13. The method of claim 1 , wherein a word is considered to be correctly recognized if the word has a confidence score higher than a threshold value.

14. A tangible computer-readable medium that stores a program which, upon execution on a processor, causes the processor to train an automatic speech recognition module, the program comprising instructions for:

training acoustic and language models using a set of transcribed data, speech recognition scores and word confidence scores for a retrieved set of un-transcribed data, wherein the set of transcribed data, speech recognition scores, and word confidence scores is generated by steps comprising:

recognizing utterances in a set of candidates for transcription using acoustic and language models trained using an initial set of transcribed data;

computing confidence scores of the utterances;

selecting a subset of utterances that have the smallest confidence scores from the set of candidates and transcribing them into a transcribed set;

adding the transcribed set to the initial set of transcribed data to produce an updated set of transcribed data, speech recognition scores, and word confidence scores;

selecting a pre-determined amount of un-transcribed data, the un-transcribed data being associated with utterances not selected in the selecting of the subset of utterances; and

applying the pre-determined amount of un-transcribed data to train the acoustic and language models; and

iteratively performing the training if word accuracy has not converged.

15. The tangible computer-readable medium of claim 14 , the program further comprising instructions for, prior to the training, training the acoustic and language models using another set of transcribed data, recognizing utterances in a set of candidates for transcription using the acoustic and language models and computing by a processor confidence scores of the utterances.

16. The tangible computer-readable medium of claim 15 , the program further comprising instructions for selecting k utterances that have the smallest confidence scores from the set of candidates and transcribing them into a first additional transcribed set, adding the first additional transcribed set to the another set of transcribed data to produce the set of transcribed data and removing the first additional transcribed set from the set of candidates and wherein the retrieved set of un-transcribed data is retrieved from the set of candidates.

17. A spoken dialog system, the system comprising:

a processor;

an automatic-speech recognition module controlling the processor to perform automatic speech recognition, the automatic speech recognition module trained using a method of training an automatic speech recognition module and stored in a memory storage device, the method comprising:

training acoustic and language models using a set of transcribed data, speech recognition scores and word confidence scores for a retrieved set of un-transcribed data, wherein the set of transcribed data, speech recognition scores, and word confidence scores is generated by steps comprising:

recognizing utterances in a set of candidates for transcription using acoustic and language models trained using an initial set of transcribed data;

computing confidence scores of the utterances;

selecting a subset of utterances that have the smallest confidence scores from the set of candidates and transcribing them into a transcribed set;

adding the transcribed set to the initial set of transcribed data to produce an updated set of transcribed data, speech recognition scores, and word confidence scores;

selecting a pre-determined amount of un-transcribed data, the un-transcribed data being associated with utterances not selected in the selecting of the subset of utterances; and

applying the pre-determined amount of un-transcribed data to train the acoustic and language models; and

iteratively performing the training if word accuracy has not converged.

18. The spoken dialog system of claim 17 , wherein the method further comprises, prior to the training, training the acoustic and language models using another set of transcribed data, recognizing utterances in a set of candidates for transcription using the acoustic and language models and computing by a processor confidence scores of the utterances.

19. The spoken dialog system of claim 18 , wherein the method further comprises selecting k utterances that have the smallest confidence scores from the set of candidates and transcribing them into a first additional transcribed set, adding the first additional transcribed set to the another set of transcribed data to produce the set of transcribed data and removing the first additional transcribed set from the set of candidates and wherein the retrieved set of un-transcribed data is retrieved from the set of candidates.

Assignments (17)
RELEASE OF SECURITY INTEREST Recorded Sep 4, 2025
From: RUNWAY GROWTH FINANCE CORP., AS AGENT
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 072802/0931 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 27, 2022
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 060445/0733 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034482/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: HAKKANI-TUR, DILEK ZEYNEP; RICCARDI, GIUSEPPE
To: AT&T CORP.
Reel/Frame 034442/0884 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2014
From: HAKKANI-TUR, DILEK ZEYNEP; RICCADI, GIUSEPPE
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 033963/0393 →