IP Library Granted Patent US 7,533,019
Granted Patent B1
US 7,533,019 · App. 10/742,854 · Granted May 12, 2009

System and method for unsupervised and active learning for automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,533,019
App. No.
10/742,854
Granted
May 12, 2009
Kind
B1
Abstract

A system and method is provided for combining active and unsupervised learning for automatic speech recognition. This process enables a reduction in the amount of human supervision required for training acoustic and language models and an increase in the performance given the transcribed and un-transcribed data.

Claims (57)

1. A method for reducing the transcription effort for training an automatic speech recognition module, the method comprising:

(1) training acoustic and language models using a first set of transcribed data;

(2) recognizing utterances in a set of candidates for transcription using the acoustic and language models;

(3) computing by a processor confidence scores of the utterances;

(4) selecting k utterances that have the smallest confidence scores from the set of candidates and transcribing them into a first additional transcribed set;

(5) adding the first additional transcribed set to the first set of transcribed data to produce a second set of transcribed data;

(6) removing the first additional transcribed set from the set of candidates;

(7) retrieving a set of un-transcribed data from the set of candidates;

(8) training the acoustic and language models using the second set of transcribed data and speech recognition and word confidence scores for the retrieved set of un-transcribed data; and

(9) returning to step (1) if word accuracy has not converged.

2. The method of claim 1 , wherein k is more than one.

3. The method of claim 1 , wherein step (4) comprises leaving out utterances with confidence scores indicating that the utterances were correctly recognized.

4. The method of claim 1 , wherein word posterior probability estimates are used for word confidence scores associated with the utterances.

5. The method of claim 1 , wherein a word is considered to be correctly recognized if it has a confidence score higher than a threshold value.

6. The method of claim 1 , wherein step (7) comprises selecting a sample of un-transcribed data.

7. A tangible computer-readable medium that stores a program for controlling a computer device to perform a method to reduce the transcription effort for training an automatic speech recognition module, the method comprising:

(1) training acoustic and language models using a first set of transcribed data;

(2) recognizing utterances in a set of candidates for transcription using the acoustic and language models;

(3) computing confidence scores of the utterances;

(4) selecting k utterances that have the smallest confidence scores from the set of candidates and transcribing them into a first additional transcribed set;

(5) adding the first additional transcribed set to the first set of transcribed data to produce a second set of transcribed data;

(6) removing the first additional transcribed set from the set of candidates;

(7) retrieving a set of un-transcribed data from the set of candidates;

(8) training the acoustic and language models using the second set of transcribed data and speech recognition and word confidence scores for the retrieved set of un-transcribed data; and

(9) returning to step (1) if word accuracy has not converged.

8. A method for training an automatic speech recognition module, the method comprising:

(1) generating a first set of transcribed data;

(2) recognizing by a processor utterances in the first set of transcribed data;

(3) augmenting said first set of transcribed data with a second set of transcribed data that include those utterances whose speech recognition has low confidence scores;

(4) retrieving a set of un-transcribed data;

(5) recognizing by a processor utterances in the set of un-transcribed data; and

(6) training one of an acoustic model and language model using the augmented set of transcribed data and speech recognition scores and word confidence scores for the retrieved set of un-transcribed data.

9. The method of claim 8 , further comprising returning to step (1) if word accurance has not converged.

10. The method of claim 8 , wherein step (4) comprises selecting a sample of un-transcribed data.

11. The method of claim 8 , wherein step (6) comprises training both said acoustic model and said language model using the augmented set of transcribed data and the set of un-transcribed data.

12. A tangible computer-readable medium that stores a program for controlling a computer device to perform a method to train an automatic speech recognition module, the method comprising:

(1) generating a first set of transcribed data;

(2) recognizing by a processor utterances in the first set of transcribed data;

(3) augmenting said first set of transcribed data with a second set of transcribed data that include those utterances whose speech recognition has low confidence scores;

(4) retrieving a set of un-transcribed data;

(5) recognizing by a processor utterances in the set of un-transcribed data; and

(6) training one of an acoustic model and a language model using the augmented set of transcribed data and speech recognition scores and word confidence scores for the retrieved set of un-transcribed data.

13. An automatic-speech recognition module trained using a method of reducing the transcription effort for training an automatic speech recognition module and stored in a memory storage device, the method comprising:

(1) generating a first set of transcribed data;

(2) recognizing by a processor utterances in the first set of transcribed data;

(3) augmenting said first set of transcribed data with a second set of transcribed data that include those utterances whose speech recognition has low confidence scores;

(4) retrieving a set of un-transcribed data;

(5) recognizing by a processor utterances in the set of un-transcribed data; and

(6) training one of an acoustic model and language model using the augmented set of transcribed data and speech recognition scores and word confidence scores for the retrieved set of un-transcribed data.

14. A spoken dialog system, comprising:

an automatic-speech recognition module trained using a method of reducing the transcription effort for training an automatic speech recognition module and stored in a memory storage device, the method comprising:

(1) generating a first set of transcribed data;

(2) recognizing by a processor utterances in the first set of transcribed data;

(3) augmenting said first set of transcribed data with a second set of transcribed data that include those utterances whose speech recognition has low confidence scores;

(4) retrieving a set of un-transcribed data;

(5) recognizing by a processor utterances in the set of un-transcribed data; and

(6) training one of an acoustic model and language model using the augmented set of transcribed data and speech recognition scores and word confidence scores for the retrieved set of un-transcribed data.

Assignments (19)
RELEASE OF SECURITY INTEREST Recorded Sep 4, 2025
From: RUNWAY GROWTH FINANCE CORP., AS AGENT
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 072802/0931 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 27, 2022
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 060445/0733 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034482/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2014
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 034448/0549 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2014
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 034447/0744 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2014
From: HAKKANI-TUR, DILEK ZEYNEP; RICCADI, GIUSEPPE
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 033963/0393 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2003
From: HAKKANI-TUR, DILEK ZEYNEP; RICCARDI, GIUSEPPE
To: AT&T CORP.
Reel/Frame 014843/0762 →