IP Library Granted Patent US 8,990,084
Granted Patent B2
US 8,990,084 · App. 14/176,439 · Granted Mar 24, 2015

Method of active learning for automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,990,084
App. No.
14/176,439
Granted
Mar 24, 2015
Kind
B2
Abstract

State-of-the-art speech recognition systems are trained using transcribed utterances, preparation of which is labor-intensive and time-consuming. The present invention is an iterative method for reducing the transcription effort for training in automatic speech recognition (ASR). Active learning aims at reducing the number of training examples to be labeled by automatically processing the unlabeled examples and then selecting the most informative ones with respect to a given cost function for a human to label. The method comprises automatically estimating a confidence score for each word of the utterance and exploiting the lattice output of a speech recognizer, which was trained on a small set of transcribed data. An utterance confidence score is computed based on these word confidence scores; then the utterances are selectively sampled to be transcribed using the utterance confidence scores.

Claims (31)

1. A method comprising:

recognizing untranscribed audio utterances that are candidates for transcription using a trained acoustic model and a trained language model;

computing confidence scores associated with an accuracy of speech recognition of the untranscribed audio utterances;

transcribing an audio utterance from the untranscribed audio utterances, the audio utterance having a lowest confidence score in the confidence scores; and

removing the audio utterance which was transcribed from a database comprising the untranscribed audio utterances.

2. The method of claim 1 , further comprising iteratively repeating the transcribing with additional audio utterances from the untranscribed audio utterances until a word accuracy converges.

3. The method of claim 1 , wherein the audio utterance comprises a plurality of audio works.

4. The method of claim 1 , further comprising leaving out from consideration for transcription utterances with confidence scores indicating that the untranscribed audio utterances were correctly recognized.

5. The method of claim 1 , wherein word posterior probability estimates are used for the confidence scores associated with the untranscribed audio utterances.

6. The method of claim 5 , wherein a word is considered to be correctly recognized when an associated posterior probability is higher than a second threshold value.

7. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

recognizing untranscribed audio utterances that are candidates for transcription using a trained acoustic model and a trained language model;

computing confidence scores associated with an accuracy of speech recognition of the untranscribed audio utterances;

transcribing an audio utterance from the untranscribed audio utterances, the audio utterance having a lowest confidence score in the confidence scores; and

removing the audio utterance which was transcribed from a database comprising the untranscribed audio utterances.

8. The system of claim 7 , the computer-readable storage medium having additional instructions stored which result in operations comprising iteratively repeating the transcribing with additional audio utterances from the untranscribed audio utterances until a word accuracy converges.

9. The system of claim 7 , wherein the audio utterance comprises a plurality of audio works.

10. The system of claim 7 , the computer-readable storage medium having additional instructions stored which result in operations comprising leaving out from consideration for transcription utterances with confidence scores indicating that the untranscribed audio utterances were correctly recognized.

11. The system of claim 7 , wherein word posterior probability estimates are used for the confidence scores associated with the untranscribed audio utterances.

12. The system of claim 11 , wherein a word is considered to be correctly recognized when an associated posterior probability is higher than a second threshold value.

13. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

recognizing untranscribed audio utterances that are candidates for transcription using a trained acoustic model and a trained language model;

computing confidence scores associated with an accuracy of speech recognition of the untranscribed audio utterances;

transcribing an audio utterance from the untranscribed audio utterances, the audio utterance having a lowest confidence score in the confidence scores; and

removing the audio utterance which was transcribed from a database comprising the untranscribed audio utterances.

14. The computer-readable storage device of claim 13 , having additional instructions stored which result in operations comprising iteratively repeating the transcribing with additional audio utterances from the untranscribed audio utterances until a word accuracy converges.

15. The computer-readable storage device of claim 13 , wherein the audio utterance comprises a plurality of audio works.

16. The computer-readable storage device of claim 13 , having additional instructions stored which result in operations comprising leaving out from consideration for transcription utterances with confidence scores indicating that the untranscribed audio utterances were correctly recognized.

17. The computer-readable storage device of claim 16 , wherein word posterior probability estimates are used for the confidence scores associated with the untranscribed audio utterances.

Assignments (23)
RELEASE OF SECURITY INTEREST Recorded Sep 4, 2025
From: RUNWAY GROWTH FINANCE CORP., AS AGENT
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 072802/0931 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 043039/0808 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060557/0636 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 27, 2022
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 060445/0733 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 29, 2017
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 043039/0808 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
CORRECTIVE ASSIGNMENT TO CORRECT THE NAME OF ASSIGNOR PREVIOUSLY RECORDED AT REEL: 035962 FRAME: 0190. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST. Recorded Jul 21, 2015
From: ORIX VENTURES, LLC
To: INTERACTIONS LLC
Reel/Frame 036143/0704 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
RELEASE OF SECURITY INTEREST Recorded Jun 24, 2015
From: INTERACTIONS LLC
To: ARES CAPITAL CORPORATION
Reel/Frame 035962/0190 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2014
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034464/0574 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: GORIN, ALLEN LOUIS; HAKKANI-TUR, DILEK Z.; RICCARDI, GIUSEPPE
To: AT&T CORP.
Reel/Frame 034435/0172 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 034435/0278 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 034435/0378 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2014
From: GORIN, ALLEN LOUIS; HAKKANI-TUR, DILEK Z.; RICCARDI, GIUSEPPE
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 033361/0390 →