IP Library Granted Patent US 9,147,394
Granted Patent B2
US 9,147,394 · App. 14/551,739 · Granted Sep 29, 2015

System and method for unsupervised and active learning for automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,147,394
App. No.
14/551,739
Granted
Sep 29, 2015
Kind
B2
Abstract

A system and method is provided for combining active and unsupervised learning for automatic speech recognition. This process enables a reduction in the amount of human supervision required for training acoustic and language models and an increase in the performance given the transcribed and un-transcribed data.

Claims (46)

1. A computer-implemented method performed by a processor, the method comprising:

identifying, in a database of utterances, transcribed utterances and un-transcribed utterances;

ordering, via the processor, transcription candidate utterances from the un-transcribed utterances based on confidence scores of the transcription candidate utterances, to yield a selectively sampled order;

transcribing, via the processor, a top n utterances from the selectively sampled order, to yield additional transcribed utterances and remainder un-transcribed utterances, wherein the remainder un-transcribed utterances are the un-transcribed utterances without the additional transcribed utterances;

receiving human-transcribed utterances, wherein the human-transcribed utterances are selected from a bottom k utterances from the selectively sampled order for human transcription based on the confidence scores;

adding the additional transcribed utterances and the human-transcribed utterances to the database of utterances; and

training acoustic and language models using the additional transcribed utterances and the human-transcribed utterances.

2. The method of claim 1 , further comprising continuing the identifying, the transcribing, and the adding until a word error rate has converged.

3. The method of claim 2 , further comprising determining the confidence scores using an acoustic model.

4. The method of claim 2 , further comprising determining the confidence scores using a language model.

5. The method of claim 1 , further comprising:

upon adding the additional transcribed utterances to the database of utterances,

removing the additional utterances from the un-transcribed utterances.

6. The method of claim 1 , wherein the confidence scores of the transcription candidate utterances are associated with an arithmetic mean of confidences scores of words contained within each transcription candidate utterance.

7. The method of claim 1 , wherein the adding of the additional transcribed utterances and the human-transcribed utterances to the database of utterances is used to create an automatic speech recognition module.

8. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, result in the processor performing operations comprising:

identifying, in a database of utterances, transcribed utterances and untranscribed utterances;

ordering, via the processor, transcription candidate utterances from the un-transcribed utterances based on confidence scores of the transcription candidate utterances, to yield a selectively sampled order;

transcribing, via the processor, a top n utterances from the selectively sampled order, to yield additional transcribed utterances and remainder un-transcribed utterances, wherein the remainder un-transcribed utterances are the un-transcribed utterances without the additional transcribed utterances;

receiving human-transcribed utterances, wherein the human-transcribed utterances are selected from a bottom k utterances from the selectively sampled order for human transcription based on the confidence scores;

adding the additional transcribed utterances and the human-transcribed utterances to the database of utterances; and

training acoustic and language models using the additional transcribed utterances and the human-transcribed utterances.

9. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising continuing the identifying, the transcribing, and the adding until a word error rate has converged.

10. The system of claim 9 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising determining the confidence scores using an acoustic model.

11. The system of claim 9 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising determining the confidence scores using a language model.

12. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising:

upon adding the additional transcribed utterances to the database of utterances,

removing the additional utterances from the un-transcribed utterances.

13. The system of claim 8 , wherein the confidence scores of the transcription candidate utterances are associated with an arithmetic mean of confidences scores of words contained within each transcription candidate utterance.

14. The system of claim 8 , wherein the adding of the additional transcribed utterances and the human-transcribed utterances to the database of utterances is used to create an automatic speech recognition module.

15. A computer-readable storage device having instructions stored which, when executed by a computing device, result in the computing device performing operations comprising:

identifying, in a database of utterances, transcribed utterances and un-transcribed utterances;

ordering, via the processor, transcription candidate utterances from the un-transcribed utterances based on confidence scores of the transcription candidate utterances, to yield a selectively sampled order;

transcribing, via the processor, a top n utterances from the selectively sampled order, to yield additional transcribed utterances and remainder un-transcribed utterances, wherein the remainder un-transcribed utterances are the un-transcribed utterances without the additional transcribed utterances;

receiving human-transcribed utterances, wherein the human-transcribed utterances are selected from a bottom k utterances from the selectively sampled order for human transcription based on the confidence scores;

adding the additional transcribed utterances and the human-transcribed utterances to the database of utterances; and

training acoustic and language models using the additional transcribed utterances and the human-transcribed utterances.

16. The computer-readable storage device of claim 15 , having additional instructions stored which, when executed by the processor, result in operations comprising continuing the identifying, the transcribing, and the adding until a word error rate has converged.

17. The computer-readable storage device of claim 16 , having additional instructions stored which, when executed by the processor, result in operations comprising determining the confidence scores using an acoustic model.

18. The computer-readable storage device of claim 16 , having additional instructions stored which, when executed by the processor, result in operations comprising determining the confidence scores using a language model.

19. The computer-readable storage device of claim 15 , having additional instructions stored which, when executed by the processor, result in operations comprising:

upon adding the additional transcribed utterances to the database of utterances,

removing the additional utterances from the un-transcribed utterances.

20. The computer-readable storage device of claim 15 , wherein the confidence scores of the transcription candidate utterances are associated with an arithmetic mean of confidences scores of words contained within each transcription candidate utterance.

Assignments (20)
RELEASE OF SECURITY INTEREST Recorded Sep 4, 2025
From: RUNWAY GROWTH FINANCE CORP., AS AGENT
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 072802/0931 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 043039/0808 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060557/0636 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 27, 2022
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 060445/0733 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 29, 2017
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 043039/0808 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034482/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2014
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 034447/0744 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2014
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 034448/0549 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: HAKKANI-TUR, DILEK ZEYNEP; RICCARDI, GIUSEPPE
To: AT&T CORP.
Reel/Frame 034442/0884 →