IP Library Granted Patent US 9,666,182
Granted Patent B2
US 9,666,182 · App. 14/874,843 · Granted May 30, 2017

Unsupervised and active learning in automatic speech recognition for call classification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,666,182
App. No.
14/874,843
Granted
May 30, 2017
Kind
B2
Abstract

Utterance data that includes at least a small amount of manually transcribed data is provided. Automatic speech recognition is performed on ones of the utterance data not having a corresponding manual transcription to produce automatically transcribed utterances. A model is trained using all of the manually transcribed data and the automatically transcribed utterances. A predetermined number of utterances not having a corresponding manual transcription are intelligently selected and manually transcribed. Ones of the automatically transcribed data as well as ones having a corresponding manual transcription are labeled. In another aspect of the invention, audio data is mined from at least one source, and a language model is trained for call classification from the mined audio data to produce a language model.

Claims (37)

1. A method comprising:

performing, via a processor, automatic speech recognition using a bootstrap model on utterance data not having a corresponding manual transcription, to produce automatically transcribed utterances, wherein the bootstrap model is based on text data mined from a website relevant to a specific domain;

selecting, via the processor, a predetermined number of utterances not having a corresponding manual transcription based on a geometrically computed n-tuple confidence score; and

generating a language model based on the automatically transcribed utterances, the predetermined number of utterances, and transcriptions of the predetermined number of utterances.

2. The method of claim 1 , wherein the transcriptions of the predetermined number of utterances are made by a human being.

3. The method of claim 1 , further comprising:

performing additional automatic speech recognition using the language model.

4. The method of claim 2 , further comprising:

iteratively repeating the performing of the automatic speech recognition using the bootstrap model, the selecting, and the performing of additional speech recognition using the language model until a word accuracy converges.

5. The method of claim 1 , wherein the predetermined number of utterances correspond to a specific number of utterances having lowest confidence scores.

6. The method of claim 1 , wherein the predetermined number of utterances used in generating the language model are equal in number to the automatically transcribed utterances.

7. The method of claim 1 , wherein the predetermined number of utterances are randomly selected.

8. The method of claim 1 , wherein the language model is further based on the bootstrap model.

9. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

performing automatic speech recognition using a bootstrap model on utterance data not having a corresponding manual transcription, to produce automatically transcribed utterances, wherein the bootstrap model is based on text data mined from a web site relevant to a specific domain;

selecting a predetermined number of utterances not having a corresponding manual transcription based on a geometrically computed n-tuple confidence score; and

generating a language model based on the automatically transcribed utterances, the predetermined number of utterances, and transcriptions of the predetermined number of utterances.

10. The system of claim 9 , wherein the transcriptions of the predetermined number of utterances are made by a human being.

11. The system of claim 9 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

performing additional automatic speech recognition using the language model.

12. The system of claim 11 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

iteratively repeating the performing of the automatic speech recognition using the bootstrap model, the selecting, and the performing of additional speech recognition using the language model until a word accuracy converges.

13. The system of claim 9 , wherein the predetermined number of utterances correspond to a specific number of utterances having lowest confidence scores.

14. The system of claim 9 , wherein the predetermined number of utterances used in generating the language model are equal in number to the automatically transcribed utterances.

15. The system of claim 9 , wherein the predetermined number of utterances are randomly selected.

16. The system of claim 9 , wherein the language model is further based on the bootstrap model.

17. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

performing automatic speech recognition using a bootstrap model on utterance data not having a corresponding manual transcription, to produce automatically transcribed utterances, wherein the bootstrap model is based on text data mined from a website relevant to a specific domain;

selecting a predetermined number of utterances not having a corresponding manual transcription based on a geometrically computed n-tuple confidence score; and

generating a language model based on the automatically transcribed utterances, the predetermined number of utterances, and transcriptions of the predetermined number of utterances.

18. The computer-readable storage device of claim 17 , wherein the transcriptions of the predetermined number of utterances are made by a human being.

19. The computer-readable storage device of claim 17 , having additional instructions stored which, when executed by the computing device, cause the computing device to perform operations comprising:

performing additional automatic speech recognition using the language model.

20. The computer-readable storage device of claim 17 , having additional instructions stored which, when executed by the computing device, cause the computing device to perform operations comprising:

iteratively repeating the performing of the automatic speech recognition using the bootstrap model, the selecting, and the performing of additional speech recognition using the language model until a word accuracy converges.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065531/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2015
From: HAKKANI-TUR, DILEK Z.; RAHIM, MAZIN G.; RICCARDI, GIUSEPPE; TUR, GOKHAN
To: AT&T CORP.
Reel/Frame 036938/0113 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2015
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 037031/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2015
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 037032/0001 →