IP Library Granted Patent US 9,026,444
Granted Patent B2
US 9,026,444 · App. 12/561,005 · Granted May 5, 2015

System and method for personalization of acoustic models for automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,026,444
App. No.
12/561,005
Granted
May 5, 2015
Kind
B2
Abstract

Disclosed herein are methods, systems, and computer-readable storage media for automatic speech recognition. The method includes selecting a speaker independent model, and selecting a quantity of speaker dependent models, the quantity of speaker dependent models being based on available computing resources, the selected models including the speaker independent model and the quantity of speaker dependent models. The method also includes recognizing an utterance using each of the selected models in parallel, and selecting a dominant speech model from the selected models based on recognition accuracy using the group of selected models. The system includes a processor and modules configured to control the processor to perform the method. The computer-readable storage medium includes instructions for causing a computing device to perform the steps of the method.

Claims (38)

1. A method comprising:

starting an automatic speech recognition session for a phone call initiated from a communication device;

determining that the communication device is associated with a plurality of users;

identifying a group of selected speech recognition models comprising a speaker independent model and a plurality of speaker dependent models;

recognizing an utterance received from a particular user in the plurality of users using each model in the group of selected speech recognition models in parallel, to yield a group of recognition results;

selecting a dominant speech model from the group of selected speech recognition models using a heuristic search algorithm, to yield a selected dominant speech model, wherein the dominant speech model is an efficient model in the group of selected speech recognition models; and

continuously using the selected dominant speech model to recognize speech received from the particular user for a remainder of the automatic speech recognition session.

2. The method of claim 1 , further comprising dropping a speech model from the group of selected speech recognition models when recognition accuracy is below a threshold.

3. The method of claim 1 , further comprising selecting the plurality of speaker dependent models based on the communication device.

4. The method of claim 3 , further comprising selecting the plurality of speaker dependent models based on the plurality of users associated with the communication device.

5. The method of claim 3 , further comprising receiving additional utterances from the communication device and clustering the additional utterances to generate a new speaker dependent model.

6. The method of claim 1 , further comprising iteratively generating the group of selected models, recognizing the utterance, and selecting the dominant speech model, each time a new automatic speech recognition session is initiated.

7. A system comprising:

a processor;

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

starting an automatic speech recognition session for a phone call initiated from a communication device;

determining that the communication device is associated with a plurality of users;

identifying a group of selected speech recognition models comprising a speaker independent model and a plurality of speaker dependent models;

recognizing an utterance received from a particular user in the plurality of users using each model in the group of selected speech recognition models in parallel, to yield a group of recognition results;

selecting a dominant speech model from the group of selected speech recognition models using a heuristic search algorithm, to yield a selected dominant speech model, wherein the dominant speech model is an efficient model in the group of selected speech recognition models; and

continuously using the selected dominant speech model to recognize speech received from the particular user for a remainder of the automatic speech recognition session.

8. The system of claim 7 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising dropping a speech model from the group of selected speech recognition models when recognition accuracy is below a threshold.

9. The system of claim 7 , the computer-readable storage medium having additional instructions which, when executed by the processor, result in operations comprising selecting the plurality of speaker dependent models based on the communication device.

10. The system of claim 9 , the computer-readable storage medium having additional instructions which, when executed by the processor, result in operations comprising selecting the plurality of speaker dependent models based on the plurality of users associated with the communication device.

11. The system of claim 9 , the computer-readable storage medium having additional instructions which, when executed by the processor, result in operations comprising:

receiving utterances from the communication device; and

clustering the utterances to generate a new speaker dependent model.

12. The system of claim 7 , the computer-readable storage medium storing additional instructions which, when executed by the processor, result in operations comprising iteratively generating the group of selected speech recognition models, recognizing the utterance, and selecting the dominant speech model, each time a new automatic speech recognition session is initiated.

13. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

starting an automatic speech recognition session for a phone call initiated from a communication device;

determining that the communication device is associated with a plurality of users;

identifying a group of selected speech recognition models comprising a speaker independent model and a plurality of speaker dependent models;

recognizing an utterance received from a particular user in the plurality of users using each model in the group of selected speech recognition models in parallel, to yield a group of recognition results;

selecting a dominant speech model from the group of selected speech recognition models using a heuristic search algorithm, to yield a selected dominant speech model, wherein the dominant speech model is an efficient model in the group of selected speech recognition models; and

continuously using the selected dominant speech model to recognize speech received from the particular user for a remainder of the automatic speech recognition session.

14. The computer-readable storage device of claim 13 , having additional instructions stored which, when executed by the computing device, result in operations comprising dropping a speech model from the group of selected speech recognition models when recognition accuracy is below a threshold.

15. The computer-readable storage device of claim 13 , having additional instructions stored which, when executed by the computing device, result in operations comprising:

selecting the plurality of speaker dependent models based on the plurality of users associated with the communication device.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065532/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →