IP Library Granted Patent US 10,699,702
Granted Patent B2
US 10,699,702 · App. 15/830,535 · Granted Jun 30, 2020

System and method for personalization of acoustic models for automatic speech recognition

Inventors: Andrej Ljolje (Morris Plains, NJ); Diamantino Antonio Caseiro (Philadelphia, PA); Alistair D. Conkie (San Jose, CA)
Assignee: NUANCE COMMUNICATIONS, INC.
G10L15/07G10L15/04G10L15/083G10L15/14G10L15/22G10L15/28G10L15/32G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,699,702
App. No.
15/830,535
Granted
Jun 30, 2020
Kind
B2
Abstract

Disclosed herein are methods, systems, and computer-readable storage media for automatic speech recognition. The method includes selecting a speaker independent model, and selecting a quantity of speaker dependent models, the quantity of speaker dependent models being based on available computing resources, the selected models including the speaker independent model and the quantity of speaker dependent models. The method also includes recognizing an utterance using each of the selected models in parallel, and selecting a dominant speech model from the selected models based on recognition accuracy using the group of selected models. The system includes a processor and modules configured to control the processor to perform the method. The computer-readable storage medium includes instructions for causing a computing device to perform the steps of the method.

Claims (36)

1. A method comprising:

recognizing received speech via each model in a group of speech recognition models, to yield recognition results;

selecting, based on the recognition results, a dominant speech model from the group of speech recognition models to yield a remainder set of dropped speech recognition models; and

continuously using only the dominant speech model, without applying the remainder set of dropped speech recognition models, to recognize additional speech received from a user, wherein the additional speech is separate from the received speech and wherein the additional speech is recognized only by the dominant speech model after the received speech is recognized by the group of speech recognition models.

2. The method of claim 1 , further comprising:

identifying, via a processor, the group of speech recognition models, wherein the group of speech recognition models comprises a speaker independent model and a speaker dependent model.

3. The method of claim 1 , wherein recognizing the received speech via each model in the group of speech recognition models is performed in parallel.

4. The method of claim 1 , wherein selecting the dominant speech model from the group of speech recognition models is performed using a heuristic search algorithm.

5. The method of claim 1 , wherein the additional speech received from the user is received during a remainder of an automatic speech recognition session.

6. The method of claim 1 , further comprising dropping a speech model from the group of speech recognition models when recognition accuracy is below a threshold.

7. The method of claim 1 , further comprising selecting at least one speech recognition model from the group of speech recognition models based on a device.

8. The method of claim 7 , further comprising selecting at least one speech recognition model from the group of speech recognition models based on a plurality of users associated with the device.

9. The method of claim 1 , further comprising receiving additional utterances from a device and clustering the additional utterances to generate a new speaker dependent model.

10. The method of claim 1 , further comprising iteratively generating a group of selected models, recognizing the received speech, and selecting the dominant speech model, each time a new automatic speech recognition session is initiated.

11. The method of claim 1 , wherein the group of speech recognition models is associated with a current location of a device for use in future speech dialogs.

12. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

recognizing received speech via each model in a group of speech recognition models, to yield recognition results;

selecting, based on the recognition results, a dominant speech model from the group of speech recognition models to yield a remainder set of dropped speech recognition models; and

continuously using only the dominant speech model, without applying the remainder set of dropped speech recognition models, to recognize additional speech received from a user, wherein the additional speech is separate from the received speech and wherein the additional speech is recognized only by the dominant speech model after the received speech is recognized by the group of speech recognition models.

13. The system of claim 12 , wherein the computer-readable storage medium stores additional instructions which, when executed by the processor, cause the processor to perform operations further comprising:

identifying the group of speech recognition models, wherein the group of speech recognition models comprises a speaker independent model and a speaker dependent model.

14. The system of claim 12 , wherein recognizing the received speech via each model in the group of speech recognition models is performed in parallel.

15. The system of claim 12 , wherein the computer-readable storage medium stores additional instructions which, when executed by the processor, cause the processor to perform operations further comprising:

selecting at least one speech recognition model from the group of speech recognition models based on a device.

16. The system of claim 12 , wherein the computer-readable storage medium stores additional instructions which, when executed by the processor, cause the processor to perform further operations comprising:

receiving additional utterances from a device and clustering the additional utterances to generate a new speaker dependent model.

17. A computer-readable storage device having instructions stored which, when executed by a computing device, result in the computing device performing operations comprising:

recognizing received speech via each model in a group of speech recognition models, to yield recognition results;

selecting, based on the recognition results, a dominant speech model from the group of speech recognition models to yield a remainder set of dropped speech recognition models; and

continuously using only the dominant speech model, without applying the remainder set of dropped speech recognition models, to recognize additional speech received from a user, wherein the additional speech is separate from the received speech and wherein the additional speech is recognized only by the dominant speech model after the received speech is recognized by the group of speech recognition models.

18. The computer-readable storage device of claim 17 , wherein the computer-readable storage device has additional instructions stored which, when executed by the computing device, result in the computing device performing operations further comprising:

identifying, via a processor, the group of speech recognition models, wherein the group of speech recognition models comprises a speaker independent model and a speaker dependent model.

19. The computer-readable storage device of claim 17 , wherein recognizing the received speech via each model in the group of speech recognition models is performed in parallel.

20. The computer-readable storage device of claim 17 , wherein selecting the dominant speech model from the group of speech recognition models is performed using a heuristic search algorithm.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065532/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2020
From: LJOLJE, ANDREJ; CASEIRO, DIAMANTINO ANTONIO; CONKIE, ALISTAIR D.
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 053585/0527 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2020
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 054156/0943 →