IP Library Granted Patent US 8,346,549
Granted Patent B2
US 8,346,549 · App. 12/631,131 · Granted Jan 1, 2013

System and method for supplemental speech recognition by identified idle resources

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,346,549
App. No.
12/631,131
Granted
Jan 1, 2013
Kind
B2
Abstract

Disclosed herein are systems, methods, and computer-readable storage media for improving automatic speech recognition performance. A system practicing the method identifies idle speech recognition resources and establishes a supplemental speech recognizer on the idle resources based on overall speech recognition demand. The supplemental speech recognizer can differ from a main speech recognizer, and, along with the main speech recognizer, can be associated with a particular speaker. The system performs speech recognition on speech received from the particular speaker in parallel with the main speech recognizer and the supplemental speech recognizer and combines results from the main and supplemental speech recognizer. The system recognizes the received speech based on the combined results. The system can use beam adjustment in place of or in combination with a supplemental speech recognizer. A scheduling algorithm can tailor a particular combination of speech recognition resources and release the supplemental speech recognizer based on increased demand.

Claims (46)

1. A method comprising:

projecting, via a processor, an expected demand for speech recognition resources;

establishing a plurality of speech recognizers based at least in part on the speech recognition resources and the expected demand;

identifying a main speech recognizer and a supplemental speech recognizer from the plurality of speech recognizers;

beginning to process a first recognition task using the main speech recognizer, to yield main results; and

upon determining that the supplemental speech recognizer is idle:

reallocating the supplemental speech recognizer to the first recognition task based on actual demand for speech recognition resources;

continuing to process the first recognition task using the supplemental speech recognizer, to yield supplemental results; and

combining the main results and the supplemental results.

2. The method of claim 1 , wherein a scheduling algorithm tailors a particular combination of the plurality of speech recognizers.

3. The method of claim 2 , wherein the scheduling algorithm releases the supplemental speech recognizer based on increased demand for speech recognition.

4. The method of claim 1 , the method further comprising establishing extra speech recognizers, tailored to speech which requires additional accuracy, upon determining that the supplemental speech recognizer is idle.

5. The method of claim 1 , the method further comprising identifying the supplemental speech recognizer as idle upon completing a second recognition task.

6. The method of claim 1 , wherein the plurality of speech recognizers comprises networked computing devices.

7. The method of claim 1 , further comprising utilizing at least one of word lattices and confusion networks to provide recognition strings as results.

8. The method of claim 1 , wherein each speech recognizer in the plurality of speech recognizers differs from each other in at least one of a spectral analysis in a front end, pronouncing dictionaries, and training algorithms.

9. A system comprising:

a processor;

a non-transitory computer-readable storage medium storing instructions which, when executed on the processor, performs a method comprising:

projecting, via a processor, an expected demand for speech recognition resources;

establishing a plurality of speech recognizers based at least in part on the speech recognition resources and the expected demand;

identifying a main speech recognizer and a supplemental speech recognizer from the plurality of speech recognizers;

beginning to process a first recognition task using the main speech recognizer, to yield main results; and

upon determining that the supplemental speech recognizer is idle:

reallocating the supplemental speech recognizer to the first recognition task based on actual demand for speech recognition resources;

continuing to process the first recognition task using the supplemental speech recognizer, to yield supplemental results; and

combining the main results and the supplemental results.

10. The system of claim 9 , wherein a scheduling algorithm tailors a particular combination of the plurality of speech recognizers.

11. The system of claim 10 , wherein the scheduling algorithm releases the supplemental speech recognizer based on increased demand for speech recognition.

12. The system of claim 9 , the non-transitory computer-readable storage medium storing additional instructions which, when executed on the processor, perform a step comprising establishing extra speech recognizers, tailored to speech which requires additional accuracy, upon determining that the supplemental speech recognizer is idle.

13. The system of claim 9 , the non-transitory computer-readable storage medium storing additional instructions which, when executed on the processor, perform a step comprising identifying the supplemental speech recognizer as idle upon completing a second recognition task.

14. The system of claim 9 , wherein the plurality of speech recognizers comprises networked computing devices.

15. The system of claim 9 , the non-transitory computer-readable storage medium storing additional instructions which, when executed on the processor, perform a step comprising utilizing at least one of word lattices and confusion networks to provide recognition strings as results.

16. The system of claim 9 , wherein each speech recognizer in the plurality of speech recognizers differs from each other in at least one of a spectral analysis in a front end, pronouncing dictionaries, and training algorithms.

17. A non-transitory computer-readable storage medium storing instructions which, when executed by a computing device, cause the computing device to perform a method comprising:

projecting, an expected demand for speech recognition resources;

establishing a plurality of speech recognizers based at least in part on the speech recognition resources and the expected demand;

identifying, a main speech recognizer and a supplemental speech recognizer from the plurality of speech recognizers;

beginning to process a first recognition task using the main speech recognizer, to yield main results; and

upon determining that the supplemental speech recognizer is idle:

reallocating the supplemental speech recognizer to the first recognition task based on actual demand for speech recognition resources;

continuing to perform the first recognition task using the supplemental speech recognizer, to yield supplemental results; and

combining the main results and the supplemental results.

18. The non-transitory computer-readable storage medium of claim 17 , wherein a scheduling algorithm tailors a particular combination of the plurality of speech recognizers.

19. The non-transitory computer-readable storage medium of claim 18 , wherein the scheduling algorithm releases the supplemental speech recognizer based on increased demand for speech recognition.

20. The non-transitory computer-readable storage medium of claim 17 , storing additional instructions which, when executed by the computing device, cause the computing device to perform a step comprising establishing extra speech recognizers, tailored to speech which requires additional accuracy, upon determining that the supplemental speech recognizer is idle.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2009
From: LJOLJE, ANDREJ; GILBERT, MAZIN
To: AT&T INTELLECTUAL PROPERTY I, LP
Reel/Frame 023606/0522 →