IP Library › Granted Patent US 8,892,439
Granted Patent B2
US 8,892,439 · App. 12/503,191 · Granted Nov 18, 2014

Combination and federation of local and remote speech recognition

Inventors: Julian J. Odell (Redmond, WA); Robert L. Chambers (Sammamish, WA)
Assignee: Microsoft Corporation
G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,892,439
App. No.
12/503,191
Granted
Nov 18, 2014
Kind
B2
Abstract

Techniques to provide automatic speech recognition at a local device are described. An apparatus may include an audio input to receive audio data indicating a task. The apparatus may further include a local recognizer component to receive the audio data, to pass the audio data to a remote recognizer while receiving the audio data, and to recognize speech from the audio data. The apparatus may further include a federation component operative to receive one or more recognition results from the local recognizer and/or the remote recognizer, and to federate a plurality of recognition results to produce a most likely result. The apparatus may further include an application to perform the task indicated by the most likely result. Other embodiments are described and claimed.

Claims (50)

1. An article comprising a computer-readable storage device and containing instructions that if executed enable a computer to:

receive audio data indicating a task at a local device;

pass the audio data to a local recognizer on the local device;

perform speech recognition on the audio data with the local recognizer;

based upon network bandwidth, delay passing the audio data to the remote recognizer until it is determined that the local recognizer cannot complete the speech recognition;

receive a recognition result from at least one of the local and the remote recognizers; and

perform the task indicated by the recognition result.

2. The article of claim 1 , further comprising instructions that if executed enable the computer to:

federate the recognition results to produce a most likely result; and

perform the task indicated by the most likely result.

3. The article of claim 2 , further comprising instructions that if executed enable the computer to:

display at least one recognition result; and

receive an operator selection of the most likely result from the displayed recognition result.

4. The article of claim 3 , further comprising instructions that if executed enable the computer to provide feedback to at least one of the local and remote recognizers in response to the operator selection and to update at least one of a usage model or an acoustic model for at least one of the local or the remote recognizer based on the feedback.

5. The article of claim 1 , further comprising instructions that if executed enable the computer to pass the audio data to both the local recognizer and the remote recognizer substantially simultaneously.

6. The article of claim 1 , further comprising instructions that if executed enable the computer to stop passing audio to the remote recognizer when the local recognizer produces a substantially unambiguous result.

7. A computer-implemented method, comprising:

receiving audio data indicating a task at a local device;

passing the audio data to a local recognizer on the local device;

recognizing speech from the audio data with the local recognizer;

in response to limited or unavailable network bandwidth, delaying passing the audio data to a remote recognizer it is determined that the local recognizer cannot complete the speech recognition;

receiving a recognition result from at least one of the local and the remote recognizers; and

displaying the recognition result.

8. The method of claim 7 , comprising:

federating the recognition results to produce a most likely result; and

performing the task indicated by the most likely result.

9. The method of claim 8 , comprising:

displaying a plurality of recognition results; and

receiving an operator selection of the most likely result from the displayed recognition results.

10. The method of claim 9 , comprising:

providing feedback to at least one of the local and remote recognizers in response to the operator selection; and

updating at least one of a usage model or an acoustic model for at least one of the local or the remote recognizer based on the feedback.

11. The method of claim 7 , comprising:

passing the audio data to both the local recognizer and the remote recognizer substantially simultaneously.

12. The method of claim 7 , comprising:

stopping the passing of audio data to the remote recognizer when the local recognizer produces a substantially unambiguous result.

13. An apparatus, comprising:

an audio input operative to receive audio data indicating a task;

a local recognizer component to receive the audio data, to recognize speech from the audio data, to pass the audio data to both a local recognizer and a remote recognizer substantially simultaneously;

in response to limited or unavailable network bandwidth, delay passing the audio data to the remote recognizer it is determined that the local recognizer cannot complete the speech recognition;

a federation component operative to receive one or more recognition results from at least one of the local and the remote recognizers and to federate a plurality of recognition results to produce a most likely result; and

an application to perform the task indicated by the most likely result as indicated by the federation component.

14. The apparatus of claim 13 , further comprising a display, and wherein the federation component is operative to display at least one recognition result, and to receive an operator selection of the most likely result from the displayed recognition result.

15. The apparatus of claim 14 , the federation component operative to provide feedback to at least one of the local and remote recognizers in response to the operator selection.

16. The apparatus of claim 15 , comprising:

updating at least one of a usage model or an acoustic model for at least one of the local or the remote recognizer based on the feedback.

17. The apparatus of claim 13 , further comprising:

a local grammar operative to be used by the local recognizer to complete a recognition task at the apparatus; and

a container grammar operative to be used by the local recognizer to partially complete a recognition task at the apparatus.

18. The apparatus of claim 17 , wherein when a recognition task is not complete according to the local grammar and is partially complete according to the container grammar, the federation component waits for a recognition result from the remote recognizer.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034564/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2009
From: ODELL, JULIAN J.; CHAMBERS, ROBERT L.
To: MICROSOFT CORPORATION
Reel/Frame 022957/0489 →
Continuity (1)
Related Publication 20110015928A1 · Jan 20, 2011