IP Library › Granted Patent US 10,049,672
Granted Patent B2
US 10,049,672 · App. 15/171,374 · Granted Aug 14, 2018

Speech recognition with parallel recognition tasks

Inventors: Brian Patrick Strope (Palo Alto, CA); Francoise Beaufays (Mountain View, CA); Olivier Siohan (New York, NY)
Assignee: Google LLC
G10L15/32G10L15/00G10L15/01G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,049,672
App. No.
15/171,374
Granted
Aug 14, 2018
Kind
B2
Abstract

The subject matter of this specification can be embodied in, among other things, a method that includes receiving an audio signal and initiating speech recognition tasks by a plurality of speech recognition systems (SRS's). Each SRS is configured to generate a recognition result specifying possible speech included in the audio signal and a confidence value indicating a confidence in a correctness of the speech result. The method also includes completing a portion of the speech recognition tasks including generating one or more recognition results and one or more confidence values for the one or more recognition results, determining whether the one or more confidence values meets a confidence threshold, aborting a remaining portion of the speech recognition tasks for SRS's that have not generated a recognition result, and outputting a final recognition result based on at least one of the generated one or more speech results.

Claims (36)

1. A computer-implemented method comprising:

providing particular audio data to each automated speech recognizer of a set of automated speech recognizers;

before all of the automated speech recognizers have output a respective hypothesis for the particular audio data, determining that a particular automated speech recognizer of the set of automated speech recognizers has output a hypothesis for the particular audio data, and that a confidence value associated with the hypothesis that is output by the particular automated speech recognizer satisfies a particular confidence value threshold; and

while at least one of the automated speech recognizers that has been provided the particular audio data is indicated as not yet finished generating a respective hypothesis for the particular audio data, and in response to determining that the particular automated speech recognizer of the set of automated speech recognizers has output the hypothesis for the particular audio data, and that the confidence value associated with the hypothesis that is output by the particular automated speech recognizer satisfies the particular confidence value threshold:

providing the hypothesis that is output by the particular automated speech recognizer, of the set of automated speech recognizers, as a top speech recognition hypothesis; and

transmitting a command to stop the at least one of the automated speech recognizers of the set of automated speech recognizers that has been provided the particular audio data and that is indicated as not yet finished generating the respective hypothesis for the particular audio data from finishing generating the respective hypothesis for the particular audio data.

2. The method of claim 1 , wherein each automated speech recognizer of the set of automated speech recognizers uses a different one of a plurality of language models.

3. The method of claim 1 , wherein information that identifies the particular automated speech recognizer from the set of automated speech recognizers is provided with the hypothesis that is output by the particular automated speech recognizer.

4. The method of claim 2 , wherein the plurality of language models are each associated with a different one of a plurality of languages.

5. The method of claim 2 , wherein the language models were each generated based on a different one of a plurality of training procedures.

6. The method of claim 1 , wherein the top speech recognition hypothesis comprises a particular recognition result from multiple recognition results generated by the particular automated speech recognizer processing of the particular audio data.

7. A system comprising:

one or more computing devices;

an interface of the one or more computing devices that is programmed to receive an audio signal;

a set of automated speech recognizers; and

a processor that is configured to perform operations comprising:

providing particular audio data to each automated speech recognizer of a set of automated speech recognizers;

before all of the automated speech recognizers have output a respective hypothesis for the particular audio data, determining that a particular automated speech recognizer of the set of automated speech recognizers has output a hypothesis for the particular audio data, and that a confidence value associated with the hypothesis that is output by the particular automated speech recognizer satisfies a particular confidence value threshold; and

while at least one of the automated speech recognizers that has been provided the particular audio data is indicated as not yet finished generating a respective hypothesis for the particular audio data, and in response to determining that the particular automated speech recognizer of the set of automated speech recognizers has output the hypothesis for the particular audio data, and that the confidence value associated with the hypothesis that is output by the particular automated speech recognizer satisfies the particular confidence value threshold:

providing the hypothesis that is output by the particular automated speech recognizer, of the set of automated speech recognizers, as a top speech recognition hypothesis; and

transmitting a command to stop the at least one of the automated speech recognizers of the set of automated speech recognizers that has been provided the particular audio data and that is indicated as not yet finished generating the respective hypothesis for the particular audio data from finishing generating the respective hypothesis for the particular audio data.

8. The system of claim 7 , wherein each automated speech recognizer of the set of automated speech recognizers uses a different one of a plurality of language models.

9. The system of claim 7 , wherein information that identifies the particular automated speech recognizer from the set of automated speech recognizers is provided with the hypothesis that is output by the particular automated speech recognizer.

10. The system of claim 8 , wherein the plurality of language models are each associated with a different one of a plurality of languages.

11. The system of claim 8 , wherein the language models were each generated based on a different one of a plurality of training procedures.

12. The system of claim 7 , wherein the top speech recognition hypothesis comprises a particular recognition result from multiple recognition results generated by the particular automated speech recognizer processing of the particular audio data.

13. A non-transitory computer-readable medium storing instructions executable by one or more processors which, upon such execution, cause the one or more processors to perform operations comprising:

providing particular audio data to each automated speech recognizer of a set of automated speech recognizers;

before all of the automated speech recognizers have output a respective hypothesis for the particular audio data, determining that a particular automated speech recognizer of the set of automated speech recognizers has output a hypothesis for the particular audio data, and that a confidence value associated with the hypothesis that is output by the particular automated speech recognizer satisfies a particular confidence value threshold; and

while at least one of the automated speech recognizers that has been provided the particular audio data is indicated as not yet finished generating a respective hypothesis for the particular audio data, and in response to determining that the particular automated speech recognizer of the set of automated speech recognizers has output the hypothesis for the particular audio data, and that the confidence value associated with the hypothesis that is output by the particular automated speech recognizer satisfies the particular confidence value threshold:

providing the hypothesis that is output by the particular automated speech recognizer, of the set of automated speech recognizers, as a top speech recognition hypothesis; and

transmitting a command to stop the at least one of the automated speech recognizers of the set of automated speech recognizers that has been provided the particular audio data and that is indicated as not yet finished generating the respective hypothesis for the particular audio data from finishing generating the respective hypothesis for the particular audio data.

14. The computer-readable medium of claim 13 , wherein each automated speech recognizer of the set of automated speech recognizers uses a different one of a plurality of language models.

15. The computer-readable medium of claim 13 , wherein information that identifies the particular automated speech recognizer from the set of automated speech recognizers is provided with the hypothesis that is output by the particular automated speech recognizer.

16. The computer-readable medium of claim 14 , wherein the plurality of language models are each associated with a different one of a plurality of languages.

17. The computer-readable medium of claim 14 , wherein the language models were each generated based on a different one of a plurality of training procedures.

Assignments (2)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2016
From: STROPE, BRIAN PATRICK; BEAUFAYS, FRANCOISE; SIOHAN, OLIVIER
To: GOOGLE INC.
Reel/Frame 038787/0688 →
Continuity (4)
Continuation 14064755 · Oct 28, 2013
Continuation 13750807 · Jan 25, 2013
Continuation 12166822 · Jul 2, 2008
Related Publication 20160275951A1 · Sep 22, 2016