IP Library Granted Patent US 10,354,650
Granted Patent B2
US 10,354,650 · App. 13/838,379 · Granted Jul 16, 2019

Recognizing speech with mixed speech recognition models to generate transcriptions

Inventors: Alexander H. Gruenstein (Sunnyvale, CA); Petar Aleksic (Jersey City, NJ)
Assignee: Google LLC
G10L15/26G10L15/18G10L15/22G10L15/32G10L15/193G10L15/197G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,354,650
App. No.
13/838,379
Granted
Jul 16, 2019
Kind
B2
Abstract

In one aspect, a method comprises accessing audio data generated by a computing device based on audio input from a user, the audio data encoding one or more user utterances. The method further comprises generating a first transcription of the utterances by performing speech recognition on the audio data using a first speech recognizer that employs a language model based on user-specific data. The method further comprises generating a second transcription of the utterances by performing speech recognition on the audio data using a second speech recognizer that employs a language model independent of user-specific data. The method further comprises determining that the second transcription of the utterances includes a term from a predefined set of one or more terms. The method further comprises, based on determining that the second transcription of the utterance includes the term, providing an output of the first transcription of the utterance.

Claims (37)

1. A computer-implemented method comprising:

accessing audio data generated by a computing device based on audio input from a user, the audio data encoding one or more user utterances;

generating a first transcription of the utterances by performing speech recognition on the audio data using a first speech recognizer, wherein the first speech recognizer employs a language model that is based on user-specific data;

generating a second transcription of the utterances by performing speech recognition on the audio data using a second speech recognizer, wherein the second speech recognizer employs a language model independent of user-specific data;

determining that the second transcription of the utterances includes a term from a predefined set of one or more terms associated with actions that are performable by the computing device; and

based on determining that the second transcription of the utterance includes the term from the predefined set of one or more terms, providing an output of the first transcription of the utterance.

2. The method of claim 1 wherein the first speech recognizer employs a grammar-based language model.

3. The method of claim 2 wherein the grammar-based language model includes a context free grammar.

4. The method of claim 1 wherein the second speech recognizer employs a statistics-based language model.

5. The method of claim 1 wherein the user-specific data includes a contact list for the user, an applications list of applications installed on the computing device, or a media list of media stored on the computing device.

6. The method of claim 1 wherein the first speech recognizer is implemented on the computing device and the second speech recognizer is implemented on one or more server devices.

7. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

accessing audio data generated by a computing device based on audio input from a user, the audio data encoding one or more user utterances;

generating a first transcription of the utterances by performing speech recognition on the audio data using a first speech recognizer, wherein the first speech recognizer employs a language model that is developed based on user-specific data;

generating a second transcription of the utterances by performing speech recognition on the audio data using a second speech recognizer, wherein the second speech recognizer employs a language model developed independent of user-specific data;

determining that the second transcription of the utterances includes a term from a predefined set of one or more terms associated with actions that are performable by the computing device; and

based on determining that the second transcription of the utterance includes the term from the predefined set of one or more terms, providing an output of the first transcription of the utterance.

8. The system of claim 7 wherein the first speech recognizer employs a grammar-based language model.

9. The system of claim 8 wherein the grammar-based language model includes a context free grammar.

10. The system of claim 7 wherein the second speech recognizer employs a statistics-based language model.

11. The system of claim 7 wherein the user-specific data includes a contact list for the user, an applications list of applications installed on the computing device, or a media list of media stored on the computing device.

12. The system of claim 7 wherein the first speech recognizer is implemented on the computing device and the second speech recognizer is implemented on one or more server devices.

13. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

accessing audio data generated by a computing device based on audio input from a user, the audio data encoding one or more user utterances;

determining a first transcription of the utterances by performing speech recognition on the audio data using a first speech recognizer, wherein the first speech recognizer employs a language model that is developed based on user-specific data;

determining a second transcription of the utterances by performing speech recognition on the audio data using a second speech recognizer, wherein the second speech recognizer employs a language model developed independent of user-specific data;

determining that the second transcription of the utterances includes a term from a predefined set of one or more terms associated with actions that are performable by the computing device; and

based on determining that the second transcription of the utterance includes the term from the predefined set of one or more terms, providing an output of the first transcription of the utterance.

14. The medium of claim 13 wherein the first speech recognizer employs a grammar-based language model.

15. The medium of claim 13 wherein the second speech recognizer employs a statistics-based language model.

16. The medium of claim 13 wherein the user-specific data includes a contact list for the user, an applications list of applications installed on the computing device, or a media list of media stored on the computing device.

17. The medium of claim 13 wherein the first speech recognizer is implemented on the computing device and the second speech recognizer is implemented on one or more server devices.

18. The method of claim 1 , further comprising determining that the second transcription represents a search query, and

wherein determining that the second transcription of the utterances includes a term from a predefined set of one or more terms is performed in response to determining that the second transcription represents the search query.

19. The method of claim 1 , further comprising determining that the second transcription represents a search query and that the first transcription represents an action, and

wherein determining that the second transcription of the utterances includes a term from a predefined set of one or more terms is performed in response to determining that the second transcription represents the search query and that the first transcription represents the action.

Assignments (2)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2013
From: GRUENSTEIN, ALEXANDER H.; ALEKSIC, PETAR
To: GOOGLE INC.
Reel/Frame 030565/0632 →
Continuity (2)
Provisional Application 61664324 · Jun 26, 2012
Related Publication 20130346078A1 · Dec 26, 2013
Cited By (15)
US 12,211,490 US 12,217,748 US 12,230,291 US 12,236,932 US 12,283,269 US 12,327,549 US 12,327,556 US 12,360,734 US 12,387,716 US 12,424,220 US 12,505,832 US 12,513,479 US 12,518,756 US 12,699,543 US 12,711,962