IP Library Granted Patent US 11,037,551
Granted Patent B2
US 11,037,551 · App. 16/417,714 · Granted Jun 15, 2021

Language model biasing system

Inventors: Petar Aleksic (Jersey City, NJ); Pedro J. Moreno Mengibar (Jersey City, NJ)
G10L15/07G10L15/187G10L15/1815G10L15/197G10L15/01G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,037,551
App. No.
16/417,714
Granted
Jun 15, 2021
Kind
B2
Abstract

Methods, systems, and apparatus for receiving audio data corresponding to a user utterance and context data, identifying an initial set of one or more n-grams from the context data, generating an expanded set of one or more n-grams based on the initial set of n-grams, adjusting a language model based at least on the expanded set of n-grams, determining one or more speech recognition candidates for at least a portion of the user utterance using the adjusted language model, adjusting a score for a particular speech recognition candidate determined to be included in the expanded set of n-grams, determining a transcription of user utterance that includes at least one of the one or more speech recognition candidates, and providing the transcription of the user utterance for output.

Claims (47)

1. A computer-implemented method comprising:

obtaining a dialog state of a user device associated with a user;

before the user speaks an utterance to the user device:

identifying, from an n-gram cache, an initial set of n-grams that represent one or more words or phrases corresponding to the dialog state of the user device; and

biasing a language model on the initial set of n-grams to increase a likelihood of output of the one or more words or phrases corresponding to the dialog state of the user device;

receiving audio data indicating the utterance of the user;

processing the audio data using the biased language model to generate a transcription of the utterance; and

providing the transcription of the utterance for output.

2. The method of claim 1 , further comprising:

receiving context data for the user device

wherein obtaining the dialog state comprises identifying the dialog state based on the context data.

3. The method of claim 2 , wherein the utterance is detected by the user device providing an interface to the user, and wherein the context data comprises data that indicates a topic corresponding to the interface.

4. The method of claim 2 , wherein the utterance is detected by the user device providing an interface to the user, and wherein the context data comprises data indicating a task to be performed using the interface.

5. The method of claim 2 , wherein the utterance is detected by the user device providing an interface to the user, and wherein the context data comprises data indicating a step for completing a portion of task to be performed using the interface.

6. The method of claim 2 , wherein the context data indicates one or more words or phrases included in a graphical user interface of the user device at a time that the utterance was spoken.

7. The method of claim 2 , wherein:

the context data comprises an application identifier or dialog state identifier; and

identifying the initial set of n-grams that represent the one or more words or phrases comprises retrieving data indicating one or more words or phrases corresponding to the application identifier or dialog state identifier.

8. The method of claim 2 , wherein identifying the initial set of n-grams that represent one or more words or phrases corresponding to the dialog state of the user device comprises:

identifying a first set of one or more n-grams from the context data; and

generating an expanded set of one or more n-grams based at least on the first set of n-grams, the expanded set of n-grams comprising one or more n-grams that are different from the n-grams in the first set of n-grams.

9. The method of claim 1 , wherein identifying the dialog state of the user device comprises identifying one of a plurality of different dialog states of the user device that each correspond to a different interface or view of an application.

10. The method of claim 9 , wherein each of the plurality of different dialog states of the user device is associated with a predetermined set of n-grams.

11. The method of claim 10 , wherein for at least one of the dialog states of the user device, one or more of the n-grams in the predetermined set of n-grams for the dialog state of the user device are not displayed by the application during the dialog state of the user device.

12. The method of claim 10 , wherein at least some of the n-grams in the predetermined set of n-grams are designated by a developer of the application.

13. The method of claim 1 , wherein receiving the audio data comprises receiving, by a server system, audio data provided by the user device over a communication network.

14. The method of claim 13 , wherein identifying the dialog state of the user device comprises identifying, by the server system, the dialog state of the user device based on additional data provided by the user device over the communication network.

15. The method of claim 14 , wherein the additional data comprises a dialog state identifier provided by the user device or an application identifier provided by the user device.

16. The method of claim 15 , wherein determining the one or more words or phrases corresponding to the dialog state of the user device comprises retrieving the one or more words or phrases from a data repository based on the dialog state identifier or the application identifier.

17. A system comprising:

one or more computers; and

one or more computer-readable media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

obtaining a dialog state of a user device associated with a user;

before the user speaks an utterance to the user device:

identifying, from an n-gram cache, an initial set of n-grams that represent one or more words or phrases corresponding to the dialog state of the user device; and

biasing a language model on the initial set of n-grams to increase a likelihood of output of the one or more words or phrases corresponding to the dialog state of the user device;

receiving audio data indicating the utterance of the user;

processing the audio data using the biased language model to generate a transcription of the utterance; and

providing the transcription of the utterance for output.

18. One or more non-transitory computer-readable media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

obtaining a dialog state of a user device associated with a user;

before the user speaks an utterance to the user device:

identifying, from an n-gram cache, an initial set of n-grams that represent one or more words or phrases corresponding to the dialog state of the user device; and

biasing a language model on the initial set of n-grams to increase a likelihood of output of the one or more words or phrases corresponding to the dialog state of the user device;

receiving audio data indicating the utterance of the user;

processing the audio data using the biased language model to generate a transcription of the utterance; and

providing the transcription of the utterance for output.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2019
From: ALEKSIC, PETAR; MENGIBAR, PEDRO J. MORENO
To: GOOGLE INC.
Reel/Frame 049240/0482 →
ENTITY CONVERSION Recorded May 21, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 049240/0619 →
Continuity (2)
Continuation 15432620 · Feb 14, 2017
Related Publication 20190341024A1 · Nov 7, 2019
Cited By (1)
US 12,272,261