IP Library Granted Patent US 11,664,021
Granted Patent B2
US 11,664,021 · App. 17/643,423 · Granted May 30, 2023

Contextual biasing for speech recognition

Inventors: Rohit Prakash Prabhavalkar (Santa Clara, CA); Golan Pundak (New York, NY); Tara N. Sainath (Jersey City, NJ); Antoine Jean Bruguier (Milpitas, CA)
Assignee: Google LLC
G10L15/187G06N20/10G10L19/04G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,664,021
App. No.
17/643,423
Granted
May 30, 2023
Kind
B2
Abstract

A method of biasing speech recognition includes receiving audio data encoding an utterance and obtaining a set of one or more biasing phrases corresponding to a context of the utterance. Each biasing phrase in the set of one or more biasing phrases includes one or more words. The method also includes processing, using a speech recognition model, acoustic features derived from the audio data and grapheme and phoneme data derived from the set of one or more biasing phrases to generate an output of the speech recognition model. The method also includes determining a transcription for the utterance based on the output of the speech recognition model.

Claims (46)

1. A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving audio data encoding one or more words of an utterance;

obtaining grapheme data and phoneme data derived from the one or more words of the utterance, the one or more words comprising a proper noun;

generating, using a grapheme encoder configured to receive the grapheme data, grapheme encodings;

generating, using a phoneme encoder configured to receive the phoneme data, phoneme encodings;

generating, using an attention module configured to receive a representation of the grapheme encodings output from the grapheme encoder and the phoneme encodings output from the phoneme encoder, attention outputs; and

processing, using a decoder, the attention outputs generated by the attention module to determine likelihoods of speech elements.

2. The computer-implemented method of claim 1 , wherein the operations further comprise generating a transcription of the utterance based on the likelihoods of speech elements.

3. The computer-implemented method of claim 1 , wherein the grapheme encoder and the phoneme encoder each comprise neural networks.

4. The computer-implemented method of claim 1 , wherein:

the grapheme encoder is configured to generate a corresponding grapheme encoding for a particular word;

the phoneme encoder is configured to generate a corresponding phoneme encoding for the particular word; and

the attention module is configured to encode a corresponding second attention output that comprises a corresponding contextual biasing vector for the particular word based on the corresponding grapheme and phoneme encodings for the particular word.

5. The computer-implemented method of claim 1 , wherein the representation of the grapheme encodings output from the grapheme encoder and the phoneme encodings output from the phoneme encoder comprises a concatenation between the grapheme encodings and the phoneme encodings.

6. The computer-implemented method of claim 5 , wherein the concatenation between the grapheme encodings and the phoneme encodings is represented by a projection vector input to the attention module.

7. The computer-implemented method of claim 1 , wherein the grapheme encoder, the phoneme encoder, the attention module, and the decoder are trained jointly to predict a sequence of graphemes.

8. The computer-implemented method of claim 1 , wherein the speech elements comprise graphemes.

9. The computer-implemented method of claim 1 , wherein the speech elements comprise words or wordpieces.

10. The computer-implemented method of claim 1 , wherein the operations further comprise determining a context of the utterance based on at least one of:

a location of a user that spoke the utterance;

one or more applications open on a user device associated with a user that spoke the utterance; or

a current date and/or time of the utterance.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving audio data encoding one or more words of an utterance;

obtaining grapheme data and phoneme data derived from the one or more words of the utterance, the one or more words comprising a proper noun;

generating, using a grapheme encoder configured to receive the grapheme data, grapheme encodings;

generating, using a phoneme encoder configured to receive the phoneme data, phoneme encodings;

generating, using an attention module configured to receive a representation of the grapheme encodings output from the grapheme encoder and the phoneme encodings output from the phoneme encoder, attention outputs; and

processing, using a decoder, the attention outputs generated by the attention module to determine likelihoods of speech elements.

12. The system of claim 11 , wherein the operations further comprise generating a transcription of the utterance based on the likelihoods of speech elements.

13. The system of claim 11 , wherein the grapheme encoder and the phoneme encoder each comprise neural networks.

14. The system of claim 11 , wherein:

the grapheme encoder is configured to generate a corresponding grapheme encoding for a particular word;

the phoneme encoder is configured to generate a corresponding phoneme encoding for the particular word; and

the attention module is configured to encode a corresponding second attention output that comprises a corresponding contextual biasing vector for the particular word based on the corresponding grapheme and phoneme encodings for the particular word.

15. The system of claim 11 , wherein the representation of the grapheme encodings output from the grapheme encoder and the phoneme encodings output from the phoneme encoder comprises a concatenation between the grapheme encodings and the phoneme encodings.

16. The system of claim 15 , wherein the concatenation between the grapheme encodings and the phoneme encodings is represented by a projection vector input to the attention module.

17. The system of claim 11 , wherein the grapheme encoder, the phoneme encoder, the attention module, and the decoder are trained jointly to predict a sequence of graphemes.

18. The system of claim 11 , wherein the speech elements comprise graphemes.

19. The system of claim 11 , wherein the speech elements comprise words or wordpieces.

20. The system of claim 11 , wherein the operations further comprise determining a context of the utterance based on at least one of:

a location of a user that spoke the utterance;

one or more applications open on a user device associated with a user that spoke the utterance; or

a current date and/or time of the utterance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2021
From: PRABHAVALKAR, ROHIT PRAKASH; PUNDAK, GOLAN; SAINATH, TARA N.; BRUGUIER, ANTOINE JEAN
To: GOOGLE LLC
Reel/Frame 058342/0251 →
Continuity (3)
Continuation 16863766 · Apr 30, 2020
Provisional Application 62863308 · Jun 19, 2019
Related Publication 20220101836A1 · Mar 31, 2022
Cited By (1)
US 12,658,191