IP Library › Granted Patent US 12,230,258
Granted Patent B2
US 12,230,258 · App. 17/659,836 · Granted Feb 18, 2025

Sub-models for neural contextual biasing

Inventors: Fadi Biadsy (Mountain View, CA); Pedro J. Moreno Mengibar (Jersey City, NJ)
Assignee: Google LLC
G10L15/183G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,230,258
App. No.
17/659,836
Granted
Feb 18, 2025
Kind
B2
Abstract

A method for contextual biasing for speech recognition includes obtaining a base automatic speech recognition (ASR) model trained on non-biased data and a sub-model trained on biased data representative of a particular domain. The method includes receiving a speech recognition request including audio data characterizing an utterance captured in streaming audio. The method further includes determining whether the speech recognition request includes a contextual indicator indicating the particular domain. When the speech recognition request does not include the contextual indicator, the method includes generating, using the base ASR model, a first speech recognition result of the utterance by processing the audio data. When the speech recognition request includes the contextual indicator the method includes biasing, using the sub-model, the base ASR model toward the particular domain and generating, using the biased base ASR model, a second speech recognition result of the utterance by processing the audio data.

Claims (50)

1. A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:

obtaining a base automatic speech recognition (ASR) model trained on non-biased data;

obtaining a sub-model trained on biased data, the biased data representative of a particular domain;

receiving a speech recognition request comprising audio data characterizing an utterance captured in streaming audio;

determining whether the speech recognition request includes a contextual indicator indicating the particular domain;

when the speech recognition request does not include the contextual indicator, generating, using the base ASR model, a first speech recognition result of the utterance by processing the audio data; and

when the speech recognition request includes the contextual indicator:

generating, using the base ASR model, an encoded output by processing the audio data;

biasing, using the sub-model, the base ASR model toward the particular domain;

generating, using the biased base ASR model, a sub-model output by processing the audio data, the sub-model output generated in parallel with the encoded output; and

generating, using a decoder of the base ASR model, a second speech recognition result of the utterance by processing the encoded output and the sub-model output, the second speech recognition result biased toward one or more terms in the particular domain.

2. The computer-implemented method of claim 1 , wherein the contextual indicator comprises a one-hot vector.

3. The computer-implemented method of claim 2 , wherein the one-hot vector indicates a particular sub-model from a plurality of sub-models to be activated, each sub-model of the plurality of sub-models associated with a different domain.

4. The computer-implemented method of claim 2 , wherein the operations further comprise projecting the one-hot vector into a phrase set embedding of an embedding space.

5. The computer-implemented method of claim 4 , wherein projecting the one-hot vector into the phrase set embedding causes the phrase set embedding to activate a portion of the sub-model.

6. The computer-implemented method of claim 1 , wherein the sub-model is disposed in a layer of the base ASR model.

7. The computer-implemented method of claim 6 , wherein:

the base ASR model comprises an encoder and the decoder; and

the sub-model is disposed in between two layers of the encoder.

8. The computer-implemented method of claim 1 , wherein one or more parameters of the base ASR model are frozen when generating the first and second speech recognition results.

9. The computer-implemented method of claim 1 , wherein:

the data processing hardware resides on a user device that captured the utterance in the streaming audio; or

the data processing hardware resides on a remote system in communication with the user device via a network.

10. The computer-implemented method of claim 1 , wherein the first speech recognition result is different from the second speech recognition result.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

obtaining a base automatic speech recognition (ASR) model trained on non-biased data;

obtaining a sub-model trained on biased data, the biased data representative of a particular domain;

receiving a speech recognition request comprising audio data characterizing an utterance captured in streaming audio;

determining whether the speech recognition request includes a contextual indicator indicating the particular domain;

when the speech recognition request does not include the contextual indicator, generating, using the base ASR model, a first speech recognition result of the utterance by processing the audio data; and

when the speech recognition request includes the contextual indicator:

generating, using the base ASR model, an encoded output by processing the audio data;

biasing, using the sub-model, the base ASR model toward the particular domain;

generating, using the biased base ASR model, a sub-model output by processing the audio data, the sub-model output generated in parallel with the encoded output; and

generating, using a decoder of the base ASR model, a second speech recognition result of the utterance by processing the audio data encoded output and the sub-model output, the second speech recognition result biased toward one or more terms in the particular domain.

12. The system of claim 11 , wherein the contextual indicator comprises a one-hot vector.

13. The system of claim 12 , wherein the one-hot vector indicates a particular sub-model from a plurality of sub-models to be activated, each sub-model of the plurality of sub-models associated with a different domain.

14. The system of claim 12 , wherein the operations further comprise projecting the one-hot vector into a phrase set embedding of an embedding space.

15. The system of claim 14 , wherein projecting the one-hot vector into the phrase set embedding causes the phrase set embedding to activate a portion of the sub-model.

16. The system of claim 11 , wherein the sub-model is disposed in a layer of the base ASR model.

17. The system of claim 16 , wherein:

the base ASR model comprises an encoder and the decoder; and

the sub-model is disposed in between two layers of the encoder.

18. The system of claim 11 , wherein one or more parameters of the base ASR model are frozen when generating the first and second speech recognition results.

19. The system of claim 11 , wherein:

the data processing hardware resides on a user device that captured the utterance in the streaming audio; or

the data processing hardware resides on a remote system in communication with the user device via a network.

20. The system of claim 11 , wherein the first speech recognition result is different from the second speech recognition result.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2022
From: BIADSY, FADI; MENGIBAR, PEDRO MORENO
To: GOOGLE LLC
Reel/Frame 059646/0257 →
Continuity (1)
Related Publication 20230335122A1 · Oct 19, 2023
References Cited (19)
US 11587569B2 · Ye · 2023 [cited by examiner]
US 20200302916A1 · Mengibar et al. · 2020 [cited by applicant]
US 20200380215A1 · Kannan · 2020 [cited by examiner]
US 20200402501A1 · Prabhavalkar et al. · 2020 [cited by applicant]
US 20220122596A1 · Jessa · 2022 [cited by examiner]
US 20230026945A1 · Friedlander · 2023 [cited by examiner]
CN 112992127A · 2021 [cited by examiner]
Katrin Tomanek, Vicky Zayats, Dirk Padfield, Kara Vaillancourt, Fadi Biadsy “Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and Accented Speech” arXiv:2109.06952v1 [cs.CL] Sep. 14, 2021 (Year: 2021… [cited by examiner]
Tsendsuren Munkhdalai, Youzheng Chen, Khe Chai Sim, Fadi Biadsy, Tara Sainath, Pedro Moreno Mengibar “Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of Large Speech Models” arXiv:2403.19709v1 [eess.… [cited by examiner]
Fadi Biadsy, Youzheng Chen, Xia Zhang, Oleg Rybakov, Andrew Rosenberg, Pedro J. Moreno “A Scalable Model Specialization Framework for Training and Inference using Submodels and its Application to Speech Model Personaliz… [cited by examiner]
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, Sylvain Gelly “Parameter-Efficient Transfer Learning for NLP” arXiv:1902.00751v2 [cs. LG] Jun.… [cited by examiner]
Asa Cooper Stickland, Iain Murray “BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning” arXiv:1902.02671v2 [cs.LG] May 15, 2019 (Year: 2019). [cited by examiner]
Asa Cooper Stickland, lain Murray. “BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning” arXiv:1902.02671v2 (Year: 2019). [cited by examiner]
International Search Report and Written Opinion for the related Application No. PCT/US2023/019030, dated Jul. 13, 2023, 60 pages. [cited by applicant]
Petar Aleksic et al: “Bringing Contextual Information to Google Speech Recognition”, INTERSPEECH 2015, Sep. 6, 2015 (Sep. 6, 2015), pp. 468-472, XP055345740, Retrieved from the Internet: URL:https://static.googleusercon… [cited by applicant]
Pundak Golan et al: “Deep Context: End-to-end Contextual Speech Recognition”, 2018 IEEE Spoken Language Technology Workshop (SLT), IEEE, Dec. 18, 2018 (Dec. 18, 2018), pp. 418-425, XP033516980, DOI: 10.1109/SLT.2018.863… [cited by applicant]
Moriokal Tsuyoshi et al: “Language Model Domain Adaptation Via Recurrent Neural Networks with Domain- Shared and Domain-Specific Representations”, 2018 IEEE International Conference on Acoustics, Speech and Signal Proce… [cited by applicant]
“Embedling, Beyond Just Words”, Mohammed Alhamid, Jun. 4, 2021. [cited by applicant]
“Theme-weighted Ranking of Keywords from Text Documents using Phrase Embeddings”, Mahata et al., 2018. [cited by applicant]