IP Library Granted Patent US 12,431,122
Granted Patent B2
US 12,431,122 · App. 18/065,692 · Granted Sep 30, 2025

Training a language model of an end-to-end automatic speech recognition model using random encoder features

Inventors: Adam Michael Stooke (San Francisco, CA); Khe Chai Sim (Dublin, CA); Mason Vijay Chua (Mountain View, CA)
Assignee: Google LLC
G10L15/063G10L15/16G10L15/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,431,122
App. No.
18/065,692
Granted
Sep 30, 2025
Kind
B2
Abstract

A method includes obtaining a training text sample, the training text sample not paired with corresponding audio data, and generating a sequence of pseudo-random encoder variables. The method also includes processing, using a decoder of a sequence transduction model, the sequence of pseudo-random encoder variables to predict a probability distribution over possible output labels. The method further includes determining a loss based metric based on the training text sample and the predicted probability distribution over possible output labels, and training the decoder based on the loss metric.

Claims (58)

1. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

obtaining a training text sample, the training text sample not paired with corresponding audio data;

generating, using a pseudo-random variable generator, a sequence of pseudo-random encoder variables for the training text sample that is not paired with corresponding audio data;

processing, using a decoder of a sequence transduction model, the sequence of pseudo-random encoder variables generated for the training text sample that is not paired with corresponding audio data to predict a probability distribution over possible output labels;

determining a loss metric based on the training text sample and the predicted probability distribution over possible output labels; and

training, based on the loss metric, the decoder with hybrid autoregressive transducer factorization to integrate a language model of the sequence transduction model,

wherein the language model comprises a neural network model comprising a stack of conformer layers or transformer layers.

2. The computer-implemented method of claim 1 , wherein training the decoder based on the loss metric comprises training the decoder based on a negative log of the probability distribution over possible output labels for the training text sample conditioned on the sequence of pseudo-random encoder variables.

3. The computer-implemented method of claim 1 , wherein:

the sequence transduction model is pre-trained on paired audio-text samples; and

training the decoder based on the loss metric comprises fine-tuning the decoder of the sequence transduction model based on the loss metric.

4. The computer-implemented method of claim 1 , wherein a number of sequence pseudo-random encoder variables generated increases as a number of text labels in the training text sample increases.

5. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

obtaining a training text sample, the training text sample not paired with corresponding audio data;

generating a sequence of pseudo-random encoder variables;

generating, by a prediction network, a corresponding text label representation for each prior sequence of non-blank symbols output by a final softmax layer of the sequence transduction model;

processing, using a decoder of the sequence transduction model, the sequence of pseudo-random encoder variables and the corresponding text label representations to predict a probability distribution over possible output labels;

determining a loss metric based on the training text sample and the predicted probability distribution over possible output labels; and

training the decoder based on the loss metric.

6. The computer-implemented method of claim 5 , wherein the decoder comprises the prediction network and a joint network, the joint network configured to receive each corresponding text label representation generated by the prediction network and each pseudo-random encoder variable in the sequence of pseudo-random encoder variables.

7. The computer-implemented method of claim 6 , wherein training the decoder comprises:

updating coefficients of the joint network;

holding coefficients of the prediction network fixed; and

holding coefficients of an encoder network of the sequence transduction model fixed.

8. The computer-implemented method of claim 6 , wherein the prediction network comprises a multi-headed attention mechanism, the multi-headed attention mechanism sharing a shared embedding matrix across each head of the multi-headed attention mechanism.

9. The computer-implemented method of claim 8 , wherein the prediction network ties a dimensionality of the shared embedding matrix to a dimensionality of an output layer of the joint network.

10. The computer-implemented method of claim 1 , wherein the sequence transduction model comprises a recurrent neural network-transducer (RNN-T) based speech recognition model.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

obtaining a training text sample, the training text sample not paired with corresponding audio data;

generating, using a pseudo-random variable generator, a sequence of pseudo-random encoder variables for the training text sample that is not paired with corresponding audio data;

processing, using a decoder of a sequence transduction model, the sequence of pseudo-random encoder variables generated for the training text sample that is not paired with corresponding audio data to predict a probability distribution over possible output labels;

determining a loss metric based on the training text sample and the predicted probability distribution over possible output labels; and

training, based on the loss metric, the decoder with hybrid autoregressive transducer factorization to integrate a language model of the sequence transduction model,

wherein the language model comprises a neural network model comprising a stack of conformer layers or transformer layers.

12. The system of claim 11 , wherein training the decoder based on the loss metric comprises training the decoder based on a negative log of the probability distribution over possible output labels for the training text sample conditioned on the sequence of pseudo-random encoder variables.

13. The system of claim 11 , wherein:

the sequence transduction model is pre-trained on paired audio-text samples; and

training the decoder based on the loss metric comprises fine-tuning the decoder of the sequence transduction model based on the loss metric.

14. The system of claim 11 , wherein a number of sequence pseudo-random encoder variables generated increases as a number of text labels in the training text sample increases.

15. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

obtaining a training text sample, the training text sample not paired with corresponding audio data;

generating a sequence of pseudo-random encoder variables;

generating, by a prediction network, a corresponding text label representation for each prior sequence of non-blank symbols output by a final softmax layer of the sequence transduction model;

processing, using a decoder of the sequence transduction model, the sequence of pseudo-random encoder variables and the corresponding text label representations to predict a probability distribution over possible output labels;

determining a loss metric based on the training text sample and the predicted probability distribution over possible output labels; and

training the decoder based on the loss metric.

16. The system of claim 15 , wherein the decoder comprises the prediction network and a joint network, the joint network configured to receive each corresponding text label representation generated by the prediction network and each pseudo-random encoder variable in the sequence of pseudo-random encoder variables.

17. The system of claim 16 , wherein training the decoder comprises:

updating coefficients of the joint network;

holding coefficients of the prediction network fixed; and

holding coefficients of an encoder network of the sequence transduction model fixed.

18. The system of claim 16 , wherein the prediction network comprises a multi-headed attention mechanism, the multi-headed attention mechanism sharing a shared embedding matrix across each head of the multi-headed attention mechanism.

19. The system of claim 18 , wherein the prediction network ties a dimensionality of the shared embedding matrix to a dimensionality of an output layer of the joint network.

20. The system of claim 11 , wherein the sequence transduction model comprises a recurrent neural network-transducer (RNN-T) based speech recognition model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2023
From: STOOKE, ADAM MICHAEL; SIM, KHE CHAI; CHUA, MASON VIJAY
To: GOOGLE LLC
Reel/Frame 062369/0554 →
Continuity (1)
Related Publication 20240203399A1 · Jun 20, 2024
References Cited (21)
US 11423312B2 · Choi · 2022 [cited by examiner]
US 11551668B1 · Baevski · 2023 [cited by examiner]
US 12087306B1 · Le · 2024 [cited by examiner]
US 12159617B2 · Chen · 2024 [cited by examiner]
US 12249317B2 · Li · 2025 [cited by examiner]
US 12249336B2 · Li · 2025 [cited by examiner]
US 20190108832A1 · Tomar et al. · 2019 [cited by applicant]
US 20190279618A1 · Yadav et al. · 2019 [cited by applicant]
US 20190349426A1 · Smith · 2019 [cited by examiner]
US 20210165976A1 · Lee · 2021 [cited by examiner]
US 20210287430A1 · Li · 2021 [cited by examiner]
US 20220310097A1 · Kim et al. · 2022 [cited by applicant]
US 20230081171A1 · Zhang · 2023 [cited by examiner]
US 20230360646A1 · Coucheiro Limeres · 2023 [cited by examiner]
US 20240153508A1 · Moritz · 2024 [cited by examiner]
US 20250094519A1 · Kol · 2025 [cited by examiner]
US 20250166622A1 · Chhugani · 2025 [cited by examiner]
Zhong Meng et al: “International Language Model Adaptation with Text-Only Data for End-to-End Speech Recognition”, arvix.org, Cornell University Library, 201 Olin Liberary Cornell University Ithaca, NY 14853, Feb. 18, 2… [cited by applicant]
Zheng Xianrui et al: “Using Synthetic Audio to Improve the Recognition of Out-of-Vocabulary Words in End-to-End Asr Systems”, ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (IC… [cited by applicant]
Stooke Adam et al: “Internal Language Model Personalization of E2E Automatic Speech Recognition Using Random Encoder Features”, 2022 IEEE Spoken Language Technology Workshop (SLT), IEEE, Jan. 9, 2023 (Jan. 9, 2023), pp.… [cited by applicant]
International Search Report and Written Opinion issued in related PCT Application No. PCT/US2023/080831 dated Feb. 16, 2024. [cited by applicant]