IP Library › Granted Patent US 12,597,417
Granted Patent B2
US 12,597,417 · App. 18/494,763 · Granted Apr 7, 2026

Exporting modular encoder features for streaming and deliberation ASR

Inventors: Rami Magdi Fahmi Botros (Mountain View, CA); Rohit Prakash Prabhavalkar (Palo Alto, CA); Johan Schalkwyk (Scarsdale, NY); Tara N. Sainath (Jersey City, NJ); Ciprian Ioan Chelba (Mountain View, CA); Francoise Beaufays (Mountain View, CA)
Assignee: Google LLC
G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,417
App. No.
18/494,763
Granted
Apr 7, 2026
Kind
B2
Abstract

A method includes obtaining a base encoder from a pre-trained model, and receiving training data comprising a sequence of acoustic frames characterizing an utterance paired with a ground-truth transcription of the utterance. At each of a plurality of output steps, the method includes: generating, by the base encoder, a first encoded representation for a corresponding acoustic frame; generating, by an exporter network configured to receive a continuous sequence of first encoded representations generated by the base encoder, a second encoded representation for a corresponding acoustic frame; generating, by an exporter decoder, a probability distribution over possible logits; and determining an exporter decoder loss based on the probability distribution over possible logits generated by the exporter decoder at the corresponding output step and the ground-truth transcription. The method also includes training the exporter network based on the exporter decoder losses while parameters of the base encoder are frozen.

Claims (58)

1 . A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

obtaining a base encoder from a pre-trained model, the pre-trained model comprising the base encoder and a decoder;

receiving training data comprising a sequence of acoustic frames characterizing an utterance paired with a ground-truth transcription of the utterance;

at each corresponding output step of a plurality of output steps:

generating, by the base encoder, a first encoded representation for a corresponding acoustic frame in the sequence of acoustic frames;

generating, by an exporter network configured to receive a continuous sequence of first encoded representations generated by the base encoder, a second encoded representation for a corresponding acoustic frame in the sequence of acoustic frames;

generating, by an exporter decoder, a probability distribution over possible logits; and

determining a corresponding exporter decoder loss based on the ground-truth transcription and the probability distribution over possible logits generated by the exporter decoder at the corresponding output step; and

training the exporter network based on the corresponding exporter decoder loss determined at each corresponding output step of the plurality of output steps while parameters of the base encoder are frozen.

2 . The method of claim 1 , wherein the operations further comprise obtaining the base encoder from a pre-trained streaming recognition model that comprises the base encoder, a prediction network, and a joint network.

3 . The method of claim 1 , wherein:

the base encoder comprises a first plurality of multi-head self-attention blocks; and

the exporter network comprises a second plurality of multi-head self-attention blocks.

4 . The method of claim 3 , wherein the second plurality of multi-head self-attention blocks of the exporter network are non-causal.

5 . The method of claim 3 , wherein the first plurality of multi-head self-attention blocks of the base encoder and the second plurality of multi-head self-attention blocks of the exporter network comprise conformer blocks.

6 . The method of claim 1 , wherein:

the exporter decoder comprises a connectionist temporal classification (CTC) decoder;

the corresponding exporter decoder loss comprises a CTC loss; and

the logits comprise sub-word units.

7 . The method of claim 6 , wherein the sub-word units comprise wordpieces, graphemes, phonemes, or triphones.

8 . The method of claim 1 , wherein the operations further comprise, at each corresponding output step of the plurality of output steps:

determining a modular encoded representation for a corresponding acoustic frame in the sequence of acoustic frames by extracting a top-k indices of logits from the probability distribution over possible logits generated by the exporter decoder at the corresponding output step;

re-embedding the modular encoded representation determined at the corresponding output step;

generating, by an importer network, an importer representation for a corresponding re-embedded modular encoded representation; and

generating, by a speech decoder, a speech recognition hypothesis for a corresponding importer representation.

9 . The method of claim 8 , wherein the speech decoder comprises a recurrent neural network-transducer (RNN-T) decoder.

10 . The method of claim 8 , wherein the speech decoder comprises a listen-attend-spell (LAS) decoder.

11 . The method of claim 10 , wherein the exporter decoder comprises a recurrent neural network-transducer (RNN-T) decoder.

12 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations comprising:

obtaining a base encoder from a pre-trained model, the pre-trained model comprising the base encoder and a decoder;

receiving training data comprising a sequence of acoustic frames characterizing an utterance paired with a ground-truth transcription of the utterance;

at each corresponding output step of a plurality of output steps:

generating, by the base encoder, a first encoded representation for a corresponding acoustic frame in the sequence of acoustic frames;

generating, by an exporter network configured to receive a continuous sequence of first encoded representations generated by the base encoder, a second encoded representation for a corresponding acoustic frame in the sequence of acoustic frames;

generating, by an exporter decoder, a probability distribution over possible logits; and

determining a corresponding exporter decoder loss based on the ground-truth transcription and the probability distribution over possible logits generated by the exporter decoder at the corresponding output step; and

training the exporter network based on the corresponding exporter decoder loss determined at each corresponding output step of the plurality of output steps while parameters of the base encoder are frozen.

13 . The system of claim 12 , wherein the operations further comprise obtaining the base encoder from a pre-trained streaming recognition model that comprises the base encoder, a prediction network, and a joint network.

14 . The system of claim 12 , wherein:

the base encoder comprises a first plurality of multi-head self-attention blocks; and

the exporter network comprises a second plurality of multi-head self-attention blocks.

15 . The system of claim 14 , wherein the second plurality of multi-head self-attention blocks of the exporter network are non-causal.

16 . The system of claim 14 , wherein the first plurality of multi-head self-attention blocks of the base encoder and the second plurality of multi-head self-attention blocks of the exporter network comprise conformer blocks.

17 . The system of claim 12 , wherein:

the exporter decoder comprises a connectionist temporal classification (CTC) decoder;

the corresponding exporter decoder loss comprises a CTC loss; and

the logits comprise sub-word units.

18 . The system of claim 17 , wherein the sub-word units comprise wordpieces, graphemes, phonemes, or triphones.

19 . The system of claim 12 , wherein the operations further comprise, at each corresponding output step of the plurality of output steps:

determining a modular encoded representation for a corresponding acoustic frame in the sequence of acoustic frames by extracting a top-k indices of logits from the probability distribution over possible logits generated by the exporter decoder at the corresponding output step;

re-embedding the modular encoded representation determined at the corresponding output step;

generating, by an importer network, an importer representation for a corresponding re-embedded modular encoded representation; and

generating, by a speech decoder, a speech recognition hypothesis for a corresponding importer representation.

20 . The system of claim 19 , wherein the speech decoder comprises a recurrent neural network-transducer (RNN-T) decoder.

21 . The system of claim 19 , wherein the speech decoder comprises a listen-attend-spell (LAS) decoder.

22 . The system of claim 21 , wherein the exporter decoder comprises a recurrent neural network-transducer (RNN-T) decoder.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING PARTY NAME IS CIPRIAN LOAN CHELBA PREVIOUSLY RECORDED AT REEL: 66399 FRAME: 647. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 22, 2024
From: BOTROS, RAMI MAGDI FAHMI; PRABHAVALKAR, ROHIT PRAKASH; SCHALKWYK, JOHAN; SAINATH, TARA N.; CHELBA, CIPRIAN LOAN; BEAUFAYS, FRANCOISE
To: GOOGLE LLC
Reel/Frame 066870/0385 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2024
From: BOTROS, RAMI MAGDI FAHRNI; PRABHAVALKAR, ROHIT; SCHALKWYK, JOHAN; SAINATH, TARA N.; CHEIBA, CIPRIAN IOAN; BEAUFAYS, FRANCOIS SIMONE
To: GOOGLE LLC
Reel/Frame 066399/0647 →
Continuity (2)
Provisional Application 63381117 · Oct 26, 2022
Related Publication 20240144917A1 · May 2, 2024
References Cited (16)
US 8751239B2 · Tian · 2014 [cited by examiner]
US 11908461B2 · Hu · 2024 [cited by examiner]
US 12002451B1 · Liu · 2024 [cited by examiner]
US 20170270919A1 · Parthasarathi · 2017 [cited by examiner]
US 20190205748A1 · Fukuda · 2019 [cited by examiner]
US 20200027444A1 · Prabhavalkar · 2020 [cited by examiner]
US 20200335091A1 · Chang · 2020 [cited by examiner]
US 20210225369A1 · Hu · 2021 [cited by examiner]
US 20210375289A1 · Zhu · 2021 [cited by examiner]
US 20210383195A1 · Gygli et al. · 2021 [cited by applicant]
US 20220351718A1 · Wu · 2022 [cited by examiner]
International Search Report and Written Opinion for related PCT Application No. PCT/US2023/035938, dated Jan. 3, 2024. [cited by applicant]
Rami Botros et al: “Lego-Features: Exporting modular encoder features for streaming and deliberation ASR”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 31, 2023 (Mar.… [cited by applicant]
Gygli Michael et al: “Towards Reusable Network Componenets by Learning Compatible Representations”, Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, No. 9, Feb. 2, 2021 (Feb. 2, 2021), pp. 7620-76… [cited by applicant]
Dalmia Siddharth et al: “LegoNN: Building Modular Encoder-Decorder Models”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, [Online] vol. 31, Jan. 1, 2023 (Jan. 1, 2023), pp. 3112-3126, XP093111303, ISS… [cited by applicant]
Hu Ke et al: “Transformer Based Deliberation for Two-Pass Speech Recognition”, 2021 IEEE Spoken Language Technology Workship (SLT), [Online] Jan. 19, 2021 (Jan. 19, 2021), pp. 68-74, XP093000484, DOI: 10.1109/SLT48900.2… [cited by applicant]