IP Library › Granted Patent US 12,417,770
Granted Patent B2
US 12,417,770 · App. 18/182,925 · Granted Sep 16, 2025

Unified cascaded encoder ASR model for dynamic model sizes

Inventors: Shaojin Ding (Mountain View, CA); Yangzhang He (Mountain View, CA); Xin Wang (Mountain View, CA); Weiran Wang (Palo Alto, CA); Trevor Strohman (Mountain View, CA); Tara N. Sainath (Jersey City, NJ); Rohit Prakash Prabhavalkar (Palo Alto, CA); Robert David (Mountain View, CA); Rina Panigrahy (Mountain View, CA); Rami Botros (Mountain View, CA); Qiao Liang (Mountain View, CA); Ian Mcgraw (Mountain View, CA); Ding Zhao (Mountain View, CA); Dongseong Hwang (Mountain View, CA)
Assignee: Google LLC
G10L15/32G10L15/16G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,770
App. No.
18/182,925
Granted
Sep 16, 2025
Kind
B2
Abstract

An automated speech recognition (ASR) model includes a first encoder, a first encoder, a second encoder, and a second decoder. The first encoder receives, as input, a sequence of acoustic frames, and generates, at each of a plurality of output steps, a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames. The first decoder receives, as input, the first higher order feature representation generated by the first encoder, and generates a first probability distribution over possible speech recognition hypotheses. The second encoder receives, as input, the first higher order feature representation generated by the first encoder, and generates a second higher order feature representation for a corresponding first higher order feature frame. The second decoder receives, as input, the second higher order feature representation generated by the second encoder, and generates a second probability distribution over possible speech recognition hypotheses.

Claims (101)

1. An automated speech recognition (ASR) model comprising:

a first encoder configured to:

receive, as input, a sequence of acoustic frames; and

generate, at each of a plurality of output steps, a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames;

a first decoder configured to:

receive, as input, the first higher order feature representation generated by the first encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a first probability distribution over possible speech recognition hypotheses;

a second encoder configured to:

receive, as input, the first higher order feature representation generated by the first encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a second higher order feature representation for a corresponding first higher order feature frame; and

a second decoder configured to:

receive, as input, the second higher order feature representation generated by the second encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a second probability distribution over possible speech recognition hypotheses,

wherein the first decoder and the second decoder each comprise a respective recurrent neural network-transducer (RNN-T) architecture having a same number of parameters.

2. The ASR model of claim 1 , wherein the first decoder is further configured to generate partial speech recognition results based on the first probability distribution over possible speech recognition hypotheses.

3. The ASR model of claim 1 , wherein the first encoder comprises a causal encoder comprising one of:

a plurality of unidirectional long short-term memory (LSTM) layers;

a plurality of conformer layers; or

a plurality of transformer layers.

4. The ASR model of claim 1 , wherein the second encoder comprises a non-causal encoder comprising one of:

a plurality of unidirectional long short-term memory (LSTM) layers;

a plurality of conformer layers; or

a plurality of transformer layers.

5. The ASR model of claim 1 , wherein the first decoder comprises:

a prediction network configured to:

receive, as input, a sequence of non-blank symbols output by a final softmax layer; and

generate, at each of the plurality of output steps, a dense representation; and

a joint network configured to:

receive, as input, the dense representation generated by the prediction network at each of the plurality of output steps and the first higher order feature representation generated by the first encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, the first probability distribution over possible speech recognition hypotheses.

6. The ASR model of claim 5 , wherein the prediction network comprises:

a long short-term memory (LSTM)-based prediction network; or

a V2 embedding look-up table.

7. The ASR model of claim 1 , wherein the second decoder comprises:

a prediction network configured to:

receive, as input, a sequence of non-blank symbols output by a final softmax layer; and

generate, at each of the plurality of output steps, a dense representation; and

a joint network configured to:

receive, as input, the dense representation generated by the prediction network at each of the plurality of output steps and the second higher order feature representation generated by the second encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, the second probability distribution over possible speech recognition hypotheses.

8. The ASR model of claim 7 , wherein the prediction network comprises:

a long short-term memory (LSTM)-based prediction network; or

a V2 embedding look-up table.

9. The ASR model of claim 1 , wherein the first encoder comprises a greater number of parameters than the second encoder.

10. An automated speech recognition (ASR) model comprising:

a first encoder configured to:

receive, as input, a sequence of acoustic frames; and

generate, at each of a plurality of output steps, a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames;

a first decoder configured to:

receive, as input, the first higher order feature representation generated by the first encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a first probability distribution over possible speech recognition hypotheses;

a second encoder configured to:

receive, as input, the first higher order feature representation generated by the first encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a second higher order feature representation for a corresponding first higher order feature frame;

a second decoder configured to:

receive, as input, the second higher order feature representation generated by the second encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a second probability distribution over possible speech recognition hypotheses; and

a third encoder configured to:

receive, as input, the second higher order feature representation generated by the second encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a third higher order feature representation for a corresponding second higher order feature representation; and

a third decoder configured to:

receive, as input, the third higher order feature representation generated by the third encoder at each of the plurality of output steps; and

generate, at each of the plurality of output steps, a third probability distribution over possible speech recognition hypotheses.

11. A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving a sequence of acoustic frames; and

generating, by a first encoder, at each of a plurality of output steps, a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames;

generating, by a second encoder, at each of the plurality of output steps, a second higher order feature representation for a corresponding first higher order feature representation;

at each of the plurality of output steps, generating, by a first prediction network of a first decoder based on a sequence of non-blank symbols output by a final softmax layer, a dense representation;

generating, by a first joint network of the first decoder, at each of the plurality of output steps, a first probability distribution over possible speech recognition hypotheses based on the dense representation generated by the first prediction network of the first decoder; and

generating, by a second decoder, at each of the plurality of output steps, a second probability distribution over possible speech recognition hypotheses,

wherein the first decoder and the second decoder each comprise a respective recurrent neural network-transducer (RNN-T) architecture having a same number of parameters.

12. The computer-implemented method of claim 11 , wherein the operations further comprise generating partial speech recognition results based on the first probability distribution over possible speech recognition hypotheses.

13. The computer-implemented method of claim 11 , wherein the first encoder comprises a causal encoder comprising one of:

a plurality of unidirectional long short-term memory (LSTM) layers;

a plurality of conformer layers; or

a plurality of transformer layers.

14. The computer-implemented method of claim 11 , wherein the second encoder comprises a non-causal encoder comprising one of:

a plurality of unidirectional long short-term memory (LSTM) layers;

a plurality of conformer layers; or

a plurality of transformer layers.

15. The computer-implemented method of claim 11 , wherein the first prediction network of the first decoder comprises:

a long short-term memory (LSTM)-based prediction network; or

a V2 embedding look-up table.

16. The computer-implemented method of claim 11 , wherein the operations further comprise, at each of the plurality of output steps:

generating, by a second prediction network of the second decoder, at each of the plurality of output steps, a dense representation;

receiving, as input to a second joint network of the second decoder, the dense representation generated by the second prediction network at each of the plurality of output steps and the second higher order feature representation generated by the second encoder at each of the plurality of output steps; and

generating, by the second joint network of the second decoder, at each of the plurality of output steps, the second probability distribution over possible speech recognition hypotheses.

17. The computer-implemented method of claim 16 , wherein the second prediction network of the second decoder comprises:

a long short-term memory (LSTM)-based prediction network; or

a V2 embedding look-up table.

18. The computer-implemented method of claim 11 , wherein the first encoder comprises a greater number of parameters than the second encoder.

19. A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving a sequence of acoustic frames;

generating, by a first encoder, at each of a plurality of output steps, a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames;

generating, by a second encoder, at each of the plurality of output steps, a second higher order feature representation for a corresponding first higher order feature representation;

generating, by a first decoder, at each of the plurality of output steps, a first probability distribution over possible speech recognition hypotheses;

generating, by a second decoder, at each of the plurality of output steps, a second probability distribution over possible speech recognition hypotheses;

receiving, as input to a third encoder, the second higher order feature representation generated by the second encoder at each of the plurality of output steps;

generating, by the third encoder, at each of the plurality of output steps, a third higher order feature representation for a corresponding second higher order feature representation;

receiving, as input to a third decoder, the third higher order feature representation generated by the third encoder at each of the plurality of output steps; and

generating, by the third decoder, at each of the plurality of output steps, a third probability distribution over possible speech recognition hypotheses.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE INVENTOR'S NAME FROM TREVOR STROHAM TO TREVOR STROHMAN PREVIOUSLY RECORDED AT REEL: 064031 FRAME: 0836. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jul 17, 2023
From: DING, SHAOJIN; HE, YANZHANG; WANG, XIN; WANG, WEIRAN; STROHMAN, TREVOR; SAINATH, TARA N; PRABHAVALKAR, ROHIT; DAVID, ROBERT; PANIGRAHY, RINA; BOTROS, RAMI; LIANG, QIAO; MCGRAW, IAN; ZHAO, DING; HWANG, DONGSEONG
To: GOOGLE LLC
Reel/Frame 064286/0548 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: DING, SHAOJIN; HE, YANZHANG; WANG, XIN; WANG, WEIRAN; STROHAM, TREVOR; SAINATH, TARA N.; PRABHAVALKAR, ROHIT; DAVID, ROBERT; PANIGRAHY, RINA; BOTROS, RAMI; LIANG, QIAO; MCGRAW, IAN; ZHAO, DING; HWANG, DONGSEONG
To: GOOGLE LLC
Reel/Frame 064031/0836 →
Continuity (2)
Provisional Application 63269703 · Mar 21, 2022
Related Publication 20230326461A1 · Oct 12, 2023
References Cited (19)
US 11810552B2 · Moritz · 2023 [cited by examiner]
US 20200234713A1 · Gowda et al. · 2020 [cited by applicant]
US 20200349922A1 · Peyser · 2020 [cited by examiner]
US 20210312294A1 · Kurata · 2021 [cited by examiner]
US 20220122622A1 · Narayanan · 2022 [cited by examiner]
US 20220208179A1 · Kurata · 2022 [cited by examiner]
US 20220310062A1 · Sainath · 2022 [cited by examiner]
US 20230017503A1 · Moritz · 2023 [cited by examiner]
US 20230109407A1 · Hu · 2023 [cited by examiner]
US 20230306958A1 · Zhang · 2023 [cited by examiner]
US 20230326461A1 · Ding · 2023 [cited by examiner]
US 20240169981A1 · Huang · 2024 [cited by examiner]
US 20240290320A1 · Huang · 2024 [cited by examiner]
Narayanan A, Sainath TN, Pang R, Yu J, Chiu CC, Prabhavalkar R, Variani E, Strohman T. Cascaded encoders for unifying streaming and non-streaming ASR. arXiv preprint arXiv:2010.14606. Oct. 2, 20207 (Year: 2020). [cited by examiner]
Gao Z, Zhang S, Lei M, McLoughlin I. Universal asr: Unifying streaming and non-streaming asr using a single encoder-decoder model. arXiv preprint arXiv:2010.14099. Oct. 27, 2020.) (Year: 2020). [cited by examiner]
Narayanan A, Sainath TN, Pang R, Yu J, Chiu CC, Prabhavalkar R, Variani E, Strohman T. Cascaded encoders for unifying streaming and non-streaming ASR. arXiv preprint arXiv:2010.14606. Oct. 27, 2020. (Year: 2020). [cited by examiner]
International Search Report and Written opinion for the related Application No. PCT/US2023/064253, dated May 8, 2023, 64 pages. [cited by applicant]
Shaojin Ding et al: “A Unified Cascaded Encoder ASR Model for Dynamic Model Sizes”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 13, 2022 (Apr. 13, 2022), XP091202853… [cited by applicant]
Narayanan Arun et al: “Cascaded Encoders for Unifying Streaming and Non-Streaming ASR”, ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Jun. 6, 2021 (Jun. 6, 202… [cited by applicant]