IP Library › Granted Patent US 12,725,607
Granted Patent B2
US 12,725,607 · App. 18/772,263 · Granted Sep 1, 2026

Efficient streaming non-recurrent on-device end-to-end model

Inventors: Tara Sainath (Jersey City, NJ); Arun Narayanan (Milpitas, CA); Rami Botros (Mountain View, CA); Yanzhang He (Mountain View, CA); Ehsan Variani (Mountain View, CA); Cyril Allauzen (Mountain View, CA); David Rybach (Munich, DE); Ruoming Pang (New York, NY); Trevor Strohman (Mountain View, CA)
Assignee: Google LLC
G10L15/063G10L15/02G10L15/22G10L15/30G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,607
App. No.
18/772,263
Granted
Sep 1, 2026
Kind
B2
Abstract

An ASR model includes a first encoder configured to receive a sequence of acoustic frames and generate a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames. The ASR model also includes a second encoder configured to receive the first higher order feature representation generated by the first encoder at each of the plurality of output steps and generate a second higher order feature representation for a corresponding first higher order feature frame. The ASR model also includes a decoder configured to receive the second higher order feature representation generated by the second encoder at each of the plurality of output steps and generate a first probability distribution over possible speech recognition hypothesis. The ASR model also includes a language model configured to receive the first probability distribution over possible speech hypothesis and generate a rescored probability distribution.

Claims (32)

1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving a sequence of acoustic frames corresponding to an utterance as a user speaks the utterance;

at a corresponding output step of a plurality of output steps each associated with a corresponding acoustic frame in the sequence of acoustic frames:

processing, using a stack of multi-headed attention layers, the corresponding acoustic frame to generate a corresponding higher order feature representation; and

generating, by a decoder configured to receive the corresponding higher order feature representation generated at the corresponding output step, a probability distribution over possible output labels; and

rescoring, by an external language model, the probability distribution over possible output labels generated by the decoder at each of the plurality of output steps to generate a transcription of the utterance.

2 . The computer-implemented method of claim 1 , wherein processing the corresponding acoustic frame commences a predefined duration after an initial acoustic frame in the sequence of acoustic frames is received.

3 . The computer-implemented method of claim 1 , wherein the stack of multi-headed attention layers comprises a stack of Transformer layers.

4 . The computer-implemented method of claim 1 , wherein the stack of multi-headed attention layers comprises a stack of Conformer layers.

5 . The computer-implemented method of claim 1 , wherein the possible output labels comprise wordpieces.

6 . The computer-implemented method of claim 1 , wherein the possible output labels comprise graphemes.

7 . The computer-implemented method of claim 1 , wherein the external language model comprises a neural language model.

8 . The computer-implemented method of claim 7 , wherein the neural language model comprises a plurality of multi-headed attention layers.

9 . The computer-implemented method of claim 1 , wherein the external language model is trained on text-only data.

10 . The computer-implemented method of claim 1 , wherein the utterance comprises a long-form utterance that comprises a plurality of sentences.

11 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

receiving a sequence of acoustic frames corresponding to an utterance as a user speaks the utterance;

at a corresponding output step of a plurality of output steps each associated with a corresponding acoustic frame in the sequence of acoustic frames:

processing, using a stack of multi-headed attention layers, the corresponding acoustic frame to generate a corresponding higher order feature representation; and

generating, by a decoder configured to receive the corresponding higher order feature representation generated at the corresponding output step, a probability distribution over possible output labels; and

rescoring, by an external language model, the probability distribution over possible output labels generated by the decoder at each of the plurality of output steps to generate a transcription of the utterance.

12 . The system of claim 11 , wherein processing the corresponding acoustic frame commences a predefined duration after an initial acoustic frame in the sequence of acoustic frames is received.

13 . The system of claim 11 , wherein the stack of multi-headed attention layers comprises a stack of Transformer layers.

14 . The system of claim 11 , wherein the stack of multi-headed attention layers comprises a stack of Conformer layers.

15 . The system of claim 11 , wherein the possible output labels comprise wordpieces.

16 . The system of claim 11 , wherein the possible output labels comprise graphemes.

17 . The system of claim 11 , wherein the external language model comprises a neural language model.

18 . The system of claim 17 , wherein the neural language model comprises a plurality of multi-headed attention layers.

19 . The system of claim 11 , wherein the external language model is trained on text-only data.

20 . The system of claim 11 , wherein the utterance comprises a long-form utterance that comprises a plurality of sentences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2024
From: SAINATH, TARA; NARAYANAN, ARUN; BOTROS, RAMI; HE, YANZHANG; VARIANI, EHSAN; ALLAUZEN, CYRIL; RYBACH, DAVID; PANG, RUOMING; STROHMAN, TREVOR
To: GOOGLE LLC
Reel/Frame 067983/0500 →
Continuity (4)
Continuation 18336211 · Jun 16, 2023
Continuation 17316198 · May 10, 2021
Provisional Application 63165068 · Mar 23, 2021
Related Publication 20240371363A1 · Nov 7, 2024
References Cited (16)
US 9502029B1 · Bell · 2016 [cited by examiner]
US 10186255B2 · Tapuhi · 2019 [cited by examiner]
US 11715458B2 · Sainath et al. · 2023 [cited by applicant]
US 12051404B2 · Sainath · 2024 [cited by examiner]
US 20200349950A1 · Yoshioka et al. · 2020 [cited by applicant]
US 20200349954A1 · Yoshioka et al. · 2020 [cited by applicant]
US 20210034966A1 · Qian · 2021 [cited by applicant]
JP 2020129015A · 2020 [cited by applicant]
WO 2019163718A1 · 2019 [cited by applicant]
USPTO. Office Action relating to U.S. Appl. No. 17/316,198, dated Dec. 8, 2022. [cited by applicant]
USPTO. Office Action relating to U.S. Appl. No. 18/336,211, dated Feb. 29, 2024. [cited by applicant]
Japanese Office Action for the related Application No. 2023-558609, dated Feb. 12, 2025. [cited by applicant]
Ke Hu et al.: “Deliberation Model Based Two-Pass End-to-End Speech Recognition”, Mar. 17, 2020 (Mar. 17, 2020), https://arxiv.org/abs/2003.07962. [cited by applicant]
Arun Narayanan et al.: “Cascaded Encoders for Unifying Streaming and Non-Streaming ASR”, Oct. 27, 2020 (Oct. 27, 2020), https://arxiv.org/abs/2010.14606. [cited by applicant]
Ehsan Variani et al.: “Hybrid Autoregressive Transducer”, Mar. 12, 2020 (Mar. 12, 2020), https://arxiv.org/abs/2003.07705. [cited by applicant]
Chinese Office Action for the related Application No. 202180096175.5 dated May 30, 2026. [cited by applicant]