IP Library Granted Patent US 12,361,927
Granted Patent B2
US 12,361,927 · App. 18/680,797 · Granted Jul 15, 2025

Emitting word timings with end-to-end models

Inventors: Tara N. Sainath (Jersey City, NJ); Basilio Garcia Castillo (Mountain View, CA); David Rybach (Munich, DE); Trevor Strohman (Mountain View, CA); Ruoming Pang (New York, NY)
Assignee: Google LLC
G10L15/063G10L25/30G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,927
App. No.
18/680,797
Granted
Jul 15, 2025
Kind
B2
Abstract

A method includes receiving a training example that includes audio data representing a spoken utterance and a ground truth transcription. For each word in the spoken utterance, the method also includes inserting a placeholder symbol before the respective word identifying a respective ground truth alignment for a beginning and an end of the respective word, determining a beginning word piece and an ending word piece, and generating a first constrained alignment for the beginning word piece and a second constrained alignment for the ending word piece. The first constrained alignment is aligned with the ground truth alignment for the beginning of the respective word and the second constrained alignment is aligned with the ground truth alignment for the ending of the respective word. The method also includes constraining an attention head of a second pass decoder by applying the first and second constrained alignments.

Claims (42)

1. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving audio data characterizing a spoken utterance;

encoding, using an encoder of a neural network speech recognition model, the audio data into a sequence of audio encodings,

processing, using a first decoder of the neural network speech recognition model, the sequence of audio encodings to generate, as output from the first decoder, a plurality of speech recognition hypotheses for the spoken utterance, each speech recognition hypothesis of the plurality of speech recognition hypotheses comprising a corresponding sequence of words;

for each respective speech recognition hypothesis in the top-K speech recognition hypotheses among the plurality of speech recognition hypotheses generated as output from the first decoder, processing, using a second decoder of the neural network speech recognition model, the corresponding sequence of words of the respective speech recognition hypothesis to compute a respective score;

selecting, using the second decoder, as a transcription for the spoken utterance, the corresponding sequence of words of the respective speech recognition hypothesis having the highest respective score; and

determining, using a constrained attention head of the neural network speech recognition model, actual word timings for each word in the corresponding sequence of words selected as the transcription for the spoken utterance.

2. The computer-implemented method of claim 1 , wherein the first decoder comprises a prediction network and a joint network.

3. The computer-implemented method of claim 1 , wherein:

the first decoder of the neural network speech recognition model comprises a recurrent neural network-transducer (RNN-T) decoder; and

the second decoder of the neural network speech recognition model comprises a Listen, Attend, and Spell (LAS) decoder.

4. The computer-implemented method of claim 1 , wherein the data processing hardware resides on a user device that captured the spoken utterance in streaming audio.

5. The computer-implemented method of claim 1 , wherein the operations further comprise displaying, on a screen in communication with the data processing hardware, the transcription of the spoken utterance, the transcription annotated with the actual word timings determined for each word in the corresponding sequence of words selected as the transcription for the spoken utterance.

6. The computer-implemented method of claim 1 , wherein the second decoder comprises a plurality of attention heads.

7. The computer-implemented method of claim 1 , wherein determining the actual word timings for each word in the corresponding sequence of words selected as the transcription for the spoken utterance comprises:

determining a time corresponding to a maximum probability at the constrained attention head of the neural network speech recognition model; and

generating a word start time or a word end time for the determined time corresponding to a maximum probability at the constrained attention head of the neural network speech recognition model.

8. The computer-implemented method of claim 1 , wherein the constrained attention head of the neural network speech recognition model is configured to generate an attention probability that indicates the actual word timings.

9. The computer-implemented method of claim 1 , wherein a training process constrains the constrained attention head of the neural network speech recognition model to generate an attention probability that indicates word timings of an output of the neural network speech recognition model.

10. The computer-implemented method of claim 1 , wherein the plurality of speech recognition hypotheses generated as output from the first decoder comprise a plurality of streaming speech recognition hypotheses.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving audio data characterizing a spoken utterance;

encoding, using an encoder of a neural network speech recognition model, the audio data into a sequence of audio encodings,

processing, using a first decoder of the neural network speech recognition model, the sequence of audio encodings to generate, as output from the first decoder, a plurality of speech recognition hypotheses for the spoken utterance, each speech recognition hypothesis of the plurality of speech recognition hypotheses comprising a corresponding sequence of words;

for each respective speech recognition hypothesis in the top-K speech recognition hypotheses among the plurality of speech recognition hypotheses generated as output from the first decoder, processing, using a second decoder of the neural network speech recognition model, the corresponding sequence of words of the respective speech recognition hypothesis to compute a respective score;

selecting, using the second decoder, as a transcription for the spoken utterance, the corresponding sequence of words of the respective speech recognition hypothesis having the highest respective score; and

determining, using a constrained attention head of the neural network speech recognition model, actual word timings for each word in the corresponding sequence of words selected as the transcription for the spoken utterance.

12. The system of claim 11 , wherein the first decoder comprises a prediction network and a joint network.

13. The system of claim 11 , wherein:

the first decoder of the neural network speech recognition model comprises a recurrent neural network-transducer (RNN-T) decoder; and

the second decoder of the neural network speech recognition model comprises a Listen, Attend, and Spell (LAS) decoder.

14. The system of claim 11 , wherein the data processing hardware resides on a user device that captured the spoken utterance in streaming audio.

15. The system of claim 11 , wherein the operations further comprise displaying, on a screen in communication with the data processing hardware, the transcription of the spoken utterance, the transcription annotated with the actual word timings determined for each word in the corresponding sequence of words selected as the transcription for the spoken utterance.

16. The system of claim 11 , wherein the second decoder comprises a plurality of attention heads.

17. The system of claim 11 , wherein determining the actual word timings for each word in the corresponding sequence of words selected as the transcription for the spoken utterance comprises:

determining a time corresponding to a maximum probability at the constrained attention head of the neural network speech recognition model; and

generating a word start time or a word end time for the determined time corresponding to a maximum probability at the constrained attention head of the neural network speech recognition model.

18. The system of claim 11 , wherein the constrained attention head of the neural network speech recognition model is configured to generate an attention probability that indicates the actual word timings.

19. The system of claim 11 , wherein a training process constrains the constrained attention head of the neural network speech recognition model to generate an attention probability that indicates word timings of an output of the neural network speech recognition model.

20. The system of claim 11 , wherein the plurality of speech recognition hypotheses generated as output from the first decoder comprise a plurality of streaming speech recognition hypotheses.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2024
From: SAINATH, TARA N.; CASTILLO, BASILIO GARCIA; RYBACH, DAVID; STROHMAN, TREVOR; PANG, RUOMING
To: GOOGLE LLC
Reel/Frame 067585/0548 →
Continuity (4)
Continuation 18167050 · Feb 9, 2023
Continuation 17204852 · Mar 17, 2021
Provisional Application 63021660 · May 7, 2020
Related Publication 20240321263A1 · Sep 26, 2024
References Cited (26)
US 3535457A · Poschenrieder · 1970 [cited by examiner]
US 6822589B1 · Dye · 2004 [cited by examiner]
US 10706840B2 · Sak · 2020 [cited by examiner]
US 20060053004A1 · Ceperkovic · 2006 [cited by examiner]
US 20140007250A1 · Stefanov · 2014 [cited by examiner]
US 20170148226A1 · Zhang · 2017 [cited by examiner]
US 20180107925A1 · Choi · 2018 [cited by examiner]
US 20190393903A1 · Mandt · 2019 [cited by examiner]
US 20200184278A1 · Zadeh · 2020 [cited by examiner]
US 20210064822A1 · Velikovich · 2021 [cited by examiner]
US 20210089863A1 · Yang · 2021 [cited by examiner]
US 20210312905A1 · Zhao · 2021 [cited by examiner]
US 20210350794A1 · Sainath · 2021 [cited by examiner]
US 20220083743A1 · Chiu · 2022 [cited by examiner]
A. Gandhe and A. Rastrow, “Audio-Attention Discriminative Language Model for ASR Rescoring,” ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2020, pp… [cited by examiner]
P. Serai, A. Stiff and E. Fosler-Lussier, “End to End Speech Recognition Error Prediction with Sequence to Sequence Learning,” ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (I… [cited by examiner]
International Search Report for the related Application No. PCT/US2021/022851, Dated Jul. 15, 2021. [cited by applicant]
Tara N . Sainath et al : “Two-Pass End-to-End Speech Recognition”, arxv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 29, 2019 (Aug. 29, 2019), XP081489070, p. 1, paragraph … [cited by applicant]
Li Bo et al: 11 Towards Fast and Accurate Streaming End-To-End ASR 11, ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (I CASSP), I EEE, May 4, 2020 (May 4, 2020), pp. 6069-6073… [cited by applicant]
Sainath Tara N et al: 11 An Attention-Based Joint Acoustic and Text on-Device End-To-End Model , ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, 2020 (May… [cited by applicant]
Sainath Tara N. et al: “Emitting Word Timings with End-to-End Models”, INTERSPEECH 2020, [Online] Oct. 25, 202 (Oct. 25, 2020), pp. 3615-3619, XP055817592, ISCA, DOI: 10.21437/Interspeech. 2020-1059. Retrieved from the … [cited by applicant]
Chinese Office Action for application No. 201880027521.2 dated Dec. 5, 2022. [cited by applicant]
P. Bell and S. Renals, “A system for automatic alignment of broadcast media captions using weighted finite-state transducers,” 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), 2015, pp. 675-6… [cited by applicant]
H. Hu, R. Zhao, J. Li, L. Lu and Y. Gong, “Exploring Pre-Training with Alignments for RNN Transducer Based End-to-End Speech Recognition,” ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal P… [cited by applicant]
T. Afouras, J. S. Chung, A. Senior, O. Vinyals and A. Zisserman, “Deep Audio-Visual Speech Recognition,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 12, pp. 8717-8727, Dec. 1, 2022, d… [cited by applicant]
USPTO. Office Action relatiing U.S. Appl. No. 18/167,050, dated Nov. 24, 2024. [cited by applicant]