IP Library › Granted Patent US 12,738,264
Granted Patent B2
US 12,738,264 · App. 18/585,020 · Granted Sep 15, 2026

Semantic segmentation with language models for long-form automatic speech recognition

Inventors: Wenqian Huang (Mountain View, CA); Hao Zhang (Jericho, NY); Shankar Kumar (New York, NY); Shuo-yiin Chang (Sunnyvale, CA); Tara N. Sainath (Jersey City, NJ)
Assignee: Google LLC
G10L15/063G06F40/30G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,264
App. No.
18/585,020
Granted
Sep 15, 2026
Kind
B2
Abstract

A joint segmenting and ASR model includes an encoder to receive a sequence of acoustic frames and generate, at each of a plurality of output steps, a higher order feature representation for a corresponding acoustic frame. The model also includes a decoder to generate based on the higher order feature representation at each of the plurality of output steps a probability distribution over possible speech recognition hypotheses, and an indication of whether the corresponding output step corresponds to an end of segment (EOS). The model is trained on a set of training samples, each training sample including audio data characterizing multiple segments of long-form speech; and a corresponding transcription of the long-form speech, the corresponding transcription annotated with ground-truth EOS labels obtained via distillation from a language model teacher that receives the corresponding transcription as input and injects the ground-truth EOS labels into the corresponding transcription between semantically complete segments.

Claims (78)

1 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to implement a joint segmenting and automated speech recognition (ASR) model comprising:

an encoder configured to:

receive, as input, a sequence of acoustic frames characterizing one or more spoken utterances; and

generate, at each output step of a plurality of output steps, a higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames; and

a decoder configured to:

receive, as input, the higher order feature representation generated by the encoder at each output step of the plurality of output steps; and

generate, at each output step of the plurality of output steps:

a probability distribution over possible speech recognition hypotheses; and

an indication of whether the output step corresponds to an end of segment,

wherein the joint segmenting and ASR model is trained on a set of training samples, each training sample in the set of training samples comprising:

audio data characterizing multiple segments of long-form speech; and

a corresponding transcription of the long-form speech, the corresponding transcription annotated with ground-truth end of segment labels obtained via distillation from a language model teacher that receives the corresponding transcription as input and injects the ground-truth end of segment labels into the corresponding transcription between semantically complete segments,

wherein:

the language model teacher is trained on a corpus of written text containing punctuation to teach the language model teacher to learn how to semantically predict ground-truth end of segment labels based on positions of punctuation in the written text;

the language model teacher processes the corresponding transcription to generate predicted punctuation; and

an augmentor inserts a ground-truth end of segment label into the corresponding transcription for each comma, period, question mark, and exclamation point predicted by the language model teacher.

2 . The system of claim 1 , wherein the language model teacher comprises a bi-directional recurrent neural network architecture.

3 . The system of claim 1 , wherein the decoder comprises:

a prediction network configured to, at each output step of the plurality of output steps:

receive, as input, a sequence of non-blank symbols output by a final Softmax layer; and

generate a hidden representation;

a first joint network configured to:

receive, as input, the hidden representation generated by the prediction network at each output step of the plurality of output steps and the higher order feature representation generated by the encoder at each output step of the plurality of output steps; and

generate, at each output step of the plurality of output steps, the probability distribution over possible speech recognition hypotheses; and

a second joint network configured to:

receive, as input, the hidden representation generated by the prediction network at each output step of the plurality of output steps and the higher order feature representation generated by the encoder at each output step of the plurality of output steps; and

generate, at each of output step the plurality of output steps, the indication of whether the output step corresponds to an end of segment.

4 . The system of claim 3 , wherein, at each output step of the plurality of output steps:

the sequence of previous non-blank symbols received as input at the prediction network comprises a sequence of N previous non-blank symbols output by the final Softmax layer; and

the prediction network is configured to generate the hidden representation by:

for each non-blank symbol of the sequence of N previous non-blank symbols, generating a respective embedding; and

generating an average embedding by averaging the respective embeddings, the average embedding comprising the hidden representation.

5 . The system of claim 3 , wherein the prediction network comprises an embedding look-up table.

6 . The system of claim 3 , wherein a training process trains the joint segmenting and ASR model on the set of training samples by:

initially training the first joint network to learn how to predict the corresponding transcription of the spoken utterance characterized by the audio data of each training sample; and

after training the first joint network, initializing the second joint network with the same parameters as the trained first joint network and using the ground-truth end of segment label inserted into the corresponding transcription of the spoken utterance characterized by the audio data of each training sample.

7 . The system of claim 1 , wherein the encoder comprises a causal encoder comprising a stack of conformer layers or transformer layers.

8 . The system of claim 1 , wherein the ground-truth end of segment labels are inserted into the corresponding transcription automatically without any human annotation.

9 . The system of claim 1 , wherein the joint segmenting and ASR model is trained to maximize a probability of emitting the ground-truth end of segment label.

10 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to implement a joint segmenting and automated speech recognition (ASR) model, the joint segmenting and ASR model comprising:

an encoder configured to:

receive, as input, a sequence of acoustic frames characterizing one or more spoken utterances; and

generate, at each output step of a plurality of output steps, a higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames; and

a decoder configured to:

receive, as input, the higher order feature representation generated by the encoder at each output step of the plurality of output steps; and

generate, at each output step of the plurality of output steps:

a probability distribution over possible speech recognition hypotheses; and

an indication of whether the output step corresponds to an end of segment, wherein the joint segmenting and ASR model is trained on a set of training samples, each training sample in the set of training samples comprising:

audio data characterizing multiple segments of long-form speech; and

a corresponding transcription of the long-form speech, the corresponding transcription annotated with ground-truth end of segment labels obtained via distillation from a language model teacher that receives the corresponding transcription as input and injects the ground-truth end of segment labels into the corresponding transcription between semantically complete segments,

wherein the language model teacher is trained on a corpus of written text containing punctuation to teach the language model teacher to learn how to semantically predict ground-truth end of segment labels based on positions of punctuation in the written text,

wherein the language model teacher processes the corresponding transcription to generate predicted punctuation, and

wherein an augmentor inserts a ground-truth end of segment label into the corresponding transcription for each comma, period, question mark, and exclamation point predicted by the language model teacher.

11 . The computer-implemented method of claim 10 , wherein the language model teacher comprises a bi-directional recurrent neural network architecture.

12 . The computer-implemented method of claim 10 , wherein the decoder comprises:

a prediction network configured to, at each output step of the plurality of output steps:

receive, as input, a sequence of non-blank symbols output by a final Softmax layer; and

generate a hidden representation;

a first joint network configured to:

receive, as input, the hidden representation generated by the prediction network at each output step of the plurality of output steps and the higher order feature representation generated by the encoder at each output step of the plurality of output steps; and

generate, at each output step of the plurality of output steps, the probability distribution over possible speech recognition hypotheses; and

a second joint network configured to:

receive, as input, the hidden representation generated by the prediction network at each output step of the plurality of output steps and the higher order feature representation generated by the encoder at each output step of the plurality of output steps; and

generate, at each output step of the plurality of output steps, the indication of whether the output step corresponds to an end of segment.

13 . The computer-implemented method of claim 12 , wherein, at each output step of the plurality of output steps:

the sequence of previous non-blank symbols received as input at the prediction network comprises a sequence of N previous non-blank symbols output by the final Softmax layer; and

the prediction network is configured to generate the hidden representation by:

for each non-blank symbol of the sequence of N previous non-blank symbols, generating a respective embedding; and

generating an average embedding by averaging the respective embeddings, the average embedding comprising the hidden representation.

14 . The computer-implemented method of claim 12 , wherein the prediction network comprises an embedding look-up table.

15 . The computer-implemented method of claim 12 , wherein a training process trains the joint segmenting and ASR model on the set of training samples by:

initially training the first joint network to learn how to predict the corresponding transcription of the spoken utterance characterized by the audio data of each training sample; and

after training the first joint network, initializing the second joint network with the same parameters as the trained first joint network and using the ground-truth end of segment labels inserted into the corresponding transcription of the spoken utterance characterized by the audio data of each training sample.

16 . The computer-implemented method of claim 10 , wherein the encoder comprises a causal encoder comprising a stack of conformer layers or transformer layers.

17 . The computer-implemented method of claim 10 , wherein the ground-truth end of segment labels are inserted into the corresponding transcription automatically without any human annotation.

18 . The computer-implemented method of claim 10 , wherein the joint segmenting and ASR model is trained to maximize a probability of emitting the ground-truth end of segment labels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2024
From: HUANG, WENQIAN; ZHANG, HAO; KUMAR, SHANKAR; CHANG, SHUO-YIIN; SAINATH, TARA N.
To: GOOGLE LLC
Reel/Frame 067594/0418 →
Continuity (2)
Provisional Application 63487600 · Feb 28, 2023
Related Publication 20240290320A1 · Aug 29, 2024
References Cited (19)
US 20200335091A1 · Chang · 2020 [cited by examiner]
US 20200349922A1 · Peyser · 2020 [cited by examiner]
US 20210350794A1 · Sainath et al. · 2021 [cited by applicant]
US 20220238101A1 · Sainath et al. · 2022 [cited by applicant]
US 20220310062A1 · Sainath et al. · 2022 [cited by applicant]
US 20220310071A1 · Botros et al. · 2022 [cited by applicant]
US 20230343332A1 · Huang · 2023 [cited by examiner]
US 20240169981A1 · Huang · 2024 [cited by examiner]
Zhou, Zhikai, Tian Tan, and Yanmin Qian. “Punctuation prediction for streaming on-device speech recognition.” ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 202… [cited by examiner]
Huang, W. Ronny, et al. “E2e segmenter: Joint segmenting and decoding for long-form asr.” arXiv preprint arXiv:2204.10749 (Apr. 2022). (Year: 2022). [cited by examiner]
Jain, “RNN-T Based ASR Systems”, [online] www.sites.cc.gatech.edu, published in 2021. (Year: 2021). [cited by examiner]
Behre, Piyush, et al. “Streaming punctuation: A novel punctuation technique leveraging bidirectional context for continuous speech recognition.” arXiv preprint arXiv:2301.03819 (2022). (Year: 2022). [cited by examiner]
Ronny Huang W et al, “E2E Segmenter: Joint Segmenting and Decoding for Long-Form AST”, arxiv.org, Cornell University Library, 201 OLIN Library Cornell University Ithaca, NY 14853, Jun. 15, 2022 (Jun. 15, 2022), XP091246… [cited by applicant]
Ronny Huang W et al, “E2E Segmentation in a Two-Pass Cascaded Encoder ASR Model”, arxiv.org, Cornell University Library, 201 OLIN Library Cornell University Ithaca, NY 14853,Nov. 28, 2022 (Nov. 28, 2022), XP091380649. [cited by applicant]
Li Bo et al, “Towards Fast and Accurate Streaming End-To-End ASR”, ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE,May 4, 2020 (May 4, 2020), p. 6069-6073, XP0337… [cited by applicant]
Chang Shuo-Yiin et al, “Joint Endpointing and Decoding with End-to-end Models”, ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 12, 2019 (May 12, 2019), p. 5… [cited by applicant]
Bo Li et al, “A Language Agnostic Multilingual Streaming On-Device ASR System”, arxiv.org, Cornell University Library, 201 OLIN Library Cornell University Ithaca, NY 14853,Aug. 29, 2022 (Aug. 29, 2022), XP091305038. [cited by applicant]
Ronny Huang W et al, “Semantic Segmentation with Bidirectional Language Models Improves Long-form ASR”, arxiv.org, Cornell University Library, 201 OLIN Library Cornell University Ithaca, NY 14853,May 28, 2023 (May 28, 2… [cited by applicant]
International Search Report and Written Opinion issued in related PCT Application No. PCT/US2024/016965, dated May 2, 2024. [cited by applicant]