IP Library › Granted Patent US 11,804,212
Granted Patent B2
US 11,804,212 · App. 17/348,118 · Granted Oct 31, 2023

Streaming automatic speech recognition with non-streaming model distillation

Inventors: Thibault Doutre (Mountain View, CA); Wei Han (Mountain View, CA); Min Ma (Mountain View, CA); Zhiyun Lu (Mountain View, CA); Chung-Cheng Chiu (Sunnyvale, CA); Ruoming Pang (New York, NY); Arun Narayanan (Santa Clara, CA); Ananya Misra (Mountain View, CA); Yu Zhang (Mountain View, CA); Liangliang Cao (Mountain View, CA)
Assignee: Google LLC
G10L15/063G06N3/045G10L15/083G10L15/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,804,212
App. No.
17/348,118
Granted
Oct 31, 2023
Kind
B2
Abstract

A method for training a streaming automatic speech recognition student model includes receiving a plurality of unlabeled student training utterances. The method also includes, for each unlabeled student training utterance, generating a transcription corresponding to the respective unlabeled student training utterance using a plurality of non-streaming automated speech recognition (ASR) teacher models. The method further includes distilling a streaming ASR student model from the plurality of non-streaming ASR teacher models by training the streaming ASR student model using the plurality of unlabeled student training utterances paired with the corresponding transcriptions generated by the plurality of non-streaming ASR teacher models.

Claims (46)

1. A computer-implemented method executed by data processing hardware that causes the data processing hardware to perform operations comprising:

receiving a plurality of unlabeled student training utterances;

for each unlabeled student training utterance of the plurality of unlabeled student training utterances, generating a transcription corresponding to the unlabeled student training utterance using a plurality of non-streaming automated speech recognition (ASR) teacher models; and

distilling a streaming ASR student model from the plurality of non-streaming ASR teacher models by training the streaming ASR student model using the plurality of unlabeled student training utterances paired with the corresponding transcriptions generated by the plurality of non-streaming ASR teacher models,

wherein:

generating the transcription corresponding to the unlabeled student training utterance comprises:

receiving, as input at the plurality of non-streaming ASR teacher models, the unlabeled student training utterance;

at each non-streaming ASR teacher model, predicting an initial transcription for the unlabeled student training utterance; and

generating the transcription for the respective unlabeled student training utterance to be output by the plurality of non-streaming ASR teacher models based on the initial transcriptions of each non-streaming ASR teacher model predicted for the respective unlabeled student training utterance;

generating the transcription for the respective unlabeled student training utterance to be output by the plurality of non-streaming ASR teacher models based on the initial transcriptions of each non-streaming ASR teacher model predicted for the respective unlabeled student training utterance comprises constructing the transcription using output voting; and

constructing the transcription using output voting comprises:

aligning the initial transcriptions from each non-streaming ASR teacher model to define a sequence of frames;

dividing each initial transcription into transcription segments, each transcription segment corresponding to a respective frame;

for each respective frame, selecting a most repeated transcription segment across all initial transcriptions; and

concatenating the most repeated transcription segment of each respective frame to form the transcription.

2. The method of claim 1 , wherein the streaming ASR student model comprises a recurrent neural network transducer (RNN-T) architecture.

3. The method of claim 1 , wherein the streaming ASR student model comprises a conformer-based encoder.

4. The method of claim 1 , wherein each non-streaming ASR teacher model comprises a connectionist temporal classification (CTC) architecture.

5. The method of claim 4 , wherein the CTC architecture comprises a language model configured to capture contextual information for a respective utterance.

6. The method of claim 1 , wherein each non-streaming ASR teacher model comprises a conformer-based encoder.

7. The method of claim 1 , wherein the plurality of non-streaming ASR teacher models comprise at least two different recurrent neural network architectures.

8. The method of claim 7 , wherein a first non-streaming ASR teacher model comprises a recurrent neural network architecture and a second non-streaming ASR teacher model comprises a connectionist temporal classification (CTC) architecture.

9. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving a plurality of unlabeled student training utterances;

for each unlabeled student training utterance of the plurality of unlabeled student training utterances, generating a transcription corresponding to the unlabeled student training utterance using a plurality of non-streaming automated speech recognition (ASR) teacher models; and

distilling a streaming ASR student model from the plurality of non-streaming ASR teacher models by training the streaming ASR student model using the plurality of unlabeled student training utterances paired with the corresponding transcriptions generated by the plurality of non-streaming ASR teacher models,

wherein:

generating the transcription corresponding to the unlabeled student training utterance comprises:

receiving, as input at the plurality of non-streaming ASR teacher models, the unlabeled student training utterance;

at each non-streaming ASR teacher model, predicting an initial transcription for the unlabeled student training utterance; and

generating the transcription for the respective unlabeled student training utterance to be output by the plurality of non-streaming ASR teacher models based on the initial transcriptions of each non-streaming ASR teacher model predicted for the respective unlabeled student training utterance;

generating the transcription for the respective unlabeled student training utterance to be output by the plurality of non-streaming ASR teacher models based on the initial transcriptions of each non-streaming ASR teacher model predicted for the respective unlabeled student training utterance comprises constructing the transcription using output voting; and

constructing the transcription using output voting comprises:

aligning the initial transcriptions from each non-streaming ASR teacher model to define a sequence of frames;

dividing each initial transcription into transcription segments, each transcription segment corresponding to a respective frame;

for each respective frame, selecting a most repeated transcription segment across all initial transcriptions; and

concatenating the most repeated transcription segment of each respective frame to form the transcription.

10. The system of claim 9 , wherein the streaming ASR student model comprises a recurrent neural network transducer (RNN-T) architecture.

11. The system of claim 9 , wherein the streaming ASR student model comprises a conformer-based encoder.

12. The system of claim 9 , wherein each non-streaming ASR teacher model comprises a connectionist temporal classification (CTC) architecture.

13. The system of claim 12 , wherein the CTC architecture comprises a language model configured to capture contextual information for a respective utterance.

14. The system of claim 9 , wherein each non-streaming ASR teacher model comprises a conformer-based encoder.

15. The system of claim 9 , wherein the plurality of non-streaming ASR teacher models comprise at least two different recurrent neural network architectures.

16. The system of claim 15 , wherein a first non-streaming ASR teacher model comprises a recurrent neural network architecture and a second non-streaming ASR teacher model comprises a connectionist temporal classification (CTC) architecture.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2021
From: DOUTRE, THIBAULT; HAN, WEI; MA, MIN; LU, ZHIYUN; CHIU, CHUNG-CHENG; PANG, RUOMING; NARAYANAN, ARUN; MISRA, ANANYA; ZHANG, ZU; CAO, LIANGLIANG
To: GOOGLE LLC
Reel/Frame 057357/0315 →
Continuity (2)
Provisional Application 63179084 · Apr 23, 2021
Related Publication 20220343894A1 · Oct 27, 2022