IP Library › Granted Patent US 12,450,480
Granted Patent B2
US 12,450,480 · App. 17/544,570 · Granted Oct 21, 2025

Self-adaptive distillation

Inventors: Isabel Leal (Mountain View, CA); Neeraj Gaur (Mountain View, CA); Parisa Haghani (Mountain View, CA); Brian Farris (Mountain View, CA); Bhuvana Ramabhadran (Mt. Kisco, NY); Manasa Prasad (Mountain View, CA); Pedro J. Moreno Mengibar (Jersey City, NJ); Yun Zhu (Mountain View, CA)
Assignee: Google LLC
G06N3/08G06N3/045G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,480
App. No.
17/544,570
Granted
Oct 21, 2025
Kind
B2
Abstract

A method for distilling one or more trained teacher automatic speech recognition (ASR) models into a multilingual student model includes receiving a plurality of teacher training examples and a plurality of student training examples. The method also includes training one or more teacher automatic speech recognition (ASR) models using the plurality of teacher training examples. Each teacher ASR model is configured to output a respective textual representation of a respective audio input. The method further includes generating a multi-lingual student ASR model by training the multi-lingual student ASR model using the plurality of student training examples and distilling the trained one or more teacher ASR models into the multilingual student ASR model using a tunable distillation loss weight. Each student ASR model is configured to receive an audio input and output a corresponding textual representation of the received audio input.

Claims (32)

1. A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:

receiving a plurality of teacher training examples and a plurality of student training examples;

training one or more teacher automatic speech recognition (ASR) models using the plurality of teacher training examples, each teacher ASR model comprising a recurrent neural network-transducer (RNN-T) architecture and configured to output a respective textual representation of a respective audio input; and

generating a multi-lingual student ASR model comprising an RNN-T architecture by:

training the multi-lingual student ASR model using the plurality of student training examples, the student ASR model configured to receive an audio input and to output a corresponding textual representation of the received audio input; and

distilling the trained one or more teacher ASR models into the multi-lingual student ASR model by applying a tunable weight to a distillation loss of the multi-lingual student ASR model to generate a tunable distillation loss weight, the tunable distillation loss weight comprising a decreasing function based on a first RNN-T loss corresponding to the one or more teacher ASR models and a second RNN-T loss corresponding to the multi-lingual student ASR model.

2. The method of claim 1 , wherein the one or more teacher ASR models are configured to collectively recognize fewer languages than the multi-lingual student ASR model.

3. The method of claim 1 , wherein:

training the multi-lingual student ASR model occurs across n number of training steps; and

the tunable distillation loss weight comprises a decreasing function decreasing based on the n number of training steps.

4. The method of claim 1 , wherein the decreasing function:

decreases the first RNN-T loss corresponding to the one or more teacher ASR models over an instance of time; and

increases the second RNN-T loss corresponding to the multi-lingual student ASR model over the instance of time.

5. The method of claim 1 , wherein each teacher ASR model of the one or more teacher ASR models corresponds to a mono-lingual teacher ASR model.

6. The method of claim 1 , wherein the one or more teacher ASR models corresponds to a single multi-lingual ASR model.

7. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a plurality of teacher training examples and a plurality of student training examples;

training one or more teacher automatic speech recognition (ASR) models using the plurality of teacher training examples, each teacher ASR model comprising a recurrent neural network-transducer (RNN-T) architecture and configured to output a respective textual representation of a respective audio input; and

generating a multi-lingual student ASR model comprising an RNN-T architecture by:

training the multi-lingual student ASR model using the plurality of student training examples, the student ASR model configured to receive an audio input and to output a corresponding textual representation of the received audio input; and

distilling the trained one or more teacher ASR models into the multi-lingual student ASR model by applying a tunable weight to a distillation loss of the multi-lingual student ASR model to generate a tunable distillation loss weight, the tunable distillation loss weight comprising a decreasing function based on a first RNN-T loss corresponding to the one or more teacher ASR models and a second RNN-T loss corresponding to the multi-lingual student ASR model.

8. The system of claim 7 , wherein the one or more teacher ASR models are configured to collectively recognize fewer languages than the multi-lingual student ASR model.

9. The system of claim 7 , wherein:

training the multi-lingual student ASR model occurs across n number of training steps; and

the tunable distillation loss weight comprises a decreasing function decreasing based on the n number of training steps.

10. The system of claim 7 , wherein the decreasing function:

decreases the first RNN-T loss corresponding to the one or more teacher ASR models over an instance of time; and

increases the second RNN-T loss corresponding to the multi-lingual student ASR model over the instance of time.

11. The system of claim 7 , wherein each teacher ASR model of the one or more teacher ASR models corresponds to a mono-lingual teacher ASR model.

12. The system of claim 7 , wherein the one or more teacher ASR models corresponds to a single multi-lingual ASR model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2022
From: LEAL, ISABEL; GAUR, NEERAJ; HAGHANI, PARISA; FARRIS, BRIAN; RAMABHADRAN, BHUVANA; PRASAD, MANASA; MENGIBAR, PEDRO J. MORENO; ZHU, YUN
To: GOOGLE LLC
Reel/Frame 058746/0393 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2021
From: LEAL, ISABEL; GAUR, NEERAJ; HAGHANI, PARISA; FARRIS, BRIAN; RAMABHADRAN, BHUVANA; PRASAD, MANASA; MENGIBAR, PEDRO J. MORENO; ZHU, YUN
To: GOOGLE LLC
Reel/Frame 058335/0080 →
Continuity (2)
Provisional Application 63166938 · Mar 26, 2021
Related Publication 20220309340A1 · Sep 29, 2022
References Cited (16)
US 20200387782A1 · Hegde · 2020 [cited by examiner]
US 20220199258A1 · Yoo · 2022 [cited by examiner]
US 20220343175A1 · Lu · 2022 [cited by examiner]
JP 2020537765A · 2020 [cited by applicant]
WO 2020242580A1 · 2020 [cited by applicant]
Xu et al., Knowledge Distillation from Multilingual and Monolingual Teachers from End-to-End Multilingual Speech Recognition, Proceedings of APSIPA Annual Summit and Conference 2019, Nov. 18-21, 2019 (Year: 2019). [cited by examiner]
Wang et al., Structure-Level Knowledge Distillation for Multilingual Sequence Labeling, https://arxiv.org/abs/2004.03846, May 4, 2020 (Year: 2020). [cited by examiner]
Mar. 28, 2022 Written Opinion (WO) of the International Searching Authority (ISA) and International Search Report (ISR) issued in International Application No. PCT/US2021/062255. [cited by applicant]
Xu Jingyi et al: “Knowledge Distillation from Multilingual and Monolingual Teachers 9-13,19, for End-to-End Multilingual Speech Recognition”, 2019 Asia-Pacific Signal and Information Processing Association Annual Summit… [cited by applicant]
Hayato Futami et al: “Distilling the Knowledge of BERT for Sequence-to-Sequence ASR”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 9, 2020 (Aug. 9, 2020), XP081737409. [cited by applicant]
Kinyu Wang et al: “Structure-Level Knowledge Distillation For Multilingual Sequence Labeling”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 8, 2020 (Apr. 8, 2020), XP… [cited by applicant]
Kevin Clark et al: “BAM1 Born-Again Multi-Task Networks for Natural Language Understanding”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jul. 10, 2019 (Jul. 10, 2019), XP… [cited by applicant]
Xinyu Wang et al: “Structural Knowledge Distillation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 10, 2020 (Oct. 10, 2020), XP081783511. [cited by applicant]
Han Zhu et al: “Domain Adaptation Using Class Similarity for Robust Speech Recognition”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 5, 2020 (Nov. 5, 2020), XP081796… [cited by applicant]
Leal Isabel et al: “Self-Adaptive Distillation for Multilingual Speech Recognition: Leveraging Student Independence”, INTERSPEECH 2021, Aug. 30, 2021 (Aug. 30, 2021), pp. 2556-2560, XP055898227. [cited by applicant]
Office Action issued in related Japanese Patent Application No. 2023-558805, dated Nov. 19, 2024. [cited by applicant]