IP Library › Granted Patent US 12,640,141
Granted Patent B1
US 12,640,141 · App. 18/649,203 · Granted May 26, 2026

Semi-supervised training of automatic speech recognition systems using iterative

Inventors: Juan Hussain (Wörth am Rhein, DE); Thai Son Nguyen (Karlsruhe, DE); Sebastian Stüker (Karlsruhe, DE); Zhirong Ye (Hangzhou, CN)
Assignee: Zoom Communications, Inc.
G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,640,141
App. No.
18/649,203
Granted
May 26, 2026
Kind
B1
Abstract

One example method includes generating teacher and student ASR models from an ASR model; generating a plurality of successive generations of the teacher ASR model, including: generating, for each unannotated audio sample of a training data set, one or more pseudo-labels corresponding to the respective unannotated audio sample using the current generation teacher ASR model; generating a plurality of successive generations of the student ASR model, including: providing a subset of the training data set as training inputs to the student ASR model, the training data set including the plurality of unannotated audio samples and the corresponding pseudo-labels, training the student ASR model using the subset of the training data set to generate a next generation of the student ASR model, and evaluating the current student ASR model to determine whether to include the current model in a subset of the generations of the student ASR model; generating a next generation of the teacher ASR model based on the subset of the generations of the student ASR model; responsive to reaching a completion condition, outputting a current generation of the teacher ASR model.

Claims (60)

1 . A method comprising:

generating a teacher ASR model and a student ASR model from an ASR model;

generating a plurality of successive generations of the teacher ASR model, comprising, for a respective teacher ASR model generation:

generating, for each unannotated audio sample of a plurality of unannotated audio samples of a training data set, one or more pseudo-labels corresponding to the respective unannotated audio sample using the current generation teacher ASR model;

generating a plurality of successive generations of the student ASR model, comprising, for a respective student ASR model generation:

providing a subset of the training data set as training inputs to the respective generation of the student ASR model, the training data comprising the plurality of unannotated audio samples and the respective corresponding pseudo-labels,

training the respective generation of the student ASR model using the subset of the training data set to generate a next generation of the student ASR model, and

evaluating the respective generation of the student ASR model to determine whether to include the respective generation in a subset of the generations of the student ASR model;

generating a next generation of the teacher ASR model based on the subset of the generations of the student ASR model; and

responsive to reaching a completion condition, outputting a current generation of the teacher ASR model.

2 . The method of claim 1 , wherein the training data set further comprises a plurality of annotated audio samples, and wherein providing the subset of the training data set further comprises providing a subset of the plurality of annotated audio samples.

3 . The method of claim 1 , wherein evaluating the current generation of the student ASR model comprises:

providing annotated audio samples from a validation data set to the respective student ASR model generation;

determining a score for the respective student ASR model generation; and

adding the respective student ASR model generation to the subset of the generations of the student ASR model based on the score.

4 . The method of claim 3 , wherein evaluating each generation of the student ASR model comprises determining whether the score for the respective generation of the student ASR model satisfies a predetermined threshold.

5 . The method of claim 1 , wherein generating the next generation of the teacher ASR model comprises averaging the subset of the generations of the student ASR model.

6 . The method of claim 5 , wherein averaging the subset of the generations of the student ASR model comprises generating a weighted average of the subset of the generations of the student ASR model.

7 . The method of claim 1 , wherein the completion condition comprises determining that the subset of the generations of the student ASR model has remained unchanged for a predetermined number of new generations of the student ASR model.

8 . The method of claim 1 , wherein the completion condition comprises generating a predetermined number of generations of the teacher ASR model.

9 . The method of claim 1 , wherein generating a next generation of the teacher ASR model is responsive to replacing a predetermined number of the subset of the generations of the student ASR model with newer generations of the student ASR model.

10 . A system comprising:

a communications interface;

a non-transitory computer-readable medium; and

one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

generate a teacher ASR model and a student ASR model from an ASR model;

generate a plurality of successive generations of the teacher ASR model, comprising executing processor-executable instructions, for a respective teacher ASR model generation, to:

generate, for each unannotated audio sample of a plurality of unannotated audio samples of a training data set, one or more pseudo-labels corresponding to the respective unannotated audio sample using the current generation teacher ASR model;

generate a plurality of successive generations of the student ASR model, comprising executing processor-executable instructions, for a respective student ASR model generation, to:

provide a subset of the training data set as training inputs to the respective generation of the student ASR model, the training data comprising the plurality of unannotated audio samples and the respective corresponding pseudo-labels,

train the respective generation of the student ASR model using the subset of the training data set to generate a next generation of the student ASR model, and

evaluate the respective generation of the student ASR model to determine whether to include the respective

generation in a subset of the generations of the student ASR model;

generate a next generation of the teacher ASR model based on the subset of the generations of the student ASR model; and

responsive to reaching a completion condition, output a current generation of the teacher ASR model.

11 . The system of claim 10 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

provide annotated audio samples from a validation data set to the respective student ASR model generation;

determine a score for the respective student ASR model generation; and

add the respective student ASR model generation to the subset of the generations of the student ASR model based on the score.

12 . The system of claim 11 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to determine whether the score for the respective generation of the student ASR model satisfies a predetermined threshold.

13 . The system of claim 10 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to average the subset of the generations of the student ASR model.

14 . The system of claim 13 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to generate a weighted average of the subset of the generations of the student ASR model.

15 . The system of claim 10 , wherein the completion condition comprises determining that the subset of the generations of the student ASR model has remained unchanged for a predetermined number of new generations of the student ASR model.

16 . The system of claim 10 , wherein the completion condition comprises generating a predetermined number of generations of the teacher ASR model.

17 . The system of claim 10 , wherein generating a next generation of the teacher ASR model is responsive to replacing a predetermined number of the subset of the generations of the student ASR model with newer generations of the student ASR model.

18 . A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:

generate a teacher ASR model and a student ASR model from an ASR model;

generate a plurality of successive generations of the teacher ASR model, comprising processor executable instructions configured to cause the one or more processors, for a respective teacher ASR model generation, to:

generate, for each unannotated audio sample of a plurality of unannotated audio samples of a training data set, one or more pseudo-labels corresponding to the respective unannotated audio sample using the current generation teacher ASR model;

generate a plurality of successive generations of the student ASR model, comprising processor executable instructions configured to cause the one or more processors, for a respective student ASR model generation, to:

provide a subset of the training data set as training inputs to the respective generation of the student ASR model, the training data comprising the plurality of unannotated audio samples and the respective corresponding pseudo-labels,

train the respective generation of the student ASR model using the subset of the training data set to generate a next generation of the student ASR model, and

evaluate the respective generation of the student ASR model to determine whether to include the respective generation in a subset of the generations of the student ASR model;

generate a next generation of the teacher ASR model based on the subset of the generations of the student ASR model; and

responsive to reaching a completion condition, output a current generation of the teacher ASR model.

19 . The non-transitory computer-readable medium of claim 18 , further comprising processor-executable instructions configured to cause the one or more processors to:

provide annotated audio samples from a validation data set to the respective student ASR model generation;

determine a score for the respective student ASR model generation; and

add the respective student ASR model generation to the subset of the generations of the student ASR model based on the score.

20 . The non-transitory computer-readable medium of claim 19 , further comprising processor-executable instructions configured to cause the one or more processors to determine whether the score for the respective generation of the student ASR model satisfies a predetermined threshold.

Assignments (2)
CHANGE OF NAME Recorded Apr 30, 2026
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 075312/0052 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2026
From: HUSSAIN, JUAN; NGUYEN, THAI SON; STÜKER, SEBASTIAN; YE, ZHIRONG
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 074534/0902 →
References Cited (8)
US 12198689B1 · Whitenack · 2025 [cited by examiner]
US 20230335122A1 · Biadsy · 2023 [cited by examiner]
US 20240013777A1 · Lu · 2024 [cited by examiner]
US 20240021190A1 · Biadsy · 2024 [cited by examiner]
US 20240290320A1 · Huang · 2024 [cited by examiner]
US 20250078842A1 · Park · 2025 [cited by examiner]
US 20250095652A1 · Chen · 2025 [cited by examiner]
WO WO2023183262A1 · 2023 [cited by examiner]