Semi-supervised training of automatic speech recognition systems using iterative
One example method includes generating teacher and student ASR models from an ASR model; generating a plurality of successive generations of the teacher ASR model, including: generating, for each unannotated audio sample of a training data set, one or more pseudo-labels corresponding to the respective unannotated audio sample using the current generation teacher ASR model; generating a plurality of successive generations of the student ASR model, including: providing a subset of the training data set as training inputs to the student ASR model, the training data set including the plurality of unannotated audio samples and the corresponding pseudo-labels, training the student ASR model using the subset of the training data set to generate a next generation of the student ASR model, and evaluating the current student ASR model to determine whether to include the current model in a subset of the generations of the student ASR model; generating a next generation of the teacher ASR model based on the subset of the generations of the student ASR model; responsive to reaching a completion condition, outputting a current generation of the teacher ASR model.
1 . A method comprising:
generating a teacher ASR model and a student ASR model from an ASR model;
generating a plurality of successive generations of the teacher ASR model, comprising, for a respective teacher ASR model generation:
generating, for each unannotated audio sample of a plurality of unannotated audio samples of a training data set, one or more pseudo-labels corresponding to the respective unannotated audio sample using the current generation teacher ASR model;
generating a plurality of successive generations of the student ASR model, comprising, for a respective student ASR model generation:
providing a subset of the training data set as training inputs to the respective generation of the student ASR model, the training data comprising the plurality of unannotated audio samples and the respective corresponding pseudo-labels,
training the respective generation of the student ASR model using the subset of the training data set to generate a next generation of the student ASR model, and
evaluating the respective generation of the student ASR model to determine whether to include the respective generation in a subset of the generations of the student ASR model;
generating a next generation of the teacher ASR model based on the subset of the generations of the student ASR model; and
responsive to reaching a completion condition, outputting a current generation of the teacher ASR model.
2 . The method of claim 1 , wherein the training data set further comprises a plurality of annotated audio samples, and wherein providing the subset of the training data set further comprises providing a subset of the plurality of annotated audio samples.
3 . The method of claim 1 , wherein evaluating the current generation of the student ASR model comprises:
providing annotated audio samples from a validation data set to the respective student ASR model generation;
determining a score for the respective student ASR model generation; and
adding the respective student ASR model generation to the subset of the generations of the student ASR model based on the score.
4 . The method of claim 3 , wherein evaluating each generation of the student ASR model comprises determining whether the score for the respective generation of the student ASR model satisfies a predetermined threshold.
5 . The method of claim 1 , wherein generating the next generation of the teacher ASR model comprises averaging the subset of the generations of the student ASR model.
6 . The method of claim 5 , wherein averaging the subset of the generations of the student ASR model comprises generating a weighted average of the subset of the generations of the student ASR model.
7 . The method of claim 1 , wherein the completion condition comprises determining that the subset of the generations of the student ASR model has remained unchanged for a predetermined number of new generations of the student ASR model.
8 . The method of claim 1 , wherein the completion condition comprises generating a predetermined number of generations of the teacher ASR model.
9 . The method of claim 1 , wherein generating a next generation of the teacher ASR model is responsive to replacing a predetermined number of the subset of the generations of the student ASR model with newer generations of the student ASR model.
10 . A system comprising:
a communications interface;
a non-transitory computer-readable medium; and
one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:
generate a teacher ASR model and a student ASR model from an ASR model;
generate a plurality of successive generations of the teacher ASR model, comprising executing processor-executable instructions, for a respective teacher ASR model generation, to:
generate, for each unannotated audio sample of a plurality of unannotated audio samples of a training data set, one or more pseudo-labels corresponding to the respective unannotated audio sample using the current generation teacher ASR model;
generate a plurality of successive generations of the student ASR model, comprising executing processor-executable instructions, for a respective student ASR model generation, to:
provide a subset of the training data set as training inputs to the respective generation of the student ASR model, the training data comprising the plurality of unannotated audio samples and the respective corresponding pseudo-labels,
train the respective generation of the student ASR model using the subset of the training data set to generate a next generation of the student ASR model, and
evaluate the respective generation of the student ASR model to determine whether to include the respective
generation in a subset of the generations of the student ASR model;
generate a next generation of the teacher ASR model based on the subset of the generations of the student ASR model; and
responsive to reaching a completion condition, output a current generation of the teacher ASR model.
11 . The system of claim 10 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:
provide annotated audio samples from a validation data set to the respective student ASR model generation;
determine a score for the respective student ASR model generation; and
add the respective student ASR model generation to the subset of the generations of the student ASR model based on the score.
12 . The system of claim 11 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to determine whether the score for the respective generation of the student ASR model satisfies a predetermined threshold.
13 . The system of claim 10 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to average the subset of the generations of the student ASR model.
14 . The system of claim 13 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to generate a weighted average of the subset of the generations of the student ASR model.
15 . The system of claim 10 , wherein the completion condition comprises determining that the subset of the generations of the student ASR model has remained unchanged for a predetermined number of new generations of the student ASR model.
16 . The system of claim 10 , wherein the completion condition comprises generating a predetermined number of generations of the teacher ASR model.
17 . The system of claim 10 , wherein generating a next generation of the teacher ASR model is responsive to replacing a predetermined number of the subset of the generations of the student ASR model with newer generations of the student ASR model.
18 . A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:
generate a teacher ASR model and a student ASR model from an ASR model;
generate a plurality of successive generations of the teacher ASR model, comprising processor executable instructions configured to cause the one or more processors, for a respective teacher ASR model generation, to:
generate, for each unannotated audio sample of a plurality of unannotated audio samples of a training data set, one or more pseudo-labels corresponding to the respective unannotated audio sample using the current generation teacher ASR model;
generate a plurality of successive generations of the student ASR model, comprising processor executable instructions configured to cause the one or more processors, for a respective student ASR model generation, to:
provide a subset of the training data set as training inputs to the respective generation of the student ASR model, the training data comprising the plurality of unannotated audio samples and the respective corresponding pseudo-labels,
train the respective generation of the student ASR model using the subset of the training data set to generate a next generation of the student ASR model, and
evaluate the respective generation of the student ASR model to determine whether to include the respective generation in a subset of the generations of the student ASR model;
generate a next generation of the teacher ASR model based on the subset of the generations of the student ASR model; and
responsive to reaching a completion condition, output a current generation of the teacher ASR model.
19 . The non-transitory computer-readable medium of claim 18 , further comprising processor-executable instructions configured to cause the one or more processors to:
provide annotated audio samples from a validation data set to the respective student ASR model generation;
determine a score for the respective student ASR model generation; and
add the respective student ASR model generation to the subset of the generations of the student ASR model based on the score.
20 . The non-transitory computer-readable medium of claim 19 , further comprising processor-executable instructions configured to cause the one or more processors to determine whether the score for the respective generation of the student ASR model satisfies a predetermined threshold.