IP Library Granted Patent US 10,643,602
Granted Patent B2
US 10,643,602 · App. 15/923,795 · Granted May 5, 2020

Adversarial teacher-student learning for unsupervised domain adaptation

Inventors: Jinyu Li (Redmond, WA); Zhong Meng (Redmond, WA); Yifan Gong (Sammamish, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/063G06N20/00G10L15/02G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,643,602
App. No.
15/923,795
Granted
May 5, 2020
Kind
B2
Abstract

Methods, systems, and computer programs are presented for training, with adversarial constraints, a student model for speech recognition based on a teacher model. One method includes operations for training a teacher model based on teacher speech data, initializing a student model with parameters obtained from the teacher model, and training the student model with adversarial teacher-student learning based on the teacher speech data and student speech data. Training the student model with adversarial teacher-student learning further includes minimizing a teacher-student loss that measures a divergence of outputs between the teacher model and the student model; minimizing a classifier condition loss with respect to parameters of a condition classifier; and maximizing the classifier condition loss with respect to parameters of a feature extractor. The classifier condition loss measures errors caused by acoustic condition classification. Further, speech is recognized with the trained student model.

Claims (55)

1. A method comprising:

training, by one or more processors, a teacher model based on teacher speech data;

initializing, by the one or more processors, a student model with parameters obtained from the trained teacher model;

training, by the one or more processors, the student model with adversarial teacher-student learning based on the teacher speech data and student speech data, training the student model with adversarial teacher-student learning further comprising:

minimizing a teacher-student loss that measures a divergence of outputs between the teacher model and the student model;

minimizing a classifier condition loss with respect to parameters of a condition classifier, the classifier condition loss measuring errors caused by acoustic condition classification; and

maximizing the classifier condition loss with respect to parameters of a feature extractor; and

recognizing speech with the trained student model.

2. The method as recited in claim 1 , wherein the condition classifier is a neural network for mapping each deep feature to an acoustic condition.

3. The method as recited in claim 1 , wherein the divergence is a Kullback-Leibler divergence that measures how an output distribution of the teacher model diverges from an output distribution of the student model.

4. The method as recited in claim 1 , wherein the student model further comprises:

a classifier to classify units of speech, the units of speech being one of a senone, a phoneme, a tri-phone, a syllable, a character, a part of a word, or a word; and

a feature extractor for extracting deep features from the student speech data.

5. The method as recited in claim 1 , wherein the teacher speech data comprises a plurality of utterances in a teacher domain, wherein the student speech data comprises the plurality of utterances in a student domain, wherein training the student model further comprises:

providing the plurality of utterances from the teacher speech data in parallel to the plurality of utterances in the student speech data.

6. The method as recited in claim 1 , wherein the teacher-student loss is calculated by:

calculating a teacher senone posterior;

calculating a student senone posterior for a deep feature; and

calculating the teacher-student loss as a difference between the teacher senone posterior and the student senone posterior.

7. The method as recited in claim 1 , wherein a condition defines characteristics of a speaker and an environment where speech is captured.

8. The method as recited in claim 1 , wherein training the student model with adversarial teacher-student learning causes the student model to recognize senones similarly to how the teacher model recognizes senones in a condition-robust fashion.

9. The method as recited in claim 1 , wherein training the student model with adversarial teacher-student learning causes the student model to lack differentiation among different conditions.

10. The method as recited in claim 1 , wherein training the student model is performed iteratively by analyzing the teacher speech data and the student speech data.

11. A system comprising:

a memory comprising instructions; and

one or more computer processors, wherein the instructions, when executed by the one or more computer processors, cause the one or more computer processors to perform operations comprising:

training a teacher model based on teacher speech data;

initializing a student model with parameters obtained from the trained teacher model;

training the student model with adversarial teacher-student learning based on the teacher speech data and student speech data, training the student model with adversarial teacher-student learning further comprising:

minimizing a teacher-student loss that measures a divergence of outputs between the teacher model and the student model;

minimizing a classifier condition loss with respect to parameters of a condition classifier, the classifier condition loss measuring errors caused by acoustic condition classification; and

maximizing the classifier condition loss with respect to parameters of a feature extractor; and

recognizing speech with the trained student model.

12. The system as recited in claim 11 , wherein the condition classifier is a neural network for mapping each deep feature to an acoustic condition.

13. The system as recited in claim 11 , wherein the student model further comprises:

a classifier to classify units of speech, the units of speech being one of a senone, a phoneme, a tri-phone, a syllable, a character, a part of a word, or a word; and

a feature extractor for extracting deep features from the student speech data.

14. The system as recited in claim 11 , wherein the teacher speech data comprises a plurality of utterances in a teacher domain, wherein the student speech data comprises the plurality of utterances in a student domain, wherein training the student model further comprises:

providing the plurality of utterances from the teacher speech data in parallel to the plurality of utterances in the student speech data.

15. The system as recited in claim 11 , wherein training the student model with adversarial teacher-student learning that causes the student model to recognize senones similarly to how the teacher model recognizes senones in a condition-robust fashion, wherein training the student model with adversarial teacher-student learning that cause the student model to lack differentiation among different conditions.

16. A machine-readable storage medium including instructions that, when executed by a machine, cause the machine to perform operations comprising:

training a teacher model based on teacher speech data:

initializing a student model with parameters obtained from the trained teacher model;

training the student model with adversarial teacher-student learning based on the teacher speech data and student speech data, training the student model with adversarial teacher-student learning further comprising:

minimizing a teacher-student loss that measures a divergence of outputs between the teacher model and the student model;

minimizing a classifier condition loss with respect to parameters of a condition classifier, the classifier condition loss measuring errors caused by acoustic condition classification; and

maximizing the classifier condition loss with respect to parameters of a feature extractor; and

recognizing speech with the trained student model.

17. The machine-readable storage medium as recited in claim 16 , wherein the condition classifier is a neural network mapping each deep feature to an acoustic condition.

18. The machine-readable storage medium as recited in claim 16 , wherein the student model further comprises:

a classifier to classify units of speech, the units of speech being one of a senone, a phoneme, a tri-phone, a syllable, a character, a part of a word, or a word; and

a feature extractor for extracting deep features from the student speech data.

19. The machine-readable storage medium as recited in claim 16 , wherein the teacher speech data comprises a plurality of utterances in a teacher domain, wherein the student speech data comprises the plurality of utterances in a student domain, wherein training the student model further comprises:

providing the plurality of utterances from the teacher speech data in parallel to the plurality of utterances in the student speech data.

20. The machine-readable storage medium as recited in claim 16 , wherein training the student model with adversarial teacher-student learning causes the student model to recognize senones similarly to how the teacher model recognizes senones in a condition-robust fashion, wherein training the student model with adversarial teacher-student learning cause the student model to lack differentiation among different conditions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2018
From: LI, JINYU; GONG, YIFAN; MENG, ZHONG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 045376/0946 →
Continuity (1)
Related Publication 20190287515A1 · Sep 19, 2019
Cited By (2)
US 12,211,491 US 12,524,677