IP Library Granted Patent US 10,635,977
Granted Patent B2
US 10,635,977 · App. 16/458,506 · Granted Apr 28, 2020

Multi-task learning using knowledge distillation

Inventors: Junyoung Chung (Gunpo-si, KR); Melvin Jose Johnson Premkumar (Sunnyvale, CA); Michael Schuster (Saratoga, CA); Wolfgang Macherey (Sunnyvale, CA)
Assignee: Google LLC
G06N3/08G06F17/289G06N3/0454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,635,977
App. No.
16/458,506
Granted
Apr 28, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media for performing multi-task learning. In one method a system obtains a respective set of training data for each of multiple machine learning tasks. For each of the machine learning tasks, the system configures a respective teacher machine learning model to perform the machine learning task by training the teacher machine learning model on the training data. The system trains a single student machine learning model to perform the multiple machine learning tasks using (i) the configured teacher machine learning models, and (ii) the obtained training data.

Claims (50)

1. A computer implemented method comprising:

obtaining a respective set of training data for each of a plurality of machine learning tasks;

for each of the machine learning tasks, configuring a respective teacher machine learning model to perform the machine learning task by training the teacher machine learning model on the training data for the task; and

training a single student machine learning model having a plurality of student machine learning model parameters to perform all of the plurality of machine learning tasks using (i) the configured teacher machine learning models, and (ii) the obtained training data, wherein training the single student machine learning model comprises:

for each of the plurality of machine learning tasks:

selecting one or more subsets from the set of training data for the machine learning task;

processing the selected subsets using the respective teacher machine learning model to generate respective teacher machine learning model outputs; and

training the single student machine learning model to perform the machine learning task using (i) the selected one or more subsets, and (ii) respective generated teacher machine learning model outputs, comprising, for each subset:

augmenting the subset with an identifier for the machine learning task;

processing the augmented subset using the student machine learning model to generate a student machine learning model output; and

adjusting values of the student machine learning model parameters to match the generated student machine learning model output to the respective generated teacher machine learning model output for the subset.

2. The method of claim 1 , wherein the teacher machine learning model outputs comprise soft target outputs.

3. The method of claim 1 , wherein the training data for each of the plurality of machine learning tasks comprises (i) an input text segment in an input language, and (ii) an output text segment in a target language that is different from the input language.

4. The method of claim 3 , wherein the plurality of machine learning tasks comprise translating an input text segment in an input language into a target language.

5. The method of claim 4 , wherein augmenting the subset with an identifier for the machine learning task comprises prepending each input text segment with a token identifying at least the target language.

6. The method of claim 3 , wherein selecting one or more subsets from the set of training data for the machine learning task comprises selecting one or more sub-word units from the input text segment.

7. The method of claim 6 , wherein each generated respective teacher machine learning model output comprises a probability distribution indicating a respective translation of the corresponding sub-word unit.

8. The method of claim 3 , wherein the training data comprises an equal distribution of text segments in different languages.

9. The method of claim 1 , wherein augmenting the subset with an identifier for the machine learning task comprises prepending the subset with a token identifier for the machine learning task.

10. The method of claim 1 , wherein the student machine learning model is smaller in size than the teacher machine learning models.

11. The method of claim 1 , wherein the student machine learning model is larger in size or the same size as the teacher machine learning models.

12. The method of claim 1 , wherein the size of each of the teacher machine learning models is independent of the student machine learning model.

13. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining a respective set of training data for each of a plurality of machine learning tasks;

for each of the machine learning tasks, configuring a respective teacher machine learning model to perform the machine learning task by training the teacher machine learning model on the training data for the task; and

training a single student machine learning model having a plurality of student machine learning model parameters to perform all of the plurality of machine learning tasks using (i) the configured teacher machine learning models, and (ii) the obtained training data, wherein training the single student machine learning model comprises:

for each of the plurality of machine learning tasks:

selecting one or more subsets from the set of training data for the machine learning task;

processing the selected subsets using the respective teacher machine learning model to generate respective teacher machine learning model outputs; and

training the single student machine learning model to perform the machine learning task using (i) the selected one or more subsets, and (ii) respective generated teacher machine learning model outputs, comprising, for each subset:

augmenting the subset with an identifier for the machine learning task;

processing the augmented subset using the student machine learning model to generate a student machine learning model output; and

adjusting values of the student machine learning model parameters to match the generated student machine learning model output to the respective generated teacher machine learning model output for the subset.

14. The system of claim 13 , wherein the teacher machine learning model outputs comprise soft target outputs.

15. The system of claim 13 , wherein the training data for each of the plurality of machine learning tasks comprises (i) an input text segment in an input language, and (ii) an output text segment in a target language that is different from the input language.

16. One or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

obtaining a respective set of training data for each of a plurality of machine learning tasks;

for each of the machine learning tasks, configuring a respective teacher machine learning model to perform the machine learning task by training the teacher machine learning model on the training data for the task; and

training a single student machine learning model having a plurality of student machine learning model parameters to perform all of the plurality of machine learning tasks using (i) the configured teacher machine learning models, and (ii) the obtained training data, wherein training the single student machine learning model comprises:

for each of the plurality of machine learning tasks:

selecting one or more subsets from the set of training data for the machine learning task;

processing the selected subsets using the respective teacher machine learning model to generate respective teacher machine learning model outputs; and

training the single student machine learning model to perform the machine learning task using (i) the selected one or more subsets, and (ii) respective generated teacher machine learning model outputs, comprising, for each subset:

augmenting the subset with an identifier for the machine learning task;

processing the augmented subset using the student machine learning model to generate a student machine learning model output; and

adjusting values of the student machine learning model parameters to match the generated student machine learning model output to the respective generated teacher machine learning model output for the subset.

17. The non-transitory computer-readable media of claim 16 , wherein the teacher machine learning model outputs comprise soft target outputs.

18. The non-transitory computer-readable media of claim 16 , wherein the training data for each of the plurality of machine learning tasks comprises (i) an input text segment in an input language, and (ii) an output text segment in a target language that is different from the input language.

19. The non-transitory computer-readable media of claim 18 , wherein the plurality of machine learning tasks comprise translating an input text segment in an input language into a target language.

20. The non-transitory computer-readable media of claim 19 , wherein augmenting the subset with an identifier for the machine learning task comprises prepending each input text segment with a token identifying at least the target language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2019
From: CHUNG, JUNYOUNG; JOHNSON PREMKUMAR, MELVIN JOSE; SCHUSTER, MICHAEL; MACHEREY, WOLFGANG
To: GOOGLE LLC
Reel/Frame 050288/0767 →
Continuity (3)
Continuation PCTUS2017069102 · Dec 29, 2017
Provisional Application 62441119 · Dec 30, 2016
Related Publication 20190325308A1 · Oct 24, 2019
Cited By (1)
US 12,651,203