IP Library Granted Patent US 11,620,515
Granted Patent B2
US 11,620,515 · App. 16/716,249 · Granted Apr 4, 2023

Multi-task knowledge distillation for language model

Inventors: Linqing Liu (Menlo Park, CA); Caiming Xiong (Menlo Park, CA)
Assignee: salesforce.com, inc.
G06N3/08G06F40/30G06F40/40G06N3/0454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,620,515
App. No.
16/716,249
Granted
Apr 4, 2023
Kind
B2
Abstract

Systems and methods are provided that employ knowledge distillation under a multi-task learning setting. In some embodiments, the systems and methods are implemented with a larger teacher model and a smaller student model, each of which comprise one or more shared layers and a plurality of task layers for performing multiple tasks. During training of the teacher model, its shared layers are initialized, and then the teacher model is multi-task refined. The teacher model predicts teacher logits. During training of the student model, its shared layers are initialized. Knowledge distillation is employed to transfer knowledge from the teacher model to the student model by the student model updating its shared layers and task layers, for example, according to the teacher logits of the teacher model. Other features are also provided.

Claims (53)

1. A method for transfer of knowledge from a teacher model to a student model, the method comprising:

initializing one or more shared layers of the teacher model;

refining multiple task layers of the teacher model, each task layer capable of performing a respective task;

randomly initializing parameters of the student model;

separating training data corresponding to multiple tasks that the teacher model has been refined for into a plurality of batches, wherein each batch is specific to at least one task from the multiple tasks;

for each task-specific batch:

predicting logits from the teacher model based on training inputs in the task-specific batch;

predicting logits from the student model based on training inputs in the task-specific batch; and

computing a task-specific distillation loss based on a difference between the logits from the teacher model and the logits from the student model; and

computing an aggregated distillation loss by summing task-specific distillation losses corresponding to the multiple tasks; and

jointly updating the student model based on the aggregated distillation loss.

2. The method of claim 1 , wherein the tasks comprise at least one natural language processing task.

3. The method of claim 1 , wherein the tasks comprise at least one of a natural language inference task, a single sentence classification task, a sentiment classification task, a semantic text similarity task, and a relevance ranking task.

4. The method of claim 1 , wherein at least one of the student model and the teacher model comprises a language representational model.

5. The method of claim 1 , wherein at least one of the student model and the teacher model comprises a transformer model natural language understanding neural network.

6. The method of claim 1 , wherein the student model comprises one or more shared layers and a plurality of task layers.

7. The method of claim 1 , wherein updating the student model comprises minimizing a mean squared error between logits predicted by the student model against the predicted logits from the teacher model.

8. A system for transfer of knowledge from a teacher model to a student model, the system comprising:

a memory storing machine executable code; and

one or more processors coupled to the memory and configurable to execute the machine executable code to cause the one or more processors to:

initialize one or more shared layers of the teacher model;

refine multiple task layers of the teacher model, each task layer capable of performing a respective task;

randomly initialize parameters of the student model;

separate training data corresponding to multiple tasks that the teacher model has been refined for into a plurality of batches, wherein each batch is specific to at least one task from the multiple tasks;

for each task-specific batch:

predict logits from the teacher model based on training inputs in the task-specific batch;

predict logits from the student model based on training inputs in the task-specific batch; and

compute a task-specific distillation loss based on a difference between the logits from the teacher model and the logits from the student model; and

compute an aggregated distillation loss by summing task-specific distillation losses corresponding to the multiple tasks; and

jointly update the student model based on the aggregated distillation loss.

9. The system of claim 8 , wherein the tasks comprise at least one natural language processing task.

10. The system of claim 8 , wherein the tasks comprise at least one of a natural language inference task, a single sentence classification task, a sentiment classification task, a semantic text similarity task, and a relevance ranking task.

11. The system of claim 8 , wherein at least one of the student model and the teacher model comprises a language representational model.

12. The system of claim 8 , wherein at least one of the student model and the teacher model comprises a transformer model natural language understanding neural network.

13. The system of claim 8 , wherein the student model comprises one or more shared layers and a plurality of task layers.

14. The system of claim 8 , wherein updating the student model comprises minimizing a mean squared error between logits predicted by the student model against the predicted logits from the teacher model.

15. A non-transitory machine-readable medium comprising executable code which when executed by one or more processors associated with a computer are adapted to cause the one or more processors to perform a method for transfer of knowledge from a teacher model to a student model comprising:

initializing one or more shared layers of the teacher model;

refining multiple task layers of the teacher model, each task layer capable of performing a respective task;

randomly initializing parameters of the student model;

separating training data corresponding to multiple tasks that the teacher model has been refined for into a plurality of batches, wherein each batch is specific to at least one task from the multiple tasks;

for each task-specific batch:

predicting logits from the teacher model based on training inputs in the task-specific batch;

predicting logits from the student model based on training inputs in the task-specific batch; and

computing a task-specific distillation loss based on a difference between the logits from the teacher model and the logits from the student model; and

computing an aggregated distillation loss by summing task-specific distillation losses corresponding to the multiple tasks; and

jointly updating the student model based on the aggregated distillation loss.

16. The non-transitory machine-readable medium of claim 15 , wherein the tasks comprise at least one natural language processing task.

17. The non-transitory machine-readable medium of claim 15 , wherein the tasks comprise at least one of a natural language inference task, a single sentence classification task, a sentiment classification task, a semantic text similarity task, and a relevance ranking task.

18. The non-transitory machine-readable medium of claim 15 , wherein at least one of the student model and the teacher model comprises a language representational model.

19. The non-transitory machine-readable medium of claim 15 , wherein at least one of the student model and the teacher model comprises a transformer model natural language understanding neural network.

20. The non-transitory machine-readable medium of claim 15 , wherein the student model comprises one or more shared layers and a plurality of task layers.

21. The non-transitory machine-readable medium of claim 15 , wherein updating the student model comprises minimizing a mean squared error between logits predicted by the student model against the predicted logits from the teacher model.

Assignments (2)
CHANGE OF NAME Recorded Dec 18, 2024
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 069717/0470 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2019
From: LIU, LINQING; XIONG, CAIMING
To: SALESFORCE.COM, INC.
Reel/Frame 051348/0859 →
Continuity (2)
Provisional Application 62932163 · Nov 7, 2019
Related Publication 20210142164A1 · May 13, 2021
Cited By (6)
US 12,367,248 US 12,367,249 US 12,608,649 US 12,613,927 US 12,675,648 US 12,705,265