IP Library › Granted Patent US 12,619,880
Granted Patent B2
US 12,619,880 · App. 17/231,514 · Granted May 5, 2026

Methods, devices and media for re-weighting to improve knowledge distillation

Inventors: Peng Lu (Montreal, CA); Ahmad Rashid (Montreal, CA); Mehdi Rezagholizadeh (Montreal, CA); Abbas Ghaddar (Montreal, CA)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06N3/088G06N3/045G06N3/084G06N3/0985
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,880
App. No.
17/231,514
Granted
May 5, 2026
Kind
B2
Abstract

Methods, devices and processor-readable media for re-weighting to improve knowledge distillation are described. A reweighting module may be used to determine relative weights to assign to a ground truth label and dark knowledge distilled from the teacher (i.e. the teacher output logits used as soft labels). A meta-reweighting method is described to optimize the weights for a given labeled data sample.

Claims (130)

1 . A method for knowledge distillation, comprising:

obtaining a batch of training data comprising one or more labeled training data samples, each labeled training data sample having a respective ground truth label;

processing the batch of training data, using a student model comprising a plurality of learnable parameters, to generate, for input data in each data sample in the batch of training data, a student prediction;

for each labeled training data sample in the batch of training data, processing the student prediction and the ground truth label to compute a respective ground truth loss;

processing the batch of training data, using a trained teacher model, to generate, for each labeled training data sample in the batch of training data, a teacher prediction;

for each labeled data sample in the batch of training data, processing the student prediction and the teacher prediction to compute a respective knowledge distillation loss;

determining a weighted loss that is a weighted sum of the knowledge distillation loss and ground truth loss for each labeled training data sample in the batch of training data, the determining including:

determining a knowledge distillation weight for the labeled training data sample, the knowledge distillation weight being used to weight the knowledge distillation loss representing a difference between the student prediction and the teacher prediction for the labeled training data sample;

determining a ground truth weight for the labeled training data sample, the ground truth weight being used to weight the ground truth loss representing a difference between the student prediction and the ground truth label for the labeled training data sample; and

computing the weighted loss as the sum of:

the knowledge distillation loss weighted by the knowledge distillation weight; and

the ground truth loss weighted by the ground truth weight;

performing gradient descent on the student model using the weighted loss to identify an adjusted set of values for the plurality of learnable parameters of the student; and

adjusting the values of the plurality of learnable parameters of the student to the adjusted set of values.

2 . The method of claim 1 , wherein the knowledge distillation weight and ground truth weight are determined based on user input.

3 . The method of claim 1 , wherein the knowledge distillation weight and ground truth weight are determined by a meta reweighting process.

4 . The method of claim 3 , wherein the meta reweighting process comprises:

for each respective learnable parameter of the plurality of learnable parameters of the student, determining an optimized value of the respective learnable parameter as a function of a knowledge distillation perturbation variable and a ground truth perturbation variable with respect to the batch of training data;

determining, for each labeled training data sample in the batch of training data, a respective estimated optimized value of the knowledge distillation perturbation variable and a respective estimated optimized value of the ground truth perturbation variable with respect to a batch of validation data; and

for each labeled training data sample in the batch of training data, using the respective estimated optimized value of the knowledge distillation perturbation variable as the knowledge distillation weight and the respective estimated optimized value of the ground truth perturbation variable as the ground truth weight.

5 . The method of claim 4 , wherein:

determining the optimized value of each respective learnable parameter as a function of the knowledge distillation weight and the ground truth weight comprises:

generating a meta student model having a plurality of learnable parameters with values equal to the values of the plurality of learnable parameters of the student model;

processing the batch of training data, using the meta student model, to generate, for each labeled training data sample in the batch of training data, a meta student prediction;

for each labeled training data sample in the batch of training data, processing the meta student prediction and the ground truth label to compute a respective meta ground truth loss;

for each labeled training data sample in the batch of training data, processing the meta student prediction and the teacher prediction to compute a respective meta knowledge distillation loss;

determining a perturbed loss as the sum of:

the meta knowledge distillation loss weighted by the knowledge distillation perturbation variable; and

the meta ground truth loss weighted by the ground truth perturbation variable;

performing gradient descent on the meta student model using the perturbed loss to identify an optimal set of values of the plurality of learnable parameters of the meta student model, such that each value in the optimal set of values is defined as a function of the knowledge distillation perturbation variable and the ground truth perturbation variable; and

adjusting the values of the plurality of learnable parameters of the meta student model to the optimal set of values based on a predetermined value of the knowledge distillation perturbation variable and a predetermined value of the ground truth perturbation variable, thereby generating an adjusted meta student model.

6 . The method of claim 5 , wherein:

the predetermined value of the knowledge distillation perturbation variable is zero; and

the predetermined value of the ground truth perturbation variable is zero, and

wherein adjusting the values of the plurality of learnable parameters of the meta student model comprises leaving the values of the plurality of learnable parameters of the meta student model unchanged.

7 . The method of claim 5 , wherein:

the batch of validation data comprises one or more labeled validation data samples, each labeled validation data sample having a respective ground truth label; and

determining, for each data sample in the batch of training data, a respective estimated optimized value of the knowledge distillation perturbation variable and a respective estimated optimized value of the ground truth perturbation variable with respect to a second batch of data comprises:

processing the batch of validation data, using the adjusted meta student model, to generate, for each labeled validation data sample in the batch of validation data, a meta student prediction;

for each labeled validation data sample in the batch of validation data, processing the adjusted meta student prediction and the ground truth label to compute a respective meta ground truth loss;

processing the batch of validation data, using the teacher model, to generate, for each labeled validation data sample in the batch of validation data, a teacher validation prediction;

for each labeled validation data sample in the batch of training data, processing the adjusted meta student prediction and the teacher validation prediction to compute a respective meta knowledge distillation loss;

determining a validation loss based on the meta knowledge distillation loss and the meta ground truth loss for each labeled validation data sample in the batch of validation data;

for each labeled training data sample in the batch of training data:

computing a gradient of the validation loss with respect to the knowledge distillation perturbation variable and the ground truth perturbation variable by computing a gradient of the validation loss with respect to the optimal set of values of the plurality of learnable parameters of the meta student model, each learnable parameter value of the optimal set of values being defined as a function of the knowledge distillation perturbation variable and the ground truth perturbation variable;

performing gradient descent to compute:

the estimated optimized value of the knowledge distillation perturbation variable; and

the estimated optimized value of the ground truth perturbation variable.

8 . The method of claim 7 , further comprising:

obtaining one or more additional batches of training data;

obtaining, for each additional batch of training data, an additional batch of validation data; and

repeating, for each additional batch of training data, generating the student predictions, computing the ground truth losses, generating the teacher predictions, computing the knowledge distillation losses, determining the weighted loss, identifying the adjusted set of values of the plurality of learnable parameters of the student, and adjusting the values of the plurality of learnable parameters of the student.

9 . The method of claim 1 , wherein:

the ground truth loss comprises a cross entropy loss; and

the trained teacher model is trained to perform a natural language processing binary classification task on training data wherein each labeled training data sample comprises input data comprising a plurality of text tokens.

10 . A device, comprising:

a processor; and

a memory having stored thereon instructions which, when executed by the processor, cause the device to:

obtain a batch of training data comprising one or more labeled training data samples, each labeled training data sample having a respective ground truth label;

process the batch of training data, using a student model comprising a plurality of learnable parameters, to generate, for input data in each data sample in the batch of training data, a student prediction;

for each labeled training data sample in the batch of training data, process the student prediction and the ground truth label to compute a respective ground truth loss;

process the batch of training data, using a trained teacher model, to generate, for each labeled training data sample in the batch of training data, a teacher prediction;

for each labeled data sample in the batch of training data, process the student prediction and the teacher prediction to compute a respective knowledge distillation loss;

determine a weighted loss based on that is a weighted sum of the knowledge distillation loss and ground truth loss for each labeled training data sample in the batch of training data by:

determining a knowledge distillation weight for the labeled training data sample, the knowledge distillation weight being used to weight the knowledge distillation loss representing a difference between the student prediction and the teacher prediction for the labeled training data sample;

determining a ground truth weight for the labeled training data sample, the ground truth weight being used to weight the ground truth loss representing a difference between the student prediction and the ground truth label for the labeled training data sample; and

computing the weighted loss as the sum of:

the knowledge distillation loss weighted by the knowledge distillation weight; and

the ground truth loss weighted by the ground truth weight;

perform gradient descent on the student model using the weighted loss to identify an adjusted set of values for the plurality of learnable parameters of the student; and

adjust the values of the plurality of learnable parameters of the student to the adjusted set of values.

11 . The device of claim 10 , wherein the knowledge distillation weight and ground truth weight are determined by a meta reweighting process comprising:

for each respective learnable parameter of the plurality of learnable parameters of the student, determining an optimized value of the respective learnable parameter as a function of a knowledge distillation perturbation variable and a ground truth perturbation variable with respect to the batch of training data;

determining, for each labeled training data sample in the batch of training data, a respective estimated optimized value of the knowledge distillation perturbation variable and a respective estimated optimized value of the ground truth perturbation variable with respect to a batch of validation data; and

for each labeled training data sample in the batch of training data, using the respective estimated optimized value of the knowledge distillation perturbation variable as the knowledge distillation weight and the respective estimated optimized value of the ground truth perturbation variable as the ground truth weight.

12 . The device of claim 11 , wherein:

determining the optimized value of each respective learnable parameter as a function of the knowledge distillation weight and the ground truth weight comprises:

generating a meta student model having a plurality of learnable parameters with values equal to the values of the plurality of learnable parameters of the student model;

processing the batch of training data, using the meta student model, to generate, for each labeled training data sample in the batch of training data, a meta student prediction;

for each labeled training data sample in the batch of training data, processing the meta student prediction and the ground truth label to compute a respective meta ground truth loss;

for each labeled training data sample in the batch of training data, processing the meta student prediction and the teacher prediction to compute a respective meta knowledge distillation loss;

determining a perturbed loss as the sum of:

the meta knowledge distillation loss weighted by the knowledge distillation perturbation variable; and

the meta ground truth loss weighted by the ground truth perturbation variable;

performing gradient descent on the meta student model using the perturbed loss to identify an optimal set of values of the plurality of learnable parameters of the meta student model, such that each value in the optimal set of values is defined as a function of the knowledge distillation perturbation variable and the ground truth perturbation variable; and

adjusting the values of the plurality of learnable parameters of the meta student model to the optimal set of values based on a predetermined value of the knowledge distillation perturbation variable and a predetermined value of the ground truth perturbation variable, thereby generating an adjusted meta student model.

13 . The device of claim 12 , wherein:

the predetermined value of the knowledge distillation perturbation variable is zero; and

the predetermined value of the ground truth perturbation variable is zero, and

wherein adjusting the values of the plurality of learnable parameters of the meta student model comprises leaving the values of the plurality of learnable parameters of the meta student model unchanged.

14 . The device of claim 12 , wherein:

the batch of validation data comprises one or more labeled validation data samples, each labeled validation data sample having a respective ground truth label; and

determining, for each data sample in the batch of training data, a respective estimated optimized value of the knowledge distillation perturbation variable and a respective estimated optimized value of the ground truth perturbation variable with respect to a second batch of data comprises:

processing the batch of validation data, using the adjusted meta student model, to generate, for each labeled validation data sample in the batch of validation data, a meta student prediction;

for each labeled validation data sample in the batch of validation data, processing the adjusted meta student prediction and the ground truth label to compute a respective meta ground truth loss;

processing the batch of validation data, using the teacher model, to generate, for each labeled validation data sample in the batch of validation data, a teacher validation prediction;

for each labeled validation data sample in the batch of training data, processing the adjusted meta student prediction and the teacher validation prediction to compute a respective meta knowledge distillation loss;

determining a validation loss based on the meta knowledge distillation loss and the meta ground truth loss for each labeled validation data sample in the batch of validation data;

for each labeled training data sample in the batch of training data:

computing a gradient of the validation loss with respect to the knowledge distillation perturbation variable and the ground truth perturbation variable by computing a gradient of the validation loss with respect to the optimal set of values of the plurality of learnable parameters of the meta student model, each learnable parameter value of the optimal set of values being defined as a function of the knowledge distillation perturbation variable and the ground truth perturbation variable;

performing gradient descent to compute:

the estimated optimized value of the knowledge distillation perturbation variable; and

the estimated optimized value of the ground truth perturbation variable.

15 . The device of claim 14 , wherein the instructions, when executed by the processor, further cause the device to:

obtain one or more additional batches of training data;

obtain, for each additional batch of training data, an additional batch of validation data; and

repeat, for each additional batch of training data, generating the student predictions, computing the ground truth losses, generating the teacher predictions, computing the knowledge distillation losses, determining the weighted loss, identifying the adjusted set of values of the plurality of learnable parameters of the student, and adjusting the values of the plurality of learnable parameters of the student.

16 . The device of claim 10 , wherein:

the ground truth loss comprises a cross entropy loss; and

the trained teacher model is trained to perform a natural language processing binary classification task on training data wherein each labeled training data sample comprises input data comprising a plurality of text tokens.

17 . A non-transitory processor-readable medium containing instructions which, when executed by a processor of a device, cause the device to:

obtain a batch of training data comprising one or more labeled training data samples, each labeled training data sample having a respective ground truth label;

process the batch of training data, using a student model comprising a plurality of learnable parameters, to generate, for input data in each data sample in the batch of training data, a student prediction;

for each labeled training data sample in the batch of training data, process the student prediction and the ground truth label to compute a respective ground truth loss;

process the batch of training data, using a trained teacher model, to generate, for each labeled training data sample in the batch of training data, a teacher prediction;

for each labeled data sample in the batch of training data, process the student prediction and the teacher prediction to compute a respective knowledge distillation loss;

determine a weighted loss that is a weighted sum of the knowledge distillation loss and ground truth loss for each labeled training data sample in the batch of training data by:

determining a knowledge distillation weight for the labeled training data sample, the knowledge distillation weight being used to weight the knowledge distillation loss representing a difference between the student prediction and the teacher prediction for the labeled training data sample;

determining a ground truth weight for the labeled training data sample, the ground truth weight being used to weight the ground truth loss representing a difference between the student prediction and the ground truth label for the labeled training data sample; and

computing the weighted loss as the sum of:

the knowledge distillation loss weighted by the knowledge distillation weight; and

the ground truth loss weighted by the ground truth weight;

perform gradient descent on the student model using the weighted loss to identify an adjusted set of values for the plurality of learnable parameters of the student; and

adjust the values of the plurality of learnable parameters of the student to the adjusted set of values.

18 . The non-transitory process-readable medium of claim 17 , wherein the knowledge distillation weight and ground truth weight are determined based on user input.

19 . The non-transitory process-readable medium of claim 17 , wherein the knowledge distillation weight and ground truth weight are determined by a meta reweighting process.

20 . The non-transitory process-readable medium of claim 19 , wherein the meta reweighting process comprises:

for each respective learnable parameter of the plurality of learnable parameters of the student, determining an optimized value of the respective learnable parameter as a function of a knowledge distillation perturbation variable and a ground truth perturbation variable with respect to the batch of training data;

determining, for each labeled training data sample in the batch of training data, a respective estimated optimized value of the knowledge distillation perturbation variable and a respective estimated optimized value of the ground truth perturbation variable with respect to a batch of validation data; and

for each labeled training data sample in the batch of training data, using the respective estimated optimized value of the knowledge distillation perturbation variable as the knowledge distillation weight and the respective estimated optimized value of the ground truth perturbation variable as the ground truth weight.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2022
From: LU, PENG; RASHID, AHMAD; REZAGHOLIZADEH, MEHDI; GHADDAR, ABBAS
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 058786/0846 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2021
From: LU, PENG; RASHID, AHMAD; REZAGHOLIZADEH, MEHDI; GHADDAR, ABBAS
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 057619/0006 →
Continuity (1)
Related Publication 20220343175A1 · Oct 27, 2022
References Cited (18)
US 20200104717A1 · Alistarh · 2020 [cited by applicant]
US 20200302295A1 · Tung et al. · 2020 [cited by applicant]
CN 111223483A · 2020 [cited by applicant]
WO WO2019240964A1 · 2019 [cited by examiner]
Ren, Mengye, et al. “Learning to reweight examples for robust deep learning.” International conference on machine learning. PMLR, 2018. (Year: 2018). [cited by examiner]
Liu, Linqing, et al. “Mkd: a multi-task knowledge distillation approach for pretrained language models.” arXiv preprint arXiv:1911.03588 (2019). (Year: 2019). [cited by examiner]
Zhang, Youcai, et al. “Prime-Aware Adaptive Distillation.” arXiv preprint arXiv:2008.01458 (2020). (Year: 2020). [cited by examiner]
Vyas, Nidhi, Shreyas Saxena, and Thomas Voice. “Learning soft labels via meta learning.” arXiv preprint arXiv:2009.09496 (2020). (Year: 2020). [cited by examiner]
Jin, Woojeong, et al. “Msd: Saliency-aware knowledge distillation for multimodal understanding.” arXiv preprint arXiv:2101.01881 (2021). (Year: 2021). [cited by examiner]
Shu, Jun, et al. “Meta-weight-net: Learning an explicit mapping for sample weighting.” Advances in neural information processing systems 32 (2019). (Year: 2019). [cited by examiner]
Heydari, A. Ali, Craig A. Thompson, and Asif Mehmood. “Softadapt: Techniques for adaptive loss weighting of neural networks with multi-part loss functions.” arXiv preprint arXiv:1912.12355 (2019). (Year: 2019). [cited by examiner]
Groenendijk, Rick, et al. “Multi-Loss Weighting with Coefficient of Variations.” arXiv e-prints (2020): arXiv-2009. (Year: 2020). [cited by examiner]
M. Ren, W. Zeng, B. Yang, and R. Urtasun. Learning to reweight examples for robust deep learning. In ICML, 2018. [cited by applicant]
Koh, Pang Wei and Liang, Percy. Understanding black-box predictions via influence functions. In Proceedings of the 34th International Conference on Machine Learning, ICML, 2017. [cited by applicant]
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. NIPS Workshop, 2014. [cited by applicant]
Jun Shu, Qi Xie, Lixuan Yi, Qian Zhao, Sanping Zhou, Zongben Xu, and Deyu Meng. Meta-weight-net: Learning an explicit mapping for sample weighting. arXiv preprint arXiv:1902.07379, 2019. [cited by applicant]
Zhenzhen Li, Jian-Yun Nie, Benyou Wang, Pan Du, Yuhan Zhang, Lixin Zou, and Dongsheng Li. Meta-learning for neural relation classification with distant supervision. In Proceedings of the 29th ACM International Conferenc… [cited by applicant]
Dong, B., Hou, J., Lu, Y., and Zhang, Z. Harvesting dark knowledge utilizing anisotropic information retrieval for overparameterized neural network. arXiv preprint arXiv:1910.01255, 2019. [cited by applicant]