IP Library › Granted Patent US 12,260,340
Granted Patent B2
US 12,260,340 · App. 18/471,866 · Granted Mar 25, 2025

Extreme language model compression with optimal sub-words and shared projections

Inventors: Yang Song (Bellevue, WA); Raghav Gupta (Mountain View, CA); Dengyong Zhou (Redmond, WA); Sanqiang Zhao (Pittsburgh, PA)
Assignee: GOOGLE LLC
G06N3/088G06F40/284G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,340
App. No.
18/471,866
Granted
Mar 25, 2025
Kind
B2
Abstract

Provided is a knowledge distillation technique for training a student language model that, relative to a larger teacher language model, has a significantly smaller vocabulary, lower embedding dimensions, and/or hidden state dimensions. Specifically, aspects of the present disclosure are directed to a dual-training mechanism that trains the teacher and student language models simultaneously to obtain optimal word embeddings for the student vocabulary. In some implementations, this approach can be combined with learning shared projection matrices that transfer layer-wise knowledge from the teacher language model to the student language model. Example experimental results have also demonstrated higher compression efficiency and accuracy when compared with other state-of-the-art compression techniques, including the ability to compress the BERT BASE model by more than 60×, with only a minor drop in downstream task metrics, resulting in a language model with a footprint of under 7 MB.

Claims (67)

1. A computing system for training a machine-learned model, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store:

a first language model comprising one or more transformer layers, wherein the first language model includes a plurality of first language model parameters, wherein each first language model parameter of the plurality of first language model parameters is associated with at least one transformer layer of the one or more transformer layers of the first language model;

a second language model comprising one or more transformer layers, wherein the second language model includes a plurality of second language model parameters, wherein each second language model parameter of the plurality of second language model parameters is associated with at least one transformer layer of the one or more transformer layers of the second language model, wherein the one or more transformer layers of the second language model are of a different dimension than the one or more transformer layers of the first language model, and

instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

projecting the first language model parameters into a shared space with the second language model parameters; and

training the second language model using a loss function based on a comparison of the projected first language model parameters and the second language model parameters.

2. The computing system of claim 1 , wherein the loss function is a mean square error loss function.

3. The computing system of claim 1 , the instructions further comprising:

evaluating a second loss function based on the first language model parameters and the second language model parameters; and

training the second language model based on the second loss function.

4. The computing system of claim 1 , wherein the first language model is a teacher language model and the second language model is a student language model.

5. The computing system of claim 4 , wherein a dimension of the one or more transformer layers of the second language model is smaller than a dimension of the one or more transformer layers of the first language model.

6. The computing system of claim 5 , wherein projecting the first language model parameters into the shared space with the second language model parameters includes down projecting the first language model parameters into the dimension of the one or more transformer layers of the second language model.

7. The computing system of claim 1 , wherein the first language model is a student language model and the second language model is a teacher language model.

8. The computing system of claim 7 , wherein a dimension of the one or more transformer layers of the second language model is larger than a dimension of the one or more transformer layers of the first language model.

9. The computing system of claim 8 , wherein projecting the first language model parameters into the shared space with the second language model parameters includes up projecting the first language model parameters into the dimension of the one or more transformer layers of the second language model.

10. A computer-implemented method for training a machine-learned model, the method comprising:

projecting a plurality of first language model parameters into a shared space with a plurality of second language model parameters, wherein:

a first language model comprises one or more transformer layers, wherein the first language model includes a plurality of first language model parameters, wherein each first language model parameter of the plurality of first language model parameters is associated with at least one transformer layer of the one or more transformer layers of the first language model;

a second language model comprises one or more transformer layers, wherein the second language model includes the plurality of second language model parameters, wherein each second language model parameter of the plurality of second language model parameters is associated with at least one transformer layer of the one or more transformer layers of the second language model, wherein the one or more transformer layers of the second language model are of a different dimension than the one or more transformer layers of the first language model; and

training the second language model using a loss function based on a comparison of the projected first language model parameters and the second language model parameters.

11. The method of claim 10 , wherein the loss function is a mean square error loss function.

12. The method of claim 10 , the method comprising:

evaluating a second loss function based on the first language model parameters and the second language model parameters; and

training the second language model based on the second loss function.

13. The method of claim 10 , wherein the first language model is a teacher language model and the second language model is a student language model.

14. The method of claim 13 , wherein a dimension of the one or more transformer layers of the second language model is smaller than a dimension of the one or more transformer layers of the first language model.

15. The method of claim 14 , wherein projecting the first language model parameters into the shared space with the second language model parameters includes down projecting the first language model parameters into the dimension of the one or more transformer layers of the second language model.

16. The method of claim 10 , wherein the first language model is a student language model and the second language model is a teacher language model.

17. The method of claim 16 , wherein a dimension of the one or more transformer layers of the second language model is larger than a dimension of the one or more transformer layers of the first language model.

18. The method of claim 17 , wherein projecting the first language model parameters into the shared space with the second language model parameters includes up projecting the first language model parameters into the dimension of the one or more transformer layers of the second language model.

19. A computing system for executing a student machine-learned model having parameters aligned with those of a teacher machine-learned model, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store:

a student language model comprising one or more transformer layers, wherein the student language model includes a plurality of student language model parameters;

wherein the student language model was trained using a loss function based on a comparison of the plurality of student language model parameters and a plurality of teacher language model parameters in a shared space, wherein the plurality of student language model parameters were projected into the shared space, the plurality of teacher language model parameters were projected into the shared space, or the plurality of student language model parameters and the plurality of teacher language model parameters were projected into the shared space; and

instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

generating an output sequence by processing an input sequence using the student machine-learned model.

20. The computing system of claim 19 , wherein the shared space was characterized by a dimension of a layer of the teacher machine-learned model, and wherein the plurality of student language model parameters were projected into the shared space.

21. A computing system for training a machine-learned model, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store:

a first language model comprising one or more layers, wherein the first language model includes a plurality of first language model parameters, wherein each first language model parameter of the plurality of first language model parameters is associated with at least one layer of the one or more layers of the first language model;

a second language model comprising one or more layers, wherein the second language model includes a plurality of second language model parameters, wherein each second language model parameter of the plurality of second language model parameters is associated with at least one layer of the one or more layers of the second language model, wherein the one or more layers of the second language model are of a different dimension than the one or more layers of the first language model, and

instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

projecting the first language model parameters into a shared space with the second language model parameters; and

training the second language model using a loss function based on a comparison of the projected first language model parameters and the second language model parameters.

22. The computing system of claim 21 , wherein the loss function is a mean square error loss function.

23. The computing system of claim 21 , the instructions further comprising:

evaluating a second loss function based on the first language model parameters and the second language model parameters; and

training the second language model based on the second loss function.

24. The computing system of claim 21 , wherein the first language model is a teacher language model and the second language model is a student language model.

25. The computing system of claim 24 , wherein a dimension of the second language model is smaller than a dimension of the first language model.

26. The computing system of claim 25 , wherein projecting the first language model parameters into the shared space with the second language model parameters includes down projecting the first language model parameters into the dimension of the second language model.

27. The computing system of claim 21 , wherein the first language model is a student language model and the second language model is a teacher language model.

28. The computing system of claim 27 , wherein a dimension of the second language model is larger than a dimension of the first language model.

29. The computing system of claim 28 , wherein projecting the first language model parameters into the shared space with the second language model parameters includes up projecting the first language model parameters into the dimension of the second language model.

30. A computing system for executing a student machine-learned model having parameters aligned with those of a teacher machine-learned model, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store:

a student language model comprising one or more layers, wherein the student language model includes a plurality of student language model parameters;

wherein the student language model was trained using a loss function based on a comparison of the plurality of student language model parameters and a plurality of teacher language model parameters in a shared space, wherein the plurality of student language model parameters were projected into the shared space, the plurality of teacher language model parameters were projected into the shared space, or the plurality of student language model parameters and the plurality of teacher language model parameters were projected into the shared space; and

instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

generating an output sequence by processing an input sequence using the student machine-learned model.

31. The computing system of claim 30 , wherein the shared space was characterized by a dimension of a layer of the teacher machine-learned model, and wherein the plurality of student language model parameters were projected into the shared space.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2023
From: SONG, YANG; GUPTA, RAGHAV; ZHOU, DENGYONG; ZHAO, SANQIANG
To: GOOGLE LLC
Reel/Frame 064994/0641 →
Continuity (2)
Continuation 16749570 · Jan 22, 2020
Related Publication 20240013059A1 · Jan 11, 2024
References Cited (54)
US 11797862B2 · Song · 2023 [cited by examiner]
US 20190205748A1 · Fukuda · 2019 [cited by examiner]
US 20190266236A1 · Battach · 2019 [cited by examiner]
US 20190266250A1 · Toplyn · 2019 [cited by applicant]
US 20200027444A1 · Prabhavalkar et al. · 2020 [cited by applicant]
US 20200111483A1 · Shafran et al. · 2020 [cited by applicant]
US 20200135174A1 · Cui · 2020 [cited by examiner]
US 20200234694A1 · Griffiths · 2020 [cited by examiner]
US 20210065699A1 · Kaushik · 2021 [cited by examiner]
US 20210141798A1 · Steedman Henderson · 2021 [cited by examiner]
US 20210182662A1 · Lai et al. · 2021 [cited by applicant]
Al-Rfou et al., “Conversational Contextual Cues: The Case of Personalization and History for Response Ranking” arXiv:1606.00372, Jun. 1, 2016, 10 pages. [cited by applicant]
Anwar et al., “Structured Pruning of Deep Convolutional Neural Networks”, ACM Journal on Emerging Technologies in Computing Systems, vol. 13, No. 3, Dec. 2015, 11 pages. [cited by applicant]
Ba et al., “Do Deep Nets Really Need to be Deep?”, Twenty-eighth Conference on Neural Information Processing Systems, Dec. 8-13, 2014, Montreal, Canada, 9 pages. [cited by applicant]
Chen et al., “Compressing Neural Networks with the Hashing Trick”, International Conference on Machine Learning, Jul. 6-11, 2015, Lille, France, 10 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv:1810.04805v1, Oct. 11, 2018, 14 pages. [cited by applicant]
Dolan et al., “Automatically Constructing a Corpus of Sentential Paraphrases”, Third International Workshop on Paraphrasing, Oct. 2005, Jeju Island, Korea, 8 pages. [cited by applicant]
Glorot et al., “Understanding the difficulty of training deep feedforward neural networks” Thirteenth International Conference on Artificial Intelligence and Statistics, May 13-15, 2010, Sardinia, Italy, 8 pages. [cited by applicant]
Gong et al., “Compressing Deep Convolutional Networks Using Vector Quantization”, arXiv:1412.6115v1, Dec. 18, 2014, 10 pages. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network”, arXiv:1503.02531v1, Mar. 9, 2015, 9 pages. [cited by applicant]
Huang et al., “Like What You Like: Knowledge Distill via Neuron Selectivity Transfer”, arXiv:1707.01219v2, Dec. 18, 2017, 9 pages. [cited by applicant]
Joulin et al., “FastText.Zip: Compressing Text Classification Models”, arXiv:1612.03651v1, Dec. 12, 2016, 13 pages. [cited by applicant]
Kim et al., “Sequence-Level Knowledge Distillation”, arXiv:1606.07947v4, Sep. 22, 2016, 11 pages. [cited by applicant]
Lai et al., “RACE: Large-scale Reading Comprehension Dataset From Examinations”, Conference on Empirical Methods in Natural Language Processing, Sep. 7-11, 2017, Copenhagen, Denmark, 10 pages. [cited by applicant]
Lample et al., Cross-lingual Language Model Pretraining, arXiv:1901.07291v1, Jan. 22, 2019, 10 pages. [cited by applicant]
Li et al., “Pruning Filters for Efficient ConvNets”, arXiv:1608.08710v2, Sep. 15, 2016, 9 pages. [cited by applicant]
Lin et al., “Fixed Point Quantization of Deep Convolutional Networks”, International Conference on Machine Learning, Jun. 19-24, 2016, New York City, NY, 10 pages. [cited by applicant]
Lin et al., “Towards Accurate Binary Convolutional Neural Network”, Thirty-first Conference on Neural Information Processing Systems, Dec. 4-9, 2017, Long Beach, CA, 9 pages. [cited by applicant]
Luo et al., “ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression”, IEEE International Conference on Computer Vision (ICCV), Oct. 22-29, 2017, Venice, Italy, 9 pages. [cited by applicant]
Mikolov et al., “Distributed Representations of Words and Phrases and their Compositionality”, Twenty-seventh Conference on Neural Information Processing Systems, Dec. 5-10, 2013, Lake Tahoe, NV, 9 pages. [cited by applicant]
Pennington et al., “GloVe: Global Vectors for Word Representation”, Conference on Empirical Methods in Natural Language Processing, Oct. 25-29, 2014, Doha, Qatar, 12 pages. [cited by applicant]
Peters et al., Deep Contextualized Word Representations. 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun. 1-6, 2018, New Orleans, Louisian… [cited by applicant]
Radford et al., Language Models are Unsupervised Multitask Learners, 2019, 24 pages. [cited by applicant]
Rajpurkar et al., “Squad: 100,000+ Questions for Machine Comprehension of Text”, arXiv:1606.05250v3, Oct. 11, 2016, 10 pages. [cited by applicant]
Romero et al., “FitNets: Hints for Thin Deep Nets”, arXiv:1412.6550v1, Dec. 19, 2014, 12 pages. [cited by applicant]
See et al., “Compression of Neural Machine Translation Models via Pruning”, 20th SIGNLL Conference on ComputationalNatural Language Learning, Aug. 7-12, 2016, Berlin, Germany, pp. 291-301. [cited by applicant]
Sennrich et al., “Neural Machine Translation of Rare Words with Subword Units”, arXiv:1508.07909v5, Jun. 10, 2016, 11 pages. [cited by applicant]
Shen et al., “Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT” arXiv:1909.05840v2, Sep. 22, 2019, 15 pages. [cited by applicant]
Sindhwani et al., “Structured Transforms for Small-Footprint Deep Learning”, Twenty-ninth Conference on Neural Information Processing Systems, Dec. 7-12, 2015, Montreal, Canada, 9 pages. [cited by applicant]
Socher et al., “Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank”, Conference on Empirical Methods in Natural Language Processing, Oct. 18-21, 2013, Seattle, WA, 12 pages. [cited by applicant]
Sun et al., “Patient Knowledge Distillation for BERT Model Compression”, arxiv:1908.09355vl, Aug. 25, 2019, 10 pages. [cited by applicant]
Tang et al., “Distilling Task-Specific Knowledge from BERT into Simple Neural Networks”, arXiv:1903.12136v1, Mar. 28, 2019, 8 pages. [cited by applicant]
Tulloch et al., “High performance ultra-low-precision convolutions on mobile devices”, arXiv:1712.02427vl, Dec. 6, 2017, 5 pages. [cited by applicant]
Wang et al., “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”, Conference on Learning Representations, May 6-9, 2019, New Orleans, Louisiana, 20 pages. [cited by applicant]
Williams et al., “A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference” arXiv:1704.05426v4, Feb. 19, 2018, 11 pages. [cited by applicant]
Wu et al., “Binarized Neural Networks on the ImageNet Classification Task”, arxiv:1604.03058v5, Nov. 19, 2016, 4 pages. [cited by applicant]
Wu et al., “Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation”, arxiv:1609.08144v2, Oct. 8, 2016, 23 pages. [cited by applicant]
Yang et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding”, arXiv:1906.08237v1, Jun. 19, 2019, 18 pages. [cited by applicant]
Yim et al., “A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning”, IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu, Hawaii, pp. 4133-4… [cited by applicant]
You et al., “Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes”, arXiv:1904.00962v3, May 24, 2019, 36 pages. [cited by applicant]
Yu et al., “On-Device Neural Language Model based Word Prediction”, 27th International Conference on Computational Linguistics: System Demonstrations, Aug. 20-26, 2018, Sante Fe, NM, pp. 128-131. [cited by applicant]
Zagoruyko et al., Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer, arXiv:1612.02928vl, Dec. 12, 2016, 12 pages. [cited by applicant]
Zhou et al., “Adaptive Quantization for Deep Neural Network”, arXiv:1712.01048v1, Dec. 4, 2017, 14 pages. [cited by applicant]
Zhu et al., “Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books”, IEEE International Conference on Computer Vision, Dec. 11-18, 2015, Santiago, Chile, pp. 19-27. [cited by applicant]