IP Library › Granted Patent US 12,346,665
Granted Patent B2
US 12,346,665 · App. 17/670,617 · Granted Jul 1, 2025

Neural architecture search of language models using knowledge distillation

Inventors: Michele Merler (New York City, NY); Aashka Trivedi (Cedar Park, TX); Rameswar Panda (Medford, MA); Bishwaranjan Bhattacharjee (Yorktown Heights, NY); Taesun Moon (Scarsdale, NY); Avirup Sil (Hopewell Junction, NY)
Assignee: International Business Machines Corporation
G06F40/40G06N3/042
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,346,665
App. No.
17/670,617
Granted
Jul 1, 2025
Kind
B2
Abstract

A neural architecture search method, system, and computer program product that determines, by a computing device, a best fit language model of a plurality of language models that is a best fit for interpretation of a corpus of natural language and interprets, by the computing device, the corpus of natural language using the best fit language model.

Claims (30)

1. A computer-implemented neural architecture search method, the method comprising:

accessing, by a computing device, a teacher model;

utilizing, by the computing device, the teacher model to distill knowledge from a corpus of natural language using knowledge distillation (KD);

determining, by a computing device, a best fit language model of a plurality of language models for language understanding that is a best fit for interpretation of the corpus of natural language, determining the best fit language model by performing Knowledge Distillation (KD)-guided Neural Architecture Search (NAS) over the plurality of language models, the NAS process using a NAS algorithm including a language model architecture search space, the Knowledge Distillation (KD)-guided Neural Architecture Search (NAS) combining knowledge distillation (KD) of knowledge distilled from the corpus of natural language by the teacher model with estimation of performance of each of the plurality of language models by the Neural Architecture Search (NAS) to select the best fit language model which maximizes accuracy, minimizes latency, optimizes knowledge distillation (KD), and meets a constraint on size to produce a pre-trained student model architecture that outperforms a non-pre-trained randomly initialized student model architecture, the pre-trained student model architecture including knowledge distilled from the teacher model and having less than half of the parameters of the teacher model; and

interpreting, by the computing device, the corpus of natural language using the pre-trained student model architecture performing language understanding.

2. The computer-implemented neural architecture search method of claim 1 , further comprising receiving, by the computing device, access to the corpus of natural language.

3. The computer-implemented neural architecture search method of claim 1 , wherein the language model architecture search space is defined to be a set of operations of a language model including deep learning-based models.

4. The computer-implemented neural architecture search method of claim 1 , wherein the student model and teacher model share a same set of operations.

5. The computer-implemented neural architecture search method of claim 1 , wherein the student model and the teacher model have a different set of operations.

6. The computer-implemented neural architecture search method of claim 1 , wherein the teacher model includes a language model, pre-trained on a text corpora.

7. The computer-implemented neural architecture search method of claim 1 , wherein the NAS algorithm includes a reward function which includes a weighted combination of an accuracy on the downstream task, the latency and the KD from the teacher model.

8. The computer-implemented neural architecture search method of claim 7 , wherein the knowledge is distilled from the teacher model in an unsupervised manner via intermediate features similarity.

9. The computer-implemented neural architecture search method of claim 7 , wherein the knowledge is distilled from the teacher model in a supervised manner via logits similarity.

10. The computer-implemented neural architecture search method of claim 7 , wherein the knowledge is distilled from the teacher model via a combination of:

an unsupervised manner via intermediate features similarity; and

a supervised manner via logits similarity.

11. The computer-implemented neural architecture search method of claim 10 , wherein a measure of a similarity between the teacher model and the student model includes a statistical measure of a distance between vectors.

12. The method of claim 1 , wherein the neural architecture search method is used to deploy high performing language models in resource-constrained environments.

13. A neural architecture search computer program product, the neural architecture search computer program product comprising a computer-readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

accessing, by a computing device, a teacher model;

utilizing, by the computing device, the teacher model to distill knowledge from a corpus of natural language using knowledge distillation;

determining, by a computing device, a best fit language model of a plurality of language models for language understanding that is a best fit for interpretation of the corpus of natural language, determining the best fit language model by performing Knowledge Distillation (KD)-guided Neural Architecture Search (NAS) over the plurality of language models, the NAS process using a NAS algorithm including a language model architecture search space, the Knowledge Distillation (KD)-guided Neural Architecture Search (NAS) combining knowledge distillation (KD) of knowledge distilled from the corpus of natural language by the teacher model with estimation of performance of each of the plurality of language models by the Neural Architecture Search (NAS) to select the best fit language model which maximizes accuracy, minimizes latency, optimizes knowledge distillation (KD), and meets a constraint on size to produce a pre-trained student model architecture that outperforms a non-pre-trained randomly initialized student model architecture, the pre-trained student model architecture including knowledge distilled from the teacher model and having less than half of the parameters of the teacher model; and

interpreting, by the computing device, the corpus of natural language using the pre-trained student model architecture performing language understanding.

14. A neural architecture search system, said neural architecture search system comprising:

a processor; and

a memory, the memory storing instructions to cause the processor to perform a method comprising:

accessing, by a computing device, a teacher model;

utilizing, by the computing device, the teacher model to distill knowledge from a corpus of natural language using knowledge distillation (KD);

determining, by a computing device, a best fit language model of a plurality of language models that is a best fit for interpretation of the corpus of natural language, determining the best fit language model by performing Knowledge Distillation (KD)-guided Neural Architecture Search (NAS) over the plurality of language models, the NAS process using a language model architecture search space, the Knowledge Distillation (KD)-guided Neural Architecture Search (NAS) combining knowledge distillation (KD) of knowledge distilled from the corpus of natural language by the teacher model with estimation of performance of each of the plurality of language models by the Neural Architecture Search (NAS) to select a best fid model which maximizes accuracy, minimizes latency, optimizes knowledge distillation (KD), and meets a constraint on size to produce a pre-trained student model architecture that outperforms a non-pre-trained randomly initialized student model architecture, the pre-trained student model architecture including knowledge distilled from the teacher model and having less than half of the parameters of the teacher model; and

interpreting, by the computing device, the corpus of natural language using the student model performing language understanding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2025
From: MERLER, MICHELE; TRIVEDI, AASHKA; PANDA, RAMESWAR; BHATTACHARJEE, BISHWARANJAN; MOON, TAESUN; SIL, AVIRUP
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 071070/0196 →
Continuity (1)
Related Publication 20230259716A1 · Aug 17, 2023
References Cited (41)
US 11030523B2 · Zoph et al. · 2021 [cited by applicant]
US 11030997B2 · Li et al. · 2021 [cited by applicant]
US 20160247061A1 · Trask et al. · 2016 [cited by applicant]
US 20210303967A1 · Bender · 2021 [cited by examiner]
US 20210357752A1 · Chen · 2021 [cited by examiner]
US 20220019880A1 · Dasgupta · 2022 [cited by examiner]
US 20220076121A1 · Choi · 2022 [cited by examiner]
US 20220156596A1 · Park · 2022 [cited by examiner]
US 20220188658A1 · Wang · 2022 [cited by examiner]
US 20220198276A1 · Wang · 2022 [cited by examiner]
US 20230020886A1 · Mahapatra · 2023 [cited by examiner]
US 20230153577A1 · Kim · 2023 [cited by examiner]
US 20230237337A1 · Carlucci · 2023 [cited by examiner]
CN 1111710331A · 2020 [cited by applicant]
Mukherjee, Subhabrata, and Ahmed Hassan Awadallah. “Distilling bert into simple neural networks with unlabeled transfer data.” arXiv preprint arXiv: 1910.01769 (2019). (Year: 2019). [cited by examiner]
Zhang, Xiaofan, et al. “Auto Distill: An end-to-end framework to explore and distill hardware-efficient language models.” arXiv preprint arXiv:2201.08539 (2022). (Year: 2022). [cited by examiner]
Mel, et al. “The NIST Definition of Cloud Computing”. Recommendations of the National Institute of Standards and Technology. Nov. 16, 2015. [cited by applicant]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, I. Polosukhin. “Attention Is All You Need”, in NeurIPS, 2017. [cited by applicant]
B. Zoph and Q. V. Le. “Neural architecture search with reinforcement learning”, in ICLR, 2017. [cited by applicant]
Bello et al. “Neural optimizer search with reinforcement learning.” International Conference on Machine Learning. PMLR, 2017. [cited by applicant]
David R. So, Chen Liang, Quoc V. Le. “The Evolved Transformer”, ICML, 2019. [cited by applicant]
G. Prato, E. Charlaix, M. Rezagholizadeh. “Fully Quantized Transformer for Improved Translation”, in NeurIPS Workshops, 2019. [cited by applicant]
Geoffrey Hinton, Oriol Vinyals, Jeff Dean. “Distilling the knowledge in a neural network”, in arXiv, 2015. [cited by applicant]
H. Liu, K. Simonyan, and Y. Yang. “Darts: Differentiable architecture search”. arXiv preprint arXiv: 1806.09055, 2018. [cited by applicant]
J. Xu, X. Tan, R. Luo, K. Song, J. Li, T. Qin, T. Liu, “NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search”, in AAAI, 2021. [cited by applicant]
Merity et al., Regularizing and optimizing LSTM language models. arXiv preprint arXiv: 1708.02182 (2017). [cited by applicant]
Pham et al. “Efficient neural architecture search via parameters sharing.” International Conference on Machine Learning. PMLR, 2018. [cited by applicant]
Pouya Bashivan, Mark Tensen, James J DiCarlo. “Teacher Guided Architecture Search”, in CVPR, 2019. [cited by applicant]
R. Tang, Y. Lu, L. Liu, L. Mou, et al. “Distilling Task-Specific Knowledge from BERT into Simple Neural Networks”, in arXiv, 2019. [cited by applicant]
S. Mukherjee, A. Awadallah, J. Gao. “XtremeDistilTransformers: Task Transfer for Task-agnostic Distillation”, in arXiv, 2021. [cited by applicant]
Subhabrata Mukherjee, Ahmed Awadallah. “XtremeDistil: Multi-stage Distillation for Massive Multilingual Models”, in ACL, 2020. [cited by applicant]
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. “Neural architecture search: A survey.”, in arXiv, 2018. [cited by applicant]
V. Sanh, L. Debut, J. Chaumond, T. Wolf. “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”, in NeurIPS 2019. [cited by applicant]
W. Wang et al. “MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers”, in NeurIPS, 2020. [cited by applicant]
William Fedus, Barret Zoph, Noam Shazeer. “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”, in arXiv, 2021. [cited by applicant]
Xiaobo Wang. “Teacher Guided Neural Architecture Search for Face Recognition”, in AAAI, 2021. [cited by applicant]
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, et al. “RoBERTa: A Robustly Optimized BERT Pretraining Approach”, in arXiv, 2019. [cited by applicant]
Y. Liu, X. Jia, M. Tan, R. Vemulapalli, Y. Zhu, B. Green, X. Wang. “Search to Distill: Pearls Are Everywhere but Not the Eyes”, in CVPR, 2020. [cited by applicant]
Yu al. “Evaluating the search phase of neural architecture search.” ar Xiv preprint arXiv: 1902.08142 (2019). [cited by applicant]
Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, Song Han. “Lite Transformer with Long-Short Range Attention”, in ICLR, 2020. [cited by applicant]
Zhiheng Huang, Wei Xu, Kai Yu. “Bidirectional LSTM-CRF Models for Sequence Tagging”, in arXiv, 2015. [cited by applicant]
Cited By (1)
US 12,626,056