IP Library › Granted Patent US 12,210,973
Granted Patent B2
US 12,210,973 · App. 16/938,098 · Granted Jan 28, 2025

Compressing neural networks for natural language understanding

Inventor: Mark Edward Johnson (Sydney, AU)
Assignee: Oracle International Corporation
G06N3/082G06F40/205G06F40/295G06F40/30G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,973
App. No.
16/938,098
Filed
Jul 24, 2020
Granted
Jan 28, 2025
Kind
B2
Examiner
HUANG, YAO D
Art Unit
2124
USPC
706/15
Abstract

A model for a natural language understanding task is generated based on labeled data generated by a labeling model. The model for the natural language understanding task is smaller than the labeling model (i.e., with lower computational and memory requirements than the combined model), but with substantially the same performance as the labeling model. In some cases, the labeling model may be generated based on a large pre-trained model.

Claims (53)

1. A computer-implemented method comprising:

obtaining first labeled conversational data comprising conversation logs annotated with data specific to a conversational task, the conversational task comprising intent determination, named entity recognition, or semantic parsing;

obtaining second unlabeled conversational data comprising data from a plurality of websites;

training a third neural network to perform a proxy task using the second unlabeled conversational data, wherein the proxy task comprises a predictive language task, and wherein the third neural network comprises a large language model;

generating a first neural network based on the third neural network, wherein generating the first neural network based on the third neural network comprises constructing the first neural network to include a portion of the third neural network, with one or more layers of the third neural network replaced with new layers;

training the first neural network using the first labeled conversational data;

obtaining first unlabeled conversational data comprising one or more of call logs or text chat logs;

executing the trained first neural network on the first unlabeled conversational data to generate one or more labels as output of an output layer of the first neural network and label the first unlabeled conversational data with the one or more labels, wherein executing the trained first neural network to generate the one or more labels comprises performing a natural language understanding task to output each label, of the one or more labels, the natural language understanding task comprising intent determination, named entity recognition, or semantic parsing, thereby generating second labeled conversational data; and

training a second neural network for performing the natural language understanding task using the second labeled conversational data.

2. The method of claim 1 , wherein training the second neural network comprises:

for a first training input, of the second labeled conversational data, outputting by the second neural network a predicted output;

computing a loss measuring an error between the predicted output and a first label associated with the first training input;

computing, based upon the loss, updated values for a first set of parameters of the second neural network; and

updating the second neural network by changing values of the first set of parameters to the updated values.

3. A non-transitory computer-readable memory storing a plurality of instructions executable by one or more processors, the plurality of instructions comprising instructions that when executed by the one or more processors cause the one or more processors to perform processing comprising:

obtaining first labeled conversational data comprising conversation logs annotated with data specific to a conversational task, the conversational task comprising intent determination, named entity recognition, or semantic parsing;

obtaining second unlabeled conversational data comprising data from a plurality of websites;

training a third neural network to perform a proxy task using the second unlabeled conversational data, wherein the proxy task comprises a predictive language task, and wherein the third neural network comprises a large language model;

generating a first neural network based on the third neural network, wherein generating the first neural network based on the third neural network comprises constructing the first neural network to include a portion of the third neural network, with one or more layers of the third neural network replaced with new layers;

training the first neural network using the first labeled conversational data;

obtaining first unlabeled conversational data comprising one or more of call logs or text chat logs;

executing the trained first neural network on the first unlabeled conversational data to generate one or more labels as output of an output layer of the first neural network and label the first unlabeled conversational data with the one or more labels, wherein executing the trained first neural network to generate the one or more labels comprises performing a natural language understanding task to output each label, of the one or more labels, the natural language understanding task comprising intent determination, named entity recognition, or semantic parsing, thereby generating second labeled conversational data; and

training a second neural network for performing the natural language understanding task using the second labeled conversational data.

4. The non-transitory computer-readable memory of claim 3 , wherein training the second neural network comprises:

for a first training input, of the second labeled conversational data, outputting by the second neural network a predicted output;

computing a loss measuring an error between the predicted output and a first label associated with the first training input;

computing, based upon the loss, updated values for a first set of parameters of the second neural network; and

updating the second neural network by changing values of the first set of parameters to the updated values.

5. A system comprising:

one or more processors; and

a memory coupled to the one or more processors, the memory storing a plurality of instructions executable by the one or more processors, the plurality of instructions comprising instructions that when executed by the one or more processors cause the one or more processors to perform processing comprising:

obtaining first labeled conversational data comprising conversation logs annotated with data specific to a conversational task, the conversational task comprising intent determination, named entity recognition, or semantic parsing;

obtaining second unlabeled conversational data comprising data from a plurality of websites;

training a third neural network to perform a proxy task using the second unlabeled conversational data, wherein the proxy task comprises a predictive language task, and wherein the third neural network comprises a large language model;

generating a first neural network based on the third neural network, wherein generating the first neural network based on the third neural network comprises constructing the first neural network to include a portion of the third neural network, with one or more layers of the third neural network replaced with new layers;

training the first neural network using the first labeled conversational data;

obtaining first unlabeled conversational data comprising one or more of call logs or text chat logs;

executing the trained first neural network on the first unlabeled conversational data to generate one or more labels as output of an output layer of the first neural network and label the first unlabeled conversational data with the one or more labels, wherein executing the trained first neural network to generate the one or more labels comprises performing a natural language understanding task to output each label, of the one or more labels, the natural language understanding task comprising intent determination, named entity recognition, or semantic parsing, thereby generating second labeled conversational data; and

training a second neural network for performing the natural language understanding task using the second labeled conversational data.

6. The system of claim 5 , wherein training the second neural network comprises:

for a first training input, of the second labeled conversational data, outputting by the second neural network a predicted output;

computing a loss measuring an error between the predicted output and a first label associated with the first training input;

computing, based upon the loss, updated values for a first set of parameters of the second neural network; and

updating the second neural network by changing values of the first set of parameters to the updated values.

7. The method of claim 1 , wherein the first neural network is a recurrent neural network having 12 or 24 layers.

8. The method of claim 7 , wherein the second neural network is a recurrent neural network having 2 or 4 layers.

9. The method of claim 1 , wherein the first neural network is trained to perform the natural language understanding task.

10. The non-transitory computer-readable memory of claim 3 , wherein the first neural network is a recurrent neural network having 12 or more layers.

11. The non-transitory computer-readable memory of claim 10 , wherein the second neural network is a recurrent neural network having 2 or 4 layers.

12. The non-transitory computer-readable memory of claim 3 , wherein the first neural network is trained to perform the natural language understanding task.

13. The system of claim 5 , wherein the first neural network is a recurrent neural network having 12 or 24 layers.

14. The system of claim 13 , wherein the second neural network is a recurrent neural network having 2 or 4 layers.

15. The system of claim 5 , wherein the first neural network is trained to perform the natural language understanding task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2020
From: JOHNSON, MARK EDWARD
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 053317/0789 →
Continuity (2)
Provisional Application 62899650 · Sep 12, 2019
Related Publication 20210081799A1 · Mar 18, 2021
References Cited (28)
US 20190180175A1 · Meteer · 2019 [cited by examiner]
US 20190197119A1 · Zhang et al. · 2019 [cited by applicant]
US 20190205748A1 · Fukuda · 2019 [cited by examiner]
US 20190228763A1 · Czarnowski · 2019 [cited by examiner]
US 20200042596A1 · Ravi · 2020 [cited by examiner]
US 20200125927A1 · Kim · 2020 [cited by examiner]
US 20210319554A1 · Yu · 2021 [cited by examiner]
CN 106875940A · 2017 [cited by applicant]
CN 109993300A · 2019 [cited by applicant]
WO 2018126213A1 · 2018 [cited by applicant]
Ter-Sarkisov et al., “Incremental Adaptation Strategies for Neural Network Language Models,” arXiv:1412.6650v4 [cs.NE] Jul. 7, 2015 (Year: 2015). [cited by examiner]
Ba et al., “Do Deep Nets Really Need to be Deep?,” Advances in Neural Information Processing Systems 27 (NIPS 2014) (Year: 2014). [cited by examiner]
Zhang et al., “Improving Clinical Named-Entity Recognition with Transfer Learning” in E. Cummings et al. (Eds.), Connecting the System to Enhance the Practitioner and Consumer Experience in Healthcare, 2018 (Year: 2018). [cited by examiner]
Ziegler et al., “Encoder-Agnostic Adaptation for Conditional Language Generation,” arXiv:1908.06938v1 [cs.CL] Aug. 19, 2019 (Year: 2019). [cited by examiner]
Chronopoulou et al., “An Embarrassingly Simple Approach for Transfer Learning from Pretrained Language Models,” arXiv:1902.10547v3 [cs.CL] May 31, 2019 (Year: 2019). [cited by examiner]
“Mean Absolute Error”, Encyclopedia of Machine Learning, 2011, 1 page. [cited by applicant]
Bromley et al., “Signature Verification using a “Siamese” Time Delay Neural Network”, Proceedings of the 6th International Conference on Neural Information Processing Systems, Nov. 1994, pp. 737-744. [cited by applicant]
Bucila et al., “Model Compression”, Proceedings of the 12th Association for Computing Machinery Special Interest Group on Knowledge Discovery and Data Mining International Conference on Knowledge Discovery and Data Mini… [cited by applicant]
Cheng et al., “Building a Neural Semantic Parser from a Domain Ontology”, Computer Science, Dec. 25, 2018, 38 pages. [cited by applicant]
Devlin et al., “Bert: Pre-Training of Deep Bidirectional Transformers for Language Understanding”, Available Online at: https://arxiv.org/pdf/1810.04805.pdf, May 24, 2019, 16 pages. [cited by applicant]
Dong et al., “Language to Logical Form with Neural Attention”, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, vol. 1, Aug. 2016, pp. 33-43. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network”, Conference on Neural Information Processing Systems Deep Learning and Representation Learning Workshop, Mar. 9, 2015, 9 pages. [cited by applicant]
Jia et al., “Data Recombination for Neural Semantic Parsing”, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, vol. 1, Aug. 2016, 11 pages. [cited by applicant]
Nielson , “3.1: The Cross-entropy Cost Function”, Available Online at: https://eng.libretexts.org/@go/page/3752, Jul. 20, 2020, 12 pages. [cited by applicant]
Peters et al., “Deep Contextualized Word Representations”, Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 1, Jun. 201… [cited by applicant]
Wong et al., “Learning for Semantic Parsing with Statistical Machine Translation”, Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, … [cited by applicant]
Zhou et al., “Ensembling Neural Networks: Many Could Be Better Than All”, Artificial intelligence, vol. 137, No. 1-2, May 2002, pp. 239-263. [cited by applicant]
International Application No. CN202010907744.7, “Office Action”, mailed Aug. 5, 2024, 6 pages. [cited by applicant]
Cited By (6)
US 12,367,425 US 12,367,426 US 12,399,907 US 12,443,620 US 12,536,045 US 12,647,326