IP Library › Granted Patent US 12,223,269
Granted Patent B2
US 12,223,269 · App. 17/664,031 · Granted Feb 11, 2025

Language-model pretraining with gradient-disentangled embedding sharing

Inventors: Pengcheng He (Sammamish, WA); Jianfeng Gao (Woodinville, WA); Weizhu Chen (Kirkland, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F40/284G06F40/295G06N3/08G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,269
App. No.
17/664,031
Granted
Feb 11, 2025
Kind
B2
Abstract

A method for training a language model comprises (a) receiving vectorized training data as input to a multitask pretraining problem; (b) generating modified vectorized training data based on the vectorized training data, according to an upstream data embedding; (c) emitting pretraining output based on the modified vectorized training data, according to a downstream data embedding equivalent to the upstream data embedding; and (d) adjusting the upstream data embedding and the downstream data embedding by computing, based on the pretraining output, a gradient of the upstream data embedding disentangled from a gradient of the downstream data embedding, thereby advancing the multitask pretraining problem toward a pretrained state.

Claims (34)

1. A computer system configured as a language-model training service, the computer system comprising:

a language model hosted in the computer system, the language model including:

an upstream encoder configured to receive vectorized training data and emit modified vectorized training data, the upstream encoder including an upstream data embedding, and

a downstream encoder configured to receive the modified vectorized training data and emit pretraining output, the downstream encoder including a downstream data embedding equivalent to the upstream data embedding; and

pretraining logic configured to adjust the upstream data embedding and the downstream data embedding by computing a gradient of the upstream data embedding disentangled from a gradient of the downstream data embedding, wherein the gradient of the upstream data embedding is computed based on a loss function of the upstream encoder and not on a loss function of the downstream encoder, and wherein the gradient of the downstream data embedding is computed based on the loss function of the upstream encoder and the loss function of the downstream encoder.

2. The training service of claim 1 wherein the vectorized training data comprises non-masked data and the modified vectorized training data comprises masked data, wherein the upstream encoder comprises a generator of the masked data, and wherein the downstream encoder comprises a discriminator operating on the masked data.

3. The training service of claim 2 wherein the pretraining output indicates whether each of a plurality of tokens of the masked data is originally present in the non-masked data or is replaced by the generator.

4. The training service of claim 2 wherein the generator is configured to generate ambiguous corruptions in the non-masked data, and wherein the discriminator is configured to distinguish the ambiguous corruptions from tokens originally present in the non-masked data.

5. The training service of claim 2 wherein each data embedding is a token embedding, and wherein the discriminator is configured to execute replaced token detection (RTD).

6. The training service of claim 2 wherein the generator and the discriminator each comprise a neural network, and wherein the generator has half a depth of the discriminator and a full width of the discriminator.

7. The training service of claim 1 wherein the pretraining logic suppresses back propagation of the gradient of the downstream data embedding into the upstream data embedding.

8. The training service of claim 1 wherein the upstream and downstream encoders are configured to execute collectively a multitask pretraining problem.

9. A language-processing service configured for natural language understanding (NLU), the language-processing service comprising:

a language model including:

an upstream sequence of transformer blocks configured to receive vectorized training data and emit modified vectorized training data during pretraining, the upstream sequence of transformer blocks including an upstream data embedding,

a downstream sequence of transformer blocks configured to receive the modified vectorized training data and emit pretraining output during the pretraining, the downstream sequence of transformer blocks including a downstream data embedding equivalent to the upstream data embedding, wherein pretraining logic operative during the pretraining is configured to adjust the upstream data embedding and the downstream data embedding by computing a gradient of the upstream data embedding disentangled from a gradient of the downstream data embedding, wherein the gradient of the upstream data embedding is computed based on a loss function of the upstream sequence of transformer blocks and not on a loss function of the downstream sequence of transformer blocks, and wherein the gradient of the downstream data embedding is computed based on the loss function of the upstream sequence of transformer blocks and the loss function of the downstream sequence of transformer blocks,

wherein the upstream and downstream sequences of transformer blocks are configured to execute collectively a multitask pretraining problem;

an input module configured to convey language input to the language-processing model; and

an output module configured to expose an output of the language-processing model.

10. The language-processing service of claim 9 wherein the upstream sequence of transformer blocks comprises an upstream encoder, and the downstream sequence of transformer blocks comprises a downstream encoder.

11. The language-processing service of claim 9 wherein the NLU includes one or more of question answering, natural language inference, and named-entity recognition.

12. The language-processing service of claim 9 wherein the at least one of the upstream and downstream sequences of transformer blocks provide disentangled attention over a plurality of encodings.

13. The language-processing service of claim 9 wherein the vectorized training data includes multilingual training data.

14. The language-processing service of claim 9 wherein the vectorized training data comprises non-masked data and the modified vectorized training data comprises masked data, wherein the upstream sequence of transformer blocks comprises a generator of the masked data, and wherein the downstream sequence of transformer blocks comprises a discriminator operating on the masked data.

15. The language-processing service of claim 14 wherein the pretraining output indicates whether each of a plurality of tokens of the masked data is originally present in the non-masked data or is replaced by the generator.

16. A method for training a language model, the method comprising:

receiving vectorized training data in an upstream encoder, as input to a multitask pretraining problem;

generating modified vectorized training data in the upstream encoder based on the vectorized training data, according to an upstream data embedding;

emitting pretraining output from a downstream encoder based on the modified vectorized training data, according to a downstream data embedding equivalent to the upstream data embedding; and

adjusting the upstream data embedding and the downstream data embedding in pretraining logic by computing a gradient of the upstream data embedding disentangled from a gradient of the downstream data embedding, thereby advancing the multitask pretraining problem toward a pretrained state, wherein the gradient of the upstream data embedding is computed based on a loss function of the upstream encoder and not on a loss function of the downstream encoder, and wherein the gradient of the downstream data embedding is computed based on the loss function of the upstream encoder and the loss function of the downstream encoder.

17. The method of claim 16 wherein the vectorized training data comprises non-masked data and the modified vectorized training data comprises masked data.

18. The method of claim 17 wherein the pretraining output indicates whether each of a plurality of tokens of the masked data is originally present in the non-masked data.

19. The method of claim 17 wherein generating the modified vectorized training data includes generating ambiguous corruptions in the non-masked data, and wherein the pretraining output distinguishes the ambiguous corruptions from tokens originally present in the non-masked data.

20. The method of claim 17 wherein each data embedding is a token embedding, and wherein the pretraining output is a product of replaced token detection (RTD).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2022
From: HE, PENGCHENG; GAO, JIANFENG; CHEN, WEIZHU
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 059951/0627 →
Continuity (2)
Provisional Application 63264163 · Nov 16, 2021
Related Publication 20230153532A1 · May 18, 2023
References Cited (69)
US 20210174784A1 · Min · 2021 [cited by examiner]
US 20220262377A1 · Park · 2022 [cited by examiner]
Clark, Kevin; Luong, Minh-Thang; Le, Quoc V.; Manning, Christopher D., “ELECTRA: Pre-training text encoders as discriminators rather than generators”, Mar. 2020, ICLR 2020 (Year: 2020). [cited by examiner]
Zhu, et al., “Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books”, In Proceedings of the IEEE International Conference on Computer Vision, Dec. 2015, pp. 19-27. [cited by applicant]
Vaswani, et al., “Attention is all you Need”, In Proceedings of 31st Conference on Neural Information Processing Systems, Dec. 2017, 11 Pages. [cited by applicant]
Wang, et al., “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”, In Proceedings of 7th International Conference on Learning Representations, Feb. 23, 2019, 20 Pages. [cited by applicant]
Wang, et al., “MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers”, In Repository of arXiv:2002.10957v1, Feb. 25, 2020, 14 Pages. [cited by applicant]
Wang, et al., “MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers”, In Repository of arXiv:2012.15828v1, Dec. 31, 2020, 10 Pages. [cited by applicant]
Wang, et al., “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems”, In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Dec. 2019, 15 Pages. [cited by applicant]
Warstadt, et al., “Neural Network Acceptability Judgments”, In Repository of arXiv:1805.12471v1, May 31, 2018, 15 Pages. [cited by applicant]
Williams, et al., “A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human … [cited by applicant]
Xue, et al., “mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolo… [cited by applicant]
Yang, et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding”, In Proceedings of 33rd Conference on Neural Information Processing Systems, Dec. 2019, 11 Pages. [cited by applicant]
Zellers, et al., “SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Oct. 31, 2018, pp. 93-104. [cited by applicant]
Zhang, et al., “ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension”, In Repository of arXiv:1810.12885v1, Oct. 30, 2018, 14 Pages. [cited by applicant]
“Microsoft / DeBERTa”, Retrieved from: https://github.com/microsoft/DeBERTa/tree/master/experiments/, Jan. 2020, 1 Page. [cited by applicant]
“Models”, Retrieved from: https://huggingface.co/models?other-deberta-v3, Retrieved on: Mar. 15, 2022, 2 Pages. [cited by applicant]
“Tensorflow / Models”, Retrieved from: https://web.archive.org/web/20180613101354/https://github.com/tensorflow/models/tree/master/research/lm_commonsense, Retrieved on: May 2, 2020, 4 Pages. [cited by applicant]
Bentivogli, et al., “The Fifth PASCAL Recognizing Textual Entailment Challenge”, In Proceedings of the Second Text Analysis Conference, Nov. 17, 2009, 15 Pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners”, In Repository of arXiv:2005.14165v1, May 18, 2020, 72 Pages. [cited by applicant]
Cer, et al., “SemEval-2017 Task 1: Semantic Textual Similarity—Multilingual and Cross-lingual Focused Evaluation”, In Repository of arXiv:1708.00055v1, Jul. 31, 2017, 14 Pages. [cited by applicant]
Chi, et al., “XLM-E: Cross-lingual Language Model Pre-training via ELECTRA”, In Repository of arXiv:2106.16138v1, Jun. 30, 2021, 12 Pages. [cited by applicant]
Clark, et al., “BooIQ: Exploring the Surprising Difficulty of Natural Yes/No Questions”, In Repository of arXiv:1905.10044v1, May 24, 2019, 13 Pages. [cited by applicant]
Clark, et al., “ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators”, In Proceedings of International Conference on Learning Representations, Mar. 11, 2020, 18 Pages. [cited by applicant]
Clark, et al., “Google-Research / Electra”, Retrieved from: https://github.com/google-research/electra, Apr. 1, 2021, 9 Pages. [cited by applicant]
Conneau, et al., “Unsupervised Cross-lingual Representation Learning at Scale”, In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 5, 2020, pp. 8440-8451. [cited by applicant]
Conneau, et al., “XNLI: Evaluating Cross-lingual Sentence Representations”, In Proceedings of Conference on Empirical Methods in Natural Language Processing, Oct. 31, 2018, pp. 2475-2485. [cited by applicant]
Dagan, et al., “The Pascal Recognising Textual Entailment Challenge”, In Proceedings of the First International Conference on Machine Learning Challenges: Evaluating Predictive Uncertainty Visual Object Classification, … [cited by applicant]
Dagan, et al., “The Second PASCAL Recognising Textual Entailment Challenge”, In Proceedings of the Second PASCAL Challenges Workshop on Recognising Textual Entailment, Apr. 2006, 9 Pages. [cited by applicant]
Dai, et al., “Transformer-XL: Attentive Language Models beyond a Fixed-Length Context”, In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 28, 2019, pp. 2978-2988. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human … [cited by applicant]
Dolan, et al., “Automatically Constructing a Corpus of Sentential Paraphrases”, In Proceedings of the Third International Workshop on Paraphrasing, Jan. 2005, pp. 9-16. [cited by applicant]
Fedus, et al., “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”, In Repository of arXiv:2101.03961v1, Jan. 11, 2021, 31 Pages. [cited by applicant]
Giampiccolo, et al., “The Third PASCAL Recognizing Textual Entailment Challenge”, In Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, Jun. 2007, 9 Pages. [cited by applicant]
Gokaslan, Aaron, “Publications”, Retrieved from: https://skylion007.github.io/, Retrieved on: Mar. 15, 2022, 2 Pages. [cited by applicant]
Hadsell, et al., “Embracing Change: Continual Learning in Deep Neural Networks”, In Journal of Trends in Cognitive Sciences, vol. 24, Issue 12, Dec. 2020, pp. 1028-1040. [cited by applicant]
HHe, et al., “Deberta: Decoding-Enhanced Bert with Disentangled Attention”, In Proceedings of International Conference on Learning Representation, Sep. 28, 2020, 21 Pages. [cited by applicant]
He, et al., “DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing”, In Repository of arXiv:2111.09543v2, Dec. 8, 2021, 16 Pages. [cited by applicant]
Huang, et al., “Music Transformer: Generating Music with Long-Term Structure”, Retrieved from: https://magenta.tensorflow.org/music-transformer#:˜:text=Generating%20long%20pieces%20of%20music,with%20improved%20long%2Dte… [cited by applicant]
Jiao, et al., “TinyBERT: Distilling BERT for Natural Language Understanding”, In Repository of arXiv:1909.10351v1, Sep. 23, 2019, 13 Pages. [cited by applicant]
Kanakarajan, et al., “Small-Bench NLP: Benchmark for Small Single GPU Trained Models in Natural Language Processing”, In Repository of arXiv:2109.10847v1, Sep. 22, 2021, 5 Pages. [cited by applicant]
Khashabi, et al., “Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple Sentences”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Ling… [cited by applicant]
Kingma, et al., “Adam: A Method for Stochastic Optimization”, In Repository of arXiv:1412.6980v1, Dec. 22, 2014, 9 Pages. [cited by applicant]
Kiyono, Shun, “Butsugiri / Homemade_Bookcorpus”, Retrieved from: https://github.com/butsugiri/homemade_bookcorpus, Jul. 22, 2018, 3 Pages. [cited by applicant]
Kudo, Taku, “Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1: Lon… [cited by applicant]
Kumar, et al., “Microsoft / DeBERTa”, Retrieved from: https://github.com/microsoft/DeBERTa, Jan. 2020, 11 Pages. [cited by applicant]
Lai, et al., “RACE: Large-scale Reading Comprehension Dataset From Examinations”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Sep. 2017, pp. 785-794. [cited by applicant]
Lan, et al., “ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations”, In Proceedings of International Conference on Learning Representations, Sep. 26, 2019, 17 Pages. [cited by applicant]
Levesque, et al., “The Winograd Schema Challenge”, In Proceedings of Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning, May 17, 2012, pp. 552-561. [cited by applicant]
Liu, et al., “On the Variance of the Adaptive Learning Rate and Beyond”, In Proceedings of International Conference on Learning Representations, Sep. 26, 2019, 13 Pages. [cited by applicant]
Loshchilov, et al., “Fixing Weight Decay Regularization in Adam”, In Proceedings of International Conference on Learning Representations, Feb. 16, 2018, 14 Pages. [cited by applicant]
Meng, et al., “COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining”, In Repository of arXiv:2102.08473v1, Feb. 16, 2021, 13 Pages. [cited by applicant]
Nagel, Sebastian, “Cc-News”, Retrieved from: https://commoncrawl.org/2016/10/news-dataset-available/, Oct. 4, 2016, 3 Pages. [cited by applicant]
Ott, et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach”, In Repository of arXiv:1907.11692v1, Jul. 26, 2019, 13 Pages. [cited by applicant]
Pilehvar, et al., “WiC: The Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguist… [cited by applicant]
Radford, et al., “Language Models are Unsupervised Multitask Learners”, In Journal of OpenAI blog, vol. 1, Issue 8, Feb. 24, 2019, 24 Pages. [cited by applicant]
Raffel, et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, In Journal of Machine Learning Research, vol. 21, Issue 140, Jun. 2020, 67 Pages. [cited by applicant]
Rajpurkar, et al., “Know What You Don't Know: Unanswerable Questions for SQuAD”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 2: Short Papers), Jul. 15, 2018, pp. 784-… [cited by applicant]
Rajpurkar, et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Nov. 2016, pp. 2383-2392. [cited by applicant]
Sennrich, et al., “Neural Machine Translation of Rare Words with Subword Units”, In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), Aug. 7, 2016, pp. 1715-1… [cited by applicant]
Shaw, et al., “Self-Attention with Relative Position Representations”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 2… [cited by applicant]
Shoeybi, et al., “Megatron-LM: Training Multi-Billion Parameter Language Models Using GPU Model Parallelism”, In Repository of arXiv:1909.08053v1, Sep. 17, 2019, 15 Pages. [cited by applicant]
Simons, et al., “The CommitmentBank: Investigating projection in naturally occurring discourse”, In Proceedings of Sinn und Bedeutung, vol. 23, Issue 2, Jul. 25, 2019, pp. 107-124. [cited by applicant]
Socher, et al., “Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Oct. 18, 2013, pp. 1631-1642. [cited by applicant]
Trinh, et al., “A Simple Method for Commonsense Reasoning”, In Repository of arXiv:1806.02847v1, Jun. 7, 2018, 12 Pages. [cited by applicant]
Pizzati, et al., “Model-based Disentanglement of Lens Occlusions”, In Repository of arXiv:2004.01071v1, Apr. 2, 2020, 17 Pages. [cited by applicant]
Guo, et al., “Phonetic Posteriorgrams based Many-to-Many Singing Voice Conversion via Adversarial Training”, In Repository of arXiv:2012.01837v1, Dec. 3, 2020, 8 Pages. [cited by applicant]
Hu, et al., “Toward Controlled Generation of Text”, In Repository of arXiv:1703.00955v4, Sep. 13, 2018, 10 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/043564”, Mailed Date: Dec. 21, 2022, 9 Pages. [cited by applicant]
Cited By (1)
US 12,450,168