IP Library Granted Patent US 12,236,206
Granted Patent B1
US 12,236,206 · App. 17/331,478 · Granted Feb 25, 2025

Pretraining a language machine-learning model

Inventors: Michael William Lewis (Seattle, WA); Marjan Ghazvini Nejad (Seattle, WA); Gargi Ghosh (Bellevue, WA); Armen Aghajanyan (Bellevue, WA); Sida Wang (Bellevue, WA); Luke Zettlemoyer (Seattle, WA)
Assignee: Meta Platforms, Inc.
G06F40/58G06F18/22G06N3/048G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,206
App. No.
17/331,478
Granted
Feb 25, 2025
Kind
B1
Abstract

In one embodiment, a method includes accessing a first document, accessing a plurality of second documents, calculating a relevance score for each of the plurality of second documents indicating a degree of relevance of the second document to the first document using an encoder of a machine-learning model, selecting a subset of the second documents based on their corresponding relevance scores, generating a target document by using the machine-learning model to process the subset of second documents and their corresponding relevance scores, and updating parameters of the machine-learning model based on a comparison between the first document and the generated target document.

Claims (46)

1. A method comprising:

accessing a first document;

accessing a plurality of second documents;

determining, for the plurality of second documents, corresponding relevance scores indicating a degree of relevance of the plurality of second documents to the first document using a machine-learning model;

selecting a subset of the plurality of second documents based on the corresponding relevance scores;

generating a target document by using the machine-learning model to process the subset of the plurality of second documents and the corresponding relevance scores; and

updating parameters of the machine-learning model based on a comparison between the first document and the generated target document.

2. The method of claim 1 , wherein the machine-learning model comprises a sequence-to-sequence machine-learning model that comprises an encoder and a decoder.

3. The method of claim 1 , wherein the plurality of second documents comprise documents published on a date that the first document is published on.

4. The method of claim 1 , wherein the plurality of second documents comprise documents written in different languages.

5. The method of claim 1 , wherein the plurality of second documents comprise documents associated with corresponding relevance scores determined using an encoder of the machine-learning model with parameter values updated in previous trainings exceed a pre-determined threshold.

6. The method of claim 1 , wherein determining a relevance score for a second document comprises:

generating a first embedding vector representing the first document using the machine-learning model;

generating a second embedding vector representing the second document using the machine-learning model; and

determining a relevance metric between the first embedding vector and the second embedding vector.

7. The method of claim 6 , wherein the relevance metric comprises a cosine similarity.

8. The method of claim 1 , wherein the selecting the subset of the plurality of second documents comprises selecting k second documents whose associated relevance scores are higher than corresponding relevance scores for other second documents among the plurality of second documents.

9. The method of claim 1 , wherein the generating the target document by using the machine-learning model to process the subset of the plurality of second documents and the corresponding relevance scores comprises:

generating, for the subset of the plurality of second documents, one or more latent representations using the machine-learning model;

concatenating the generated one or more latent representations, wherein a corresponding relevance score for generated embedding vectors is used to bias cross-attention from a decoder of the machine-learning model to an encoder of the machine-learning model; and

generating the target document by using the decoder of the machine-learning model to process the generated one or more latent representations.

10. The method of claim 1 , wherein the updating parameters of the machine-learning model is performed as a backpropagation procedure.

11. The method of claim 1 , wherein the machine-learning model, after being trained, is used for at least one task.

12. The method of claim 11 , wherein the at least one task comprises a paraphrasing of a document, or a translation of a document.

13. The method of claim 11 , wherein the at least one task comprises a multi-document summarization, wherein a plurality of documents and associated pre-determined relevance scores are processed by the machine-learning model to generate a document summarizing the plurality of documents.

14. The method of claim 13 , wherein the pre-determined corresponding relevance scores are identical to each other.

15. The method of claim 11 , wherein the at least one task comprises an information retrieval, wherein k documents among a large number of documents that are more relevant to a given document are selected based on their corresponding relevance scores determined by an encoder of the machine-learning model.

16. The method of claim 11 , wherein the at least one task comprises a document classification.

17. The method of claim 16 , wherein an encoder of the machine-learning model is connected to a classifier that is trained to determine a class of an input document based on a latent representation that the encoder generates based on the input document.

18. The method of claim 16 , wherein a decoder of the machine-learning model is re-trained to generate a word string indicating a class of an input document.

19. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

access a first document;

access a plurality of second documents;

determine, for the plurality of second documents, corresponding relevance scores indicating a degree of relevance of the plurality of second documents to the first document using a machine-learning model;

select a subset of the plurality of second documents based on the corresponding relevance scores;

generate a target document by using the machine-learning model to process the subset of the plurality of second documents and the corresponding relevance scores; and

update parameters of the machine-learning model based on a comparison between the first document and the generated target document.

20. A system comprising:

one or more processors; and

a non-transitory memory coupled to the one or more processors comprising instructions executable by the one or more processors, the one or more processors operable when executing the instructions to:

access a first document;

access a plurality of second documents;

determine, for the plurality of second documents, corresponding relevance scores indicating a degree of relevance of the plurality of second documents to the first document using a machine-learning model;

select a subset of the plurality of second documents based on the corresponding relevance scores;

generate a target document by using the machine-learning model to process the subset of the plurality of second documents and the corresponding relevance scores; and

update parameters of the machine-learning model based on a comparison between the first document and the generated target document.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2021
From: LEWIS, MICHAEL WILLIAM; GHAZVINI NEJAD, MARJAN; GHOSH, GARGI; AGHAJANYAN, ARMEN; WANG, SIDA; ZETTLEMOYER, LUKE
To: FACEBOOK, INC.
Reel/Frame 056687/0510 →
References Cited (64)
US 8095544B2 · Boone · 2012 [cited by examiner]
US 8533148B1 · Feuersanger · 2013 [cited by examiner]
US 10083229B2 · Eyres · 2018 [cited by examiner]
US 10324936B2 · Feuersänger · 2019 [cited by examiner]
US 11232358B1 · Ramezani · 2022 [cited by examiner]
US 11410072B2 · Burstein · 2022 [cited by examiner]
US 11436419B2 · Li · 2022 [cited by examiner]
US 11921728B2 · Ahmed · 2024 [cited by examiner]
US 20130006954A1 · Nikoulina · 2013 [cited by examiner]
US 20130103390A1 · Fujita · 2013 [cited by examiner]
US 20130212090A1 · Sperling · 2013 [cited by examiner]
US 20140350914A1 · Andrade Silva · 2014 [cited by examiner]
US 20160098456A1 · Contreras · 2016 [cited by examiner]
US 20160155067A1 · Dubnov · 2016 [cited by examiner]
US 20170228434A1 · Beller · 2017 [cited by examiner]
US 20190163817A1 · Milenova · 2019 [cited by examiner]
US 20200210523A1 · Aghajanyan · 2020 [cited by examiner]
US 20210133498A1 · Zhang · 2021 [cited by examiner]
US 20210142210A1 · Cheng · 2021 [cited by examiner]
US 20220075945A1 · Zhang · 2022 [cited by examiner]
US 20220083744A1 · Li · 2022 [cited by examiner]
US 20220198144A1 · Yang · 2022 [cited by examiner]
US 20220245161A1 · Ahmed · 2022 [cited by examiner]
Shen et al., title={Zero-shot cross-lingual neural headline generation}, journal={IEEE/ACM Transactions on Audio, Speech, and language Processing}, volume={26}, number={12}, pages={2319-2327}, year={2018}, publisher=IEE… [cited by examiner]
Title={Zero-shot paraphrase generation with multilingual language models}, author={Guo, Yinpeng and Liao, Yi and Jiang, Xin and Zhang, Qing and Zhang, Yibo and Liu, Qun}, journal={arXiv preprint arXiv: 1911.03597}, year… [cited by examiner]
Artetxe M., et al., “Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond,” Transactions of the Association for Computational Linguistics, 2019, vol. 7, pp. 597-610. [cited by applicant]
Artetxe M., et al., “Unsupervised Neural Machine Translation,” arXiv preprint arXiv:1710.11041, 2017, 11 pages. [cited by applicant]
Clark K., et al., “ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators,” arXiv preprint arXiv:2003.10555, 2020, 18 pages. [cited by applicant]
Conneau A., et al., “Unsupervised Cross-lingual Representation Learning at Scale,” arXiv preprint arXiv:1911.02116, 2019, 12 pages. [cited by applicant]
Devlin J., et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” arXiv preprint arXiv:1810.04805, 2018, 14 pages. [cited by applicant]
Dong L., et al., “Unified Language Model Pre-training for Natural Language Understanding and Generation,” Microsoft Research, arXiv preprint arXiv:1905.03197, 2019, 14 pages. [cited by applicant]
Fan A., et al., “Controllable Abstractive Summarization,” arXiv preprint arXiv:1711.05217, 2017, 10 pages. [cited by applicant]
Guu K., et al., “Generating Sentences by Editing Prototypes,” Transactions of the Association for Computational Linguistics, 2018, vol. 6, pp. 437-450. [cited by applicant]
Guu K., et al., “REALM: Retrieval-Augmented Language Model Pre-Training,” arXiv preprint arXiv:2002.08909, 2020, 12 pages. [cited by applicant]
Hu J., et al., “Xtreme: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization,” arXiv preprint arXiv:2003.11080, 2020, 20 pages. [cited by applicant]
Johnson J., et al., “Billion-Scale Similarity Search with GPUs,” IEEE Transactions on Big Data, Jul.-Sep. 2021, vol. 7, No. 3, pp. 535-547. [cited by applicant]
Johnson M., et al., “Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation,” Transactions of the Association for Computational Linguistics, 2017, vol. 5, pp. 339-351. [cited by applicant]
Joulin A., et al., “FastText.zip: Compressing Text Classification Models,” arXiv preprint arXiv:1612.03651, 2016, 13 pages. [cited by applicant]
Kaplan J., et al., “Scaling Laws for Neural Language Models,” arXiv preprint arXiv:2001.08361, 2020, 30 pages. [cited by applicant]
Khandelwal U., et al., “Generalization Through Memorization: Nearest Neighbor Language Models,” arXiv preprint arXiv:1911.00172, 2019, 13 pages. [cited by applicant]
Lample G., et al., “Cross-lingual Language Model Pretraining,” arXiv preprint arXiv:1901.07291, 2019, 10 pages. [cited by applicant]
Lample G., et al., “Unsupervised Machine Translation Using Monolingual Corpora Only,” arXiv preprint arXiv:1711.00043, 2017, 12 pages. [cited by applicant]
Lewis M., et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,” arXiv preprint arXiv:1910.13461, 2019, 10 pages. [cited by applicant]
Lewis M., et al., “Pre-training via Paraphrasing,” arXiv preprint arXiv:2006.15020v1 [cs.CL], Jun. 26, 2020, 14 pages. [cited by applicant]
Lewis P., et al., “MLQA: Evaluating Cross-lingual Extractive Question Answering,” arXiv preprint arXiv:1910.07475, 2019, 14 pages. [cited by applicant]
Lewis P., et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” arXiv preprint arXiv:2005.11401, 2020, 19 pages. [cited by applicant]
Li Z., et al., “Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers,” arXiv preprint arXiv:2002.11794, 2020, 14 pages. [cited by applicant]
Liu P.J., et al., “Generating Wikipedia by Summarizing Long Sequences,” arXiv preprint arXiv:1801.10198, 2018, 18 pages. [cited by applicant]
Liu Y., et al., “Multilingual Denoising Pre-Training for Neural Machine Translation,” arXiv preprint arXiv:2001.08210, 2020, 17 pages. [cited by applicant]
Liu Y., et al., “ROBERTa: A Robustly Optimized BERT Pretraining Approach,” arXiv preprint arXiv:1907.11692, 2019, 13 pages. [cited by applicant]
McCann B., et al., “Learned in Translation: Contextualized Word Vectors,” In Advances in Neural Information Processing Systems, 2017, pp. 6294-6305. [cited by applicant]
Miculicich L., et al., “Document-Level Neural Machine Translation with Hierarchical Attention Networks,” arXiv preprint arXiv:1809.01576, 2018, 8 pages. [cited by applicant]
Post M., “A Call for Clarity in Reporting BLEU Scores,” arXiv preprint arXiv:1804.08771, 2018, 6 pages. [cited by applicant]
Raffel C., et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” arXiv preprint arXiv:1910.10683, 2019, 53 pages. [cited by applicant]
Rajpurkar P., et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” arXiv preprint arXiv:1606.05250, 2016, 10 pages. [cited by applicant]
Rogers A., et al., “A Primer in BERTology: What We Know About How BERT Works,” arXiv preprint arXiv:2002.12327, 2020, 23 pages. [cited by applicant]
Schwenk H., et al., “CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web,” arXiv preprint arXiv:1911.04944, 2019, 13 pages. [cited by applicant]
Scialom T., et al., “MLSUM: The Multilingual Summarization Corpus,” arXiv preprint arXiv:2004.14900, 2020, 16 pages. [cited by applicant]
Siddhant A., et al., “Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine Translation,” arXiv preprint arXiv:1909.00437, 2019, 13 pages. [cited by applicant]
Vaswani A., et al., “Attention is All You Need,” Advances in Neural Information Processing Systems, 2017, pp. 5998-6008. [cited by applicant]
Wieting J., et al., “No Training Required: Exploring Random Encoders for Sentence Classification,” arXiv preprint arXiv:1901.10444, 2019, 16 pages. [cited by applicant]
Yang Y., et al., “Paws-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification,” arXiv preprint arXiv:1908.11828, 2019, 6 pages. [cited by applicant]
Yang Z., et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding,” arXiv preprint arXiv:1906.08237, 2019, 18 pages. [cited by applicant]
Zweigenbaum P., et al., “Overview of the Third BUCC Shared Task: Spotting Parallel Sentences in Comparable Corpora,” In Proceedings of 11th Workshop on Building and Using Comparable Corpora, 2018, pp. 39-42. [cited by applicant]