IP Library › Granted Patent US 12,223,273
Granted Patent B2
US 12,223,273 · App. 18/473,386 · Granted Feb 11, 2025

Learned evaluation model for grading quality of natural language generation outputs

Inventors: Thibault Sellam (New York City, NY); Dipanjan Das (Jersey City, NJ); Ankur Parikh (New York City, NY)
Assignee: Google LLC
G06F40/289G06F40/205G06F40/47G06F40/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,273
App. No.
18/473,386
Filed
Sep 25, 2023
Granted
Feb 11, 2025
Kind
B2
Art Unit
2657
USPC
704/9
Abstract

Systems and methods for automatic evaluation of the quality of NLG outputs. In some aspects of the technology, a learned evaluation model may be pretrained first using NLG model pretraining tasks, and then with further pretraining tasks using automatically generated synthetic sentence pairs. In some cases, following pretraining, the evaluation model may be further fine-tuned using a set of human-graded sentence pairs, so that it learns to approximate the grades allocated by the human evaluators.

Claims (30)

1. A method of training a neural network, comprising:

generating, by one or more processors, a first training signal of a plurality of training signals based on whether a given synthetic sentence pair was generated using backtranslation; and

generating, by the one or more processors, one or more second training signals of the plurality of training signals based on a prediction from a textual entailment model regarding a likelihood that a modified passage of text of the given synthetic sentence pair entails or contradicts an original passage of text of the given synthetic sentence pair;

pretraining, by the one or more processors, the neural network to predict the plurality of training signals for the given synthetic sentence pair; and

fine-tuning, by the one or more processors, the neural network to predict a grade allocated to a graded sentence pair.

2. The method of claim 1 , wherein the pretraining further includes pretraining the neural network to predict a mask token in one or more masked language modeling tasks.

3. The method of claim 2 , further comprising generating the one or more masked language modeling tasks.

4. The method of claim 1 , wherein the pretraining further includes pretraining the neural network to predict, for a next-sentence prediction task, whether a second passage of text of the next-sentence prediction task directly follows a first passage of text of the next-sentence prediction task.

5. The method of claim 4 , further comprising generating the next-sentence prediction task.

6. The method of claim 1 , further comprising generating the given synthetic sentence pair.

7. The method of claim 6 , wherein generating the given synthetic sentence pair comprises translating the original passage of text from a first language into a second language, to create a translated passage of text.

8. The method of claim 7 , wherein generating the given synthetic sentence pair further comprises translating the translated passage of text from the second language into the first language, to create the modified passage of text of the given synthetic sentence pair.

9. The method of claim 1 , wherein the grade allocated to a graded-sentence pair is based on one or more human-graded sentence pairs.

10. The method of claim 1 , further comprising generating, for the given synthetic sentence pair, a third training signal of the plurality of training signals based on one or more scores generated by comparing the original passage of text to the modified passage of text.

11. The method of claim 10 , wherein comparing the original passage of text to the modified passage of text is performed using one or more automatic metrics.

12. The method of claim 11 , wherein the one or more automatic metrics includes at least one of the BLEU metric, the ROUGE metric, or the BERTscore metric.

13. The method of claim 1 , further comprising generating, for the given synthetic sentence pair, a third training signal of the plurality of training signals based on a prediction from a backtranslation prediction model regarding a likelihood that one of the original passage of text or the modified passage of text of the given synthetic sentence pair could have been generated by backtranslating the other one of the original passage of text or the modified passage of text of the given synthetic sentence pair.

14. A processing system comprising:

memory; and

one or more processors operatively coupled to the memory and configured to:

generate a first training signal of a plurality of training signals based on whether a given synthetic sentence pair was generated using backtranslation; and

generate one or more second training signals of the plurality of training signals based on a prediction from a textual entailment model regarding a likelihood that a modified passage of text of the given synthetic sentence pair entails or contradicts an original passage of text of the given synthetic sentence pair;

pretrain a neural network to predict the plurality of training signals for the given synthetic sentence pair; and

fine-tune the neural network to predict a grade allocated to a graded sentence pair.

15. The processing system of claim 14 , wherein the one or more processors are further configured to pretrain the neural network to predict a mask token in one or more masked language modeling tasks.

16. The processing system of claim 14 , wherein the one or more processors are further configured to pretrain the neural network to predict, for a next-sentence prediction task, whether a second passage of text of the next-sentence prediction task directly follows a first passage of text of the next-sentence prediction task.

17. The processing system of claim 14 , wherein the one or more processors are further configured to generate the given synthetic sentence pair by translation of the original passage of text from a first language into a second language, to create a translated passage of text.

18. The processing system of claim 17 , wherein generation of the given synthetic sentence pair further comprises translation of the translated passage of text from the second language into the first language, to create the modified passage of text of the given synthetic sentence pair.

19. The processing system of claim 14 , wherein the one or more processors are further configured to generate, for the given synthetic sentence pair, a third training signal of the plurality of training signals based on one or more scores generated by comparison of the original passage of text to the modified passage of text.

20. The processing system of claim 14 , wherein the one or more processors are further configured to generate, for the given synthetic sentence pair, a third training signal of the plurality of training signals based on a prediction from a backtranslation prediction model regarding a likelihood that one of the original passage of text or the modified passage of text of the given synthetic sentence pair could have been generated by backtranslation of the other one of the original passage of text or the modified passage of text of the given synthetic sentence pair.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: SELLAM, THIBAULT; DAS, DIPANJAN; PARIKH, ANKUR
To: GOOGLE LLC
Reel/Frame 065008/0284 →
Continuity (3)
Continuation 18079148 · Dec 12, 2022
Continuation 17003572 · Aug 26, 2020
Related Publication 20240012999A1 · Jan 11, 2024
References Cited (80)
US 11551002B2 · Sellam et al. · 2023 [cited by applicant]
US 20120101804A1 · Roth et al. · 2012 [cited by applicant]
US 20160117316A1 · Le · 2016 [cited by examiner]
US 20170060855A1 · Song · 2017 [cited by examiner]
US 20200210772A1 · Bojar · 2020 [cited by examiner]
US 20210174204A1 · Yin · 2021 [cited by examiner]
US 20210182662A1 · Lai et al. · 2021 [cited by applicant]
US 20210365837A1 · Kashihara · 2021 [cited by examiner]
US 20220067285A1 · Sellam et al. · 2022 [cited by applicant]
US 20220067309A1 · Sellam et al. · 2022 [cited by applicant]
Smith, Ronnie W. , et al., Spoken Natural Language Dialog Systems, A Practical Approach, Oxford University Press (excerpts), 1994, pp. 1-37. [cited by applicant]
Stanojevic, Milos , et al., BEER: BEtter Evaluation as Ranking, Proceedings of the Ninth Workshop on Statistical Machine Translation, pp. 414-419,2014. [cited by applicant]
Sutskever, Ilya , et al., Sequence to Sequence Learning with Neural Networks, arXiv:1409.3215v3, Dec. 14, 2014. [cited by applicant]
Tian, Ran , et al., Sticking to The Facts: Confident Decoding for Faithful Data-To-Text Generation, arXiv:1910.08684v1, Oct. 19, 2019, pp. 1-12. [cited by applicant]
Tian, Ran , et al., Sticking to The Facts: Confident Decoding for Faithful Data-To-Text Generation, arXiv:1910.08684v2, Nov. 15, 2019, pp. 1-16. [cited by applicant]
Tomar, Gaurav Singh , et al., Neural Paraphrase Identification of Questions with Noisy Pretraining, Proceedings of the First Workshop on Subword and Character Level Models in NLP, 2017, pp. 142-147. [cited by applicant]
Turc, Iulia , et al., Well-Read Students Learn Better: On The Importance of Pre-Training Compact Models, Google Research, Sep. 25, 2019, pp. 1-13, arXiv:1908.08962v2 [cs.CL]. [cited by applicant]
Turc, Iulia , et al., Well-Read Students Learn Better: The Impact of Student Initialization on Knowledge Distillation, arXiv:1908.08962v1 [cs.CL] Sep. 23, 2019, pp. 1-12. [cited by applicant]
Vaswani, Ashish , et al., Attention is All You Need, 31st Conference on Neural Information Processing Systems (NIPS 2017), pp. 1-11. [cited by applicant]
Mnyals, Orio , et al., A Neural Conversational Model, arXiv:1506.05869v3, Jul. 22, 2015, pp. 1-8. [cited by applicant]
Wang, Alex, et al., GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp. 353-… [cited by applicant]
Wang, Alex , et al., GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, Published as a conference paper at ICLR 2019, arXiv:1804.07461v3, Feb. 22, 2019, pp. 1-20. [cited by applicant]
Wieting, John , et al., Towards Universal Paraphrastic Sentence Embeddings, Published as a conference paper at ICLR 2016, arXiv:1511.08198v3 Mar. 4, 2016, pp. 1-19. [cited by applicant]
Williams, Adina , et al., A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference, Proceedings of NAACL-HLT 2018, pp. 1112-1122. [cited by applicant]
Wiseman, Sam , et al., Challenges in Data-to-Document Generation, Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2253-2263. [cited by applicant]
Xenouleas, Stratos , et al., SUM-QE: a BERT-based Summary Quality Estimation Model, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Na… [cited by applicant]
Zhang, Tianyi, et al., Bertscore: Evaluating Text Generation with Bert, arXiv:1904.09675v1, Apr. 21, 2019. [cited by applicant]
Zhang, Tianyi , et al., Bertscore: Evaluating Text Generation With Bert, arXiv:1904.09675v3, Feb. 24, 2020, pp. 1-43. [cited by applicant]
Zhang, Tianyi , et al., Bertscore: Evaluating Text Generation With Bert, arXiv:1904.09675v2, Oct. 1, 2019, pp. 1-41. [cited by applicant]
Zhao, Wei , et al., “MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th In… [cited by applicant]
Bahdanau, Dzmitry , et al., Neural Machine Translation, By Jointly Learning to Align and Translate, arXiv:1409.0473v7, May 19, 2016, pp. 1-15. [cited by applicant]
Bannard, Colin , et al., Paraphrasing with Bilingual Parallel Corpora, Proceedings of the 43rd Annual Meeting of the ACL, pp. 597-604, Ann Arbor, Jun. 2005, Association for Computational Linguistics. [cited by applicant]
Belinkov, Yonatan , et al., Synthetic and Natural Noise Both Break Neural Machine Translation, Published as a conference paper at ICLR 2018, pp. 1-13. [cited by applicant]
Belz, Anja, et al., Comparing Automatic and Human Evaluation of NLG Systems, 11th Conference of the European Chapter of the Association for Computational Linguistics, www.aclweb.org, 2006, pp. 1-8. [cited by applicant]
Bojar, Ondrej , Results of the WMT16 Metrics Shared Task, Proceedings of the First Conference on Machine Translation, vol. 2: Shared Task Papers, 2016, pp. 199-231. [cited by applicant]
Bojar, Ondrej , et al., Results of the WMT17 Metrics Shared Task, Proceedings of the Conference on Machine Translation (WMT), vol. 2: Shared Task Papers, 2017, pp. 489-513. [cited by applicant]
Bowman, Samuel , et al., A large annotated corpus for learning natural language inference, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 632-642, 2015. [cited by applicant]
Callison-Burch, Chris , et al., Re-evaluating the Role of BLEU in Machine Translation Research, 11th Conference of the European Chapter of the Association for Computational Linguistics, https://www.aclweb.org/anthology/… [cited by applicant]
Celikyilmaz, Asli , et al., Evaluation of Text Generation: A Survey, arXiv:2006.14799v1, Jun. 26, 2020, pp. 1-58. [cited by applicant]
Chaganty, Arun Tejasvi, et al., The price of debiasing automatic metrics in natural language evaluation, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers), pp. 643-653… [cited by applicant]
Chen, Qian , et al., Enhanced LSTM for Natural Language Inference, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, pp. 1657-1668, 2017. [cited by applicant]
Chopra, Sumit , et al., Abstractive Sentence Summarization with Attentive Recurrent Neural Networks, Proceedings of NAACL-HLT 2016, pp. 93-98. [cited by applicant]
Clark, Elizabeth , et al., Sentence Mover's Similarity: Automatic Evaluation for Multi-Sentence Texts, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 28-Aug. 2, 2019, pp. 1… [cited by applicant]
Clark, Elizabeth , et al., Sentence Mover's Similarity: Automatic Evaluation for Multi-Sentence Texts, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 28-Aug. 2, 2019, pp. 2… [cited by applicant]
Devlin, Jacob , et al., Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding, Proceedings of NAACL-HLT 2019, pp. 4171-4186. [cited by applicant]
Dusek, Ondrej , et al., Automatic Quality Estimation for Natural Language Generation: Ranting (Jointly Rating and Ranking), In Proceedings of INLG, Tokyo, Japan, Oct. 2019, arXiv:1910.04731v1, pp. 1-9. [cited by applicant]
Dusek, Ondrej , et al., Automatic Quality Estimation for Natural Language Generation: Ranting (Jointly Rating and Ranking), Proceedings of The 12th International Conference on Natural Language Generation, pp. 369-376, 2… [cited by applicant]
Dusek, Ondrej , et al., Referenceless Quality Estimation for Natural Language Generation, arXiv:1708.01759v1, Aug. 5, 2017, pp. 1-9. [cited by applicant]
Eyal, Matan , et al., Question Answering as an Automatic Evaluation Metric for News Article Summarization, Proceedings of NAACL-HLT 2019, pp. 3938-3948. [cited by applicant]
Fang, Hao , et al., From Captions to Visual Concepts and Back, CVPR 2015, pp. 1-10. [cited by applicant]
Ganitkevitch, Juri , et al., PPDB: The Paraphrase Database, Proceedings of NAACL-HLT 2013, pp. 758-764. [cited by applicant]
Gardent, Claire , et al., The WebNLG Challenge: Generating Text from RDF Data, Proceedings of The 10th International Natural Language Generation conference, 2017, pp. 124-133. [cited by applicant]
Goodrich, Ben , et al., Assessing The Factual Accuracy of Generated Text, Research Track Paper, KDD '19, Aug. 4-8, 2019, Anchorage, AK, USA, pp. 166-175. [cited by applicant]
Hinton, Geoffrey , et al., “Distilling the Knowledge in a Neural Network,” arXiv:1503.02531v1 [stat.ML] Mar. 9, 2015, pp. 1-9. [cited by applicant]
Iyyer, Mohit , et al., Adversarial Example Generation with Syntactically Controlled Paraphrase Networks, Proceedings of NAACL-HLT 2018, pp. 1875-1885. [cited by applicant]
Jia, Robin , et al., Adversarial Examples for Evaluating Reading Comprehension Systems, Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2021-2031. [cited by applicant]
Koehn, Philipp , Statistical Machine Translation, Cambridge University Press, 2009, pp. 1-20. [cited by applicant]
Kukich, Karen , Design of a Knowledge-Based Report Generator, University of Pittsburgh, Bell Telephone Laboratories, 1983, pp. 145-150. [cited by applicant]
Lin, Chin-Yew , Rouge: A Package for Automatic Evaluation of Summaries, In Proceedings of Workshop on Text Summarization Branches Out, Post-Conference Workshop of ACL 2004, pp. 1-10. [cited by applicant]
Liu, Yinhan , ROBERTa: A Robustly Optimized BERT Pretraining Approach, arXiv:1907.11692v1, Jul. 26, 2019, pp. 1-13. [cited by applicant]
Liu, Chia-Wei , et al., How Not to Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Proceedings of the 2016 Conference on Empirical Methods in Natura… [cited by applicant]
Lo, Chi-Kiu , “YiSi—A unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources,” Proceedings of the Fourth Conference on Machine Translation (WMT), vol. 2: … [cited by applicant]
Ma, Qingsong , et al., Blend: a Novel Combined MT Metric Based on Direct Assessment, CASICT-DCU submission to WMT17 Metrics Task, Proceedings of the Conference on Machine Translation (WMT), vol. 2: Shared Task Papers, p… [cited by applicant]
Ma, Qingsong , et al., Results of the WMT18 Metrics Shared Task, Proceedings of the Third Conference on Machine Translation (WMT), vol. 2: Shared Task Papers, pp. 671-688, 2018. [cited by applicant]
Ma, Qingsong , et al., Results of the WMT19 Metrics Shared Task: Segment-Level and Strong MT Systems Pose Big Challenges, Proceedings of the Fourth Conference on Machine Translation (WMT), vol. 2: Shared Task Papers (Da… [cited by applicant]
Mani, Inderjeet , et al., Advances in Automatic Text Summarization, The MIT Press, Copyright 1999, pp. 1-32. [cited by applicant]
Mathur, Nitika , et al., Putting Evaluation in Context: Contextual Embeddings improve Machine Translation Evaluation, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 2799-280… [cited by applicant]
Mathur, Nitika , et al., “Tangled up in BLEU: Reevaluating the Evaluation of Automatic MachineTranslation Evaluation Metrics”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul… [cited by applicant]
Novikova, Jekaterina, et al., Why We Need New Evaluation Metrics for NLG, Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2241-2252. [cited by applicant]
Papineni, Kishore , et al., BLEU: a Method for Automatic Evaluation of Machine Translation, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL), Philadelphia, Jul. 2002, pp. 311… [cited by applicant]
Ribeiro, Marco Tulio, et al., Semantically Equivalent Adversarial Rules for Debugging NLP Models, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers), pp. 856-865, 2018. [cited by applicant]
Sellam, Thibault, Bleurt: Learning Robust Metrics for Text Generation, arXiv:2004.04696v4, May 20, 2020, pp. 1-12. [cited by applicant]
Sellam, Thibault, et al., Bleurt: Learning Robust Metrics for Text Generation, arXiv:2004.04696v1, Apr. 9, 2020, pp. 1-12. [cited by applicant]
Sellam, Thibault, et al., Bleurt: Learning Robust Metrics for Text Generation, arXiv:2004.04696v2, May 11, 2020, pp. 1-12. [cited by applicant]
Sellam, Thibault , et al., Bleurt: Learning Robust Metrics for Text Generation, arXiv:2004.04696v3, May 14, 2020, pp. 1-12. [cited by applicant]
Sellam, Thibault , et al., Bleurt: Learning Robust Metrics for Text Generation, arXiv:2004.04696v5, May 21, 2020, pp. 1-12. [cited by applicant]
Sellam, Thibault , et al., Evaluating Natural Language Generation with Bleurt, Google AI Blog: Evaluating Natural Language Generation with Bleurt, May 26, 2020, https://ai.googleblog.com/2020/05/evaluating-natural-langu… [cited by applicant]
Sennrich, Rico , et al., Improving Neural Machine Translation Models with Monolingual Data, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, pp. 86-96, 2016. [cited by applicant]
Shimanaka, Hiroki , et al., RUSE: Regressor Using Sentence Embeddings for Automatic Machine Translation Evaluation, Proceedings of the Third Conference on Machine Translation (WMT), vol. 2: Shared Task Papers, pp. 751-7… [cited by applicant]
Shimorina, Anastasia , et al., WebNLG Challenge: Human Evaluation Results, Jan. 2018, pp. 1-16. [cited by applicant]