IP Library › Granted Patent US 12,536,388
Granted Patent B2
US 12,536,388 · App. 18/386,343 · Granted Jan 27, 2026

Learning self-evaluation to improve selective prediction in LLMs

Inventors: Jinsung Yoon (San Jose, CA); Jiefeng Chen (Madison, WI); Sayna Ebrahimi (Los Altos Hills, CA); Sercan Omer Arik (San Francisco, CA)
Assignee: Google LLC
G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,388
App. No.
18/386,343
Granted
Jan 27, 2026
Kind
B2
Abstract

Aspects of the disclosure are directed to methods, systems, and computer readable media for adaptation with self-evaluation to improve selective prediction in large language models (LLMs), generally referred to as ASPIRE. ASPIRE includes training LLMs on a portion of training data from a question answering task to learn self-evaluation, e.g., learn to distinguish whether a generated answer is correct or not. ASPIRE further includes a selection score that combines a likelihood of that generated answer is correct with a self-evaluation score for selective prediction. ASPIRE demonstrates improved selective prediction performance with less computational cost.

Claims (51)

1 . A method for selective prediction, comprising:

training, by one or more processors, a large language model (LLM) to a task by adjusting first adaptable parameters to the task using training data;

generating, by one or more processors, a plurality of outputs associated with the task using the LLM with the adjusted first adaptable parameters;

training, by the one or more processors, the LLM on self-evaluation by:

freezing the adjusted first adaptable parameters;

determining whether each of the plurality of outputs is correct using an evaluation metric; and

adjusting second adaptable parameters for the LLM based on the determination; and

generating, by the one or more processors, a prediction for the task using the LLM based on the adjusted first adaptable parameters and adjusted second adaptable parameters, the prediction comprising a self-evaluation score.

2 . The method of claim 1 , wherein the LLM is a pretrained LLM.

3 . The method claim 1 , wherein training the LLM to the task and on self-evaluation comprises fine-tuning the LLM using soft prompt tuning.

4 . The method of claim 1 , wherein training the LLM to the task further comprises:

freezing model parameters of the LLM; and

adding and iteratively updating the first adaptable parameters for the LLM.

5 . The method of claim 1 , wherein training the LLM on self-evaluation further comprises:

freezing model parameters of the LLM; and

adding and iteratively updating the second adaptable parameters for the LLM.

6 . The method of claim 5 , wherein determining whether each of the plurality of outputs is correct further comprises labeling each of the plurality of outputs as correct or wrong.

7 . The method of claim 5 , wherein the evaluation metric compares a similarity of an output of the plurality of outputs to a reference output.

8 . The method of claim 7 , wherein determining whether each of the plurality of outputs is correct further comprises determining whether the output is within a threshold of the reference output.

9 . The method of claim 8 , wherein the threshold is a Rouge threshold having a value large enough where outputs that are wrong are not determined to be correct.

10 . A system comprising:

one or more processors; and

one or more storage devices coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations for selective prediction, the operations comprising:

training a large language model (LLM) to a task by adjusting first adaptable parameters to the task using training data;

generating a plurality of outputs associated with the task using the LLM with the adjusted first adaptable parameters;

training the LLM on self-evaluation by:

freezing the adjusted first adaptable parameters;

determining whether each of the plurality of outputs is correct using an evaluation metric; and

adjusting second adaptable parameters for the LLM based on the determination; and

generating a prediction for the task using the LLM based on the adjusted first adaptable parameters and adjusted second adaptable parameters, the prediction comprising a self-evaluation score.

11 . The system of claim 10 , wherein the LLM is a pretrained LLM.

12 . The system of claim 10 , wherein training the LLM to the task and on self-evaluation comprises fine-tuning the LLM using soft prompt tuning.

13 . The system of claim 10 , wherein training the LLM to the task further comprises:

freezing model parameters of the LLM; and

adding and iteratively updating the first adaptable parameters for the LLM.

14 . The system of claim 10 , wherein training the LLM on self-evaluation further comprises:

freezing model parameters of the LLM; and

adding and iteratively updating the second adaptable parameters for the LLM.

15 . The system of claim 14 , wherein determining whether each of the plurality of outputs is correct further comprises labeling each of the plurality of outputs as correct or wrong.

16 . The system of claim 14 , wherein the evaluation metric compares a similarity of an output of the plurality of outputs to a reference output.

17 . The system of claim 16 , wherein determining whether each of the plurality of outputs is correct further comprises determining whether the output is within a threshold of the reference output.

18 . The system of claim 17 , wherein the threshold is a Rouge threshold having a value large enough where outputs that are wrong are not determined to be correct.

19 . A non-transitory computer readable medium for storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for selective prediction, the operations comprising:

training a large language model (LLM) to a task by adjusting first adaptable parameters to the task using training data;

generating a plurality of outputs associated with the task using the LLM with the adjusted first adaptable parameters;

training the LLM on self-evaluation by:

freezing the adjusted first adaptable parameters;

determining whether each of the plurality of outputs is correct using an evaluation metric; and

adjusting second adaptable parameters for the LLM based on the determination; and

generating a prediction for the task using the LLM based on the adjusted first adaptable parameters and adjusted second adaptable parameters, the prediction comprising a self-evaluation score.

20 . The non-transitory computer readable medium of claim 19 , wherein training the LLM to the task and on self-evaluation comprises fine-tuning the LLM using soft prompt tuning.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2023
From: YOON, JINSUNG; CHEN, JIEFENG; EBRAHIMI, SAYNA; ARIK, SERCAN OMER
To: GOOGLE LLC
Reel/Frame 065435/0619 →
Continuity (2)
Provisional Application 63521930 · Jun 20, 2023
Related Publication 20240428015A1 · Dec 26, 2024
References Cited (33)
US 20210209513A1 · Torres · 2021 [cited by examiner]
US 20230342559A1 · Bhardwaj · 2023 [cited by examiner]
US 20240296294A1 · Imani · 2024 [cited by examiner]
US 20240419912A1 · Somech · 2024 [cited by examiner]
Chen et al. “Adaptation with Self-Evaluation to Improve Selective Prediction in LLMs”. arXiv:2310.11689v2 [cs.CL] Nov. 11, 2023 (Year: 2023). [cited by examiner]
Huang et al. “Large Language Models Can Self-Improve”. arXiv:2210.11610v2 [cs.CL] Oct. 25, 2022 (Year: 2022). [cited by examiner]
Amayuelas et al. (“Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models”. arXiv:2305.13712v1 [cs.CL] May 23, 2023 (Year: 2023). [cited by examiner]
Madaan et al. (“Self-Refine: Iterative Refinement with Self-Feedback”. arXiv:2303.17651v2 [cs.CL] May 25, 2023 (Year: 2023). [cited by examiner]
Wang et al. “Self-Instruct: Aligning Language Models with Self-Generated Instructions”. arXiv:2212.10560v2 [cs.CL] May 25, 2023 (Year: 2023). [cited by examiner]
Madaan et al. “Self-Refine: Iterative Refinement with Self-Feedback”. rXiv:2303.17651v2 [cs.CL] May 25, 2023 (Year: 2023). [cited by examiner]
Adina Williams, Nikita Nangia, and Samuel R Bowman. 2017. A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426. 12 pages. [cited by applicant]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30. 15 pag… [cited by applicant]
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 3045-3059. [cited by applicant]
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. 2022. Prompting gpt-3 to be reliable. arXiv preprint arXiv:2210.09150. 24 pages. [cited by applicant]
Chin-Yew Lin and Franz Josef Och. 2004. Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In Proceedings of the 42nd Annual Meeting of the Association for C… [cited by applicant]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685. 26 pages. [cited by applicant]
Elizabeth Bondi, Raphael Koster, Hannah Sheahan, Martin Chadwick, Yoram Bachrach, Taylan Cemgil, Ulrich Paquet, and Krishnamurthy Dvijotham. 2022. Role of human-ai interaction in selective prediction. In Proceedings of … [cited by applicant]
Gabriel Poesia, Oleksandr Polozov, Vu Le, Ashish Tiwari, Gustavo Soares, Christopher Meek, and Sumit Gulwani. 2022. Synchromesh: Reliable code generation from pre-trained language models. arXivpreprint arXiv:2201.11227.… [cited by applicant]
Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021. The art of abstention: Selective prediction and error regularization for natural language processing. In Proceedings of the 59th Annual Meeting of the Association … [cited by applicant]
Jie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu. 2022. Out-of-distribution detection and selective generation for conditional language models. arXiv preprint arXi… [cited by applicant]
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, et al. 2023. Towards expert-level medical question answering with large language … [cited by applicant]
Liyan Tang, Zhaoyi Sun, Betina Idnay, Jordan G Nestor, Ali Soroush, Pierre A Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin Rousseau, et al. 2023. Evaluating large language models on medical evidence summarization. m… [cited by applicant]
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Represen… [cited by applicant]
Neeraj Varshney, Swaroop Mishra, and Chitta Baral. 2022. Investigating selective prediction approaches across several tasks in iid, ood, and adversarial settings. In Findings of the Association for Computational Linguis… [cited by applicant]
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654. 17 pages. [cited by applicant]
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXi… [cited by applicant]
Shun Zhang, Zhenfang Chen, Yikang Shen, Mingyu Ding, Joshua B Tenenbaum, and Chuang Gan. 2023a. Planning with large language models for code generation. arXiv preprint arXiv:2303.05510. 28 pages. [cited by applicant]
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B Hashimoto. 2023b. Benchmarking large language models for news summarization. arXiv preprint arXiv:2301.13848. 14 pages. [cited by applicant]
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2021a. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXivpreprint arXiv:2110.0… [cited by applicant]
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021b. Gpt understands, too. arXiv preprint arXiv:2103.10385. 10 pages. [cited by applicant]
Kuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171. 24 pages. [cited by applicant]
Yonatan Geifman and Ran El-Yaniv. 2017. Selective classification for deep neural networks. Advances in neural information processing systems, 30. 10 pages. [cited by applicant]
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational … [cited by applicant]