IP Library › Granted Patent US 12,711,317
Granted Patent B2
US 12,711,317 · App. 18/140,658 · Granted Aug 18, 2026

Interacting with a language model using external knowledge and feedback

Inventors: Baolin Peng (Issaquah, WA); Michel Galley (Seattle, WA); Hao Cheng (Kirkland, WA); Pengcheng He (Sammamish, WA); Nguyen Hung Bach (Redmond, WA); Weizhu Chen (Kirkland, WA); Jianfeng Gao (Woodinville, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F40/40G06F16/3325
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,317
App. No.
18/140,658
Granted
Aug 18, 2026
Kind
B2
Abstract

A technique supplements a language model with knowledge information retrieved from external sources. The technique operates by: receiving a query; receiving knowledge information based on the query; generating original model-input information that includes the query and the knowledge information; and presenting the original model-input information to the language model. The technique further includes: receiving an original response from the language model; generating a usefulness measure that identifies usefulness of the original response; and determining whether the usefulness measure satisfies a prescribed test. Upon determining that the usefulness measure does not satisfy the test, the technique includes: generating revised model-input information that includes feedback information; presenting the revised model-input information to the language model; and receiving a revised response from the language model. According to some implementations, the technique eliminates or reduces artificial hallucination exhibited by the language model.

Claims (54)

1 . A computer-implemented method for interacting with a machine-trained language model, comprising:

receiving an input query;

providing knowledge information based on the input query;

generating original model-input information that includes the input query and the knowledge information, and presenting the original model-input information to the language model;

receiving an original response from the language model, the language model being a generative model trained to auto-regressively predict tokens to follow an initial set of tokens;

programmatically generating a usefulness measure that identifies usefulness of the original response;

in response to determining that the usefulness measure does not satisfy a prescribed test, generating revised model-input information that includes feedback information, presenting the revised model-input information to the language model, and receiving a revised response from the language model in response to the revised model-input information; and

generating and presenting one or more revised instances of model-input information until it is determined that the language model has generated a response that satisfies the prescribed test.

2 . The method of claim 1 , wherein the language model includes weights that are produced in a pre-training operation, and wherein the weights of the language model remain fixed during training of other machine-trained logic used by the method.

3 . The method of claim 1 , wherein the language model includes attention logic for assessing relevance to be given to a part of input information fed to the attention logic when interpreting another part of the input information.

4 . The method of claim 1 , wherein the generating of the revised model-input information is performed upon receiving a user request to generate the revised response.

5 . The method of claim 1 , wherein different actions performed by the method are chosen by a state machine based on state information and a policy, the state information describing aspects of a current dialogue state, and the policy expressing logic for mapping different instances of state information to the different actions.

6 . The method of claim 5 , wherein the state information describes aspects of a current dialogue turn, including at least:

the query;

the knowledge information; and

a last-received response from the language model.

7 . The method of claim 6 , wherein the state information also describes a history of previous dialogue turns, prior to the current dialogue turn.

8 . The method of claim 5 , wherein the policy is chosen to maximize attainment of an objective, and wherein an extent to which an action advances the objective is expressed by a reward signal.

9 . The method of claim 1 , wherein the providing knowledge information comprises:

retrieving initial knowledge information that matches the input query, from one or more knowledge sources;

identifying a chain of evidence based on the initial knowledge; and

validating the chain of evidence, to produce final knowledge information.

10 . The method of claim 1 , wherein the generating of the usefulness measure includes assessing an extent of overlap between the original response and the knowledge information.

11 . The method of claim 1 , further comprising generating the feedback information by retrieving pre-generated prompt information from a data store.

12 . The method of claim 1 , further comprising generating the feedback information using another generative machine-trained model, based on state information that describes aspects of a current dialogue state.

13 . The method of claim 1 , further comprising generating the feedback information using the language model, based on state information that describes aspects of a current dialogue state.

14 . A computing system for interacting with a machine-trained language model, comprising:

computer-readable storage media including:

an instruction data store for storing computer-readable instructions; and

a state data store for storing state information, the state information describing aspects of a current dialogue state; and

a processing system including one or more processors for executing the computer-readable instructions based on the state information in the state data store, to perform operations including:

receiving an input query;

providing knowledge information based on the input query;

generating original model-input information including the input query and the knowledge information, and presenting the original model-input information to the language model;

receiving an original response from the language model, the language model being a generative model trained to auto-regressively predict tokens to follow an initial set of tokens;

programmatically generating a usefulness measure that identifies usefulness of the original response;

in response to determining that the usefulness measure does not satisfy a prescribed test, generating revised model-input information that includes feedback information, presenting the revised model-input information to the language model, and receiving a revised response from the language model in response to the revised model-input information; and

generating and presenting one or more revised instances of model-input information until it is determined that the language model has generated a response that satisfies the prescribed test.

15 . The computing system of claim 14 , wherein the processing system implements a state machine for performing different actions based on the state information and a policy, the policy expressing logic for mapping different instances of state information to the different actions.

16 . The computing system of claim 14 , wherein the operations further include generating the feedback information by retrieving pre-generated prompt information from a data store.

17 . The computing system of claim 14 , wherein the operations further include generating the feedback information using another generative machine-trained model, based on the state information.

18 . The computing system of claim 14 , wherein the operations further include generating the feedback information using the language model, based on the state information.

19 . A computer-readable storage medium for storing computer-readable instructions, a processing system executing the computer-readable instructions to perform operations, the operations comprising:

receiving an input query;

providing knowledge information based on the input query;

generating original model-input information that includes the input query and the knowledge information, and presenting the original model-input information to machine-trained a language model;

receiving an original response from the language model, the language model being a generative model trained to auto-regressively predict tokens to follow an initial set of tokens;

programmatically generating a usefulness measure that identifies usefulness of the original response;

in response to determining that the usefulness measure does not satisfy a prescribed test, generating feedback information;

generating revised model-input information that includes the feedback information;

presenting the revised model-input information to the language model;

receiving a revised response from the language model in response to the revised model-input information; and

generating and presenting one or more revised instances of model-input information until it is determined that the language model has generated a response that satisfies the prescribed test.

20 . The computer-readable storage medium of claim 19 , wherein the operations further include generating the feedback information using the language model, based on state information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2023
From: PENG, BAOLIN; GALLEY, MICHEL; CHENG, HAO; HE, PENGCHENG; BACH, NGUYEN HUNG; CHEN, WEIZHU; GAO, JIANFENG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063474/0320 →
Continuity (1)
Related Publication 20240362418A1 · Oct 31, 2024
References Cited (69)
US 9454733B1 · Purpura · 2016 [cited by applicant]
US 12067366B1 · Heller · 2024 [cited by examiner]
US 20140379353A1 · Boies · 2014 [cited by examiner]
US 20180004729A1 · Qiu · 2018 [cited by examiner]
US 20210217408A1 · Hakkani-Tur et al. · 2021 [cited by applicant]
US 20220147861A1 · Oltramari et al. · 2022 [cited by applicant]
US 20220189474A1 · Sharifi · 2022 [cited by applicant]
US 20220374479A1 · Xiong et al. · 2022 [cited by applicant]
US 20230112921A1 · Cai · 2023 [cited by applicant]
US 20240273345A1 · Bharadwaj · 2024 [cited by applicant]
US 20240354319A1 · Dinu · 2024 [cited by examiner]
US 20240362422A1 · Callegari · 2024 [cited by applicant]
Gupta, “Compression of Deep Learning Models for Text: A Survey,” arXiv, arXiv:2008.05221v4 [cs.CL], Jun. 13, 2021, 53 pages. [cited by applicant]
Chen, et al., “Stabilized In-Context Learning with Pre-trained Language Models for Few Shot Dialogue State Tracking,” arXiv, arXiv:2302.05932v1 [cs.CL], Feb. 12, 2023, 14 pages. [cited by applicant]
Santra, et al., “Frugal Prompting for Dialog Models,” arXiv, arXiv:2305.14919v1 [cs.CL], May 24, 2023, 22 pages. [cited by applicant]
Banerjee, et al., “METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation… [cited by applicant]
Bender, et al., “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?,” in FAccT '21: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, Mar. 2021, pp. 610-623. [cited by applicant]
Bommasani, et al., “On the Opportunities and Risks of Foundation Models,” in arXiv, Cornell University, arXiv:2108.07258v3 [cs.LG], Jul. 12, 2022, 214 pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners,” in 34th Conference on Neural Information Processing Systems (NeurIPS 2020), 2020, 25 pages. [cited by applicant]
Chowdhery, et al., “PaLM: Scaling Language Modeling with Pathways,” in arXiv, Cornell University, arXiv:2204.02311v5 [cs.CL], Oct. 5, 2022, 87 pages. [cited by applicant]
Dinan, et al., “Wizard of Wikipedia: Knowledge-Powered Conversational agents,” in arXiv, Cornell University, arXiv:1811.01241v2 [cs.CL], Feb. 21, 2019, 18 pages. [cited by applicant]
Eric, et al., “MultiWOZ 2.1: A Consolidated Multi-Domain Dialogue Dataset with State Corrections and State Tracking Baselines,” in arXiv, Cornell University, arXiv:1907.01669v4 [cs.CL], Dec. 3, 2019, 7 pages. [cited by applicant]
Galley, et al., “Grounded Response Generation Task at DSTC7,” in Proceedings of the AAAI-19 Workshop on Dialog System Technology Challenges, 2018, 5 pages. [cited by applicant]
Scao, et al., “Bloom: A 176B-Parameter Open-Access Multilingual Language Model,” arXiv, Cornell University, arXiv:2211.05100v2 [cs.CL], Dec. 11, 2022, 62. [cited by applicant]
Gao, et al., “Neural Approaches to Conversational AI,” in arXiv, Cornell University, arXiv:1809.08267v3 [cs.CL], Sep. 10, 2019, 95 pages. [cited by applicant]
Ghazvininejad, et al., “A Knowledge-Grounded Neural Conversation Model,” in arXiv, Cornell University, arXiv:1702.01932v2 [cs.CL], Nov. 15, 2018, 9 pages. [cited by applicant]
Glaese, et al., “Improving alignment of dialogue agents via targeted human judgements,” ArXiv, in arXiv, Cornell University, arXiv:2209.14375v1 [cs.LG], Sep. 28, 2022, 77 pages. [cited by applicant]
Guu, et al., “REALM: Retrieval-Augmented Language Model Pre-Training,” in arXiv, Cornell University, arXiv:2002.08909v1 [cs.CL], Feb. 10, 2020, 12 pages. [cited by applicant]
He, et al., “Rethinking with Retrieval: Faithful Large Language Model Inference,” in arXiv, Cornell University, arXiv:2301.00303v1 [cs.CL], Dec. 31, 2022, 15 pages. [cited by applicant]
Karpukhin, et al., “Dense Passage Retrieval for Open-Domain Question Answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 6769-6781. [cited by applicant]
Kim, et al., “Beyond Domain APIs: Task-oriented Conversational Modeling with Unstructured Knowledge Access,” in Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue, Jul. 2020, … [cited by applicant]
“DSTC11 Track 5—Task-oriented Conversational Modeling with Subjective Knowledge,” available at https://github.com/alexa/dstc11-track5, Github, accessed on Mar. 23, 2023, 5 pages. [cited by applicant]
Lazaridou, et al., “Internet augmented language models through few-shot prompting for open-domain question answering,” in arXiv, Cornell University, arXiv:2203.05115v2 [cs.CL], May 23, 2022, 20 pages. [cited by applicant]
Lewis, et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in arXiv, Cornell University, arXiv:2005.11401v4 [cs.CL], Apr. 12, 2021, 19 pages. [cited by applicant]
Lian, et al., “Learning to Select Knowledge for Response Generation in Dialog Systems,” in arXiv, Cornell University, arXiv:1902.04911v2 [cs.CL], May 21, 2019, 7 pages. [cited by applicant]
Lin, et al., “Rouge: A Package for Automatic Evaluation of Summaries,” in Text Summarization Branches Out, Association for Computational Linguistics, Jul. 2004, 8 pages. [cited by applicant]
Ma, et al., “Open-domain Question Answering via Chain of Reasoning over Heterogeneous Knowledge,” in arXiv, Cornell University, arXiv:2210.12338v1 [cs.CL], Oct. 22, 2022, 15 pages. [cited by applicant]
Madaan, et al., “MemPrompt: Memory-assisted Prompt Editing with User Feedback,” in arXiv, Cornell University, arXiv:2201.06009v7 [cs.CL], Feb. 18, 2023, 30 pages. [cited by applicant]
Nakano, et al., “WebGPT: Browser-assisted question answering with human feedback,” in arXiv, Cornell University, arXiv:2112.09332v3 [cs.CL], Jun. 1, 2022, 32 pages. [cited by applicant]
Ouyang, et al., “Training language models to follow instructions with human feedback,” in arXiv, Cornell University, arXiv:2203.02155v1 [cs.CL], Mar. 4, 2022, 68 pages. [cited by applicant]
Papineni, et al., “BLEU: a Method for Automatic Evaluation of Machine Translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL), Jul. 2002, 8 pages. [cited by applicant]
Peng, et al., “GODEL: Large-Scale Pre-Training for Goal-Directed Dialog,” in arXiv, Cornell University, arXiv:2206.11309v1 [cs.CL], Jun. 22, 2022, 14 pages. [cited by applicant]
Popovic, Maja, “CHRF: character n-gram F-score for automatic MT evaluation,” in Proceedings of the Tenth Workshop on Statistical Machine Translation, Sep. 2015, pp. 392-395. [cited by applicant]
Radford, et al., “Improving Language Understanding by Generative Pre-Training,” available at https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf, OpenAI, Jun. 11, 2018, 12 pages. [cited by applicant]
Schick, et al., “Toolformer: Language Models Can Teach Themselves to Use Tools,” in arXiv, Cornell University, arXiv:2302.04761v1 [cs.CL], Feb. 9. 2023, 17 pages. [cited by applicant]
Sellam, et al., “BLEURT: Learning Robust Metrics for Text Generation,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, pp. 7881-7892. [cited by applicant]
Shi, et al., “REPLUG: Retrieval-Augmented Black-Box Language Models,” in arXiv, Cornell University, arXiv:2301.12652v2 [cs.CL], Feb. 1, 2023, 12 pages. [cited by applicant]
Shuster, et al., “Retrieval Augmentation Reduces Hallucination in Conversation,” in arXiv, Cornell University, arXiv:2104.07567v1 [cs.CL], Apr. 15, 2021, 21 pages. [cited by applicant]
Shuster, et al., “BlenderBot 3: a deployed conversational agent that continually learns to responsibly engage,” in arXiv, Cornell University, arXiv:2208.03188v3 [cs.CL], Aug. 10, 2022, 38 pages. [cited by applicant]
“Mesh Transformer JAX,” available at https://github.com/kingoflolz/mesh-transformer-jax, Github, accessed on Mar. 23, 2023, 7 pages. [cited by applicant]
Weidinger, et al., “Ethical and social risks of harm from Language Models,” in arXiv, Cornell University, arXiv:2112.04359v1 [cs.CL], Dec. 8, 2021, 64 pages. [cited by applicant]
Williams, Ronald J., “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning,” in Machine Learning, May 1992, 27 pages. [cited by applicant]
Yeh, et al., “A Comprehensive Assessment of Dialog Evaluation Metrics,” in arXiv, Cornell University arXiv:2106.03706v4 [cs.CL], Jul. 7, 2021, 23 pages. [cited by applicant]
Yuan, et al., “BARTScore: Evaluating Generated Text as Text Generation,” in arXiv, Cornell University, arXiv:2106.11520v2 [cs.CL], Oct. 27, 2021, 18 pages. [cited by applicant]
Zhang, et al., “OPT: Open Pre-trained Transformer Language Model,” in arXiv, Cornell University, arXiv:2205.01068v4 [cs.CL], Jun. 21, 2022, 30 pages. [cited by applicant]
Zhang, et al., “BERTScore: Evaluating Text Generation with BERT,” in arXiv, Cornell University, arXiv:1904.09675v3 [cs.CL], Feb. 24, 2020, 43 pages. [cited by applicant]
Zhang, et al., “RetGen: A Joint Framework for Retrieval and Grounded Text Generation Modeling,” in The Thirty-Sixth AAAI Conference on Artificial Intelligence (AAAI-22), Feb. 2022, p. 11739-11747. [cited by applicant]
Zhong, et al., “Training Language Models with Memory Augmentation,” in arXiv, Cornell University, arXiv:2205.12674v3 [cs.CL], Nov. 29, 2022, 17 pages. [cited by applicant]
Peng, et al., “Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback,” in arXiv, Cornell University, arXiv:2302.12813v3 [cs.CL], Mar. 8, 2023, 15 pages. [cited by applicant]
Shinn, et al., “Reflexion: an autonomous agent with dynamic memory and self-reflection,” in arXiv, Cornell University, arXiv:2303.11366v1 [cs.AI], Mar. 20, 2023, 17 pages. [cited by applicant]
Wood, Thomas, “What is F-Score,” available at https://deepai.org/machine-learning-glossary-and-terms/f-score, DeepAI, accessed on Mar. 29, 2023, 10 pages. [cited by applicant]
Sutton, et al., “Reinforcement Learning: An Introduction,” draft version available at progresshttps://web.stanford.edu/class/psych209/Readings/SuttonBartolPRLBook2ndEd.pdf, 2nd Edition, 2015, MIT Press, 352 pages. [cited by applicant]
“Introducing The World's Largest Open Multilingual Language Model: BLOOM,” available at https://bigscience.huggingface.co/blog/bloom, accessed on Feb. 13, 2023, 2 pages. [cited by applicant]
Galley, et al., “LLM-Augmenter,” accessible at https://github.com/pengbaolin/LLM-Augmenter, Github, Feb. 24, 2023, 3 pages. [cited by applicant]
Gao, et al., “Neural approaches to conversational information retrieval,” arXiv, Cornell University, arXiv:2201.05176v1, Jan. 13, 2022, 150 pages. [cited by applicant]
Madaan, et al., “Memory-assisted prompt editing to improve GPT-3 after deployment”, arXiv, Cornell University, arXiv:2201.06009v1, Jan. 16, 2022, 14 pages. [cited by applicant]
Driess, et al., “Palm-e: An embodied multimodal language model”, arXiv:2303.03378v1, Mar. 6, 2023, 18 pages. [cited by applicant]
Final Office Action mailed on Oct. 1, 2025, in U.S. Appl. No. 18/322,524, 22 pages. [cited by applicant]
Peng, et al., “Check your facts and try again: Improving large language models with external knowledge and automated feedback”, arXiv:2302.12813v3, Mar. 8, 2023, 15 pages. [cited by applicant]