IP Library › Granted Patent US 12,579,174
Granted Patent B1
US 12,579,174 · App. 18/621,221 · Granted Mar 17, 2026

Evaluating retrieval system for language model processing

Inventors: Nicolaas Jedema (Santa Barbara, CA); Leonardo Filipe Rodrigues Ribeiro (Kirkland, WA); Alessandro Moschitti (Playa de Rey, CA); Matteo Gabburo (Bassano del Grappa, CA); Siddhant Garg (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G06F16/3329G06F16/338G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,174
App. No.
18/621,221
Filed
Mar 29, 2024
Granted
Mar 17, 2026
Kind
B1
Art Unit
2166
USPC
707/722
Abstract

Techniques for determining whether retrieved content is sufficient for a language model to generate a response to an input are described. In some embodiments, a system may determine a retrieval complexity (RC) metric based on a question, a reference answer and document results retrieved from a retrieval component (e.g., a search engine, a knowledge search, etc.). The RC metric may be based on whether one of the document results corresponds to the reference answer and whether the reference answer can be determined from the set of the document results (e.g., using two or more document results). The RC metric can be used to determine when retrieved content is to be used by a language model for responding to an input corresponding to the question.

Claims (115)

1 . A computer-implemented method comprising:

receiving first data representing a first question;

receiving second data representing a reference answer to the first question;

sending a first request including the first question to a first component configured to provide data usable by a language model to generate a response;

receiving, in response to the first request and from the first component, a first plurality of document results;

processing the first data, the second data, and the first plurality of document results using a first trained model to determine a first value representing whether at least one document result of the first plurality of document results can be used to determine an answer that corresponds to the reference answer;

determining that the first value fails to satisfy a condition;

receiving third data representing a user query;

determining, using the third data, that the user query corresponds to the first question;

based on the first value failing to satisfy the condition, sending, to a second component, a second request including the first question, the second component configured to provide data usable by the language model to generate a response;

receiving, from the second component, a second plurality of document results;

processing, using the language model, the third data and the second plurality of document results to generate a response to the user query; and

causing presentation of the response.

2 . The computer-implemented method of claim 1 , wherein the second data includes a first set of sentences and the method further comprises:

processing, using the first trained model, a first document result of the first plurality of document results and the second data to determine a second value representing that the first document result can be used to determine the reference answer, the second value being based on a first portion of the first set of sentences being semantically similar to a second set of sentences in the first document result;

processing, using the first trained model, a second document result of the first plurality of document results and the reference answer to determine a third value representing that the second document result can be used to determine the reference answer, the third value being based on a second portion of the first set of sentences being semantically similar to a third set of sentences in the second document result; and

determining the first value based on the second value and the third value.

3 . The computer-implemented method of claim 1 , further comprising:

processing, using the first trained model, the first plurality of document results and the reference answer to determine:

a second value representing that a first document result of the first plurality of document results includes a first set of tokens semantically similar to a first portion of the reference answer, and

a third value representing that a second document result of the first plurality of document results includes a second set of tokens semantically similar to a second portion of the reference answer; and

determining the first value based on the second value and third value.

4 . The computer-implemented method of claim 1 , further comprising:

receiving fourth data including a second question and a corresponding second answer;

receiving, from the first component, a third plurality of document results based on a third request including the second question;

processing, using the first trained model, the fourth data and the third plurality of document results to determine a second value representing the third plurality of document results can be used to determine the second answer;

determining that the second value satisfies the condition;

based on the second value satisfying the condition, storing, in a data storage, fifth data including the third plurality of document results, the fifth data associated with the second question;

receiving a second user query corresponding to the second question;

determining, from the data storage, the fifth data;

determining a prompt including the second user query and the fifth data;

processing, using the language model, the prompt to generate a second response to the second user query; and

causing presentation of the second response.

5 . A computer-implemented method comprising:

receiving first question data representing a first question;

receiving first answer data representing a first answer corresponding to the first question;

receiving a first plurality of document results based on a first request to a first component, the first request including the first question data;

processing, using a trained model, the first plurality of document results, the first question data and the first answer data to determine a first value representing whether at least one of the first plurality of document results can be used to determine a first response that corresponds to the first answer;

determining that the first value fails to satisfy a first condition; and

based on the first value failing to satisfy the first condition, causing a language model to use a second plurality of document results to generate the first response, wherein the second plurality of document results are determined using a second component.

6 . The computer-implemented method of claim 5 , further comprising:

determining, using the trained model, a second value representing that a first document result of the first plurality of document results corresponds to the first answer;

and

determining the first value based on the second value.

7 . The computer-implemented method of claim 5 , further comprising:

determining, using the trained model, a second value representing that a first document result of the first plurality of document results corresponds to a first portion of the first answer;

determining, using the trained model, a third value representing that a second document result of the first plurality of document results corresponds to a second portion of the first answer, the second portion being different than the first portion; and

determining the first value based on the second value and the third value.

8 . The computer-implemented method of claim 5 , further comprising:

processing, using the trained model, the second plurality of document results, the first question data and the first answer data to determine a second value representing that at least one of the second plurality of document results can be used to determine the first response corresponding to the first answer;

after determining that the second value satisfies the first condition, storing data representing an association between the second plurality of document results and the first question data;

receiving a user query corresponding to the first question;

based on the user query corresponding to the first question, retrieving the second plurality of document results;

determining a prompt based on the user query and including the second plurality of document results; and

processing, using the language model, the prompt to determine the first response to the user query.

9 . The computer-implemented method of claim 5 , further comprising:

receiving second question data representing a second question;

receiving second answer data representing a second answer corresponding to the second question;

receiving a third plurality of document results based on a second request to the first component, the second request including the second question data;

processing, using the trained model, the third plurality of document results, the second question data and the second answer data to determine a second value representing whether at least one of the second plurality of document results can be used to determine a second response that corresponds to the second answer;

determining that the second value fails to satisfy the first condition; and

based on the second value failing to satisfy the first condition, causing the language model to generate a third response to an input corresponding to the second question, the third response indicating that the language model is unable to respond to the input.

10 . The computer-implemented method of claim 5 , further comprising:

receiving second answer data representing a second answer generated by the language model to the first question;

processing, using the trained model, the first plurality of document results, the first question data and the second answer data to determine a second value representing that at least one of the first plurality of document results can be used to determine the second answer;

determining the second value fails to satisfy the first condition; and

based on the second value failing to satisfy the first condition, determining that the language model is to be further trained with respect to the first question.

11 . The computer-implemented method of claim 5 , further comprising:

storing data representing an association between the first question data and the second component, the association indicating that the second component to be used to retrieve content for an input similar to the first question data;

receiving user input data that is semantically similar to the first question data;

determining, based on the stored data, that the second component is to be used;

determining a prompt based on the user input data and including an identifier for the second component; and

processing, using the language model, the prompt to generate a request to the second component.

12 . A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

receive first question data representing a first question;

receive first answer data representing a first answer corresponding to the first question;

receive a first plurality of document results based on a first request to a first component, the first request including the first question data;

process, using a trained model, the first plurality of document results, the first question data and the first answer data to determine a first value representing that at least one of the first plurality of document results can be used to determine a first response that corresponds to the first answer;

determine that the first value fails to satisfy a first condition; and

based on the first value failing to satisfy the first condition, cause a language model to use a second plurality of document results to generate the first response, wherein the second plurality of document results are determined using a second component.

13 . The system of claim 12 , wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to:

determine, using the trained model, a second value representing that a first document result of the first plurality of document results corresponds to the first answer;

and

determine the first value based on the second value.

14 . The system of claim 12 , wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to:

determine, using the trained model, a second value representing that a first document result of the first plurality of document results corresponds to a first portion of the first answer;

determine, using the trained model, a third value representing that a second document result of the first plurality of document results corresponds to a second portion of the first answer, the second portion being different than the first portion; and

determine the first value based on the second value and the third value.

15 . The system of claim 12 , wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to:

process, using the trained model, the second plurality of document results, the first question data and the first answer data to determine a second value representing that at least on of the second plurality of document results can be used to determine the first response corresponding to the first answer;

after determining that the second value satisfies the first condition, store data representing an association between the second plurality of document results and the first question data;

receive a user query corresponding to the first question;

based on the user query corresponding to the first question, retrieve the second plurality of document results;

determine a prompt based on the user query and including the second plurality of document results; and

process, using the language model, the prompt to determine the first response to the user query.

16 . The system of claim 12 , wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to:

receive second question data representing a second question;

receive second answer data representing a second answer corresponding to the second question;

receive a third plurality of document results based on a second request to the first component, the second request including the second question data;

process, using the trained model, the third plurality of document results, the second question data and the second answer data to determine a second value representing whether at least one of the second plurality of document results can be used to determine a second response that corresponds to the second answer;

determine that the second value fails to satisfy the first condition; and

based on the second value failing to satisfy the first condition, cause the language model to generate a third response to an input corresponding to the second question, the third response indicating that the language model is unable to respond to the input.

17 . The system of claim 12 , wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to:

receive second answer data representing a second answer generated by the language model to the first question;

process, using the trained model, the first plurality of document results, the first question data and the second answer data to determine a second value representing that at least one of the first plurality of document results can be used to determine the second answer;

determine the second value fails to satisfy the first condition; and

based on the second value failing to satisfy the first condition, determine that the language model is to be further trained with respect to the first question.

18 . The system of claim 12 , wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to:

store data representing an association between the first question data and the second component, the association indicating that the second component to be used to retrieve content for an input similar to the first question data;

receive user input data that is semantically similar to the first question data;

determine, based on the stored data, that the first component is to be used;

determine a prompt based on the user input data and including an identifier for the second component; and

process, using the language model, the prompt to generate a request to the second component.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2024
From: JEDEMA, NICOLAAS; RODRIGUES RIBEIRO, LEONARDO FILIPE; MOSCHITTI, ALESSANDRO; GABBURO, MATTEO; GARG, SIDDHANT
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 068378/0928 →
References Cited (96)
US 10331402B1 · Spector et al. · 2019 [cited by applicant]
US 10642577B2 · Spector et al. · 2020 [cited by applicant]
US 10699215B2 · Devarakonda · 2020 [cited by examiner]
US 11080336B2 · Van Dusen · 2021 [cited by examiner]
US 11314819B2 · Katzman · 2022 [cited by examiner]
US 11449556B2 · Jawagal · 2022 [cited by examiner]
US 11740863B2 · Spector et al. · 2023 [cited by applicant]
US 12008473B2 · Lazaridou · 2024 [cited by examiner]
US 12405985B1 · Kanagovi · 2025 [cited by examiner]
US 20170293638A1 · He · 2017 [cited by examiner]
US 20180081906A1 · Katz · 2018 [cited by examiner]
US 20180082187A1 · Katz · 2018 [cited by examiner]
US 20180165580A1 · Boyer · 2018 [cited by examiner]
US 20210390418A1 · Mass · 2021 [cited by examiner]
US 20220138432A1 · Galitsky · 2022 [cited by examiner]
US 20220222440A1 · Chowdhury · 2022 [cited by examiner]
US 20220230061A1 · Singh · 2022 [cited by examiner]
US 20220335046A1 · Oshio · 2022 [cited by examiner]
US 20220335231A1 · Oshio · 2022 [cited by examiner]
US 20220343903A1 · Mostafazadeh · 2022 [cited by examiner]
US 20240069860A1 · Spector et al. · 2024 [cited by applicant]
US 20240070434A1 · Garg · 2024 [cited by examiner]
US 20240176980A1 · Lee · 2024 [cited by examiner]
US 20240281487A1 · Bathwal · 2024 [cited by examiner]
US 20240362286A1 · He · 2024 [cited by examiner]
US 20240370479A1 · Hudetz · 2024 [cited by examiner]
US 20250086392A1 · Han · 2025 [cited by examiner]
US 20250291827A1 · Chong · 2025 [cited by examiner]
US 20250299059A1 · Ahmed · 2025 [cited by examiner]
US 20250307658A1 · Wang · 2025 [cited by examiner]
Asai, Akari, et al. “Self-RAG: Learning to retrieve, generate, and critique through self-reflection.” arXiv preprint arXiv:2310.11511 (2023). [cited by applicant]
Baumgärtner, Tim, et al. “Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-Ranking.” arXiv preprint arXiv:2210.10695 (2022). [cited by applicant]
Baumgärtner, Tim, et al. “UKP-SQUARE: An Online Platform for Question Answering Research.” arXiv preprint arXiv:2203.13693 (2022). [cited by applicant]
Borgeaud, Sebastian, et al. “Improving language models by retrieving from trillions of tokens.” International conference on machine learning. PMLR, 2022. [cited by applicant]
Bulian, Jannis, et al. “Tomayto, tomahto. beyond token-level answer equivalence for question answering evaluation.” arXiv preprint arXiv:2202.07654 (2022). [cited by applicant]
Chen, Wenhu, Xinyi Wang, and William Yang Wang. “A dataset for answering time-sensitive questions.” arXiv preprint arXiv:2108.06314 (2021). [cited by applicant]
Clark, Kevin, et al. “Electra: Pre-training text encoders as discriminators rather than generators.” arXiv preprint arXiv:2003.10555 (2020). [cited by applicant]
Crestani, Fabio, et al. ““Is this document relevant?. . . probably” a survey of probabilistic models in information retrieval.” ACM Computing Surveys (CSUR) 30.4 (1998): 528-552. [cited by applicant]
Dao, Xuan-Quy. “Performance comparison of large language models on vnhsge english dataset: Openai chatgpt, microsoft bing chat, and google bard.” arXiv preprint arXiv:2307.02288 (2023). [cited by applicant]
Deutsch, Daniel, and Dan Roth. “Benchmarking Answer Verification Methods for Question Answering-Based Summarization Evaluation Metrics.” arXiv preprint arXiv:2204.10206 (2022). [cited by applicant]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. “BERT: Pre-training of deep bidirectional transformers for language understanding.” Proceedings of NAACL-HLT. vol. 1. 2019. [cited by applicant]
Dua, Dheeru, et al. “DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs.” arXiv preprint arXiv:1903.00161 (2019). [cited by applicant]
Fan, Yixing, et al. “Modeling diverse relevance patterns in ad-hoc retrieval.” The 41st international ACM SIGIR conference on research & development in information retrieval. 2018. [cited by applicant]
Gabburo, Matteo, et al. “SQUARE: Automatic question answering evaluation using multiple positive and negative references.” arXiv preprint arXiv:2309.12250 (2023). [cited by applicant]
Gabburo, Matteo, et al. “Knowledge transfer from answer ranking to answer generation.” arXiv preprint arXiv:2210.12865 (2022). [cited by applicant]
Geva, Mor, et al. “Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies.” Transactions of the Association for Computational Linguistics 9 (2021): 346-361. [cited by applicant]
Hsu, Chao-Chun, et al. “Answer generation for retrieval-based question answering systems.” arXiv preprint arXiv:2106.00955 (2021). [cited by applicant]
Huang, Hao, et al. “Understand before answer: Improve temporal reading comprehension via precise question understanding.” Proceedings of the 2022 Conference of the North American Chapter of the Association for Computati… [cited by applicant]
Huang, Lei, et al. “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.” arXiv preprint arXiv:2311.05232 (2023). [cited by applicant]
Jiang, Albert Q., et al. “Mistral 7B.” arXiv preprint arXiv:2310.06825 (2023). [cited by applicant]
Karpukhin, Vladimir, et al. “Dense passage retrieval for open-domain question answering.” arXiv preprint arXiv:2004.04906 (2020). [cited by applicant]
Khattab, Omar, and Matei Zaharia. “Colbert: Efficient and effective passage search via contextualized late interaction over bert.” Proceedings of the 43rd International ACM SIGIR conference on research and development i… [cited by applicant]
Kwiatkowski, Tom, et al. “Natural questions: a benchmark for question answering research.” Transactions of the Association for Computational Linguistics 7 (2019): 453-466. [cited by applicant]
Lewis, Mike, et al. “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.” arXiv preprint arXiv:1910.13461 (2019). [cited by applicant]
Lewis, Patrick, et al. “Retrieval-augmented generation for knowledge-intensive nlp tasks.” Advances in Neural Information Processing Systems 33 (2020): 9459-9474. [cited by applicant]
Lin, Jimmy, and Xueguang Ma. “A few brief notes on deepimpact, coil, and a conceptual framework for information retrieval techniques.” arXiv preprint arXiv:2106.14807 (2021). [cited by applicant]
Liu, Yinhan, et al. “RoBERTa: A Robustly Optimized BERT Pretraining Approach.” arXiv preprint arXiv:1907.11692 (2019). [cited by applicant]
Luo, Haoran, et al. “Chatkbqa: A generate-then-retrieve framework for knowledge base question answering with fine-tuned large language models.” arXiv preprint arXiv:2310.08975 (2023). [cited by applicant]
Mallia, Antonio, et al. “Learning passage impacts for inverted indexes.” Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2021. [cited by applicant]
Mavi, Vaibhav, Anubhav Jangra, and Adam Jatowt. “A survey on multi-hop question answering and generation.” arXiv preprint arXiv:2204.09140 (2022). [cited by applicant]
Min, Sewon, et al. “Compositional questions do not necessitate multi-hop reasoning.” arXiv preprint arXiv:1906.02900 (2019). [cited by applicant]
Nguyen, Tri, et al. “Ms marco: A human-generated machine reading comprehension dataset.” (2016). [cited by applicant]
Papineni, Kishore, et al. “Bleu: a method for automatic evaluation of machine translation.” Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 2002. [cited by applicant]
Perez, Ethan, et al. “Unsupervised question decomposition for question answering.” arXiv preprint arXiv:2002.09758 (2020). [cited by applicant]
Petroni, Fabio, et al. “KILT: a benchmark for knowledge intensive language tasks.” arXiv preprint arXiv:2009.02252 (2020). [cited by applicant]
Radford, Alec, et al. “Improving language understanding by generative pre-training.” (2018). [cited by applicant]
Colin, Raffel. “Exploring the limits of transfer learning with a unified text-to-text transformer.” JMLR 21.140 (2020): 1. [cited by applicant]
Rajpurkar, Pranav, et al. “Squad: 100,000+ questions for machine comprehension of text.” arXiv preprint arXiv:1606.05250 (2016). [cited by applicant]
Rogers, Anna, Matt Gardner, and Isabelle Augenstein. “Qa dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension.” ACM Computing Surveys 55.10 (2023): 1-45. [cited by applicant]
Saxena, Apoorv, Soumen Chakrabarti, and Partha Talukdar. “Question answering over temporal knowledge graphs.” arXiv preprint arXiv:2106.01515 (2021). [cited by applicant]
Shang, Chao, et al. “Improving time sensitivity for question answering over temporal knowledge graphs.” arXiv preprint arXiv:2203.00255 (2022). [cited by applicant]
Sharma, Aditya, et al. “TwiRGCN: Temporally weighted graph convolution for question answering over temporal knowledge graphs.” arXiv preprint arXiv:2210.06281 (2022). [cited by applicant]
Su, Ming-Hsiang, et al. “RoBERTa-based traditional chinese medicine named entity recognition model.” Proceedings of the 34th Conference on Computational Linguistics and Speech Processing (ROCLING 2022). 2022. [cited by applicant]
Sun, Tianxiang, et al. “BERTScore is unfair: On social bias in language model-based metrics for text generation.” arXiv preprint arXiv:2210.07626 (2022). [cited by applicant]
Talmor, Alon, and Jonathan Berant. “The web as a knowledge-base for answering complex questions.” arXiv preprint arXiv:1803.06643 (2018). [cited by applicant]
Trivedi, Harsh, et al. “ MuSiQue: Multihop Questions via Single-hop Question Composition.” Transactions of the Association for Computational Linguistics 10 (2022): 539-554. [cited by applicant]
Vaswani, Ashish, et al. “Attention is all you need.” Advances in neural information processing systems 30 (2017). [cited by applicant]
Vu, Thuy, and Alessandro Moschitti. “AVA: an automatic evaluation approach to question answering systems.” arXiv preprint arXiv:2005.00705 (2020). [cited by applicant]
Vu, Tu, et al. “Freshllms: Refreshing large language models with search engine augmentation.” arXiv preprint arXiv:2310.03214 (2023). [cited by applicant]
Wang, Zizhen, et al. “Match [cited by applicant]
Wei, Jason, et al. “Chain-of-thought prompting elicits reasoning in large language models.” Advances in neural information processing systems 35 (2022): 24824-24837. [cited by applicant]
Yan, Yiming, et al. “BLEURT has universal translations: An analysis of automatic metrics by minimum risk training.” arXiv preprint arXiv:2307.03131 (2023). [cited by applicant]
Yang, Zhilin, et al. “HotpotQA: A dataset for diverse, explainable multi-hop question answering.” arXiv preprint arXiv:1809.09600 (2018). [cited by applicant]
Yoran, Ori, et al. “Answering questions by meta-reasoning over multiple chains of thought.” arXiv preprint arXiv:2304.13007 (2023). [cited by applicant]
Khashabi, Daniel, et al. “GooAQ: Open question answering with diverse answer types.” arXiv preprint arXiv:2104.08727 (2021). [cited by applicant]
Khot, Tushar, et al. “Qasc: A dataset for question answering via sentence composition.” Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34. No. 05. 2020. [cited by applicant]
Abujabal, Abdalghani, et al. “Comqa: A community-sourced dataset for complex factoid question answering with paraphrase clusters.” arXiv preprint arXiv:1809.09528 (2018). [cited by applicant]
BehnamGhader, Parishad, Santiago Miret, and Siva Reddy. “Can retriever-augmented language models reason? the blame game between the retriever and the language model.” arXiv preprint arXiv:2212.09146 (2022). [cited by applicant]
Patel, Pruthvi, et al. “Is a question decomposition unit all we need?.” arXiv preprint arXiv:2205.12538 (2022). [cited by applicant]
Dua, Dheeru, et al. “Successive prompting for decomposing complex questions.” arXiv preprint arXiv:2212.04092 (2022). [cited by applicant]
Min, Sewon, et al. “Multi-hop reading comprehension through question decomposition and rescoring.” arXiv preprint arXiv:1906.02916 (2019). [cited by applicant]
Radhakrishnan, Ansh, et al. “Question decomposition improves the faithfulness of model-generated reasoning.” arXiv preprint arXiv:2307.11768 (2023). [cited by applicant]
Daull, Xavier, et al. “Complex QA and language models hybrid architectures, Survey.” arXiv preprint arXiv:2302.09051 (2023). [cited by applicant]
Borji, Ali. “A categorical archive of chatgpt failures.” arXiv preprint arXiv:2302.03494 (2023). [cited by applicant]
Kazemnejad, Amirhossein, et al. “Measuring the knowledge acquisition-utilization gap in pretrained language models.” arXiv preprint arXiv:2305.14775 (2023). [cited by applicant]
Lauriola, Ivano, Kevin Small, and Alessandro Moschitti. “Building a Dataset for Automatically Learning to Detect Questions Requiring Clarification.” Proceedings of the Thirteenth Language Resources and Evaluation Confer… [cited by applicant]