IP Library › Granted Patent US 12,748,793
Granted Patent B2
US 12,748,793 · App. 19/003,593 · Granted Sep 29, 2026

Systems and methods for providing improved retrieval augmented generation (RAG) for query response

Inventors: Diana Meditz (New York, NY); Yikai Feng (New York, NY)
Assignee: The Bank of New York Mellon
G06F16/3347G06F16/33295
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,793
App. No.
19/003,593
Granted
Sep 29, 2026
Kind
B2
Abstract

Systems and methods for providing improved retrieval augmented generation for query response, retrieve key-value pairs from a key-value pairs database; index the key-value pairs into vector stores, wherein a first vector store contains the keys from the key-value pairs, a second vector store contains the values from the key-value pairs, and a third vector store contains value chunks of the values; process a query through the first and third vector stores to generate a list of keys and a list of value chunks from their respective stores; retrieve, from the second vector store, corresponding values to the keys in the list of keys from the first vector store; compose, an augmented prompt based on an aggregation of the query, the list of keys, the corresponding values, and at least a portion of the value chunks; and generate a proposed response by feeding the augmented prompt into a large language model.

Claims (46)

1 . A method for providing improved retrieval augmented generation (RAG) for query response, comprising:

retrieving, by a processor, a plurality of key-value pairs from a key-value pairs database;

indexing, by the processor, the plurality of key-value pairs into a plurality of vector stores, wherein a first vector store contains the keys from the plurality of key-value pairs, a second vector store contains the values from the plurality of key-value pairs, and a third vector store contains a plurality of value chunks of the values;

processing, by the processor, a query through the first vector store and the third vector store to generate a list of keys from the first vector store and a list of value chunks from the third vector store;

retrieving, from the second vector store, corresponding values to the keys in the list of keys from the first vector store;

composing, by the processor, an augmented prompt based on an aggregation of the query, the list of keys, the corresponding values, and at least a portion of the value chunks from the list of value chunks;

generating, by the processor, a proposed response by feeding the augmented prompt into a large language model (LLM);

retrieving, by the processor, from the second vector store, a revised list of values based on the proposed response;

comparing, by the processor, a highest ranked value from the revised list of values with the proposed response; and

based on a similarity threshold, outputting, by the processor, one of the highest ranked value from the revised list of values or the proposed response as a final response.

2 . The method as in claim 1 , further comprising altering the revised list of values based on one or more additional criteria for reranking values in the revised list of values.

3 . The method as in claim 1 , wherein the key-value pairs are stored in an external data store.

4 . The method as in claim 1 , wherein the processor is configured to retrieve as least one of keys, values, or value chunks based on respective cosine similarity scores.

5 . The method as in claim 1 , wherein the value chunks are generated by breaking down the values from the plurality of key-value pairs into shorter chunks of data.

6 . The method as in claim 1 , wherein the plurality of key-value pairs comprises a plurality of question-and-answer pairs; and wherein the key-value pairs are stored in a question-and-answer pairs store.

7 . A method for providing improved retrieval augmented generation (RAG) for query response, comprising:

retrieving, by a processor, a plurality of question-and-answer pairs from a question-and-answer pairs database;

indexing, by the processor, the plurality of question-and-answer pairs into a plurality of vector databases, wherein a first vector database contains the questions from the plurality of question-and-answer pairs, a second vector database contains the answers from the plurality of question-and-answer pairs, and a third vector database contains a plurality of answer chunks of the answers;

processing, by the processor, a query through the first vector database and the third vector database to generate a list of questions from the first vector database and a list of answer chunks from the third vector database;

retrieving, from the second vector database, corresponding answers to the questions in the list of questions from the first vector database;

composing, by the processor, an augmented prompt based on an aggregation of the query, the list of questions, the corresponding answers, and at least a portion of the answer chunks from the list of answer chunks;

generating, by the processor, a proposed response by feeding the augmented prompt into a large language model (LLM);

retrieving, by the processor, from the second database, a revised list of answers based on the proposed response;

comparing, by the processor, a highest ranked answer from the revised list of answers with the proposed response; and

based on a similarity threshold, outputting, by the processor, one of the highest ranked answer from the revised list of answers or the proposed response as a final response.

8 . The method as in claim 7 , further comprising altering the revised list of answers based on one or more additional criteria for reranking answers in the revised list of answers.

9 . The method as in claim 7 , wherein the question-and-answer pairs database is an external database.

10 . The method as in claim 7 , wherein the processor is configured to retrieve as least one of questions, answers, or answer chunks based on respective cosine similarity scores.

11 . The method as in claim 7 , wherein the answer chunks are generated by breaking down the answers from the plurality of question-and-answer pairs into shorter chunks of text.

12 . A system for providing improved retrieval augmented generation (RAG) for query response, comprising:

memory storing computer program instructions; and

one or more processors configured to execute the computer program instructions to:

retrieve a plurality of key-value pairs from a key-value pairs database;

index the plurality of key-value pairs into a plurality of vector stores, wherein a first vector store contains the keys from the plurality of key-value pairs, a second vector store contains the values from the plurality of key-value pairs, and a third vector store contains a plurality of value chunks of the values;

process a query through the first vector store and the third vector store to generate a list of keys from the first vector store and a list of value chunks from the third vector store;

retrieve, from the second vector store, corresponding values to the keys in the list of keys from the first vector store;

compose, an augmented prompt based on an aggregation of the query, the list of keys, the corresponding values, and at least a portion of the value chunks from the list of value chunks;

generate a proposed response by feeding the augmented prompt into a large language model (LLM);

retrieve, from the second vector store, a revised list of values based on the proposed response;

compare a highest ranked value from the revised list of values with the proposed response; and

based on a similarity threshold, output one of the highest ranked value from the revised list of values or the proposed response as a final response.

13 . The system as in claim 12 , further configured to alter the revised list of values based on one or more additional criteria for reranking values in the revised list of values.

14 . The system as in claim 12 , wherein the key-value pairs are stored in an external data store.

15 . The system as in claim 12 , wherein the processor is configured to retrieve as least one of keys, values, or value chunks based on respective cosine similarity scores.

16 . The system as in claim 12 , wherein the value chunks are generated by breaking down the values from the plurality of key-value pairs into shorter chunks of data.

17 . The system as in claim 12 , wherein the plurality of key-value pairs comprises a plurality of question-and-answer pairs; and wherein the key-value pairs are stored in a question-and-answer pairs store.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2024
From: MEDITZ, DIANA; FENG, YIKAI
To: THE BANK OF NEW YORK MELLON
Reel/Frame 069690/0820 →
Continuity (1)
Related Publication 20260187122A1 · Jul 2, 2026
References Cited (36)
US 9002773B2 · Bagchi et al. · 2015 [cited by applicant]
US 9348900B2 · Alkov et al. · 2016 [cited by applicant]
US 9697477B2 · Oh et al. · 2017 [cited by applicant]
US 9720981B1 · Boguraev et al. · 2017 [cited by applicant]
US 10255273B2 · Chakraborty et al. · 2019 [cited by applicant]
US 10706090B2 · Sun et al. · 2020 [cited by applicant]
US 11016966B2 · Kim et al. · 2021 [cited by applicant]
US 11106664B2 · Mcelvain et al. · 2021 [cited by applicant]
US 11157564B2 · Prakash et al. · 2021 [cited by applicant]
US 11748577B1 · Aberle · 2023 [cited by applicant]
US 11769017B1 · Gray et al. · 2023 [cited by applicant]
US 11861320B1 · Gajek et al. · 2024 [cited by applicant]
US 11966704B1 · Magary et al. · 2024 [cited by applicant]
US 11971914B1 · Watson et al. · 2024 [cited by applicant]
US 11990123B1 · Rosser · 2024 [cited by applicant]
US 20200034722A1 · Oh et al. · 2020 [cited by applicant]
US 20220237213A1 · Karlberg et al. · 2022 [cited by applicant]
US 20240012842A1 · Kislal et al. · 2024 [cited by applicant]
US 20240020538A1 · Socher et al. · 2024 [cited by applicant]
US 20240038226A1 · Nouri et al. · 2024 [cited by applicant]
US 20240095460A1 · Xu et al. · 2024 [cited by applicant]
US 20240119075A1 · Mallick et al. · 2024 [cited by applicant]
US 20240126794A1 · Cook · 2024 [cited by applicant]
US 20240160902A1 · Padgett et al. · 2024 [cited by applicant]
US 20240346256A1 · Qin · 2024 [cited by applicant]
US 20240346566A1 · Qin · 2024 [cited by applicant]
US 20250086215A1 · Kumar · 2025 [cited by examiner]
US 20250238470A1 · Penta · 2025 [cited by examiner]
US 20250272506A1 · Sassak, Jr. · 2025 [cited by examiner]
CN 117112755A · 2023 [cited by applicant]
CN 117216222A · 2023 [cited by applicant]
CN 117371457A · 2024 [cited by applicant]
CN 117648417A · 2024 [cited by applicant]
CN 117787417A · 2024 [cited by applicant]
International Search Report & Written Opinion of the International Searching Authority dated Mar. 10, 2026, issued in corresponding International Application No. PCT/US2025/054339 (11 pgs.). [cited by applicant]
Zhi Jing et al., “When Large Language Models Meet Vector Databases: A Survey”, arXiv:2402.01763v2, Feb. 2024 URL: https://arxiv.org/html/2402.01763v2 9 pgs. [cited by applicant]