Systems and methods for providing improved retrieval augmented generation (RAG) for query response
Systems and methods for providing improved retrieval augmented generation for query response, retrieve key-value pairs from a key-value pairs database; index the key-value pairs into vector stores, wherein a first vector store contains the keys from the key-value pairs, a second vector store contains the values from the key-value pairs, and a third vector store contains value chunks of the values; process a query through the first and third vector stores to generate a list of keys and a list of value chunks from their respective stores; retrieve, from the second vector store, corresponding values to the keys in the list of keys from the first vector store; compose, an augmented prompt based on an aggregation of the query, the list of keys, the corresponding values, and at least a portion of the value chunks; and generate a proposed response by feeding the augmented prompt into a large language model.
1 . A method for providing improved retrieval augmented generation (RAG) for query response, comprising:
retrieving, by a processor, a plurality of key-value pairs from a key-value pairs database;
indexing, by the processor, the plurality of key-value pairs into a plurality of vector stores, wherein a first vector store contains the keys from the plurality of key-value pairs, a second vector store contains the values from the plurality of key-value pairs, and a third vector store contains a plurality of value chunks of the values;
processing, by the processor, a query through the first vector store and the third vector store to generate a list of keys from the first vector store and a list of value chunks from the third vector store;
retrieving, from the second vector store, corresponding values to the keys in the list of keys from the first vector store;
composing, by the processor, an augmented prompt based on an aggregation of the query, the list of keys, the corresponding values, and at least a portion of the value chunks from the list of value chunks;
generating, by the processor, a proposed response by feeding the augmented prompt into a large language model (LLM);
retrieving, by the processor, from the second vector store, a revised list of values based on the proposed response;
comparing, by the processor, a highest ranked value from the revised list of values with the proposed response; and
based on a similarity threshold, outputting, by the processor, one of the highest ranked value from the revised list of values or the proposed response as a final response.
2 . The method as in claim 1 , further comprising altering the revised list of values based on one or more additional criteria for reranking values in the revised list of values.
3 . The method as in claim 1 , wherein the key-value pairs are stored in an external data store.
4 . The method as in claim 1 , wherein the processor is configured to retrieve as least one of keys, values, or value chunks based on respective cosine similarity scores.
5 . The method as in claim 1 , wherein the value chunks are generated by breaking down the values from the plurality of key-value pairs into shorter chunks of data.
6 . The method as in claim 1 , wherein the plurality of key-value pairs comprises a plurality of question-and-answer pairs; and wherein the key-value pairs are stored in a question-and-answer pairs store.
7 . A method for providing improved retrieval augmented generation (RAG) for query response, comprising:
retrieving, by a processor, a plurality of question-and-answer pairs from a question-and-answer pairs database;
indexing, by the processor, the plurality of question-and-answer pairs into a plurality of vector databases, wherein a first vector database contains the questions from the plurality of question-and-answer pairs, a second vector database contains the answers from the plurality of question-and-answer pairs, and a third vector database contains a plurality of answer chunks of the answers;
processing, by the processor, a query through the first vector database and the third vector database to generate a list of questions from the first vector database and a list of answer chunks from the third vector database;
retrieving, from the second vector database, corresponding answers to the questions in the list of questions from the first vector database;
composing, by the processor, an augmented prompt based on an aggregation of the query, the list of questions, the corresponding answers, and at least a portion of the answer chunks from the list of answer chunks;
generating, by the processor, a proposed response by feeding the augmented prompt into a large language model (LLM);
retrieving, by the processor, from the second database, a revised list of answers based on the proposed response;
comparing, by the processor, a highest ranked answer from the revised list of answers with the proposed response; and
based on a similarity threshold, outputting, by the processor, one of the highest ranked answer from the revised list of answers or the proposed response as a final response.
8 . The method as in claim 7 , further comprising altering the revised list of answers based on one or more additional criteria for reranking answers in the revised list of answers.
9 . The method as in claim 7 , wherein the question-and-answer pairs database is an external database.
10 . The method as in claim 7 , wherein the processor is configured to retrieve as least one of questions, answers, or answer chunks based on respective cosine similarity scores.
11 . The method as in claim 7 , wherein the answer chunks are generated by breaking down the answers from the plurality of question-and-answer pairs into shorter chunks of text.
12 . A system for providing improved retrieval augmented generation (RAG) for query response, comprising:
memory storing computer program instructions; and
one or more processors configured to execute the computer program instructions to:
retrieve a plurality of key-value pairs from a key-value pairs database;
index the plurality of key-value pairs into a plurality of vector stores, wherein a first vector store contains the keys from the plurality of key-value pairs, a second vector store contains the values from the plurality of key-value pairs, and a third vector store contains a plurality of value chunks of the values;
process a query through the first vector store and the third vector store to generate a list of keys from the first vector store and a list of value chunks from the third vector store;
retrieve, from the second vector store, corresponding values to the keys in the list of keys from the first vector store;
compose, an augmented prompt based on an aggregation of the query, the list of keys, the corresponding values, and at least a portion of the value chunks from the list of value chunks;
generate a proposed response by feeding the augmented prompt into a large language model (LLM);
retrieve, from the second vector store, a revised list of values based on the proposed response;
compare a highest ranked value from the revised list of values with the proposed response; and
based on a similarity threshold, output one of the highest ranked value from the revised list of values or the proposed response as a final response.
13 . The system as in claim 12 , further configured to alter the revised list of values based on one or more additional criteria for reranking values in the revised list of values.
14 . The system as in claim 12 , wherein the key-value pairs are stored in an external data store.
15 . The system as in claim 12 , wherein the processor is configured to retrieve as least one of keys, values, or value chunks based on respective cosine similarity scores.
16 . The system as in claim 12 , wherein the value chunks are generated by breaking down the values from the plurality of key-value pairs into shorter chunks of data.
17 . The system as in claim 12 , wherein the plurality of key-value pairs comprises a plurality of question-and-answer pairs; and wherein the key-value pairs are stored in a question-and-answer pairs store.