IP Library › Granted Patent US 12,405,979
Granted Patent B2
US 12,405,979 · App. 19/056,496 · Granted Sep 2, 2025

Method and system for optimizing use of retrieval augmented generation pipelines in generative artificial intelligence applications

Inventors: Vijay Madisetti (Alpharetta, GA); Arshdeep Bahga (Chandigarh, IN)
Assignee: Vijay Madisetti
G06F16/3329G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,405,979
App. No.
19/056,496
Granted
Sep 2, 2025
Kind
B2
Abstract

A system and method of improving performance of LLMs including receiving a query, retrieving relevant documents for the query to form a combined context, partitioning the combined context in context partitions, generating intermediate results from the context partitions using a mapper prompt, and generating a final result from the intermediate results using a reducer prompt. The final result is transmitted to a user.

Claims (152)

1. A method of processing large contexts in a retrieval-augmented generation (RAG) system comprising:

receiving a query from a user;

retrieving a plurality of relevant documents from at least one database based on the query, wherein the relevant documents form a combined context;

partitioning the combined context into a plurality of context partitions;

generating a plurality of intermediate analysis results by:

processing each context partition of the plurality of context partitions using a mapper prompt; and

sending an output of the mapper prompt from each context partition to one or more large language models (LLMs);

generating a final response by processing the plurality of intermediate analysis results using a reducer prompt sent to one or more LLMs; and

transmitting the final response to the user.

2. The method of claim 1 wherein retrieving the plurality of relevant documents comprises:

identifying a first set of documents by performing a keyword search in a relational database;

identifying a second set of documents by performing a vector search in a vector database; and

combining the first and second sets of documents to form the plurality of relevant documents.

3. The method of claim 1 wherein processing each context partition comprises:

processing the plurality of context partitions in parallel using a plurality of LLM instances; and

generating the plurality of intermediate analysis results simultaneously.

4. The method of claim 1 wherein combining the plurality of intermediate analysis results comprises:

identifying one or more common themes across the plurality of intermediate analysis results;

resolving any conflicts between the plurality of intermediate analysis results;

synthesizing the plurality of intermediate analysis results into a coherent response; and

validating the coherent response for at least one of accuracy and completeness.

5. The method of claim 4 further comprising processing the coherent response for at least one of:

compliance with one or more system policies;

protection from unauthorized access;

protection from privacy breaches; or

protection from potential jailbreaking attempts.

6. The method of claim 1 further comprising:

storing the plurality of intermediate analysis results in a cache;

receiving a subsequent query related to the original query;

retrieving the stored plurality of intermediate analysis results from the cache; and

generating a new final response using the stored plurality of intermediate analysis results.

7. The method of claim 1 wherein processing each context partition comprises:

extracting key information from the context partition;

generating a summary of the context partition;

identifying one or more relevant facts and relationships within the context partition; and

storing the extracted key information, the summary, and the identified one or more relevant facts and relationships as part of the intermediate analysis results.

8. The method of claim 1 wherein combining the plurality of intermediate analysis results comprises:

assigning a confidence score to each intermediate analysis result of the plurality of intermediate analysis results;

generating a weighted plurality of intermediate analysis results by weighting each intermediate analysis result of the plurality of intermediate analysis results based on their respective confidence scores; and

generating the final response based on the weighted plurality of intermediate analysis results.

9. The method of claim 1 further comprising:

monitoring system resources during processing;

dynamically adjusting a size of the context partitions of the plurality of context partitions based on available system resources; and

scheduling the processing of the plurality of context partitions to optimize resource utilization.

10. The method of claim 1 wherein:

each context partition of the plurality of context partitions is processed using a first LLM optimized for analysis and extraction;

the plurality of intermediate analysis results are combined using a second LLM optimized for synthesis and summarization; and

the final response is validated using a third LLM that is optimized for fact-checking and verification.

11. The method of claim 1 wherein the mapper prompt comprises at least one of:

instructions for analyzing the context partition;

tags for structuring the analysis of the context partition;

guidelines for extracting and formatting relevant information;

a context section containing the context partition; and

a question section containing the query.

12. The method of claim 1 wherein the reducer prompt comprises at least one of:

instructions for combining the intermediate analysis results;

tags for structuring the final response;

guidelines for generating related questions;

a context section containing the combined intermediate analysis results;

a conversation history section; and

a question section containing the query.

13. The method of claim 1 wherein processing each context partition using the mapper prompt comprises at least one of:

identifying key information relevant to the query;

extracting relevant quotes with their sources;

identifying significant details, dates, and relationships;

identifying insights from the context partition; and

generating the intermediate analysis results within specified tags.

14. The method of claim 1 wherein processing the plurality of intermediate analysis results using the reducer prompt comprises at least one of:

generating a detailed answer addressing the query;

generating a list of related questions for future exploration;

formatting the final response with specified tags;

including source documentation references; and

validating the response against provided guidelines.

15. The method of claim 1 wherein the mapper prompt and reducer prompt each include one or more guidelines comprising at least one of:

guidelines for validating document relevance based on mentioned entities;

guidelines for providing relevant excerpts with source attribution;

guidelines for ignoring irrelevant documents;

guidelines for declining to answer when no relevant context is found;

guidelines for maintaining consistent terminology; and

guidelines for acknowledging limitations of generating the final response.

16. The method of claim 1 further comprising:

maintaining a conversation history related to the user;

adding the query to the conversation history;

including the conversation history in the reducer prompt; and

generating the final response based upon each of the intermediate analysis results and the conversation history.

17. The method of claim 1 wherein the mapper prompt is sent to a plurality of instances of a single LLM.

18. The method of claim 1 wherein the mapper prompt is sent to a plurality of LLMs.

19. A retrieval-augmented generation (RAG) system configured to process large contexts comprising:

a processor;

a communication device positioned in communication with the processor and configured to send and receive transmissions across a digital network; and

a storage medium having stored thereon a document database;

a non-transitory computer-readable storage medium having stored thereon software that, when executed by the processor, is operable to:

receive a query from a user;

retrieve a plurality of relevant documents from the document database based on the query, wherein the relevant documents form a combined context;

partition the combined context into a plurality of context partitions;

generate a plurality of intermediate analysis results by:

processing each context partition of the plurality of context partitions using a mapper prompt; and

sending an output of the mapper prompt from each context partition to a first large language model (LLM);

generate a final response by processing the plurality of intermediate analysis results using a reducer prompt sent to at least one of the first LLM and a second LLM; and

transmit the final response to the user.

20. The RAG system of claim 19 wherein the software is further operable to, when executed by the processor, retrieve the plurality of relevant documents by:

identifying a first set of documents by performing a keyword search in a relational database;

identifying a second set of documents by performing a vector search in a vector database; and

combining the first and second sets of documents to form the plurality of relevant documents.

21. The RAG system of claim 19 wherein the software is further operable to, when executed by the processor, process each context partition by:

processing the plurality of context partitions in parallel using a plurality of LLM instances; and

generating the plurality of intermediate analysis results simultaneously.

22. The RAG system of claim 19 wherein the software is further operable to, when executed by the processor, combine the plurality of intermediate analysis results by:

identifying one or more common themes across the plurality of intermediate analysis results;

resolving any conflicts between the plurality of intermediate analysis results;

synthesizing the plurality of intermediate analysis results into a coherent response; and

validating the coherent response for at least one of accuracy and completeness.

23. The RAG system of claim 22 wherein the software is further operable to, when executed by the processor, process the coherent response for at least one of:

compliance with one or more system policies;

protection from unauthorized access;

protection from privacy breaches; or

protection from potential jailbreaking attempts.

24. The RAG system of claim 19 wherein the software is further operable to, when executed by the processor:

store the plurality of intermediate analysis results in a cache;

receive a subsequent query related to the query;

retrieve the stored plurality of intermediate analysis results from the cache; and

generate a new final response using the stored plurality of intermediate analysis results.

25. The RAG system of claim 19 wherein the software is further operable to, when executed by the processor, process each context partition by:

extracting key information from the context partition;

generating a summary of the context partition;

identifying one or more relevant facts and relationships within the context partition; and

storing the extracted key information, the summary, and the identified one or more relevant facts and relationships as part of the intermediate analysis results.

26. The RAG system of claim 19 wherein the software is further operable to, when executed by the processor, combine the plurality of intermediate analysis results by:

assigning a confidence score to each intermediate analysis result of the plurality of intermediate analysis results;

generating a weighted plurality of intermediate analysis results by weighting each intermediate analysis result of the plurality of intermediate analysis results based on their respective confidence scores; and

generating the final response based on the weighted plurality of intermediate analysis results.

27. The RAG system of claim 19 wherein the software is further operable to, when executed by the processor:

monitor system resources during processing;

dynamically adjust a size of the context partitions of the plurality of context partitions based on available system resources; and

schedule the processing of the plurality of context partitions to optimize resource utilization.

28. The RAG system of claim 19 wherein the software is further operable to, when executed by the processor:

process each context partition of the plurality of context partitions using the first LLM, the first LLM being optimized for analysis and extraction;

combine the plurality of intermediate analysis results using the second LLM, the second LLM being optimized for synthesis and summarization; and

validate the final response using a third LLM that is optimized for fact-checking and verification.

29. The RAG system of claim 19 wherein the software comprises instructions for the mapper prompt to comprise at least one of:

instructions for analyzing the context partition;

tags for structuring the analysis of the context partition;

guidelines for extracting and formatting relevant information;

a context section containing the context partition; and

a question section containing the query.

30. The RAG system of claim 19 wherein the software comprises instructions for the reducer prompt to comprise at least one of:

instructions for combining the intermediate analysis results;

tags for structuring the final response;

guidelines for generating related questions;

a context section containing the combined intermediate analysis results;

a conversation history section; and

a question section containing the query.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2025
From: BAHGA, ARSHDEEP
To: MADISETTI, VIJAY
Reel/Frame 070332/0062 →
Continuity (16)
Continuation In Part 19040471 · Jan 29, 2025
Continuation In Part 18921852 · Oct 21, 2024
Continuation In Part 18812707 · Aug 22, 2024
Continuation In Part 18470487 · Sep 20, 2023
Continuation 18348692 · Jul 7, 2023
Provisional Application 63742792 · Jan 7, 2025
Provisional Application 63693351 · Sep 11, 2024
Provisional Application 63647092 · May 14, 2024
Provisional Application 63607647 · Dec 8, 2023
Provisional Application 63607112 · Dec 7, 2023
Provisional Application 63535118 · Aug 29, 2023
Provisional Application 63534974 · Aug 28, 2023
Provisional Application 63529177 · Jul 27, 2023
Provisional Application 63469571 · May 30, 2023
Provisional Application 63463913 · May 4, 2023
Related Publication 20250190461A1 · Jun 12, 2025
References Cited (28)
US 5410475A · Lu et al. · 1995 [cited by applicant]
US 8775436B1 · Zhou · 2014 [cited by examiner]
US 11194868B1 · Paka · 2021 [cited by examiner]
US 11928569B1 · Douthit · 2024 [cited by applicant]
US 11995411B1 · Qadrud-Din · 2024 [cited by examiner]
US 12242503B1 · Kelsey et al. · 2025 [cited by applicant]
US 20070244937A1 · Flynn, Jr. et al. · 2007 [cited by applicant]
US 20080077569A1 · Lee · 2008 [cited by examiner]
US 20100138531A1 · Kashyap · 2010 [cited by applicant]
US 20120023073A1 · Dean et al. · 2012 [cited by applicant]
US 20120290521A1 · Frank et al. · 2012 [cited by applicant]
US 20140297845A1 · Tamura · 2014 [cited by applicant]
US 20180095845A1 · Sanakkayala et al. · 2018 [cited by applicant]
US 20190081959A1 · Yadav et al. · 2019 [cited by applicant]
US 20190130006A1 · Raviv · 2019 [cited by examiner]
US 20190130902A1 · Itoh et al. · 2019 [cited by applicant]
US 20220124013A1 · Chitalia et al. · 2022 [cited by applicant]
US 20240289560A1 · Kelly et al. · 2024 [cited by applicant]
US 20240354130A1 · Cadoni et al. · 2024 [cited by applicant]
US 20240354490A1 · Chauvin et al. · 2024 [cited by applicant]
US 20240404687A1 · Bell et al. · 2024 [cited by applicant]
US 20240412029A1 · Yang et al. · 2024 [cited by applicant]
US 20250036878A1 · Marwah · 2025 [cited by examiner]
US 20250045256A1 · Gottlob · 2025 [cited by examiner]
US 20250068665A1 · Chandel · 2025 [cited by examiner]
Non Final Office Action received in related U.S. Appl. No. 19/040,471 issued on Apr. 9, 2025; 11 pages. [cited by applicant]
Notice of Allowance received in related U.S. Appl. No. 18/812,707 issued on May 20, 2025; 26 pages. [cited by applicant]
Notice of Allowance received in related U.S. Appl. No. 18/921,852 issued Jun. 20, 2025; 25 pages. [cited by applicant]