IP Library › Granted Patent US 12,711,164
Granted Patent B2
US 12,711,164 · App. 19/287,291 · Granted Aug 18, 2026

Method and system for optimizing use of retrieval augmented generation pipelines in generative artificial intelligence applications

Inventors: Vijay Madisetti (Alpharetta, GA); Arshdeep Bahga (Chandigarh, IN)
Assignee: Vijay Madisetti
G06F16/3329G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,164
App. No.
19/287,291
Filed
Jul 31, 2025
Granted
Aug 18, 2026
Kind
B2
Examiner
YEN, ERIC L
Art Unit
2658
USPC
704/9
Abstract

Systems and methods of adaptive context partitioning in a retrieval-augmented generation system, including receiving a query from a user; forming a combined context by retrieving relevant documents from a document database based on the query, determining characteristics of the combined context, monitoring current system resources including processor availability and memory utilization, dynamically determining a partition size based on the current system resources and the one or more characteristics of the combined context, partitioning the combined context into context partitions according to the partition size, generating intermediate analysis results by processing each context partition using a mapper prompt large language models (LLMs), generating a final response by processing the intermediate analysis results using a reducer prompt through the LLMs, and transmitting the final response to the user.

Claims (83)

1 . A method of adaptive context partitioning in a retrieval-augmented generation system, comprising:

receiving a query from a user;

forming a combined context by retrieving a plurality of relevant documents from at least one document database based on the query;

monitoring current system resources including processor availability and memory utilization;

dynamically determining a partition size based on the current system resources;

partitioning the combined context into a plurality of context partitions according to the partition size while preserving document boundaries to maintain semantic coherence;

generating a plurality of intermediate analysis results by processing each context partition of the plurality of context partitions using a mapper prompt through one or more large language models (LLMs);

generating a final response by processing the plurality of intermediate analysis results using a reducer prompt through the one or more LLMs; and

transmitting the final response to the user.

2 . The method of claim 1 further comprising:

detecting a performance degradation during processing; and

dynamically adjusting the partition size for remaining unprocessed content comprised by the combined context responsive to the performance degradation.

3 . The method of claim 1 further comprising redistributing processing tasks across additional available computational resources.

4 . The method of claim 1 wherein processing each context partition of the plurality of context partitions further comprises:

processing the plurality of context partitions in parallel using a plurality of LLM instances; and

generating the plurality of intermediate analysis results simultaneously.

5 . The method of claim 1 wherein partitioning the combined context preserves document boundaries by ensuring that each context partition of the plurality of context partitions comprises one or more of complete documents or complete document segments.

6 . The method of claim 1 further comprising:

storing the plurality of intermediate analysis results as a stored plurality of intermediate analysis results in a cache;

receiving a subsequent query related to the query;

retrieving the stored plurality of intermediate analysis results from the cache; and

generating a new final response using the stored plurality of intermediate analysis results.

7 . The method of claim 1 further comprising:

assigning a confidence score to each intermediate analysis result of the plurality of intermediate analysis results;

generating a weighted plurality of intermediate analysis results by weighting each intermediate analysis result of the plurality of intermediate analysis results based on their respective confidence scores; and

generating the final response based on the weighted plurality of intermediate analysis results.

8 . The method of claim 1 wherein generating the final response is further based on a conversation history associated with the user.

9 . A system for adaptive context partitioning in a retrieval-augmented generation process, comprising:

a processor;

a communication device positioned in communication with the processor and configured to send and receive transmissions across a digital network; and

a non-transitory computer-readable storage medium having stored thereon software that, when executed by the processor, is operable to:

receive a query from a user;

retrieve a plurality of relevant documents that form a combined context;

monitor available system resources;

dynamically adjust a partition size for context partitions based on the available system resources;

partition the combined context into a plurality of context partitions according to the partition size;

implement a map phase operable to generate a plurality of intermediate analysis results, wherein each context partition of the plurality of context partitions is processed in parallel;

implement a reduce phase operable to synthesize the plurality of intermediate analysis results into a coherent response; and

transmit the coherent response to the user.

10 . The system of claim 9 , wherein the software is further operable to, when executed by the processor:

implement a feedback loop operable to refine partitioning the combined context based on processing results;

check a cache for previously processed similar queries before partitioning the combined context; and

apply one or more guardrails to the coherent response.

11 . The system of claim 9 , wherein the map phase comprises:

distributing the plurality of context partitions to a plurality of mapper instances such that each context partition of the plurality of context partitions has an assigned mapper instance of the plurality of mapper instances;

processing each context partition of the plurality of context partitions independently with its assigned mapper instance of the plurality of mapper instances; and

shuffling and sorting the plurality of intermediate analysis results before the reduce phase.

12 . The system of claim 9 , wherein dynamically adjusting the partition size comprises:

increasing the partition size when system resources are above a first threshold; and

decreasing the partition size when system resources are below a second threshold.

13 . The system of claim 9 wherein the software, when executed by the processor, implements the reduce phase by:

assigning a confidence score to each intermediate analysis result of the plurality of intermediate analysis results;

generating a weighted plurality of intermediate analysis results by weighting each intermediate analysis result of the plurality of intermediate analysis results based on their respective confidence scores; and

generating the coherent response based on the weighted plurality of intermediate analysis results.

14 . The system of claim 9 wherein generating the coherent response is further based on a conversation history associated with the user.

15 . The system of claim 9 wherein the software is further operable to, when executed by the processor, partition the combined context to preserve document boundaries by ensuring that each context partition of the plurality of context partitions comprises one or more of complete documents or complete document segments.

16 . A system for adaptive context partitioning in a retrieval-augmented generation system, comprising:

means for receiving a query from a user;

means for forming a combined context by retrieving a plurality of relevant documents from at least one document database based on the query;

means for monitoring current system resources including processor availability and memory utilization;

means for dynamically determining a partition size based on the current system resources;

means for partitioning the combined context into a plurality of context partitions according to the partition size while preserving document boundaries to maintain semantic coherence;

means for generating a plurality of intermediate analysis results by processing each context partition of the plurality of context partitions using a mapper prompt through one or more large language models (LLMs);

means for generating a final response by processing the plurality of intermediate analysis results using a reducer prompt through the one or more LLMs; and

means for transmitting the final response to the user.

17 . The system of claim 16 further comprising:

means for detecting a performance degradation during processing; and

means for dynamically adjusting the partition size for remaining unprocessed content comprised by the combined context responsive to the performance degradation.

18 . The system of claim 16 further comprising means for redistributing processing tasks across additional available computational resources.

19 . The system of claim 16 wherein the means for generating a plurality of intermediate analysis results by processing each context partition of the plurality of context partitions further is operable to:

process the plurality of context partitions in parallel using a plurality of LLM instances; and

generate the plurality of intermediate analysis results simultaneously.

20 . The system of claim 16 wherein the means for partitioning the combined context is operable to preserve document boundaries by ensuring that each context partition of the plurality of context partitions comprises one or more of complete documents or complete document segments.

21 . The system of claim 16 further comprising:

means for storing the plurality of intermediate analysis results as a stored plurality of intermediate analysis results in a cache;

means for receiving a subsequent query related to the query;

means for retrieving the stored plurality of intermediate analysis results from the cache; and

means for generating a new final response using the stored plurality of intermediate analysis results.

22 . The system of claim 16 further comprising:

means for assigning a confidence score to each intermediate analysis result of the plurality of intermediate analysis results;

means for generating a weighted plurality of intermediate analysis results by weighting each intermediate analysis result of the plurality of intermediate analysis results based on their respective confidence scores; and

means for generating the final response based on the weighted plurality of intermediate analysis results.

23 . The system of claim 16 wherein the means for generating the final response is operable to generate the final response further based on a conversation history associated with the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2025
From: BAHGA, ARSHDEEP
To: MADISETTI, VIJAY
Reel/Frame 071980/0613 →
Continuity (17)
Continuation 19056496 · Feb 18, 2025
Continuation In Part 19040471 · Jan 29, 2025
Continuation In Part 18921852 · Oct 21, 2024
Continuation In Part 18812707 · Aug 22, 2024
Continuation In Part 18470487 · Sep 20, 2023
Continuation 18348692 · Jul 7, 2023
Provisional Application 63742792 · Jan 7, 2025
Provisional Application 63693351 · Sep 11, 2024
Provisional Application 63647092 · May 14, 2024
Provisional Application 63607647 · Dec 8, 2023
Provisional Application 63607112 · Dec 7, 2023
Provisional Application 63535118 · Aug 29, 2023
Provisional Application 63534974 · Aug 28, 2023
Provisional Application 63529177 · Jul 27, 2023
Provisional Application 63469571 · May 30, 2023
Provisional Application 63463913 · May 4, 2023
Related Publication 20250363142A1 · Nov 27, 2025
References Cited (23)
US 5410475A · Lu · 1995 [cited by applicant]
US 7809697B1 · Kanefsky · 2010 [cited by examiner]
US 10261973B2 · Simhon · 2019 [cited by applicant]
US 11928569B1 · Douthit · 2024 [cited by applicant]
US 12033618B1 · Wei · 2024 [cited by examiner]
US 12242503B1 · Kelsey · 2025 [cited by applicant]
US 12361038B2 · Zang · 2025 [cited by examiner]
US 12405985B1 · Kanagovi · 2025 [cited by examiner]
US 12505148B2 · Hintz · 2025 [cited by examiner]
US 12511482B1 · Maragakis · 2025 [cited by applicant]
US 20120023073A1 · Dean · 2012 [cited by applicant]
US 20120290521A1 · Frank · 2012 [cited by applicant]
US 20190130902A1 · Itoh · 2019 [cited by applicant]
US 20240256592A1 · O'Neill · 2024 [cited by examiner]
US 20240265041A1 · Rennie · 2024 [cited by examiner]
US 20240386015A1 · Crabtree · 2024 [cited by examiner]
US 20240404687A1 · Bell · 2024 [cited by applicant]
US 20250292110A1 · Jain et al. · 2025 [cited by applicant]
CN 118839021B · 2025 [cited by examiner]
Google Translation of CN118839021B, 2025, https://patents.google.com/patent/CN118839021B/en?oq=CN118839021 (Year: 2025). [cited by examiner]
Notice of Allowance received in related U.S. Appl. No. 19/306,379 issued on Mar. 5, 2026; 21 pages. [cited by applicant]
Bhaskarjit Sarmah, Benika Hall, Rohan Rao, Sunil Patel, Stefano Pasquali, Dhagash Mehta, “HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction”, 2024, ht… [cited by applicant]
Non-Final Office Action received in related U.S. Appl. No. 19/646,057 issued on Jul. 17, 2026; 27 pages. [cited by applicant]