IP Library › Granted Patent US 12,602,412
Granted Patent B2
US 12,602,412 · App. 19/306,379 · Granted Apr 14, 2026

Method and system for optimizing use of retrieval augmented generation pipelines in generative artificial intelligence applications

Inventors: Vijay Madisetti (Alpharetta, GA); Arshdeep Bahga (Chandigarh, IN)
Assignee: Vijay Madisetti
G06F16/3329G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,412
App. No.
19/306,379
Filed
Aug 21, 2025
Granted
Apr 14, 2026
Kind
B2
Examiner
YEN, ERIC L
Art Unit
2658
USPC
704/9
Abstract

Systems and methods of processing domain-specific content in a generative AI system including receiving a prompt, tokenizing the prompt, identifying an identified domain of the tokenized prompt identifying domain-specific functions within the identified domain, generating domain-specific sub-functions from the domain-specific functions according to a hierarchical mapping, generating H-Tokens, each encapsulating one of a domain-specific function or a domain-specific sub-function and relationships domain-specific functions and the domain-specific sub-functions, implementing the of H-Tokens, assembling a response using the implemented of H-Tokens, and transmitting the response to the user.

Claims (99)

1 . A method of processing domain-specific content in a generative artificial intelligence system, comprising:

receiving a prompt from a user;

generating a tokenized prompt by performing an initial tokenization of the prompt;

identifying an identified domain comprised by a plurality of supported domains that is associated with the tokenized prompt by analyzing the tokenized prompt;

identifying one or more domain-specific functions within the identified domain by performing a domain-specific analysis on the tokenized prompt;

generating a plurality of domain-specific sub-functions from the one or more domain-specific functions according to a hierarchical mapping of the identified domain;

generating a plurality of hierarchical tokens (H-Tokens), each H-Token of the plurality of H-Tokens being configured to encapsulate one of a domain-specific function of the one or more domain-specific functions or a domain-specific sub-function of the plurality of domain-specific sub-functions, and relationships between the one or more domain-specific functions and the plurality of domain-specific sub-functions;

implementing the plurality of H-Tokens, resulting in an implemented plurality of H-Tokens, by mapping each H-Token of the plurality of H-Tokens to each of a corresponding function, event, and implementation method within the identified domain;

assembling a response using the implemented plurality of H-Tokens; and

transmitting the response to the user.

2 . The method of claim 1 , wherein implementing the plurality H-Tokens comprises:

identifying a plurality of functions as one of domain-specific tasks that lead to events or domain-specific processes that lead to events;

identifying a plurality of events as results of the plurality of functions; and

identifying a plurality of implementation methods to accomplish the plurality of functions.

3 . The method of claim 2 wherein each implementation method of the plurality of implementation methods is represented as an H-Token.

4 . The method of claim 1 , further comprising:

assembling a context using the implemented plurality of H-Tokens;

selecting a processing path from one of:

expanding the plurality H-Tokens into regular tokens for processing; or

generating processed content by directly processing the plurality of H-Tokens without expansion;

generating a processed context by processing the context consistent with the processing path; and

integrating the processed context into the response.

5 . The method of claim 1 , further comprising integrating the implemented plurality of H-Tokens with a retrieval-augmented generation (RAG) system by:

assembling an H-Token context comprising at least one of a location context, a duration context, and a preference context;

generating a processed context by processing the H-Token context through one of:

a token expansion path operable to expand the plurality of H-Tokens to regular tokens for traditional RAG processing; or

a direct H-Token processing path for H-Token-aware response generation; and

producing the response based on the processed context.

6 . The method of claim 1 wherein assembling the response using the implemented plurality of H-Tokens comprises:

assembling the implemented plurality of H-Tokens;

forming a preliminary response responsive to assembling the implemented plurality of H-Tokens; and

generating the response by performing a quality verification on the preliminary response.

7 . A system for processing domain-specific content in a generative artificial intelligence system, comprising:

a processor;

a network communication device operably coupled to the processor and configured to transmit and receive digital transmissions across a network to removed computerized devices; and

a non-transitory computer-readable storage medium having stored thereon software that, when executed by the processor, is configured to:

receive a prompt from a user;

generate a tokenized prompt by performing an initial tokenization of the prompt;

identify an identified domain comprised by a plurality of supported domains that is associated with the tokenized prompt by analyzing the tokenized prompt;

identify one or more domain-specific functions within the identified domain by performing a domain-specific analysis on the tokenized prompt;

generate a plurality of domain-specific sub-functions from the one or more domain-specific functions according to a hierarchical mapping of the identified domain;

generate a plurality of hierarchical tokens (H-Tokens), each H-Token of the plurality of H-Tokens being configured to encapsulate one of a domain-specific function of the one or more domain-specific functions or a domain-specific sub-function of the plurality of domain-specific sub-functions, and relationships between the one or more domain-specific functions and the plurality of domain-specific sub-functions;

implement the plurality of H-Tokens, resulting in an implemented plurality of H-Tokens, by mapping each H-Token of the plurality of H-Tokens to each of a corresponding function, event, and implementation method within the identified domain;

assemble a response using the implemented plurality of H-Tokens; and

transmit the response to the user.

8 . The system of claim 7 , wherein the software, when executed by the processor, is operable to implement the plurality H-Tokens by:

identifying a plurality of functions as one of domain-specific tasks that lead to events or domain-specific processes that lead to events;

identifying a plurality of events as results of the plurality of functions; and

identifying a plurality of implementation methods to accomplish the plurality of functions.

9 . The system of claim 8 wherein each implementation method of the plurality of implementation methods is represented as an H-Token.

10 . The system of claim 7 , wherein the software, when executed by the processor, is further operable to:

assemble a context using the implemented plurality of H-Tokens;

select a processing path from one of:

expand the plurality H-Tokens into regular tokens for processing; or

generate processed content by directly processing the plurality of H-Tokens without expansion;

generate a processed context by processing the context consistent with the processing path; and

integrate the processed context into the response.

11 . The system of claim 7 wherein the software, when executed by the processor, is further operable to integrate the implemented plurality of H-Tokens with a retrieval-augmented generation (RAG) system by:

assembling an H-Token context comprising at least one of a location context, a duration context, and a preference context;

generating a processed context by processing the H-Token context through one of:

a token expansion path operable to expand the plurality of H-Tokens to regular tokens for traditional RAG processing; or

a direct H-token processing path for H-Token-aware response generation; and

producing the response based on the processed context.

12 . The system of claim 7 wherein the software, when executed by the processor, is operable to assemble the response using the implemented plurality of H-Tokens by:

assembling the implemented plurality of H-Tokens;

forming a preliminary response responsive to assembling the implemented plurality of H-Tokens; and

generating the response by performing a quality verification on the preliminary response.

13 . A system for processing domain-specific content in a generative artificial intelligence system, comprising:

means for receiving a prompt from a user;

means for generating a tokenized prompt by performing an initial tokenization of the prompt;

means for identifying an identified domain comprised by a plurality of supported domains that is associated with the tokenized prompt by analyzing the tokenized prompt;

means for identifying one or more domain-specific functions within the identified domain by performing a domain-specific analysis on the tokenized prompt;

means for generating a plurality of domain-specific sub-functions from the one or more domain-specific functions according to a hierarchical mapping of the identified domain;

means for generating a plurality of hierarchical tokens (H-Tokens), each H-Token of the plurality of H-Tokens being configured to encapsulate one of a domain-specific function of the one or more domain-specific functions or a domain-specific sub-function of the plurality of domain-specific sub-functions, and relationships between the one or more domain-specific functions and the plurality of domain-specific sub-functions;

means for implementing the plurality of H-Tokens, resulting in an implemented plurality of H-Tokens, by mapping each H-Token of the plurality of H-Tokens to each of a corresponding function, event, and implementation method within the identified domain;

means for assembling a response using the implemented plurality of H-Tokens; and

means for transmitting the response to the user.

14 . The system of claim 13 , wherein the means for implementing the plurality H-Tokens is operable to:

identify a plurality of functions as one of domain-specific tasks that lead to events or domain-specific processes that lead to events;

identify a plurality of events as results of the plurality of functions; and

identify a plurality of implementation methods to accomplish the plurality of functions.

15 . The system of claim 14 wherein each implementation method of the plurality of implementation methods is represented as an H-Token.

16 . The system of claim 13 , further comprising:

means for assembling a context using the implemented plurality of H-Tokens;

means for selecting a processing path from one of:

expanding the plurality H-Tokens into regular tokens for processing; or

generating processed content by directly processing the plurality of H-Tokens without expansion;

means for generating a processed context by processing the context consistent with the processing path; and

means for integrating the processed context into the response.

17 . The system of claim 13 , further comprising means for integrating the implemented plurality of H-Tokens with a retrieval-augmented generation (RAG) system by:

assembling an H-Token context comprising at least one of a location context, a duration context, and a preference context;

generating a processed context by processing the H-Token context through one of:

a token expansion path operable to expand the plurality of H-Tokens to regular tokens for traditional RAG processing; or

a direct H-token processing path for H-Token-aware response generation; and

producing the response based on the processed context.

18 . The system of claim 13 wherein the means for assembling the response using the implemented plurality of H-Tokens is operable to:

assemble the implemented plurality of H-Tokens;

form a preliminary response responsive to assembling the implemented plurality of H-Tokens; and

generate the response by performing a quality verification on the preliminary response.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2025
From: BAHGA, ARSHDEEP
To: MADISETTI, VIJAY
Reel/Frame 072098/0858 →
Continuity (18)
Continuation 19287291 · Jul 31, 2025
Continuation 19056496 · Feb 18, 2025
Continuation In Part 19040471 · Jan 29, 2025
Continuation In Part 18921852 · Oct 21, 2024
Continuation In Part 18812707 · Aug 22, 2024
Continuation In Part 18470487 · Sep 20, 2023
Continuation 18348692 · Jul 7, 2023
Provisional Application 63742792 · Jan 7, 2025
Provisional Application 63693351 · Sep 11, 2024
Provisional Application 63647092 · May 14, 2024
Provisional Application 63607647 · Dec 8, 2023
Provisional Application 63607112 · Dec 7, 2023
Provisional Application 63535118 · Aug 29, 2023
Provisional Application 63534974 · Aug 28, 2023
Provisional Application 63529177 · Jul 27, 2023
Provisional Application 63469571 · May 30, 2023
Provisional Application 63463913 · May 4, 2023
Related Publication 20260023764A1 · Jan 22, 2026
References Cited (17)
US 5410475A · Lu · 1995 [cited by examiner]
US 7809697B1 · Kanefsky et al. · 2010 [cited by applicant]
US 10261973B2 · Simhon · 2019 [cited by examiner]
US 11928569B1 · Douthit · 2024 [cited by examiner]
US 12033618B1 · Wei et al. · 2024 [cited by applicant]
US 12242503B1 · Kelsey · 2025 [cited by examiner]
US 12361038B2 · Zang et al. · 2025 [cited by applicant]
US 12405985B1 · Kanagovi et al. · 2025 [cited by applicant]
US 12505148B2 · Hintz et al. · 2025 [cited by applicant]
US 12511482B1 · Maragakis · 2025 [cited by examiner]
US 20120023073A1 · Dean · 2012 [cited by examiner]
US 20120290521A1 · Frank · 2012 [cited by examiner]
US 20190130902A1 · Itoh · 2019 [cited by examiner]
US 20240404687A1 · Bell · 2024 [cited by examiner]
CN 118839021B · 2025 [cited by applicant]
Google Translation of CN 118839021 B, 2025, https://patents.google.com/patent/CN118839021 B/en?oq=CN118839021 (Year: 2025). [cited by applicant]
Non Final Office Action received in related U.S. Appl. No. 19/287,291 issued on Jan. 21, 2026; 25 pages. [cited by applicant]