IP Library › Granted Patent US 12,614,035
Granted Patent B2
US 12,614,035 · App. 18/478,613 · Granted Apr 28, 2026

Retrieval augmented generation

Inventors: Jolene Huang (Olympia, WA); Pankaj Rastogi (Fremont, CA); Spriha Awasthi (Superior, CO); Yiran Huang (San Diego, CA)
Assignee: Intuit Inc.
G06F40/295G06F16/345G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,035
App. No.
18/478,613
Filed
Sep 29, 2023
Granted
Apr 28, 2026
Kind
B2
Art Unit
2653
USPC
704/9
Abstract

Certain aspects of the present disclosure provide techniques retrieval augmented generation of language model responses using an embedding database. Embeddings for data is stored in an embedding database. When a prompt related to the data is received, relevant embeddings may be retrieved from the database and used generate an augmented prompt based on the initial prompt and the retrieved embeddings from the database. The augmented prompt can be input into a machine learning model. Although the model may be unaware of the data from which the embeddings of the embedding database were generated, the augmented prompt enables the model to use the data to improve breadth and depth of responses.

Claims (76)

1 . A method, comprising:

receiving a prompt from a user at an endpoint;

preprocessing the prompt by:

performing summarization on the prompt;

performing entity extraction on the prompt; and

performing classification on the prompt;

providing the prompt to an embedding module to generate an embedding for the prompt;

retrieving one or more embeddings from a vector database based on the embedding for the prompt and based on a result of the preprocessing of the prompt;

generating an augmented prompt based on the one or more embeddings and the embedding for the prompt;

generating a response to the augmented prompt using a machine learning model and using a context of the one or more embeddings as a context for the machine learning model; and

providing the response to the user at the endpoint.

2 . The method of claim 1 , further comprising:

receiving a corpus of data;

chunking the corpus of data to produce a plurality of data chunks;

generating a plurality of embeddings for the plurality of data chunks; and

storing the plurality of embeddings in the vector database.

3 . The method of claim 2 , further comprising:

determining a segment of the corpus of data corresponds to a conversation; and

determining a context of the conversation; wherein

chunking the corpus of data comprises

generating a data chunk corresponding to the segment, and

generating the plurality of embeddings by generating an embedding for the segment that includes metadata defining the context of the conversation as a context for the embedding.

4 . The method of claim 2 , wherein the corpus of text is chunked using a configurable algorithmic delimiter.

5 . The method of claim 1 , further comprising:

performing, on an embedding of the one or more embeddings, at least one of an insertion operation, an update operation, or a deletion operation.

6 . The method of claim 1 , wherein retrieving one or more embeddings from a vector database comprises:

performing a semantic search using a similarity algorithm and the embedding for the prompt to identify one or more similar embeddings in the vector database, the similar embeddings being similar to the embedding for the prompt; and

combining the similar embeddings and the embedding for the prompt to create an augmented embedding.

7 . The method of claim 6 , wherein the semantic search is performed using a filter or limited search space.

8 . The method of claim 1 , wherein the prompt comprises a selection from a test data set and the method further comprising validating the response using the test data set.

9 . The method of claim 1 , wherein the prompt is received from a user via a user interface of the endpoint, and wherein the response is displayed to the user on a display associated with the user interface.

10 . A system comprising:

a memory having executable instructions stored thereon;

an endpoint having a user interface associated with a display; and

one or more processors configured to execute the executable instructions to cause the system to perform a method comprising:

receiving a prompt from a user at the user interface;

providing the prompt to an embedding module to generate an embedding for the prompt;

retrieving one or more embeddings from a vector database based on the embedding for the prompt, wherein the retrieving the one or more embeddings comprises performing a semantic search using a similarity algorithm and the embedding for the prompt to identify one or more similar embeddings in the vector database, the similar embeddings being similar to the embedding for the prompt;

generating an augmented prompt based on the one or more embeddings and the embedding for the prompt, wherein the generating of the augmented prompt comprises combining the similar embeddings and the embedding for the prompt to create an augmented embedding;

generating a response to the augmented prompt using a machine learning model and using a context of the one or more embeddings as a context for the machine learning model; and

providing the response to the user on the display.

11 . The system of claim 10 , the method further comprising:

receiving a corpus of data;

chunking the corpus of data to produce a plurality of data chunks;

generating a plurality of embeddings for the plurality of data chunks; and

storing the plurality of embeddings in the vector database.

12 . The system of claim 11 , the method further comprising:

determining a segment of the corpus of data corresponds to a conversation; and

determining a context of the conversation; wherein

chunking the corpus of data comprises

generating a data chunk corresponding to the segment, and

generating the plurality of embeddings by generating an embedding for the segment that includes metadata defining the context of the conversation as a context for the embedding.

13 . The system of claim 11 , wherein the corpus of text is chunked using a configurable algorithmic delimiter.

14 . The system of claim 10 , the method further comprising:

preprocessing the prompt by

performing summarization on the prompt;

performing entity extraction on the prompt; and

performing classification on the prompt; and

retrieving the embedding from the vector database based on a result of preprocessing the prompt.

15 . The system of claim 10 , the method further comprising:

performing, on an embedding of the one or more embeddings, at least one of an insertion operation, an update operation, or a deletion operation.

16 . The system of claim 10 , wherein the semantic search is performed using a filter or limited search space.

17 . The system of claim 10 , wherein the prompt comprises a selection from a test data set and the method further comprising validating the response using the test data set.

18 . A non-transitory computer readable storage medium comprising instructions, that when executed by one or more processors of a computing system, cause the computing system to perform a method comprising:

receiving a corpus of data;

determining a segment of the corpus of data corresponds to a conversation;

determining a context of the conversation;

chunking the corpus of data to produce a plurality of data chunks, wherein the chunking of the corpus of data comprises generating a data chunk corresponding to the segment;

generating a plurality of embeddings for the plurality of data chunks, wherein the generating of the plurality of embeddings comprises generating an embedding for the segment that includes metadata defining the context of the conversation as a context for the embedding;

storing the plurality of embeddings in a vector database;

receiving a prompt from a user at an endpoint;

providing the prompt to an embedding module to generate an embedding for the prompt;

retrieving one or more embeddings from the vector database based on the embedding for the prompt;

generating an augmented prompt based on the one or more embeddings and the embedding for the prompt;

generating a response to the augmented prompt using a machine learning model and using a context of the one or more embeddings as a context for the machine learning model; and

providing the response to the user at the endpoint.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2023
From: HUANG, JOLENE; RASTOGI, PANKAJ; AWASTHI, SPRIHA; HUANG, YIRAN
To: INTUIT, INC.
Reel/Frame 065087/0624 →
Continuity (1)
Related Publication 20250111159A1 · Apr 3, 2025
References Cited (3)
US 11468238B2 · Tiwari · 2022 [cited by examiner]
US 20230037077A1 · Gutta · 2023 [cited by examiner]
US 20250181899A1 · Ajmera · 2025 [cited by examiner]