Retrieval augmented generation
Certain aspects of the present disclosure provide techniques retrieval augmented generation of language model responses using an embedding database. Embeddings for data is stored in an embedding database. When a prompt related to the data is received, relevant embeddings may be retrieved from the database and used generate an augmented prompt based on the initial prompt and the retrieved embeddings from the database. The augmented prompt can be input into a machine learning model. Although the model may be unaware of the data from which the embeddings of the embedding database were generated, the augmented prompt enables the model to use the data to improve breadth and depth of responses.
1 . A method, comprising:
receiving a prompt from a user at an endpoint;
preprocessing the prompt by:
performing summarization on the prompt;
performing entity extraction on the prompt; and
performing classification on the prompt;
providing the prompt to an embedding module to generate an embedding for the prompt;
retrieving one or more embeddings from a vector database based on the embedding for the prompt and based on a result of the preprocessing of the prompt;
generating an augmented prompt based on the one or more embeddings and the embedding for the prompt;
generating a response to the augmented prompt using a machine learning model and using a context of the one or more embeddings as a context for the machine learning model; and
providing the response to the user at the endpoint.
2 . The method of claim 1 , further comprising:
receiving a corpus of data;
chunking the corpus of data to produce a plurality of data chunks;
generating a plurality of embeddings for the plurality of data chunks; and
storing the plurality of embeddings in the vector database.
3 . The method of claim 2 , further comprising:
determining a segment of the corpus of data corresponds to a conversation; and
determining a context of the conversation; wherein
chunking the corpus of data comprises
generating a data chunk corresponding to the segment, and
generating the plurality of embeddings by generating an embedding for the segment that includes metadata defining the context of the conversation as a context for the embedding.
4 . The method of claim 2 , wherein the corpus of text is chunked using a configurable algorithmic delimiter.
5 . The method of claim 1 , further comprising:
performing, on an embedding of the one or more embeddings, at least one of an insertion operation, an update operation, or a deletion operation.
6 . The method of claim 1 , wherein retrieving one or more embeddings from a vector database comprises:
performing a semantic search using a similarity algorithm and the embedding for the prompt to identify one or more similar embeddings in the vector database, the similar embeddings being similar to the embedding for the prompt; and
combining the similar embeddings and the embedding for the prompt to create an augmented embedding.
7 . The method of claim 6 , wherein the semantic search is performed using a filter or limited search space.
8 . The method of claim 1 , wherein the prompt comprises a selection from a test data set and the method further comprising validating the response using the test data set.
9 . The method of claim 1 , wherein the prompt is received from a user via a user interface of the endpoint, and wherein the response is displayed to the user on a display associated with the user interface.
10 . A system comprising:
a memory having executable instructions stored thereon;
an endpoint having a user interface associated with a display; and
one or more processors configured to execute the executable instructions to cause the system to perform a method comprising:
receiving a prompt from a user at the user interface;
providing the prompt to an embedding module to generate an embedding for the prompt;
retrieving one or more embeddings from a vector database based on the embedding for the prompt, wherein the retrieving the one or more embeddings comprises performing a semantic search using a similarity algorithm and the embedding for the prompt to identify one or more similar embeddings in the vector database, the similar embeddings being similar to the embedding for the prompt;
generating an augmented prompt based on the one or more embeddings and the embedding for the prompt, wherein the generating of the augmented prompt comprises combining the similar embeddings and the embedding for the prompt to create an augmented embedding;
generating a response to the augmented prompt using a machine learning model and using a context of the one or more embeddings as a context for the machine learning model; and
providing the response to the user on the display.
11 . The system of claim 10 , the method further comprising:
receiving a corpus of data;
chunking the corpus of data to produce a plurality of data chunks;
generating a plurality of embeddings for the plurality of data chunks; and
storing the plurality of embeddings in the vector database.
12 . The system of claim 11 , the method further comprising:
determining a segment of the corpus of data corresponds to a conversation; and
determining a context of the conversation; wherein
chunking the corpus of data comprises
generating a data chunk corresponding to the segment, and
generating the plurality of embeddings by generating an embedding for the segment that includes metadata defining the context of the conversation as a context for the embedding.
13 . The system of claim 11 , wherein the corpus of text is chunked using a configurable algorithmic delimiter.
14 . The system of claim 10 , the method further comprising:
preprocessing the prompt by
performing summarization on the prompt;
performing entity extraction on the prompt; and
performing classification on the prompt; and
retrieving the embedding from the vector database based on a result of preprocessing the prompt.
15 . The system of claim 10 , the method further comprising:
performing, on an embedding of the one or more embeddings, at least one of an insertion operation, an update operation, or a deletion operation.
16 . The system of claim 10 , wherein the semantic search is performed using a filter or limited search space.
17 . The system of claim 10 , wherein the prompt comprises a selection from a test data set and the method further comprising validating the response using the test data set.
18 . A non-transitory computer readable storage medium comprising instructions, that when executed by one or more processors of a computing system, cause the computing system to perform a method comprising:
receiving a corpus of data;
determining a segment of the corpus of data corresponds to a conversation;
determining a context of the conversation;
chunking the corpus of data to produce a plurality of data chunks, wherein the chunking of the corpus of data comprises generating a data chunk corresponding to the segment;
generating a plurality of embeddings for the plurality of data chunks, wherein the generating of the plurality of embeddings comprises generating an embedding for the segment that includes metadata defining the context of the conversation as a context for the embedding;
storing the plurality of embeddings in a vector database;
receiving a prompt from a user at an endpoint;
providing the prompt to an embedding module to generate an embedding for the prompt;
retrieving one or more embeddings from the vector database based on the embedding for the prompt;
generating an augmented prompt based on the one or more embeddings and the embedding for the prompt;
generating a response to the augmented prompt using a machine learning model and using a context of the one or more embeddings as a context for the machine learning model; and
providing the response to the user at the endpoint.