Verification and citation for language model outputs
A user provides a question to be answered from detailed, dense or otherwise complex documents to a processing system that converts the question to a structured query language query and generates an embedding from the question, augmented by temporal data, synopses, themes, or other relevant information or data. The embedding is compared to embeddings generated from documents of a knowledge base to identify documents that are relevant to the question, and to rank such documents for their relevance. Highly ranking documents are combined with the query and provided to a language model that returns an answer to the question. A source for the answer is identified in at least one of the documents. The answer and the identified documents are presented to the user.
1 . A system operating on a provider network, wherein the system comprises:
one or more computer processors;
at least one data store having a knowledge base comprising a plurality of documents stored thereon;
a language model executed by the one or more computer processors, wherein the language model is trained based at least in part on a corpus of data, and wherein the corpus of data does not include any of the plurality of documents; and
a service executed by the one or more computer processors, wherein the service is configured to perform operations comprising:
receiving at least a free-form description of a question from a user;
converting the free-form description of the question to a first query;
determining a comparison of an embedding generated based at least in part on the first query to each of a plurality of embeddings, wherein each of the plurality of embeddings is generated based on one of a plurality of text chunks extracted from at least one document in a selected domain and augmented with temporal metadata, and wherein the embedding and each of the plurality of embeddings is in a common vector space;
selecting at least a subset of the plurality of text chunks based at least in part on a similarity analysis;
generating, using a large language model, a first response to the question based at least in part on the subset of the plurality of text chunks, wherein the first response comprises a first set of data points identified based at least in part on the temporal metadata;
identifying one of the plurality of documents including at least one of the first set of data points identified based at least in part on the temporal metadata; and
providing the first response to the question and the one of the plurality of documents to the user.
2 . The system of claim 1 , wherein the operations further comprise:
partitioning text extracted from the plurality of documents into the plurality of text chunks, wherein each one of the plurality of documents is in the selected domain;
augmenting the plurality of text chunks with the temporal metadata; and
indexing the plurality of text chunks augmented with the temporal metadata,
wherein each of the plurality of embeddings is generated based on one of the indexed plurality of text chunks augmented with the temporal metadata.
3 . The system of claim 1 , wherein the operations further comprise:
generating a second query based at least in part on a portion of the first response to the question comprising at least one of the first set of data points;
generating a second response to the second query, wherein the second response identifies a document including a second set of data points;
determining a comparison of the second set of data points to the first set of data points; and
verifying an accuracy of the first set of data points based at least in part on the comparison.
4 . A method comprising:
receiving at least a free-form description of a question;
converting the free-form description of the question into a query;
determining comparisons of an embedding generated based at least in part on the query to each of a plurality of embeddings, wherein each of the plurality of embeddings represents at least one of a plurality of text chunks extracted from a plurality of documents and augmented with temporal metadata, and wherein the embedding and each of the plurality of embeddings is in a common vector space;
selecting at least a subset of the plurality of text chunks based at least in part on the comparisons;
generating, using a large language model, a response to the question based at least in part on the subset of the plurality of text chunks, wherein the response comprises a set of data points identified based at least in part on the temporal metadata;
identifying one of the plurality of documents including at least one data point of the set of data points; and
providing the response to the question and at least an identifier of the one of the plurality of documents.
5 . The method of claim 4 , further comprising:
identifying a portion of the response including the at least one data point;
generating a structured query language query based at least in part on the portion of the response;
executing the structured query language query on the plurality of documents;
identifying one of the plurality of documents based at least in part on a response to the structured query language query; and
determining that a portion of the one of the plurality of documents includes the at least one data point.
6 . The method of claim 4 , further comprising:
identifying a portion of the response including one of the set of data points;
generating a structured query language query based at least in part on the portion of the response;
executing the structured query language query on the plurality of documents, wherein at least the portion of the at least one of the plurality of documents is identified based at least in part on a response to the structured query language query;
determining that the at least one data point differs from the one of the set of data points; and
substituting the at least one data point for the one of the set of data points in the response.
7 . The method of claim 4 , wherein identifying the one of the plurality of documents comprises:
executing a regular expression function on the response and the plurality of documents,
wherein the one of the plurality of documents is identified by the regular expression function.
8 . The method of claim 4 , wherein receiving at least the free-form description comprises:
receiving at least the free-form description and at least one document,
wherein generating the response to the query comprises:
providing at least the query and the at least one document as inputs to the large language model,
wherein the response is generated based at least in part on an output received from the large language model in response to the inputs.
9 . The method of claim 4 , further comprising:
converting the free-form description of the question to the query in a structured query language by a query generator module.
10 . The method of claim 4 , wherein receiving at least the query comprises:
calculating at least one quality score based at least in part on the query;
determining that the query is not optimal based at least in part on the at least one quality score;
prompting a user to provide additional information regarding the query;
receiving the additional information regarding the query; and
updating the query to include at least some of the additional information,
wherein the subset of the plurality of text chunks is identified based at least in part on the updated query.
11 . The method of claim 4 , further comprising:
determining contextual information regarding the query; and
augmenting the query to include at least some of the contextual information,
wherein the contextual information comprises at least one of:
a time associated with the query;
a theme associated with the query; or
a summary of the query, and
wherein the embedding is generated based at least in part on the augmented query.
12 . The method of claim 4 , wherein the one of the plurality of documents comprises at least one of:
a financial record;
a legal opinion; or
a medical report.
13 . The method of claim 4 , wherein determining the comparisons comprises:
generating the plurality of embeddings, wherein each one of the plurality of embeddings is generated based on one of the plurality of text chunks; and
performing a similarity analysis on each one of the plurality of embeddings based at least in part on the embedding generated based at least in part on the query,
wherein each one of the comparisons is determined based at least in part on the similarity analysis.
14 . The method of claim 13 , further comprising:
determining a ranking of at least the plurality of text chunks based at least in part on the comparisons,
wherein the subset of the plurality of text chunks is identified based at least in part on the ranking.
15 . The method of claim 4 , further comprising:
partitioning text extracted from the plurality of documents into the plurality of text chunks;
augmenting the plurality of text chunks with the temporal metadata; and
indexing the augmented plurality of text chunks in a database.
16 . The method of claim 4 , wherein at least one of the plurality of documents is associated with a user, and
wherein the response to the question and the identifier of the one of the plurality of documents are provided to the user.
17 . A non-transitory computer-readable storage medium having executable instructions stored thereon that, when executed by one or more processors of a computer system, cause the computer system to at least:
receive at least a query;
generate a response to the query based at least in part on a plurality of text chunks extracted from a corpus of documents, wherein the plurality of text chunks are augmented with temporal metadata and indexed in a database, and wherein the response includes a first set of quantitative data points identified based at least in part on the temporal metadata;
identify a second set of quantitative data points based at least in part on the response, wherein each one of the second set of quantitative data points corresponds to one of the first set of quantitative data points;
compare a first quantitative data point of the first set of quantitative data points to a second quantitative data point of the second set of quantitative data points; and
provide at least a portion of the response and an identifier of a source of a quantitative data point in the portion of the response,
wherein the quantitative data point in the portion of the response is the first quantitative data point and the source is one of the corpus of documents including the first quantitative data point if the first quantitative data point equals the second quantitative data point, and
wherein the quantitative data point in the portion of the response is the second quantitative data point and the source is one of the corpus of documents including the second quantitative data point if the first quantitative data point does not equal the second quantitative data point.
18 . The non-transitory computer-readable storage medium of claim 17 , wherein the executable instructions, when executed by the one or more processors of the computer system, further cause the computer system to at least:
generate a structured query language query based at least in part on the portion of the response; and
identify a response to the structured query language query based at least in part on the corpus of documents, wherein the response to the structured query language query includes the second set of quantitative data points.
19 . The non-transitory computer-readable storage medium of claim 17 , wherein the executable instructions, when executed by the one or more processors of the computer system, further cause the computer system to at least:
generate an embedding based at least in part on the query;
identify a plurality of embeddings, wherein each one of the plurality of embeddings is generated based on one of the plurality of text chunks; and
perform a similarity analysis on each of the plurality of embeddings based at least in part on the embedding,
wherein the response to the query is generated based at least in part on a subset of the plurality of text chunks identified based at least in part on the similarity analysis.
20 . The non-transitory computer-readable storage medium of claim 17 , wherein the executable instructions, when executed by the one or more processors of the computer system, further cause the computer system to at least:
receive at least one document; and
provide at least the query and the at least one document as inputs to a large language model,
wherein the response is generated based at least in part on an output received from the large language model in response to the inputs.