Reference and training data validation for machine learning models
Systems and methods are provided for a prompt and content generation service to validate generated references and training data sets of large language models (LLMs). The prompt and content generation service may determine if a reference generated by an LLM is valid by searching a training dataset for tokens using a prompt, generated answer, and generated reference for content and references in the training data set. If the prompt and content generation service finds content or a valid reference with a certain confidence level, the prompt and content generation service may indicate that the generated content and reference is valid. The prompt and content generation service may additionally validate a saved checkpoint during training of an LLM by executing use case scenarios against the LLM checkpoint and determining if expected answers match generated answers.
1 . A computing device for validating content and references generated by large language models (LLMs), the computing device comprising:
computer-readable memory storing executable instructions; and
at least one computing device in communication with the computer-readable memory and programmed by the executable instructions to:
receive a prompt with a natural language question, image, or sound;
generate, via an LLM, content and a reference in response to the prompt, wherein the LLM was trained using training data, and wherein the training data comprise content and predetermined references associated with the content;
determine a characterization of matching by searching the training data to identify matches between the (i) content and the reference and (ii) the training data, wherein the characterization of matching comprises a full match, partial match, and no match; and
if there was a full match:
return the generated content and the reference, as an output to the prompt, wherein the output comprises a notification that the reference is associated with attribution data of the training data;
if there was a partial match:
determine a confidence score based on a degree of matching between the (i) content and the reference and (ii) the training data;
return the content, the reference, and the confidence score as an output to the prompt, wherein the output comprises a notification that the reference was a partial match; and
if there was no match:
return the generated content and reference as an output to the prompt, wherein the output comprises a notification that the generated content and reference could not be determined to be associated with attribution data of the training data.
2 . The computing device of claim 1 , wherein the at least one computing device is further programmed by the executable instructions to:
if there was no match:
regenerate the content and reference up to a predetermined number of times until a full or partial match is found, wherein if a full match or partial match is found:
return the regenerated content and reference as an output to the prompt, wherein the output comprises a notification that the reference was a partial match or full match.
3 . The computing device of claim 1 , wherein the predetermined references are identified by identifier tokens within the training data, and wherein an identifier token of the identifier tokens is positioned directly before, or directly after, a predetermined reference within the training data.
4 . The computing device of claim 1 , wherein the content is identified by identifier tokens within the training data, and wherein an identifier token of the identifier tokens is positioned directly before, or directly after, content within the training data.
5 . A computer-implemented method comprising:
receiving a prompt;
causing generation, via a large language model (LLM), of output in response to the prompt;
identifying, in the output, a first portion corresponding to content and a second portion corresponding to a reference;
determining a degree of matching between the reference included within the output and one or more predetermined references included within training data for the LLM by searching the training data to identify a degree of matching between the reference included within the output and the one or more predetermined references included within the training data;
generating attribution information indicating a confidence value for the output representing the degree of matching between the reference included within the output and one or more predetermined references included within training data for the LLM; and
returning the output, including the content and the reference, with an indication of the attribution information indicating the confidence value for the output representing the degree of matching between the reference included within the output and one or more predetermined references included within training data for the LLM in response to the prompt.
6 . The computer-implemented method of claim 5 , wherein the matching is characterized as at least one of a full match, partial match, or no match.
7 . The computer-implemented method of claim 5 , wherein the generated attribution information is associated with a uniform resource locator (URL) of a website, and wherein the output is generated using content from the website.
8 . The computer-implemented method of claim 5 , wherein returning the output with the indication of the attribution information includes prior to returning the output, replacing at least a portion of the output with at least a portion of the training data based on a threshold matching.
9 . The computer-implemented method of claim 5 , wherein the degree of matching is determined based on a search for identifier tokens within the training data.
10 . The computer-implemented method of claim 5 , wherein the training data comprises content associated with identifier tokens and content not associated with identifier tokens.
11 . The computer-implemented method of claim 5 , wherein the content are identified by identifier tokens within the training data, and wherein an identifier token is positioned directly before or directly after content within the training data.
12 . The computer-implemented method of claim 5 , further comprising presenting the output on a user interface (UI), wherein different graphical visual indications on the user interface are based on a full match, partial match, or no match.
13 . The computer-implemented method of claim 12 , wherein the different graphical visual indications comprise a color attribute representative of a type of matching.
14 . The computer-implemented method of claim 12 , wherein the different graphical visual indications comprise a color attribute representative of a degree of matching.
15 . One or more non-transitory computer-readable media storing specific computer-executable instructions that, when executed by a computing system comprising a processor, cause the computing system to at least:
identify, in an output from a large language model (LLM), a first portion corresponding to content and a second portion corresponding to a reference;
determine a degree of matching between the reference included within the output and one or more predetermined references included within training data for the LLM by searching the training data to identify a degree of matching between the reference included within the output and the one or more predetermined references included within the training data;
generate, attribution information indicating a confidence value for the output representing the degree of matching between the reference included within the output and one or more predetermined references included within training data for the LLM; and
return the output, including the content and the reference, with an indication of the attribution information indicating the confidence value for the output representing the degree of matching between the reference included within the output and one or more predetermined references included within training data for the LLM.
16 . The one or more non-transitory computer-readable media of claim 15 , wherein the matching is characterized as at least one of a full match, partial match, or no match.
17 . The one or more non-transitory computer-readable media of claim 15 , wherein the attribution information comprises of (i) a uniform resource locator (URL) of a website containing content used to generate output and (ii) a date when the URL was accessed.
18 . The one or more non-transitory computer-readable media of claim 15 , wherein prior to returning the output, at least a portion of the output is replaced by the training data based on a threshold matching.
19 . The one or more non-transitory computer-readable media of claim 15 , wherein the output is matched to the one or more predetermined references included within the training data based on a search for identifier tokens within the training data.