Enriching language model input with contextual data
Various embodiments discussed herein are directed to improving existing technologies by providing a corpus data supplement as input into a model, such as a Large Language Model (LLM). Consequently, the model can generate accurate scores or data for predictions because the model is better able to distinguish between a general understanding of natural language concepts and domain-specific concepts.
1 . A system comprising:
at least one computer processor; and
one or more computer storage media storing computer-useable instructions that, when used by the at least one computer processor, cause the at least one computer processor to perform operations comprising:
receiving a corpus of text;
determining contextual data associated with the corpus of text;
based on the contextual data, determining a corpus data supplement, the corpus data supplement being at least one of: data to be added within the corpus of text as input into a trained machine learning model or data that supplements the corpus of text as input into the trained machine learning model, wherein determining the corpus data supplement comprises determining a first portion, among other portions, of the contextual data to be included in the corpus data supplement by determining contextual data from more relevant to less relevant until a token threshold is met, the token threshold corresponding to an input size constraint of the trained machine learning model, wherein the other portions are excluded from the corpus data supplement based on the token threshold having been met and being less relevant relative to the first portion;
based on the determining of the corpus data supplement, providing the corpus of text and the corpus data supplement as input into the trained machine learning model; and
based on the providing of the corpus of text and the corpus data supplement as input into the trained machine learning model, causing the machine learning model to perform an output that is based on the corpus data supplement.
2 . The system of claim 1 , wherein the corpus of text includes a meeting transcript that includes natural language characters indicating content spoken in a meeting, and wherein the determining of the contextual data comprises:
determining metadata associated with the meeting, and wherein the metadata includes at least one of: attendees of the meeting, date of the meeting, and agenda of the meeting.
3 . The system of claim 1 , wherein the corpus of text includes a file, and wherein the determining of the contextual data comprises at least one of:
identifying a network graph or user profile of a person who authored or commented about the file; and
determining a file name or file type associated with the file.
4 . The system of claim 1 , wherein the corpus of text includes one or more messages, and wherein the determining of the contextual data comprises at least one of:
determining a date that a message was sent or received;
determining a recipient or sender of the message; and
determining an attachment associated with the message.
5 . The system of claim 1 , wherein the corpus data supplement comprises at least one of: a portion of the contextual data or entity data enrichment data that is data associated with an entity.
6 . The system of claim 1 , wherein the determining of the corpus data supplement comprises:
detecting an entity within the corpus of text by scanning the corpus of text;
classifying the entity into an entity type;
based on the entity type, determining entity data enrichment for the entity based on the contextual data; and
adding or associating the entity data enrichment with the entity or the corpus of text as the corpus data supplement.
7 . The system of claim 1 , wherein the determining of the corpus data supplement comprises:
detecting a plurality of entities within the corpus of text; and
scoring each entity, of the plurality of entities, according to a relevance of the entity to the contextual data.
8 . The system of claim 7 , wherein the operations further comprise:
for each entity, of the plurality of entities, whose score exceeds a relevance threshold:
determining an entity type;
determining whether to enrich the entity based on the entity type and a size of at least one of: the corpus of text, a portion of the contextual data, and an entity data enrichment; and
based on determining to enrich the entity, determining an entity data enrichment data based on the contextual data, the entity data enrichment data being included in the corpus data supplement.
9 . The system of claim 1 , wherein the operations further comprise:
receiving or determining an input size constraint of the machine learning model, wherein the determining of the corpus data supplement is based on the input size constraint.
10 . The system of claim 1 , wherein the output includes at least one of: sentiment analysis, answering one or more questions, automatic summarization, text generation, machine translation, or document classification.
11 . The system of claim 1 , wherein the machine learning model is pre-trained without having been fine-tuned.
12 . The system of claim 1 , wherein the operations further comprise:
receiving a user query;
in response to the receiving of the user query, scanning the corpus of text and the corpus data supplement;
based on the scanning and the providing of the corpus of text and the corpus data supplement as input into the trained machine learning model, causing the query to be executed such that one or more results for the query are returned.
13 . A computer-implemented method comprising:
receiving a corpus of text;
determining contextual data associated with the corpus of text, wherein determining the corpus data supplement comprises:
detecting a plurality of entities within the corpus of text by scanning the corpus of text;
scoring each entity of the plurality of entities according to a relevance of the entity to the contextual data;
ranking each entity of the plurality of entities based on the score; and
for each entity of the plurality of entities:
based on the ranking, determining an entity type;
determining whether to enrich the entity based on the entity type and a size of at least one of: the corpus of text, a portion of the contextual data, or an entity data enrichment; and
based on determining to enrich the entity, determining an entity data enrichment data based on the contextual data;
based on the contextual data, determining a corpus data supplement;
subsequent to the determining of the corpus data supplement, receiving a user query;
in response to the receiving of the user query, providing the user query, the corpus of text, and the corpus data supplement as input into a machine learning model; and
based on the providing, causing the user query to be executed such that one or more results for the query are returned, wherein the one or more results are based at least on the corpus data supplement.
14 . The computer-implemented method of claim 13 , wherein the corpus of text includes at least one of: a set of chat messages or a meeting transcript that includes natural language characters indicating content spoken during a meeting, and wherein the determining of the contextual data comprises:
determining metadata associated with the meeting or the set of chats and wherein the metadata includes at least one of: attendees of the meeting, date of the meeting, agenda of the meeting, participants in a chat session associated with the set of chat messages, and a date that the chat messages were sent.
15 . The computer-implemented method of claim 13 , wherein the determining of the corpus data supplement comprises:
detecting an entity within the corpus of text;
classifying the entity into an entity type;
based on the entity type, determining entity data enrichment for the entity based on the contextual data; and
adding or associating the entity data enrichment with the entity or the corpus of text as the corpus data supplement.
16 . The computer-implemented method of claim 13 , wherein the machine learning model is pre-trained without having been fine-tuned.
17 . One or more computer storage media having computer-executable instructions embodied thereon that, when executed, by one or more processors, cause the one or more processors to perform operations comprising:
receiving a corpus of text;
determining contextual data associated with the corpus of text;
based on the contextual data, determining a corpus data supplement, the corpus data supplement being at least one of: data to be added within the corpus of text or data to be supplemented with the corpus of text as input into a language model, wherein determining the corpus data supplement comprises determining a first portion, among other portions, of the contextual data to be included in the corpus data supplement by determining contextual data from more relevant to less relevant until a token threshold is met, the token threshold corresponding to an input size constraint of the trained machine learning model, wherein the other portions are excluded from the corpus data supplement based on the token threshold having been met and being less relevant relative to the first portion;
based on the determining of the corpus data supplement, causing the corpus of text and the corpus data supplement to be used as input into the language model; and
receiving, from the language model, an output that is based on the corpus data supplement.
18 . The one or more computer storage media of claim 17 , wherein the operations further comprising:
based on the corpus of text and the corpus data supplement being used as input into the language model, causing the language model to perform at least one of: sentiment analysis, answering one or more questions, automatic summarization, text generation, machine translation, or document classification.
19 . The one or more computer storage media of claim 17 , wherein the corpus of text includes a meeting transcript that includes natural language characters indicating content spoken in a meeting.
20 . The one or more computer storage media of claim 19 , wherein the determining of the contextual data comprises:
determining metadata associated with the meeting, wherein the metadata includes at least one of: attendees of the meeting, a date of the meeting, or an agenda of the meeting.