Distilling language models
One example method includes selecting an entry from a validation set based on similarity to a first communication record, wherein the entry from the validation set includes a second communication record and a corresponding summary of the second communication record; generating a prompt that includes the first communication record, the second communication record, and the corresponding summary of the second communication record; inputting the prompt to a first language model to obtain a generated summary of the first communication record; and training a second language model based on the first communication record and the generated summary.
1 . A method comprising:
fine-tuning a small language model, comprising:
accessing a set of communication records and a set of validation records;
for each communication record in the set of communication records, generating a first embedding representing the respective communication record;
for each validation record in the set of communication records, generating a second embedding representing the respective validation record;
for each communication record:
determining a similarity between the respective first embedding and each generated second embedding;
selecting a validation record from the set of validation records based on the determined similarities; and
generating, using a large language model (“LLM”), a summary of the respective communication record comprising inputting the respective communication record and the validation record into the LLM; and
training the small language model comprising prompting the small language model to generate a summary of a first communication record and providing the first communication record and the generated summary as a prompt/completion pair.
2 . The method of claim 1 , wherein the similarity is determined based on a cosine distance between the respective first embedding and the respective second embedding.
3 . The method of claim 1 , further comprising generating an unannotated training data set based on similarities between communication records in a superset of communication records and communication records in the set of validation records comprising:
for each communication record in the set of validation records:
determining a similarity between the communication record in the set of validation records and each communication record in the superset of communication records, and
selecting the set of communication records from the superset of communication records based on the determined similarities; and
wherein the respective communication record is obtained from the unannotated training data set.
4 . The method of claim 3 , wherein the similarity between the respective communication record in the set of validation records and each communication record in the superset of communication records is determined based on a cosine distance between an embedding generated from the respective communication record and embeddings generated from the communication records in the superset of communication records.
5 . The method of claim 3 , further comprising:
generating a summary of each communication record in the unannotated training data set based on similarities to one or more entries in the set of validation records; and
generating an annotated training set based on the generated summaries and the corresponding communication records.
6 . The method of claim 5 , further comprising
determining a Shannon score for each generated summary based on the respective communication record; and
removing one or more communication records and corresponding summaries from the annotated training set based on generated summaries having Shannon scores that do not satisfy a Shannon score threshold to generate a filtered annotated training set.
7 . The method of claim 6 , further comprising training the small language model using the annotated training set.
8 . The method of claim 1 , further comprising:
selecting a second validation record from the set of validation records based on similarity to the respective communication record, wherein the second validation record from the set of validation records includes a third communication record and a corresponding summary of the third communication record; and
wherein generating, using the LLM, the summary further comprises inputting the third communication record and the corresponding summary of the third communication record to the LLM.
9 . A system comprising:
a non-transitory computer-readable medium; and
one or more processors communicatively coupled to the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:
fine-tune a small language model, the one or more processors configured to execute processor-executable instructions configured to cause one or more processors to:
access a set of communication records and a set of validation records;
for each communication record in the set of communication records, generate a first embedding representing the respective communication record;
for each validation record in the set of communication records, generate a second embedding representing the respective validation record;
for each communication record:
determine a similarity between the respective first embedding and each generated second embedding;
select a validation record from the set of validation records based on the determined similarities; and
generate, using a large language model (“LLM”), a summary of the respective communication record comprising inputting the respective communication record and the validation record into the LLM;
and
train the small language model comprising prompting the small language model to generate a summary of a first communication record and providing the first communication record and the generated summary as a prompt/completion pair.
10 . The system of claim 9 , wherein the similarity is determined based on a cosine distance between the respective first embedding and the respective second embedding.
11 . The system of claim 9 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:
generate an unannotated training data set based on similarities between communication records in a superset of communication records and communication records in the set of validation records;
for each communication record in the set of validation records:
determine a similarity between the communication record in the set of validation records and each communication record in the set of communication records, and
select the set of communication records from the superset of communication records based on the determined similarities; and
wherein the respective communication record is obtained from the unannotated training data set.
12 . The system of claim 11 , wherein the similarity between the respective communication record in the set of validation records and each communication record in the superset of communication records is determined based on a cosine distance between an embedding generated from the respective communication record and embeddings generated from the communication records in the superset of communication records.
13 . The system of claim 11 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:
generate a summary of each communication record in the unannotated training data set based on similarities to one or more entries in the set of validation records; and
generate an annotated training set based on the generated summaries and the corresponding communication records.
14 . The system of claim 13 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:
determine a Shannon score for each generated summary based on the respective communication record; and
remove one or more communication records and corresponding summaries from the annotated training set based on generated summaries having Shannon scores that do not satisfy a Shannon score threshold to generate a filtered annotated training set.
15 . The system of claim 14 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to train the small language model using the annotated training set.
16 . The system of claim 9 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:
select a second validation record from the set of validation records based on similarity to the respective communication record, wherein the second validation record from the set of validation records includes a third communication record and a corresponding summary of the third communication record; and
wherein generating, using the LLM, the summary further comprises inputting the third communication record and the corresponding summary of the third communication record to the LLM.
17 . A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:
fine-tune a small language model, comprising processor-executable instructions configured to cause one or more processors to:
access a set of communication records and a set of validation records;
for each communication record in the set of communication records, generate a first embedding representing the respective communication record;
for each validation record in the set of communication records, generate a second embedding representing the respective validation record;
for each communication record:
determine a similarity between the respective first embedding and each generated second embedding;
select a validation record from the set of validation records based on the determined similarities; and
generate, using a large language model (“LLM”), a summary of the respective communication record comprising inputting the respective communication record and the validation record into the LLM; and
train the small language model comprising prompting the small language model to generate a summary of a first communication record and providing the first communication record and the generated summary as a prompt/completion pair.
18 . The non-transitory computer-readable medium of claim 17 , further comprising processor-executable instructions configured to cause the one or more processors to:
generate an unannotated training data set based on similarities between communication records in a superset of communication records and communication records in the set of validation records;
for each communication record in the set of validation records:
determine a similarity between the communication record in the set of validation records and each communication record in the set of communication records, and
select the set of communication records from the superset of communication records based on the determined similarities; and
wherein the respective communication record is obtained from the unannotated training data set.
19 . The non-transitory computer-readable medium of claim 18 , further comprising processor-executable instructions configured to cause the one or more processors to:
generate a summary of each communication record in the unannotated training data set based on similarities to one or more entries in the set of validation records; and
generate an annotated training set based on the generated summaries and the corresponding communication records.
20 . The non-transitory computer-readable medium of claim 19 , further comprising processor-executable instructions configured to cause the one or more processors to:
determine a Shannon score for each generated summary based on the respective communication record; and
remove one or more communication records and corresponding summaries from the annotated training set based on generated summaries having Shannon scores that do not satisfy a Shannon score threshold to generate a filtered annotated training set.