Real-time summarization large language model
The present disclosure generally relates to systems and methods for generating and evaluating a transcript summary based on a transcript of an active meeting. A summary generation system may generate a summary based on an active meeting while minimizing latency and the number of calls to a large language model. The summary generation system may provide each transcript chunk of a transcript as input into an LLM for summarization and append to a transcript summary. In addition, the summary generation system may determine whether the transcript summary of the transcript has reached a threshold, such as a length or word count threshold. If the threshold has been reached, the summary generation system may input the transcript summary into the LLM to generate a revised transcript summary. The revised transcript summary may replace the existing transcript summary.
1 . A system, comprising:
a memory to store computer-executable instructions; and
a processor in communication with the memory, wherein the computer-executable instructions, when executed by the processor, cause the processor to at least:
receive a transcript chunk of a transcript of an active meeting;
provide the transcript chunk and a first prompt as input into a machine learning model;
determine, using the machine learning model, a chunk summary corresponding to the transcript chunk;
append the chunk summary to a transcript summary of the transcript of the active meeting to form a revised transcript summary;
determine that a summary threshold has been reached based on the revised transcript summary;
provide the revised transcript summary and a second prompt as input into the machine learning model in response to the summary threshold being reached;
determine, using the machine learning model, a new transcript summary; and
replace the revised transcript summary with the new transcript summary.
2 . The system of claim 1 , wherein the transcript of the active meeting comprises a text-based transcript of audio, a conversation, a video conference, an audio conference, a phone call, a video, an audio file, a recording, or a communication.
3 . The system of claim 1 , wherein the machine learning model includes a large language model.
4 . The system of claim 1 , wherein the first prompt includes instructions to generate the chunk summary of the transcript chunk based on a word count, a sentence count, a character count, a file size, a topic, or a criteria.
5 . The system of claim 1 , wherein the second prompt includes instructions to generate the new transcript summary based on a word count, a sentence count, a character count, a file size, a topic, or a criteria.
6 . The system of claim 1 , wherein the summary threshold includes one of a word count, a sentence count, or a character count of the revised transcript summary.
7 . The system of claim 1 , wherein the transcript chunk is a portion of the transcript of the active meeting up to a checkpoint.
8 . The system of claim 7 , wherein the checkpoint corresponds to a word count of the transcript, a character count of the transcript, a sentence count of the transcript, a timestamp, or a time interval.
9 . A computer implemented method comprising:
receiving a transcript chunk of a transcript;
providing the transcript chunk and a first prompt as input into a machine learning model;
determining, using the machine learning model, a chunk summary corresponding to the transcript chunk;
appending the chunk summary to a transcript summary to form a revised transcript summary;
determining that a summary threshold has been reached based on the revised transcript summary;
providing the revised transcript summary and a second prompt as input into the machine learning model in response to the summary threshold being reached;
determining, using the machine learning model, a new transcript summary; and
replacing the revised transcript summary with the new transcript summary.
10 . The method of claim 9 , wherein the transcript comprises a text-based transcript of a meeting, a conversation, a video conference, an audio conference, a phone call, a video, an audio file, a recording, or a communication.
11 . The method of claim 9 , wherein the machine learning model includes a large language model.
12 . The method of claim 9 , wherein the first prompt includes instructions to generate the chunk summary of the transcript chunk based on a word count, a sentence count, a character count, a file size, a topic, or a criteria.
13 . The method of claim 9 , wherein the second prompt includes a request to summarize the transcript summary based on a word count, a sentence count, a character count, a file size, a topic, or a criteria.
14 . The method of claim 9 , wherein the summary threshold includes one of a word count, a sentence count, or a character count of the revised transcript summary.
15 . The method of claim 9 , wherein the transcript chunk is a portion of the transcript of an active meeting up to a checkpoint.
16 . The method of claim 15 , wherein the checkpoint corresponds to a word count of the transcript, a character count of the transcript, a sentence count of the transcript, a timestamp, or a time interval.
17 . A non-transitory computer-readable medium storing computer-executable instructions that, when executed by a processor of a computing device, cause the computing device to at least:
provide a portion of a transcript as input into a machine learning model;
determine, using the machine learning model, a summary corresponding to the portion of the transcript;
determine that a summary threshold associated with the summary and a second summary generated subsequent to the summary has been reached;
provide the summary and the second summary as input into the machine learning model;
determine, using the machine learning model, a new transcript summary; and
replace the summary with the new transcript summary.
18 . The non-transitory computer-readable medium of claim 17 , wherein the transcript comprises a text-based transcript of a meeting, a conversation, a video conference, an audio conference, a phone call, a video, an audio file, a recording, or a communication.
19 . The non-transitory computer-readable medium of claim 17 , wherein the processor is further to provide the portion of the transcript and a prompt into the machine learning model, wherein the prompt includes a request to summarize the portion of the transcript based on a word count, a sentence count, a character count, a file size, a topic, or a criteria.
20 . The non-transitory computer-readable medium of claim 17 , wherein the processor is further to apply the summary, the second summary, and a prompt into the machine learning model, wherein the prompt includes a request to summarize the summary and the second summary based on a word count, a sentence count, a character count, a file size, a topic, or a criteria.