Summarization based on timing data
A method performed by a computing system comprises generating text from audio data and determining an end portion of the text to include in a summarization of the text based on a length of a portion of the audio data from which the text was generated and which ends with a proposed end portion and a time value associated with the proposed end portion, the proposed end portion including a word from the text.
1 . A method performed by a computing system, the method comprising:
generating text from audio data received by the computing system via a microphone;
determining a value associated with an end portion of the text to include in a summary of the text based on:
a length of a portion of the audio data from which the text was generated and which ends with the end portion; and
a time value associated with the end portion, the end portion including a word from the text;
determining, based on the value satisfying a condition, to summarize the portion of the audio data from which the text was generated and ends with the end portion; and
generating, by a model, the summary of the text based on the portion of the audio data from which the text was generated and ends with the end portion.
2 . The method of claim 1 , wherein the length of the audio data from which the text was generated and which ends with the end portion includes a time duration of the portion of the audio data.
3 . The method of claim 1 , wherein the length of the audio data from which the text was generated and which ends with the end portion includes a number of words included in the text transcribed from the portion of the audio data.
4 . The method of claim 1 , wherein the length of the audio data from which the text was generated and which ends with the end portion is based on the text transcribed from the portion of the audio data.
5 . The method of claim 1 , wherein the time value associated with the end portion includes a duration of a pause after the end portion.
6 . The method of claim 1 , wherein the time value associated with the end portion includes a duration of time between the end portion and a subsequent portion of the text that immediately follows the end portion.
7 . The method of claim 1 , wherein the determination of the value is further based on a punctuation mark included in the text, the punctuation mark immediately following the end portion.
8 . The method of claim 1 , wherein the determination of the value is further based on a determination that the end portion was spoken by a first person, and a subsequent portion that immediately follows the end portion was spoken by a second person, the second person being different than the first person.
9 . The method of claim 1 , wherein the determination of the value is further based on a determination that the text that is unsummarized and ends with the end portion is related to a first topic and that text that is subsequent to the end portion is related to a second topic, the first topic being different than the second topic.
10 . The method of claim 1 , wherein the determination of the value is further based on a low confidence level of transcribing speech subsequent to the text that is unsummarized and ends with the end portion.
11 . The method of claim 1 , wherein the computing system is a head-mounted device.
12 . The method of claim 1 , further comprising presenting the summarized text on a display.
13 . The method of claim 1 , further comprising:
determining, by the model, a specific term based on a general term and contextual data, the general term being included in the summary of the text, the contextual data including information associated with a user other than the text generated based on the audio data, the specific term including additional information than was included in the text; and
generating an enhanced summary based on the summary of the text and the specific term.
14 . The method of claim 1 , wherein the model is a sequence-to-sequence generative large learning model.
15 . The method of claim 1 , wherein the summary of the text includes a hyperlink that includes an address of a webpage that presents information about a person, place, or thing referred to by the text.
16 . A method performed by a computing system, the method comprising:
generating text from audio data received by the computing system via a microphone;
determining, by a model, that a proposed end portion of the text is an end portion of the text based on a duration of a pause after the proposed end portion satisfying a pause duration threshold, the pause duration threshold being less for greater lengths of the text that end with the proposed end portion and the pause duration threshold being greater for lesser lengths of the text that end with the proposed end portion; and
generating, by the model, a summary of the text based on the portion of the audio data from which the text was generated and ends with the end portion, the summary of the text including a reference to a webpage that presents information about a person, place, or thing referred to by the text.
17 . The method of claim 16 , wherein the text that ends with the proposed end portion is unsummarized.
18 . The method of claim 16 , wherein:
the determining includes determining that the proposed end portion of the text is the end portion of the text based on the duration of the pause after the proposed end portion satisfying the pause duration threshold, and the method further includes the model summarizing the text based on the determining that the proposed end portion of the text is the end portion of the text.
19 . A method performed by a computing system, the method comprising:
storing contextual data, the contextual data including textual information associated with a user;
generating text based on audio data received by the computing system via a microphone associated with the user, the text including a general term;
searching the contextual data for a specific term to replace the general term, the specific term including additional information than was included in the text; and
generating, by a model, a summary based on the text and the specific term.
20 . The method of claim 19 , wherein the summary includes fewer words than the text.