Systems for controllable summarization of content
A method of generating summaries of content items using one or more large language models (LLMs) is disclosed. A first content item is identified. The first content item includes a set of sub-content items. A level of abstraction is determined for the content item. A prompt is automatically engineered for providing to the one or more LLMs. The prompt includes a reference to the first content item and the level of the abstraction for the first content item. A response to the prompt is received from the LLM. The response includes a second content item. The second content item includes a representation of the first content item that is generated by the LLM. The representation omits or simplifies one or more of the set of sub-content items based on the level of abstraction. The representation is used to control an output that is communicated to a target device.
1 . A system comprising:
one or more computer processors;
one or more computer memories;
a set of instructions stored in the one or more computer memories, the set of instructions configuring the one or more computer processors to perform operations, the operations comprising:
receiving a content corpus to be summarized;
generating a first summarized content by applying a large language model (LLM) or an appropriate processing model to the content corpus to create an initial summary at a first predefined abstraction level;
recursively generating subsequent summarized content by applying the large language model or the appropriate processing model to previously summarized content at progressively greater abstraction levels, wherein each subsequent summary is derived from last summarized content without reprocessing the content corpus; and
storing each of the subsequent summarized content at a plurality of abstraction levels, thereby facilitating efficient retrieval and use based on user-defined summarization needs.
2 . The system of claim 1 , wherein the content corpus comprises one or more of text, audio, video, and image content.
3 . The system of claim 1 , wherein the recursive generating of the subsequent summarized content utilizes previously summarized outputs to perform further summarizations.
4 . The system of claim 1 , wherein the plurality of abstraction levels for generating subsequent summarized content includes a plurality of specific percentages to support a tiered approach to content reduction.
5 . The system of claim 1 , the operations further comprising calculating a cost savings of the recursive generating of the subsequent summarized content compared to a technique where the content corpus is reprocessed for each summarization.
6 . The system of claim 1 , wherein the LLM or the appropriate processing model is configured to adjust a depth of summarization dynamically based on a complexity or length of the content corpus.
7 . The system of claim 1 , wherein the storing of each of the subsequent summarized content comprises indexing each of the subsequent summarized content based on its abstraction level to increase a speed of retrieval in one or more applications that require different levels of content detail.
8 . A method comprising:
receiving a content corpus to be summarized;
generating a first summarized content by applying a large language model (LLM) or an appropriate processing model to the content corpus to create an initial summary at a first predefined abstraction level;
recursively generating subsequent summarized content by applying the large language model or the appropriate processing model to previously summarized content at progressively greater abstraction levels, wherein each subsequent summary is derived from last summarized content without reprocessing the content corpus; and
storing each of the subsequent summarized content at a plurality of abstraction levels, thereby facilitating efficient retrieval and use based on user-defined summarization needs.
9 . The method of claim 8 , wherein the content corpus comprises one or more of text, audio, video, and image content.
10 . The method of claim 8 , wherein the recursive generating of the subsequent summarized content utilizes previously summarized outputs to perform further summarizations.
11 . The method of claim 8 , wherein the plurality of abstraction levels for generating subsequent summarized content includes a plurality of specific percentages to support a tiered approach to content reduction.
12 . The method of claim 8 , further comprising calculating a cost savings of the recursive generating of the subsequent summarized content compared to a technique where the content corpus is reprocessed for each summarization.
13 . The method of claim 8 , wherein the LLM or the appropriate processing model is configured to adjust a depth of summarization dynamically based on a complexity or length of the content corpus.
14 . The method of claim 8 , wherein the storing of each of the subsequent summarized content comprises indexing each of the subsequent summarized content based on its abstraction level to increase a speed of retrieval in one or more applications that require different levels of content detail.
15 . A non-transitory computer-readable storage medium storing a set of instructions that, when executed by one or more computer processors, causes the one or more computer processors to perform operations, the operations comprising:
receiving a content corpus to be summarized;
generating a first summarized content by applying a large language model (LLM) or an appropriate processing model to the content corpus to create an initial summary at a first predefined abstraction level;
recursively generating subsequent summarized content by applying the large language model or the appropriate processing model to previously summarized content at progressively greater abstraction levels, wherein each subsequent summary is derived from last summarized content without reprocessing the content corpus; and
storing each of the subsequent summarized content at a plurality of abstraction levels, thereby facilitating efficient retrieval and use based on user-defined summarization needs.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein the content corpus comprises one or more of text, audio, video, and image content.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein the recursive generating of the subsequent summarized content utilizes previously summarized outputs to perform further summarizations.
18 . The non-transitory computer-readable storage medium of claim 15 , wherein the plurality of abstraction levels for generating subsequent summarized content includes a plurality of specific percentages to support a tiered approach to content reduction.
19 . The non-transitory computer-readable storage medium of claim 15 , further comprising calculating a cost savings of the recursive generating of the subsequent summarized content compared to a technique where the content corpus is reprocessed for each summarization.
20 . The non-transitory computer-readable storage medium of claim 15 , wherein the LLM or the appropriate processing model is configured to adjust a depth of summarization dynamically based on a complexity or length of the content corpus.