IP Library › Granted Patent US 12,626,071
Granted Patent B2
US 12,626,071 · App. 18/962,575 · Granted May 12, 2026

Generating summaries of texts using large language models selected based on a minimization of a classification score, processing and reading times and words in a deny list

Inventors: Walter Bender (Newton, MA); Nithi Vivatrat (Washington, DC); Richard Graves (Washington, DC); Tomá Valena (Basel, CH); David McMinn (East Kilbride, GB)
Assignee: Sorcero, Inc.
G06F40/42G06F18/241G06F3/0484G06F16/332G06F16/34G06F40/157G06F40/205G06N5/01G06N5/04G06Q10/103
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,071
App. No.
18/962,575
Filed
Nov 27, 2024
Granted
May 12, 2026
Kind
B2
Art Unit
2653
USPC
704/9
Abstract

Systems and methods for generating summaries from text using a generative model are disclosed. The system is configured to access an article; identify section; provide, to one or more generative models, a prompt including instructions to generate a section summary; generate an article summary based on the section summary; determine, from the article summary, a first concept found in the article summary that is missing from the article; determine, using a classifier, for a first sentence included in the article summary, a confidence score; and provide, for presentation at a client device, a document including the article summary.

Claims (103)

1 . A system, comprising:

one or more processors coupled to memory, the one or more processors configured to:

access an article including a title and a body;

identify, from the article, a plurality of sections;

provide, to one or more generative pre-training transformer (GPT) models, for each section of the plurality of sections of the article, a prompt including instructions to generate a section summary for the section based on one or more sentences included in the section, wherein providing the prompt to the one or more GPT models comprises the one or more processors to select the one or more GPT models from a plurality of GPT models based on the selected one or more GPT models providing higher metrics than non-selected GPT models, the metrics comprising at least two of:

a minimization of a classification score indicating a likelihood that the one or more sentences included in the section correspond to the section summary;

a minimization of processing time;

a minimization of a score indicating a duration a user associated with a client device is reading;

a minimization of a period of time associated with a reading time, and a minimization of words on a deny list included within the section summaries;

generate an article summary based on the section summary generated for each section;

determine, using the one or more GPT models, from the article summary, a first concept found in the article summary that is missing from the article;

determine, using a classifier, for a first sentence included in the article summary, a confidence score indicating a likelihood that the first sentence is not supported by the article;

and provide, for presentation at the client device, a document including the article summary, a first indicator corresponding to the first concept and a second indicator corresponding to the first sentence.

2 . The system of claim 1 , wherein the classifier is a first classifier, and comprising the one or more processors to:

determine, using the first classifier responsive to providing the prompt to generate the section summary for the section, that at least one sentence included in a section summary of a first section of the plurality of sections of the article has a classification score indicating that the at least one sentence belongs to a second section of the plurality of sections of the article;

remove the at least one sentence from the section summary of the first section to generate an updated section summary of the first section; and

generate the article summary based on the section summary generated for each section and the updated section summary of the first section.

3 . The system of claim 1 , comprising the one or more processors to:

determine, responsive to determining the confidence score indicating the likelihood that the first sentence is not supported by evidence in the article, a score indicating a reading level for the article summary; and

provide, for presentation at the client device, the document including the score.

4 . The system of claim 1 , comprising the one or more processors to:

determine, responsive to determining the confidence score indicating the likelihood that the first sentence is not supported by evidence in the article, a period of time indicating a duration a user associated with the client device is reading for each of the article summary and the article; and

provide, for presentation at the client device, the document including the period of time.

5 . The system of claim 1 , wherein providing the document comprises the one or more processors to:

identify a deny list of words associated with the article;

parse the article summary to identify a first word of the deny list appearing within the article summary; and

replace the first word appearing within the article summary with a second word based on an index mapping the deny list of words to an allow list of words.

6 . The system of claim 1 , comprising the one or more processors to:

determine, responsive to determining a section identifier and a confidence score for each sentence, that a first confidence score for a sentence is above a threshold confidence score;

associate the sentence with the section identifier corresponding to the first confidence score;

determine that a second confidence score for a second sentence is at or below the threshold confidence score; and

associate the second sentence with a second section identifier different than the section identifier associated with the first confidence score.

7 . The system of claim 1 , wherein the classifier comprises a first classifier and a second classifier, and comprising the one or more processors to:

train the first classifier using a plurality of bodies from a plurality of articles to recognize a plurality of section identifiers; and

train the second classifier using a plurality of sentences of a plurality of section summaries and the plurality of bodies to compare the plurality of sentences to the plurality of bodies.

8 . The system of claim 1 , comprising the one or more processors to:

determine that the article satisfies a format comprising the body and the title; and

access the article responsive to the determination that the article satisfies the format.

9 . The system of claim 1 , comprising the one or more processors to:

determine that a second article does not satisfy a format comprising the body and the title; and

provide, for presentation at the client device, an indication of the second article not satisfying the format.

10 . The system of claim 1 , comprising the one or more processors to provide, for presentation at the client device, the document comprising a comparison of a first score indicating a reading level associated with the article summary and a second score indicating a reading level associated with the article.

11 . The system of claim 1 , wherein to identify, from the article, a plurality of sections, the one or more processors configured to provide, to a GPT model, a prompt to cause the GPT model to output portions of the body of the article in respective sections of the plurality of sections.

12 . The system of claim 1 , wherein the classifier is a first classifier, wherein the body includes a plurality of sentences, and wherein to identify, from the article, a plurality of sections, the one or more processors configured to determine, by inputting each sentence of the one or more sentences of the article into the first classifier, a section identifier and a confidence score for each sentence, the confidence score indicating a likelihood of its respective sentence corresponding to its respective section identifier.

13 . The system of claim 1 , wherein to generate the article summary, the one or more processors configured to iteratively prompt the one or more GPT models based on a threshold associated with the article summary.

14 . A method, comprising:

accessing, by one or more processors coupled to memory, an article including a title and a body;

identifying, by the one or more processors, from the article, a plurality of sections;

providing, by the one or more processors, to one or more GPT models, for each section of the plurality of sections of the article, a prompt including instructions to generate a section summary for the section based on one or more sentences included in the section, wherein providing the prompt to the one or more GPT models comprises selecting the one or more GPT models from a plurality of GPT models based on the selected one or more GPT models providing higher metrics than non-selected GPT models, the metrics comprising at least two of:

a minimization of a classification score indicating a likelihood that the one or more sentences included in the section correspond to the section summary;

a minimization of processing time;

a minimization of a score indicating a duration a user associated with a client device is reading;

a minimization of a period of time associated with a reading time; and

a minimization of words on a deny list included within the section summaries;

generating, by the one or more processors, an article summary based on the section summary generated for each section;

determining, by the one or more processors using the one or more GPT models, from the article summary, a first concept found in the article summary that is missing from the article;

determining, by the one or more processors, using a classifier, for a first sentence included in the article summary, a confidence score indicating a likelihood that the first sentence is not supported by the article; and

providing, by the one or more processors, for presentation at the client device, a document including the article summary, a first indicator corresponding to the first concept and a second indicator corresponding to the first sentence.

15 . The method of claim 14 , wherein the classifier is a first classifier, and the method comprising:

determining, by the one or more processors responsive to providing the prompt to generate the section summary, using the first classifier, that at least one sentence included in a section summary of a first section of the plurality of sections of the article has a classification score indicating that the at least one sentence belongs to a second section of the plurality of sections of the article;

removing, by the one or more processors, the at least one sentence from the section summary of the first section to generate an updated section summary of the first section; and

generating, by the one or more processors, the article summary based on the section summary generated for each section and the updated section summary of the first section.

16 . The method of claim 14 , comprising:

determining, by the one or more processors, responsive to determining the confidence score indicating the likelihood that the first sentence is not supported by evidence in the article, a score indicating a reading level for the article summary; and

providing, by the one or more processors, for presentation at the client device, the document including the score.

17 . The method of claim 14 , comprising:

determining, by the one or more processors, responsive to determining the confidence score indicating the likelihood that the first sentence is not supported by evidence in the article, a period of time indicating a duration a user associated with the client device is reading for each of the article summary and the article; and

providing, by the one or more processors, for presentation at the client device, the document including the period of time.

18 . The method of claim 14 , comprising:

determining, by the one or more processors, responsive to determining a section identifier and a confidence score for each sentence, that a first confidence score for a sentence is above a threshold confidence score;

associating, by the one or more processors, the sentence with the section identifier corresponding to the first confidence score;

determining, by the one or more processors, that a second confidence score for a second sentence is at or below the threshold confidence score; and

associating, by the one or more processors, the second sentence with a second section identifier different than the section identifier associated with the first confidence score.

19 . The method of claim 14 , comprising:

determining, by the one or more processors, that the article satisfies a format comprising the body and the title; and

accessing, by the one or more processors, the article responsive to the determination that the article satisfies the format.

20 . The method of claim 14 , comprising:

determining, by the one or more processors, that a second article does not satisfy a format comprising the body and the title; and

providing, by the one or more processors, for presentation at the client device, an indication of the second article not satisfying the format.

21 . The method of claim 14 , wherein identifying, from the article, the plurality of sections comprises providing, to a GPT model, a prompt to cause the GPT model to output portions of the body of the article in respective sections of the plurality of sections.

22 . The method of claim 14 , wherein the classifier is a first classifier, and wherein identifying, from the article, the plurality of sections comprises determining, by inputting each sentence of the one or more sentences of the article into the first classifier, a section identifier and a confidence score for each sentence, the confidence score indicating a likelihood of its respective sentence corresponding to its respective section identifier.

23 . The method of claim 14 , wherein generating the article summary comprises iteratively prompting the one or more GPT models based on a threshold associated with the article summary.

24 . A system, comprising:

one or more processors coupled to memory, the one or more processors configured to:

access an article including a title and a body;

identify, from the article, a plurality of sections;

provide, to one or more GPT models, for each section of the plurality of sections of the article, a prompt including instructions to generate a section summary for the section based on one or more sentences included in the section, wherein providing the prompt to the one or more GPT models comprises the one or more processors to select the one or more GPT models from a plurality of GPT models based on the selected one or more GPT models providing higher metrics than non-selected GPT models, the metrics comprising at least two of:

a minimization of a classification score indicating a likelihood that the one or more sentences included in the section correspond to the section summary;

a minimization of processing time;

a minimization of a score indicating a duration a user associated with a client device is reading;

a minimization of a period of time associated with a reading time; and

a minimization of words on a deny list included within the section summaries;

generate an article summary based on the section summary generated for each section;

determine, using the one or more GPT models, that a first portion of the article summary does not satisfy a validation criteria; and

provide, for presentation at the client device, a document including the article summary and a first indicator corresponding to a first portion of the article.

25 . The system of claim 24 , comprising the one or more processors to:

determine, using a first classifier responsive to providing the prompt to generate the section summary for the section, that at least one sentence included in a section summary of a first section of the plurality of sections of the article has a classification score indicating that the at least one sentence belongs to a second section of the plurality of sections of the article;

remove the at least one sentence from the section summary of the first section to generate an updated section summary of the first section; and

generate the article summary based on the section summary generated for each section and the updated section summary of the first section.

26 . The system of claim 24 , wherein to determine that the first portion does not satisfy the validation criteria, the one or more processors are configured to determine, from the article summary, a first concept found in the article summary that is missing from the article; and

wherein the first indicator corresponds to the first concept.

27 . The system of claim 24 , wherein to determine that the first portion does not satisfy the validation criteria, the one or more processors are configured to determine, using a second classifier, for a first sentence included in the article summary, a confidence score indicating a likelihood that the first sentence is not supported by the article; and

wherein the first indicator corresponds to the first sentence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2025
From: BENDER, WALTER; VIVATRAT, NITHI; GRAVES, RICHARD; VALENA, TOMÁŠ; MCMINN, DAVID
To: SORCERO, INC.
Reel/Frame 071972/0753 →
Continuity (2)
Provisional Application 63603944 · Nov 29, 2023
Related Publication 20250173523A1 · May 29, 2025
References Cited (10)
US 8738414B1 · Nagar · 2014 [cited by examiner]
US 10496272B1 · Lonkar · 2019 [cited by examiner]
US 20200065387A1 · Matthews · 2020 [cited by examiner]
US 20230107640A1 · Pang et al. · 2023 [cited by applicant]
US 20230252225A1 · Zhelezniak · 2023 [cited by examiner]
US 20240412004A1 · Manikandan · 2024 [cited by examiner]
US 20250005276A1 · Bhat · 2025 [cited by examiner]
CN 117194653 · 2013 [cited by examiner]
KR 2640471B1 · 2023 [cited by examiner]
KR 102640449 · 2023 [cited by examiner]