IP Library › Granted Patent US 12,524,626
Granted Patent B2
US 12,524,626 · App. 18/299,503 · Granted Jan 13, 2026

Systems and methods for question-driven pretraining for controllable summarization

Inventors: Artidoro Pagnoni (Seattle, WA); Alexander Fabbri (New York, NY); Wojciech Kryscinski (Palo Alto, CA)
Assignee: Salesforce, Inc.
G06F40/40G06F16/345
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,626
App. No.
18/299,503
Granted
Jan 13, 2026
Kind
B2
Abstract

Embodiments described herein provide a task-targeted training paradigm which involves reformulating the pretraining objective to match the downstream task, e.g., query-focused summarization with or without guidance, without additional supervision. Specifically, query-focused summarization is foremost a task in which information is selected and compressed. A training objective is computed based on asking questions about sentences considered to be informative about the text that naturally incorporates a form of guidance of the generation process.

Claims (80)

1 . A method of training a pretrained language model for

question-driven summarization, the method comprising:

receiving, via a data interface, a training dataset containing at least one unlabeled document;

selecting, by a selection model implemented on one or more processors, a plurality of sentences from the unlabeled document depending on an information overlap metric with a rest of the unlabeled document;

generating, by a question generation model connected to the selection model implemented on the one or more processors, a plurality of questions in response to an input of the plurality of sentences and the unlabeled document as context;

generating a masked document by masking the plurality of sentences from the unlabeled document with mask tokens;

generating, by the pretrained model as a summarization model implemented on the one or more processors, predicted questions in response to an input of the masked document;

training the summarization model based on a first loss obtained by a comparison between the predicted questions and the plurality of questions;

generating, by the summarization model, a predicted summary containing one or more sentences in response to an input of the masked document prepended with the plurality of questions;

training again the summarization model based on a second loss obtained by a comparison between the predicted one or more sentences and the plurality of sentences;

in response to receiving a user query relating to the article from a user device, generating, by the trained summarization model, a query-driven summary of the article according to the user query; and

causing the query-driven summary to be displayed at a reading interface of a web browser at the user device at a time when the user is accessing the article through the reading interface.

2 . The method of claim 1 , wherein the plurality of sentences are selected by a metric that measure information overlap of each of the plurality of sentences with a rest of the input document.

3 . The method of claim 1 , further comprising:

generating, by the summarization model implemented on the one or more processors, predicted questions in response to an input of the masked document;

computing a third loss comparing the predicted questions and the plurality of questions and comparing the one or more predicted sentences and the plurality of sentences; and

training the summarization model by updating parameters of the summarization model based on the computed third loss via backpropagation.

4 . The method of claim 1 , further comprising:

generating, by the summarization model, reconstructed sentences in response to an input of the masked document;

computing a third loss comparing the reconstructed sentences and the plurality of sentences; and

training the summarization model by updating parameters of the summarization model based on the computed third loss via backpropagation.

5 . The method of claim 1 , further comprising:

prepending a dedicated token to the masked document before inputting the masked document to the summarization model; and

determining a type of training output depending on the dedicated token.

6 . The method of claim 1 , further comprising:

receiving a testing document and a user query relating to a content of the testing document; and

generating, by the trained summarization model, a question-driven summary of the testing document according to the user query.

7 . A system of training a pretrained language model for question-driven summarization, the system comprising:

a communication interface that receives a training dataset containing at least one unlabeled document;

a memory storing a selection model, a question generation model, a summarization model and a plurality of processor-executable instructions; and

one or more processors executing the plurality of processor-executable instructions to perform operations comprising:

selecting, by the selection model, a plurality of sentences from the unlabeled document depending on an information overlap metric with a rest of the unlabeled document;

generating, by the question generation model connected to the selection model implemented on the one or more processors, a plurality of questions in response to an input of the plurality of sentences and the unlabeled document as context;

generating a masked document by masking the plurality of sentences from the unlabeled document with mask tokens;

generating, by the pretrained model as the summarization model implemented on the one or more processors, predicted questions in response to an input of the masked document;

training the summarization model based on a first loss obtained by a comparison between the predicted questions and the plurality of questions;

generating, by the summarization model, a predicted summary containing one or more sentences in response to an input of the masked document prepended with the plurality of questions;

training again the summarization model based on a second loss obtained by a comparison between the predicted one or more sentences and the plurality of sentences;

in response to receiving a user query relating to the article from a user device, generating, by the trained summarization model, a query-driven summary of the article according to the user query; and

causing the query-driven summary to be displayed at a reading interface of a web browser at the user device at a time when the user is accessing the article through the reading interface.

8 . The system of claim 7 , wherein the plurality of sentences are selected by a metric that measure information overlap of each of the plurality of sentences with a rest of the input document.

9 . The system of claim 7 , wherein the operations further comprise:

generating, by the summarization model implemented on the one or more processors, predicted questions in response to an input of the masked document;

computing a third loss comparing the predicted questions and the plurality of questions; and

training the summarization model by updating parameters of the summarization model based on the computed third loss via backpropagation.

10 . The system of claim 7 , wherein the operations further comprise:

computing a fourth loss comparing the predicted questions and the plurality of questions and comparing the one or more predicted sentences and the plurality of sentences; and

training the summarization model by updating parameters of the summarization model based on the computed fourth loss via backpropagation.

11 . The system of claim 7 , wherein the operations further comprise:

generating, by the summarization model, reconstructed sentences in response to an input of the masked document;

computing a third loss comparing the reconstructed sentences and the plurality of sentences; and

training the summarization model by updating parameters of the summarization model based on the computed third loss via backpropagation.

12 . The system of claim 7 , wherein the operations further comprise:

prepending a dedicated token to the masked document before inputting the masked document to the summarization model; and

determining a type of training output depending on the dedicated token.

13 . The system of claim 7 , wherein the operations further comprise:

receiving a testing document and a user query relating to a content of the testing document; and

generating, by the trained summarization model, a question-driven summary of the testing document according to the user query.

14 . A non-transitory processor-readable storage medium storing a plurality of processor-executable instructions for training a pretrained language model for question-driven summarization, the instructions being executed by one or more processors to perform operations comprising:

receiving, via a data interface, a training dataset containing at least one unlabeled document;

selecting, by a selection model implemented on one or more processors, a plurality of sentences from the unlabeled document depending on an information overlap metric with a rest of the unlabeled document;

generating, by a question generation model connected to the selection model implemented on the one or more processors, a plurality of questions in response to an input of the plurality of sentences and the unlabeled document as context;

generating a masked document by masking the plurality of sentences from the unlabeled document with mask tokens;

generating, by the pretrained model as a summarization model implemented on the one or more processors, predicted questions in response to an input of the masked document;

training the summarization model based on a first loss obtained by a comparison between the predicted questions and the plurality of questions;

generating, by the summarization model, a predicted summary containing one or more sentences in response to an input of the masked document prepended with the plurality of questions;

training again the summarization model based on a second loss obtained by a comparison between the predicted one or more sentences and the plurality of sentences;

in response to receiving a user query relating to the article from a user device, generating, by the trained summarization model, a query-driven summary of the article according to the user query; and

causing the query-driven summary to be displayed at a reading interface of a web browser at the user device at a time when the user is accessing the article through the reading interface.

15 . The non-transitory processor-readable storage medium of claim 14 , wherein the plurality of sentences are selected by a metric that measure information overlap of each of the plurality of sentences with a rest of the input document.

16 . The non-transitory processor-readable storage medium of claim 14 , wherein the operations further comprise:

generating, by the summarization model implemented on the one or more processors, predicted questions in response to an input of the masked document;

computing a third loss comparing the predicted questions and the plurality of questions; and

training the summarization model by updating parameters of the summarization model based on the computed third loss via backpropagation.

17 . The non-transitory processor-readable storage medium of claim 15 , wherein the operations further comprise:

prepending a dedicated token to the masked document before inputting the masked document to the summarization model; and

determining a type of training output depending on the dedicated token.

18 . The non-transitory processor-readable storage medium of claim 14 , wherein the operations further comprise:

receiving a testing document and a user query relating to a content of the testing document; and

generating, by the trained summarization model, a question-driven summary of the testing document according to the user query.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2023
From: PAGNONI, ARTIDORO; FABBRI, ALEXANDER R.; KRYSCINSKI, WOJCIECH
To: SALESFORCE, INC.
Reel/Frame 063428/0736 →
Continuity (2)
Provisional Application 63387663 · Dec 15, 2022
Related Publication 20240202461A1 · Jun 20, 2024
References Cited (7)
WO WO2022015730A1 · 2022 [cited by examiner]
Lin, Chun Hung. “Automatic Question Generation with Pre-trained Masked Language Models.” (2020). (Year: 2020). [cited by examiner]
Kag, Anil, and Venkatesh Saligrama. “Training recurrent neural networks via forward propagation through time.” In International Conference on Machine Learning, pp. 5189-5200. PMLR, 2021. (Year: 2021). [cited by examiner]
Liu, Sen, Libin Yang, and Xiaoyan Cai. “SEASum: Syntax-enriched abstractive summarization.” Expert Systems with Applications 199 (Aug. 2022): 116819. (Year: 2022). [cited by examiner]
Kwon, Sunjae, Cheongwoong Kang, Jiyeon Han, and Jaesik Choi. “Why Do Neural Language Models Still Need Commonsense Knowledge to Handle Semantic Variations in Question Answering ?. ” arXiv preprint arXiv:2209.00599 (Sep.… [cited by examiner]
Sinclair, Neil. “Self-supervised text sentiment transfer with rationale predictions and pretrained transformers.” (Feb. 2022). (Year: 2022). [cited by examiner]
Chen, Hanjie, et al. “Explaining neural network predictions on sentence pairs via learning word-group masks.” arXiv preprint arXiv:2104.04488 (2021). (Year: 2021). [cited by examiner]