IP Library › Granted Patent US 12,204,847
Granted Patent B2
US 12,204,847 · App. 17/938,572 · Granted Jan 21, 2025

Systems and methods for text summarization

Inventors: Alexander R. Fabbri (New York, NY); Prafulla Kumar Choubey (San Jose, CA); Jesse Vig (Los Altos, CA); Chien-Sheng Wu (Mountain View, CA); Caiming Xiong (Menlo Park, CA)
Assignee: Salesforce, Inc.
G06F40/166G06F40/284G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,204,847
App. No.
17/938,572
Granted
Jan 21, 2025
Kind
B2
Abstract

Embodiments described herein provide a method for text summarization. The method includes receiving a training dataset having at least an uncompressed text, a compressed text, and one or more information entities accompanying the compressed text. The method also includes generating, using a perturber model, a perturbed text with the one or more information entities being inserted into the compressed text. The method further includes training the perturber model based on a first training objective, and generating, using the trained perturber model, a perturbed summary in response to an input of a reference summary. The method further includes generating, via an editor model, a predicted summary by removing information from the perturbed summary conditioned on a source document of the reference summary, and training the editor model based on a second training objective.

Claims (51)

1. A method for text summarization, the method comprising:

receiving, via a communication interface, a training dataset comprising at least an uncompressed text, a compressed text, and one or more information entities that are not mentioned in the compressed text but included in the uncompressed text;

generating, by a perturber model, a perturbed text based on inserting at least a first information entity from the one or more information entities into the compressed text;

training the perturber model based on a first training objective comparing the perturbed text generated from the compressed text and the at least first information entity, and the uncompressed text including the one or more information entities;

generating, by the trained perturber model, a perturbed summary in response to an input of a reference summary and at least a second information entity randomly selected from a corresponding source document of the reference summary;

generating, via an editor model, a predicted summary by removing information from the perturbed summary conditioned on the source document of the reference summary; and

training the editor model based on a second training objective of a cross-entropy loss between a predicted token distribution of the predicted summary conditioned on the source document and a token distribution of the reference summary.

2. The method of claim 1 , wherein the one or more information entities are not mentioned in the compressed text.

3. The method of claim 1 , wherein the information removed from the predicted summary comprises a plurality of information entities not mentioned in the source document.

4. The method of claim 1 , wherein the first training objective is a cross-entropy loss between a predicted token distribution of the perturbed text and a token distribution of the uncompressed text.

5. The method of claim 1 , wherein the perturbed summary is generated by inserting information into the reference summary, and wherein at least a portion of the inserted information entities is randomly selected from the source document but not present in the reference summary.

6. The method of claim 1 , wherein the perturbed summary has a format of a sequence of tokens with hashtags identifying inserted tokens corresponding to inserted one or more information entities that do not belong to the reference summary.

7. The method of claim 1 , wherein the training the editor model further comprises updating the editor model based on the cross-entropy loss via backpropagation.

8. The method of claim 1 , further comprising:

receiving, at the trained editor model, a testing summary and a test source document; and

generating, by the trained editor model, a pruned summary conditioned on the test source document.

9. The method of claim 1 , wherein:

the perturber model comprises a denoising autoencoder; and

the editor model comprises a denoising autoencoder.

10. The method of claim 1 , wherein the one or more information entities are randomly selected from a pool of information entities that is overlapping with the uncompressed text and non-overlapping with the compressed text.

11. The method of claim 1 , wherein the generating of the perturbed text with the one or more information entities comprises the inserting of one or more information entities into the compressed text without swapping with the one or more information entities.

12. The method of claim 1 , wherein the generating of the predicted summary comprises the removing of the information entities from the perturbed summary without swapping the information.

13. The method of claim 1 , wherein:

training data of the perturber model comprises a set of sentence-compression data, and the set of sentence-compression data comprises one or more pairs of a compressed text, a corresponding uncompressed, and one or more perturber information entities; and

the uncompressed text contains the compressed text and the one or more information entities.

14. The method of claim 1 , wherein a number of tokens in the uncompressed text is more a number of tokens in the compressed text by up to one third of the number of tokens in the compressed text.

15. A system for text summarization, the system comprising:

a communication interface that receives a training dataset comprising at least an uncompressed text, a compressed text, and one or more information entities that are not mentioned in the compressed text;

a memory storing a plurality of processor-executable instructions; and

a processor executing the instructions to perform operations comprising:

generating, by a perturber model, a perturbed text based on inserting at least a first information entity from the one or more information entities into the compressed text;

training the perturber model based on a first training objective comparing the perturbed text generated from the compressed text and the at least first information entity and the uncompressed text including the one or more information entities;

generating, by the trained perturber model, a perturbed summary in response to an input of a reference summary and at least a second information entity randomly selected from a corresponding source document of the reference summary;

generating, via an editor model, a predicted summary by removing information from the perturbed summary conditioned on the source document of the reference summary; and

training the editor model based on a second training objective of a cross-entropy loss between a predicted token distribution of the predicted summary conditioned on the source document and a token distribution of the reference summary.

16. The system of claim 15 , wherein the one or more information entities are not mentioned in the compressed text.

17. The system of claim 15 , wherein the information removed from the predicted summary comprises a plurality of information entities not mentioned in the source document.

18. The system of claim 15 , wherein the first training objective is a cross-entropy loss between a predicted token distribution of the perturbed text and a token distribution of the uncompressed text.

19. The system of claim 15 , wherein the perturbed summary is generated by inserting information into the reference summary, and wherein at least a portion of the inserted information entities is randomly selected from the source document but not present in the reference summary.

20. The system of claim 15 , wherein the perturbed summary has a format of a sequence of tokens with hashtags identifying inserted tokens corresponding to inserted one or more information entities that do not belong to the reference summary.

21. The system of claim 15 , wherein the training the editor model further comprises updating the editor model based on the cross-entropy loss via backpropagation.

22. The system of claim 15 , further comprising:

receiving, at the trained editor model, a testing summary and a test source document; and

generating, by the trained editor model, a pruned summary conditioned on the test source document.

23. A non-transitory processor-readable storage medium storing a plurality of processor-executable instructions for code program synthesis for a target problem, the instructions being executed by one or more hardware processors to perform operations comprising:

receiving, via a communication interface, a training dataset comprising at least an uncompressed text, a compressed text, and one or more information entities that are not mentioned in the compressed text;

generating, by a perturber model, a perturbed text based on inserting at least a first information entity from the one or more information entities into the compressed text;

training the perturber model based on a first training objective comparing the perturbed text generated from the compressed text and the at least first information entity and the uncompressed text including the one or more information entities;

generating, by the trained perturber model, a perturbed summary in response to an input of a reference summary and at least a second information entity randomly selected from a corresponding source document of the reference summary;

generating, via an editor model, a predicted summary by removing information from the perturbed summary conditioned on the source document of the reference summary; and

training the editor model based on a second training objective of a cross-entropy loss between a predicted token distribution of the predicted summary conditioned on the source document and a token distribution of the reference summary.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2022
From: FABBRI, ALEXANDER R; CHOUBEY, PRAFULLA KUMAR; VIG, JESSE; WU, CHIEN-SHENG; XIONG, CAIMING
To: SALESFORCE, INC.
Reel/Frame 061340/0254 →
Continuity (2)
Provisional Application 63355323 · Jun 24, 2022
Related Publication 20230419017A1 · Dec 28, 2023
References Cited (37)
US 10943181B2 · Shteingart · 2021 [cited by examiner]
US 20110029477A1 · Tengli · 2011 [cited by examiner]
US 20130282837A1 · Mayala · 2013 [cited by examiner]
US 20160062981A1 · Dogrultan · 2016 [cited by examiner]
US 20200081909A1 · Li · 2020 [cited by examiner]
US 20200272940A1 · Sun · 2020 [cited by examiner]
US 20210256160A1 · Hachey · 2021 [cited by examiner]
US 20210375269A1 · Yavuz · 2021 [cited by examiner]
US 20210390127A1 · Fox · 2021 [cited by examiner]
US 20220036232A1 · Patel · 2022 [cited by examiner]
US 20220229983A1 · Zohrevand · 2022 [cited by examiner]
US 20220237373A1 · Singh Bawa · 2022 [cited by examiner]
US 20220318593A1 · Chen · 2022 [cited by examiner]
US 20220382793A1 · Hwang · 2022 [cited by examiner]
US 20230004589A1 · Wu · 2023 [cited by examiner]
US 20230079879A1 · Gunasekara · 2023 [cited by examiner]
US 20230119109A1 · Choubey · 2023 [cited by examiner]
US 20230376677A1 · Choubey · 2023 [cited by examiner]
Nguyen Thanh Nguyen; Improving Abstractive Summarization By Understanding Hidden Representations And Guidance On Semantic Meaning; Ulsan National Institute of Science and Technology; 2021; pp. 1-25. [cited by examiner]
Chen, Mingda, et al. “Summscreen: A dataset for abstractive screenplay summarization.” arXiv preprint arXiv:2104.07091 (2021) (Year: 2021). [cited by examiner]
Stiennon, Nisan, et al. “Learning to summarize with human feedback.” Advances in Neural Information Processing Systems 33 (2020): 3008-3021 (Year: 2020). [cited by examiner]
Miyato, Takeru, Andrew M. Dai, and Ian Goodfellow. “Adversarial training methods for semi-supervised text classification.” arXiv preprint arXiv:1605.07725 (2016) (Year: 2016). [cited by examiner]
Cao, Ziqiang, et al. “Faithful to the original: Fact aware neural abstractive summarization.” Proceedings of the AAAI Conference on Artificial Intelligence. vol. 32. No. 1. 2018 (Year: 2018). [cited by examiner]
Adam, Griffin et al. “Learning to Revise References for Faithful Summarization”, Apr. 13, 2022, pp. 1-20, arXiv:2204.10290v1 [cs.CL]. [cited by applicant]
Cao et al., “Factual error correction for abstractive summarization models”, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov. 16-20, 2020, pp. 6251-6258. [cited by applicant]
Dong et al., “Multi-fact correction in abstractive text summarization”, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov. 2020, pp. 9320-9331. [cited by applicant]
Filippova et al., “Overcoming the lack of parallel data in sentence compression”, Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Oct. 2013, Seattle, Washington, pp. 1481-1491. [cited by applicant]
Gehrmann et al., “Bottom-up abstractive summarization”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Oct.-Nov. 2018, Brussels, Belgium, pp. 4098-4109. [cited by applicant]
Goyal et al., “Annotating and modeling fine-grained factuality in summarization”, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technol… [cited by applicant]
Hermann et al., “Teaching machines to read and comprehend”, Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, Dec. 7-12, 2015, Montreal, Quebec, Canad… [cited by applicant]
Kryscinski et al., “Evaluating the factual consistency of abstractive text summarization”, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov. 2020, pp. 9332-9346. [cited by applicant]
Lee et al., “Factual error correction for abstractive summaries using entity retrieval”, Proceedings of the 2nd Workshop on Natural Language Generation, Evaluation, and Metrics (GEM), Dec. 2022, pp. 439-444. [cited by applicant]
Lin, “ROUGE: A package for automatic evaluation of summaries”, Text Summarization Branches Out, Barcelona, Spain, pp. 74-81. [cited by applicant]
Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach”, http://arxiv.org/abs/1907.11692v1, 13 pages. [cited by applicant]
Narayan et al., “Don't give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,… [cited by applicant]
Wan et al., “Factpegasus: Factuality-aware pre-training and fine-tuning for abstractive summarization”, arXiv:2205.07830v1 [cs.CL] May 16, 2022, 19 pages. [cited by applicant]
Warstadt et al., “Neural network acceptability judgments”, Transactions of the Association for Computational Linguistics, 2019, vol. 7, pp. 625-641. [cited by applicant]