IP Library Granted Patent US 11,699,026
Granted Patent B2
US 11,699,026 · App. 17/589,675 · Granted Jul 11, 2023

Systems and methods for explainable and factual multi-document summarization

Inventors: Jered McInerney (Brighton, MA); Wojciech Kryscinski (Palo Alto, CA); Nazneen Rajani (Mountain View, CA)
Assignee: Salesforce, Inc.
G06F40/166G06F40/20G06F40/40G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,699,026
App. No.
17/589,675
Granted
Jul 11, 2023
Kind
B2
Abstract

Embodiments described herein provide methods and systems for summarizing multiple documents. A system receives a plurality of documents and generates embeddings of the sentences from the plurality of documents. The embedded sentences are clustered in a representation space. Sentences from a reference summary are embedded and aligned with the closest cluster. Sentences from each cluster are summarized with the aligned reference sentences as a target. A loss is computed based on the summarized sentences and the aligned references, and the natural language processing model is updated based on the loss. Sentences may be masked from being used in the summarization by identifying sentences that are contradicted by other sentences within the plurality of documents.

Claims (60)

1. A method for training a multi-document summarization model, comprising:

receiving, via a communication interface, a plurality of documents and a reference summary associated with the plurality of documents;

generating embeddings of sentences from the plurality of documents, wherein the embeddings indicate a relationship between the sentences across the plurality of documents;

clustering, based on the embeddings, the sentences from the plurality of documents into a plurality of clusters;

aligning one or more reference sentences in the reference summary with the plurality of clusters into a plurality of aligned reference sentence clusters, respectively;

masking a first sentence from one of the plurality of documents based on a determination that the first sentence is contradicted by a second sentence of the plurality of documents

generating, by a natural language processing model without using the first sentence based on the masking, a plurality of cluster-wise summaries corresponding to the plurality of clusters, respectively;

comparing the plurality of cluster-wise summaries and the plurality of aligned reference sentence clusters to compute a loss; and

updating the natural language processing model based on the loss.

2. The method of claim 1 , wherein the clustering further comprises:

clustering, based on the embeddings, the sentences from the plurality of documents into a plurality of clusters using K-means clustering.

3. The method of claim 1 , wherein the clustering further comprises:

clustering, based on the embeddings, the sentences from the plurality of documents into a plurality of clusters using Spectral clustering.

4. The method of claim 1 , wherein the masking is based on a determination that the first sentence is contradicted by a plurality of sentences of the plurality of documents.

5. The method of claim 1 , further comprising:

generating a composite summary based on the plurality of cluster-wise summaries.

6. The method of claim 5 , wherein the composite summary is generated with information about how the composite summary was formed,

including at least one of:

information associated with the clustering;

information associated with the masking; and

information associated with the generating the cluster-wise summaries or the composite summary.

7. A system for training a multi-document summarization model, the system comprising:

a memory that stores the multi-document summarization model;

a communication interface that receives a plurality of documents and a reference summary associated with the plurality of documents; and

one or more hardware processors that:

generates embeddings of sentences from the plurality of documents, wherein the embeddings indicate a relationship between the sentences across the plurality of documents;

clusters, based on the embeddings, the sentences from the plurality of documents into a plurality of clusters;

aligns one or more reference sentences in the reference summary with the plurality of clusters into a plurality of aligned reference sentence clusters, respectively;

masks a first sentence from one of the plurality of documents based on a determination that the first sentence is contradicted by a second sentence of the plurality of documents

generates, by a natural language processing model without using the first sentence based on the masking, a plurality of cluster-wise summaries corresponding to the plurality of clusters, respectively;

compares the plurality of cluster-wise summaries and the plurality of aligned reference sentence clusters to compute a loss; and

updates the natural language processing model based on the loss.

8. The system of claim 7 , wherein the clustering further comprises:

clustering, based on the embeddings, the sentences from the plurality of documents into a plurality of clusters using K-means clustering.

9. The system of claim 7 , wherein the clustering further comprises:

clustering, based on the embeddings, the sentences from the plurality of documents into a plurality of clusters using Spectral clustering.

10. The system of claim 7 , wherein the masking is based on a determination that the first sentence is contradicted by a plurality of sentences of the plurality of documents.

11. The system of claim 10 , wherein the one or more hardware processors further:

generates a composite summary based on the plurality of cluster-wise summaries.

12. The system of claim 11 , wherein the composite summary is generated with information about how the composite summary was formed,

including at least one of:

information associated with the clustering;

information associated with the masking; and

information associated with the generating the cluster-wise summaries or the composite summary.

13. A processor-readable non-transitory storage medium storing a plurality of processor-executable instructions for training a multi-document summarization model, the instructions being executed by a processor to perform operations comprising:

receiving, via a communication interface, a plurality of documents and a reference summary associated with the plurality of documents;

generating embeddings of sentences from the plurality of documents, wherein the embeddings indicate a relationship between the sentences across the plurality of documents;

clustering, based on the embeddings, the sentences from the plurality of documents into a plurality of clusters;

aligning one or more reference sentences in the reference summary with the plurality of clusters into a plurality of aligned reference sentence clusters, respectively;

masking a first sentence from one of the plurality of documents based on a determination that the first sentence is contradicted by a second sentence of the plurality of documents

generating, by a natural language processing model without using the first sentence based on the masking, a plurality of cluster-wise summaries corresponding to the plurality of clusters, respectively;

comparing the plurality of cluster-wise summaries and the plurality of aligned reference sentence clusters to compute a loss; and

updating the natural language processing model based on the loss.

14. The processor-readable non-transitory storage medium of claim 13 , wherein the clustering further comprises:

clustering, based on the embeddings, the sentences from the plurality of documents into a plurality of clusters using K-means clustering.

15. The processor-readable non-transitory storage medium of claim 13 , wherein the clustering further comprises:

clustering, based on the embeddings, the sentences from the plurality of documents into a plurality of clusters using Spectral clustering.

16. The processor-readable non-transitory storage medium of claim 13 , wherein the masking is based on a determination that the first sentence is contradicted by a plurality of sentences of the plurality of documents.

17. The processor-readable non-transitory storage medium of claim 13 , the operations further comprising:

generating a composite summary based on the plurality of cluster-wise summaries.

Assignments (2)
CHANGE OF NAME Recorded Dec 18, 2024
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 069717/0638 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2022
From: MCINERNEY, JERED; KRYSCINSKI, WOJCIECH; RAJANI, NAZNEEN
To: SALESFORCE.COM, INC.
Reel/Frame 061519/0731 →
Continuity (2)
Provisional Application 63240814 · Sep 3, 2021
Related Publication 20230070497A1 · Mar 9, 2023
Cited By (1)
US 12,537,852