IP Library › Granted Patent US 11,269,965
Granted Patent B2
US 11,269,965 · App. 16/670,631 · Granted Mar 8, 2022

Extractive query-focused multi-document summarization

Inventors: Odellia Boni (Giva'at Ela, IL); Guy Feigenblat (Givataym, IL); David Konopnicki (Haifa, IL); Haggai Roitman (Yoknea'm Elit, IL)
Assignee: International Business Machines Corporation
G06F16/9535G06F16/3334G06F16/345G06F16/9038G06F16/90332G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,269,965
App. No.
16/670,631
Granted
Mar 8, 2022
Kind
B2
Abstract

A method, computer system, and computer program product for generating a multi-document summary is provided. The embodiment may include receiving a query statement, one or more documents, one or more summary constraints, and quality goals. The embodiment may include identifying one or more keywords within the query statement. The embodiment may include performing a sentence selection from the one or more documents based on the one or more identified keywords. The embodiment may include generating a plurality of candidate summaries of the one or more documents based on the performed sentence selection, the goals, and a cross entropy method. The embodiment may include calculating a quality score for each of the plurality of generated candidate summaries using a plurality of quality features. The embodiment may include selecting a candidate summary from the plurality of generated candidate summaries with the highest calculated quality score that also satisfies a quality score threshold.

Claims (54)

1. A processor-implemented method for generating a multi-document summary, the method comprising:

performing a sentence selection from one or more documents based on one or more keywords within a query statement;

generating a plurality of candidate summaries of the one or more documents based on the performed sentence selection, one or more goals, and a fully-polynomial randomized approximation scheme (FPRAS) cross entropy method;

calculating a quality score for each of the plurality of generated candidate summaries using a plurality of quality features; and

selecting a candidate summary from the plurality of generated candidate summaries with the highest calculated quality score that also satisfies a quality score threshold.

2. The method of claim 1 , further comprising:

generating an expanded query statement using a plurality of query expansion techniques, wherein the expanded query statement comprises the one or more identified keywords and one or more other keywords related to the received query statement, and wherein the expanded query statement is used in the sentence selection.

3. The method of claim 1 , further comprising:

restructuring a sentence structure of the selected candidate summary using a plurality of natural language processing techniques.

4. The method of claim 1 , further comprising:

presenting the selected candidate summary to a user, wherein the selected candidate summary is presented on a display screen of a user device through a graphical user interface.

5. The method of claim 1 , further comprising:

extracting each candidate summary based on the quality score associated with each candidate summary satisfying a filter threshold;

identifying one or more frequently appearing sentences within plurality of filtered candidate summaries; and

updating an algorithm used by the cross entropy method based on the one or more identified frequently appearing sentences.

6. The method of claim 1 , wherein the plurality of quality features are selected from a group consisting of measuring a Bhattacharyya similarity between a unigram language model (LM) of the received query statement and a unigram LM of a candidate summary within the plurality of candidate summaries, measuring a relative mass that the candidate summary devotes to the received query statement, measuring to what extent the candidate summary generally covers the one or more documents, measuring a sentence diversity of each candidate summary by calculating a bigram LM entropy for each candidate summary, biasing the sentence selection towards one or more sentences that appear earlier in a containing document, and biasing the sentence selection towards one or more longer summaries that still satisfy a length constraint that contain few long sentences rather than a plurality of candidate summaries that contain many short sentences.

7. The method of claim 1 , further comprising:

removing one or more candidate summaries from the plurality of generated candidate summaries that do not satisfy one or more summary constraints.

8. A computer system for generating a multi-document summary, the computer system comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored on at least one of the one or more tangible storage media for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method comprising:

performing a sentence selection from one or more documents based on one or more keywords within a query statement;

generating a plurality of candidate summaries of the one or more documents based on the performed sentence selection, one or more goals, and a fully-polynomial randomized approximation scheme (FPRAS) cross entropy method;

calculating a quality score for each of the plurality of generated candidate summaries using a plurality of quality features; and

selecting a candidate summary from the plurality of generated candidate summaries with the highest calculated quality score that also satisfies a quality score threshold.

9. The computer system of claim 8 , further comprising:

generating an expanded query statement using a plurality of query expansion techniques, wherein the expanded query statement comprises the one or more identified keywords and one or more other keywords related to the received query statement, and wherein the expanded query statement is used in the sentence selection.

10. The computer system of claim 8 , further comprising:

restructuring a sentence structure of the selected candidate summary using a plurality of natural language processing techniques.

11. The computer system of claim 8 , further comprising:

presenting the selected candidate summary to a user, wherein the selected candidate summary is presented on a display screen of a user device through a graphical user interface.

12. The computer system of claim 8 , further comprising:

extracting each candidate summary based on the quality score associated with each candidate summary satisfying a filter threshold;

identifying one or more frequently appearing sentences within plurality of filtered candidate summaries; and

updating an algorithm used by the cross entropy method based on the one or more identified frequently appearing sentences.

13. The computer system of claim 8 , wherein plurality of quality features are selected from a group consisting of measuring a Bhattacharyya similarity between a unigram language model (LM) of the received query statement and a unigram LM of a candidate summary within the plurality of candidate summaries, measuring a relative mass that the candidate summary devotes to the received query statement, measuring to what extent the candidate summary generally covers the one or more documents, measuring a sentence diversity of each candidate summary by calculating a bigram LM entropy for each candidate summary, biasing the sentence selection towards one or more sentences that appear earlier in a containing document, and biasing the sentence selection towards one or more longer summaries that still satisfy a length constraint that contain few long sentences rather than a plurality of candidate summaries that contain many short sentences.

14. The computer system of claim 8 , further comprising:

removing one or more candidate summaries from the plurality of generated candidate summaries that do not satisfy one or more summary constraints.

15. A computer program product for generating a multi-document summary, the computer program product comprising:

one or more computer-readable tangible storage media and program instructions stored on at least one of the one or more tangible storage media, the program instructions executable by a processor of a computer to perform a method, the method comprising:

performing a sentence selection from one or more documents based on one or more keywords within a query statement;

generating a plurality of candidate summaries of the one or more documents based on the performed sentence selection, one or more goals, and a fully-polynomial randomized approximation scheme (FPRAS) cross entropy method;

calculating a quality score for each of the plurality of generated candidate summaries using a plurality of quality features; and

selecting a candidate summary from the plurality of generated candidate summaries with the highest calculated quality score that also satisfies a quality score threshold.

16. The computer program product of claim 15 , further comprising:

generating an expanded query statement using a plurality of query expansion techniques, wherein the expanded query statement comprises the one or more identified keywords and one or more other keywords related to the received query statement, and wherein the expanded query statement is used in the sentence selection.

17. The computer program product of claim 15 , further comprising:

restructuring a sentence structure of the selected candidate summary using a plurality of natural language processing techniques.

18. The computer program product of claim 15 , further comprising:

presenting the selected candidate summary to a user, wherein the selected candidate summary is presented on a display screen of a user device through a graphical user interface.

19. The computer program product of claim 15 , further comprising:

extracting each candidate summary based on the quality score associated with each candidate summary satisfying a filter threshold;

identifying one or more frequently appearing sentences within plurality of filtered candidate summaries; and

updating an algorithm used by the cross entropy method based on the one or more identified frequently appearing sentences.

20. The computer program product of claim 15 , wherein the plurality of quality features are selected from a group consisting of measuring a Bhattacharyya similarity between a unigram language model (LM) of the received query statement and a unigram LM of a candidate summary within the plurality of candidate summaries, measuring a relative mass that the candidate summary devotes to the received query statement, measuring to what extent the candidate summary generally covers the one or more documents, measuring a sentence diversity of each candidate summary by calculating a bigram LM entropy for each candidate summary, biasing the sentence selection towards one or more sentences that appear earlier in a containing document, and biasing the sentence selection towards one or more longer summaries that still satisfy a length constraint that contain few long sentences rather than a plurality of candidate summaries that contain many short sentences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2019
From: BONI, ODELLIA; FEIGENBLAT, GUY; KONOPNICKI, DAVID; ROITMAN, HAGGAI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 050883/0966 →
Continuity (4)
Continuation 16005815 · Jun 12, 2018
Continuation 15843993 · Dec 15, 2017
Continuation 15660034 · Jul 26, 2017
Related Publication 20200065346A1 · Feb 27, 2020