IP Library Granted Patent US 10,042,924
Granted Patent B2
US 10,042,924 · App. 15/019,646 · Granted Aug 7, 2018

Scalable and effective document summarization framework

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,042,924
App. No.
15/019,646
Granted
Aug 7, 2018
Kind
B2
Abstract

Systems, methods, and apparatuses are disclosed for adaptively generating a summary of web-based content based on an attribute of a mobile communication device having transmitted a request for the web-based content. By adaptively generating the summary based on an attribute of the mobile communication device such as an amount of visual space available or a number of characters permitted in the interface, a display of the web-based content may be controlled on the mobile communication device in a way that was not previously available. This enables control of displaying web-based content that has been adaptively generated to be displayed on limited display screens based on a learned attribute of the mobile communication device requesting the web-based content.

Claims (82)

1. A summarization engine comprising:

a memory configured to store a document including textual information;

an interface configured to receive a viewing request from a communication device, the viewing request corresponding to the document; and

a processor configured to:

communicate with the memory and interface;

in response to receiving the viewing request, determine a target summary length, wherein the target summary length identifies a targeted length for a generated summary;

extract the textual information from the document;

parse the textual information;

identify a plurality of sentence structures from the textual information based on the parsing;

assign each sentence structure a weighted score based on the target summary length;

generate, in accordance with a summarization policy, candidate summaries to include one or more sentence structures from the plurality of sentence structures;

determine a candidate score for each candidate summary as a linear function of the scores assigned to each sentence structure included in the respective candidate summary;

learn coefficients of the linear function from a training dataset of documents and authored summaries via a predetermined learning algorithm; and

select, from the generated candidate summaries, a generated summary determined to have the highest candidate score under the linear function.

2. The summarization engine of claim 1 , wherein the processor is configured to generate the weighted score by:

determining a position of each sentence structure within the document; and

generating the weighted score of each sentence structure based on the determined position of each corresponding sentence structure.

3. The summarization engine of claim 1 , wherein the processor is configured to generate the weighted score by:

determining a content of each sentence structure within the document; and

generating the weighted score of each sentence structure based on the determined content of each corresponding sentence structure.

4. The summarization engine of claim 1 , wherein the processor is configured to generate the weighted score by:

identifying one or more tokens from each sentence structure, wherein each identified token corresponds to a token type;

determining a token type of each identified token; and

generating the weighted score of each sentence structure based on the determined token type of each identified token of each corresponding sentence structure.

5. The summarization engine of claim 1 , wherein the processor is configured to generate the weighted score by:

identifying one or more lexical cues from each sentence structure; and

generating the weighted score of each sentence structure based on the identified lexical cues of each corresponding sentence structure.

6. The summarization engine of claim 1 , wherein the target summary length is included in the viewing request; and

wherein the processor is configured to determine the target summary length by extracting the target summary length from the viewing request.

7. The summarization engine of claim 6 , wherein the target summary length corresponds to an attribute of the communication device transmitting the viewing request.

8. A method for generating a summary of a document, the method comprising:

receiving, through an interface, a viewing request from a communication device, the viewing request corresponding to a document including textual information stored on a memory;

in response to receiving the viewing request, determining a target summary length;

extracting the textual information from the document;

parsing the textual information;

identifying a plurality of sentence structures from the textual information based on the parsing;

assigning each sentence structure a weighted score based on the target summary length;

generating, in accordance with a summarization policy, candidate summaries to include one or more sentence structures from the plurality of sentence structures;

determining a candidate score for each candidate summary as a linear function of the scores assigned to each sentence structure included in the respective candidate summary;

learning coefficients of the linear function from a training dataset of documents and authored summaries via a predetermined learning algorithm; and

selecting, from the generated candidate summaries, a generated summary determined to have the highest candidate score under the linear function.

9. The method of claim 8 , further comprising generating the weighted score by:

determining a position of each sentence structure within the document; and

generating the weighted score of each sentence structure based on the determined position of each corresponding sentence structure.

10. The method of claim 8 , further comprising generating the weighted score by:

determining a content of each sentence structure within the document; and

generating the weighted score of each sentence structure based on the determined content of each corresponding sentence structure.

11. The method of claim 8 , further comprising generating the weighted score by:

identifying one or more tokens from each sentence structure, wherein each identified token corresponds to a token type;

determining a token type of each identified token; and

generating the weighted score of each sentence structure based on the determined token type of each identified token of each corresponding sentence structure.

12. The method of claim 8 , wherein assigning each sentence structure the weighted score based on the analysis and the target summary length comprises:

scoring each sentence structure as it appears in the document; and

scoring each sentence structure as it appears in a partial summary including one or more sentence structures already selected from the document.

13. The method of claim 8 , wherein the predetermined learning algorithm is a structured perception.

14. The method of claim 8 , wherein generating, in accordance with the summarization policy, the candidate summaries comprises:

considering each sentence structure included in the document;

revising each sentence score corresponding to a considered sentence structure; and

adding a highest-scoring sentence structure to the candidate summary such that the target summary length is not exceeded.

15. The method of claim 8 , wherein generating, in accordance with the summarization policy, the candidate summaries comprises:

considering each sentence structure in the document that appears after every sentence structure in the partial summary;

revising each sentence score corresponding to a considered sentence structure; and

adding a highest-scoring sentence structure to the candidate summary such that the target summary length is not exceeded.

16. The method of claim 8 , wherein generating, in accordance with the summarization policy, the candidate summaries comprises:

considering each sentence structure in the document that appears at a beginning or an end of a paragraph in the document;

revising each sentence score corresponding to a considered sentence structure; and

adding a highest-scoring sentence to the candidate summary such that the target summary length is not exceeded.

17. A method for generating a summary of a document, the method comprising:

receiving, through an interface, a viewing request from a communication device, the viewing request corresponding to a document including textual information stored on a memory;

in response to receiving the viewing request, determining a target summary length;

extracting the textual information from the document;

parsing the textual information;

identifying a plurality of sentence structures from the textual information based on the parsing;

analyzing each sentence structure from the plurality of sentence structures;

assigning each sentence structure a weighted score based on the analysis and the target summary length by:

scoring each sentence structure as it appears in the document; and

scoring each sentence structure as it appears in a partial summary including of one or more sentence structures already selected from the document:

choosing sentence structures to form a candidate summary in accordance with a summarization policy;

modeling a combination of scores for each sentence structure in the candidate summary as a linear function of the scores for each sentence structure;

learning coefficients of the linear function from a training dataset of documents and authored summaries via a predetermined learning algorithm;

selecting the candidate summary with the highest score under the linear function as the generated summary; and

generating a summary based on the selected candidate summary in accordance to a summarization strategy, wherein the summarization strategy is adaptable based on the target summary length.

Assignments (6)
PATENT SECURITY AGREEMENT (FIRST LIEN) Recorded Sep 29, 2022
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 061571/0773 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 058982/0282 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2018
From: YAHOO HOLDINGS, INC.
To: OATH INC.
Reel/Frame 045240/0310 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2017
From: YAHOO! INC.
To: YAHOO HOLDINGS, INC.
Reel/Frame 042963/0211 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2016
From: BILLAWALA, YOUSSEF; MEHDAD, YASHAR; RADEV, DRAGOMIR; STENT, AMANDA; THADANI, KAPIL
To: YAHOO! INC.
Reel/Frame 037710/0895 →
Cited By (1)
US 12,430,395