IP Library Granted Patent US 8,402,369
Granted Patent B2
US 8,402,369 · App. 12/199,104 · Granted Mar 19, 2013

Multiple-document summarization using document clustering

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,402,369
App. No.
12/199,104
Granted
Mar 19, 2013
Kind
B2
Abstract

Systems and methods are disclosed for summarizing multiple documents by generating a model of the documents as a mixture of document clusters, each document in turn having a mixture of sentences, wherein the model simultaneously representing summarization information and document cluster structure; and determining a loss function for evaluating the model and optimizing the model.

Claims (72)

1. A method for summarizing multiple documents, comprising:

a. generating a model BUV T of the documents as a mixture of document clusters, each document in turn having a mixture of sentences, wherein U is a corresponding sentence-topic matrix, V is a corresponding document-topic matrix, and B is a corresponding term-sentence matrix;

b. simultaneously representing summarization information and document cluster structure in the model BUV T ;

c. determining a loss function l, where the loss function comprises

l ( U,V )= KL ( A∥BUV T )− In Pr ( U,V ),

where KL is the Kullback-Leibler divergence, and where A is a corresponding term-document matrix;

d. evaluating and optimizing the model by translating a summarization and clustering problem into minimizing the loss l between the given documents and model reconstructed terms using the maximum likelihood estimation task comprising

U

,

V

=

arg

min

U

,

V

l

(

U

,

V

)

;

e. generating a summary of the documents at the same time as clustering documents into a given size of targeted summarization based on the model BUV T and the maximum likelihood estimation task.

2. The method of claim 1 , comprising receiving a document language model for each document.

3. The method of claim 2 , wherein the document language model comprises a unigram language model.

4. The method of claim 1 , comprising extracting sentence candidates from the document and receiving a sentence language model for each sentence candidate.

5. The method of claim 4 , wherein the sentence language model comprises a unigram language model.

6. The method of claim 1 , comprising determining model parameters U and V from a document language model and a sentence language model.

7. The method of claim 1 , comprising generating a matrix A of features of the documents or a matrix B of the features of the sentences.

8. The method of claim 1 , wherein each column of the model BUV T comprises features of a corresponding document generated by the model with parameters U and V.

9. The method of claim 1 , comprising formulating the model BUV T to approximate the document language model.

10. A method for summarizing multiple documents, comprising:

a. receiving a document language model for each document, wherein the document language models are used to generate a corresponding term-document matrix A;

b. extracting sentence candidates from the documents and receiving a sentence language model for each sentence candidate, wherein the sentence language models are used to generate a corresponding term-sentence matrix B;

c. determining model parameters U and V from the document language models and the sentence language models, wherein U is a corresponding sentence-topic matrix, and V is a corresponding document-topic matrix;

d. simultaneously representing summarization information and document cluster structure in a model BUV T of the documents as a mixture of document clusters, each document in turn having a mixture of sentences;

e. minimizing loss according to a loss function l between the documents and model reconstructed terms according to a maximum likelihood estimation task comprising

U

,

V

=

arg

min

l

(

U

,

V

)

U

,

V

;

and

f. summarizing the documents at the same time as clustering documents into a given size of targeted summarization based on the model BUV T and the maximum likelihood estimation task.

11. The method of claim 10 , wherein the document or sentence language model comprises a unigram language model.

12. The method of claim 11 , wherein A is a corresponding matrix of features of the documents.

13. The method of claim 11 , wherein B is a corresponding matrix of features of the sentences.

14. The method of claim 10 , wherein each column BUV T of the model comprises features of a corresponding document generated by the model with parameters U and V.

15. The method of claim 10 , wherein the loss function l comprises a Kullback-Leibler divergence function or a Frobenius matrix norm.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE 8538896 AND ADD 8583896 PREVIOUSLY RECORDED ON REEL 031998 FRAME 0667. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 30, 2017
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 042754/0703 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2014
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 031998/0667 →