IP Library Granted Patent US 7,607,083
Granted Patent B2
US 7,607,083 · App. 09/817,591 · Granted Oct 20, 2009

Test summarization using relevance measures and latent semantic analysis

Assignee: NEC Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,607,083
App. No.
09/817,591
Granted
Oct 20, 2009
Kind
B2
Abstract

Text summarizers using relevance measurement technologies and latent semantic analysis techniques provide accurate and useful summarization of the contents of text documents. Generic text summaries may be produced by ranking and extracting sentences from original documents; broad coverage of document content and decreased redundancy may simultaneously be achieved by constructing summaries from sentences that are highly ranked and different from each other. In one embodiment, conventional Information Retrieval (IR) technologies may be applied in a unique way to perform the summarization; relevance measurement, sentence selection, and term elimination may be repeated in successive iterations. In another embodiment, a singular value decomposition technique may be applied to a terms-by-sentences matrix such that all the sentences from the document may be projected into the singular vector space; a text summarizer may then select sentences having the largest index values with the most important singular vectors as part of the text summary.

Claims (49)

1. A method of creating a generic text summary of a document; said method comprising:

obtaining the document;

creating a weighted document term-frequency vector for said document;

for each sentence in said document, creating a weighted sentence term-frequency vector;

computing a score for each said weighted sentence term-frequency vector in accordance with relevance to said weighted document term-frequency vector;

selecting a sentence for inclusion in said generic text summary in accordance with said computing, wherein the selected sentence has the computed score representing high degree of relevance of the corresponding weighted sentence term-frequency vector to said weighted document term-frequency vector;

deleting said selected sentence from said document and eliminating terms in said selected sentence from said document; and

generating the generic text summary based on the selected sentence.

2. The method of claim 1 further comprising:

recreating said weighted document term-frequency vector in accordance with said deleting and said eliminating; and

selectively repeating said computing, said selecting, said deleting, said eliminating, and said recreating.

3. The method of claim 2 wherein said selectively repeating is terminated when a predetermined number of sentences has been selected.

4. The method of claim 1 wherein said computing comprises calculating an inner product of said weighted sentence term-frequency vector and said weighted document term-frequency vector.

5. The method of claim 1 wherein said creating a weighted sentence term-frequency vector comprises implementing a local weighting function and implementing a global weighting function.

6. The method of claim 5 wherein said creating a weighted sentence term-frequency vector comprises normalizing each said weighted sentence term-frequency vector by dividing the weighted sentence term-frequency vector by a magnitude of the weighted sentence term-frequency vector.

7. The method of claim 1 wherein said creating a weighted document term-frequency vector comprises implementing a local weighting function and implementing a global weighting function.

8. The method of claim 7 wherein said creating a weighted document term-frequency vector comprises normalizing said weighted document term-frequency vector by dividing the weighted document term-frequency vector by a magnitude of the weighted document term-frequency vector.

9. A system for creating a generic text summary of a document; said system comprising:

a computer comprising at least a CPU and a memory;

an interface for obtaining the document;

a display for displaying said generic text summary; and

summarizer program code, operable on said computer, for analyzing and summarizing said document; said summarizer program code comprising:

a vector generator for creating a weighted document term-frequency vector for said document and creating a weighted sentence term-frequency vector for each sentence in said document;

a scoring engine for computing a score for each said weighted sentence term-frequency vector in accordance with relevance to said weighted document term-frequency vector;

a selector for selecting a sentence for inclusion in said generic text summary in accordance with output results from said scoring engine;

a document editor for deleting said selected sentence from said document and for eliminating terms in said selected sentence from said document and;

a generic summary generator for generating the generic text summary based on the selected sentence.

10. The system of claim 9 wherein said vector generator recreates said weighted document term-frequency vector in accordance with output results from said document editor.

11. The system of claim 10 wherein said summarizer further comprises a loop routine for generating iterative sequential operations of said vector generator, said scoring engine, said selector, and said document editor.

12. The system of claim 11 wherein said loop routine is responsive to a predetermined limit such that said generic text summary is of a predetermined number of sentences.

13. A method of creating a generic text summary of a document; said method comprising:

obtaining the document;

decomposing said document into individual sentences;

forming a candidate sentence set from said individual sentences;

for each of said individual sentences in said candidate sentence set, creating a weighted sentence term-frequency vector;

creating a weighted document term-frequency vector for said document;

for each of said individual sentences in said candidate sentence set, computing a relevance score for said weighted sentence term-frequency vector relative to said weighted document term-frequency vector;

selecting a sentence for inclusion in said generic text summary in accordance with said computing, wherein the selected sentence has the computed relevance score representing a high degree of relevance of the corresponding weighted sentence term-frequency vector to said weighted document term-frequency vector;

deleting said selected sentence from said candidate sentence set;

eliminating terms in said selected sentence from said document;

recreating said weighted document term-frequency vector in accordance with said deleting and said eliminating; and

generating the generic text summary based on the selected sentence.

14. The method of claim 13 further comprising: selectively repeating said computing, said selecting, said deleting, said eliminating, and said recreating.

15. The method of claim 14 wherein said selectively repeating is terminated when a predetermined number of sentences has been selected.

16. The method of claim 13 wherein said computing comprises calculating an inner product of said weighted sentence term-frequency vector and said weighted document term-frequency vector.

17. The method of claim 13 wherein said creating a weighted sentence term-frequency vector comprises implementing a local weighting function and implementing a global weighting function.

18. The method of claim 17 wherein said creating a weighted sentence term-frequency vector comprises normalizing each said weighted sentence term- frequency vector by dividing the weighted sentence term-frequency vector by a magnitude of the weighted sentence term-frequency vector.

19. The method of claim 13 wherein said creating a weighted document term-frequency vector comprises implementing a local weighting function and implementing a global weighting function.

20. The method of claim 19 wherein said creating a weighted document term-frequency vector comprises normalizing said weighted document term-frequency vector by dividing the weighted document term-frequency vector by a magnitude of the weighted document term-frequency vector.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2003
From: NEC USA, INC.
To: NEC CORPORATION
Reel/Frame 013926/0288 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 16, 2001
From: GONG, YIHONG; LIU, XIN
To: NEC USA, INC.
Reel/Frame 011989/0465 →
Continuity (2)
Provisional Application 6025453500 · Dec 12, 2000
Related Publication 20020138528A1 · Sep 26, 2002