IP Library › Granted Patent US 7,280,957
Granted Patent B2
US 7,280,957 · App. 10/321,420 · Granted Oct 9, 2007

Method and apparatus for generating overview information for hierarchically related information

Assignee: Palo Alto Research Center, Incorporated
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,280,957
App. No.
10/321,420
Granted
Oct 9, 2007
Kind
B2
Abstract

A method is provided for digesting the content of hierarchically related information. The method, which obtains relatively short overviews, selects a proportion of representative nodes and then extracts and organizes one or more sentences from the text associated with each selected node. For text trees representing archived discussions, the selection of nodes and sentences is from comment/response sequences drawn from lexically central nodes which will capture those aspects of the discussion considered most important to discussion participants.

Claims (48)

1. A method of forming an overview for hierarchically related information, comprising:

forming a vector for each document of hierarchically related information that comprises one position for each unique word contained in the document;

representing each document as a node in a cluster that initially comprises the node for only that document;

successively combining pairs of the clusters by matching the vectors for the documents of the nodes in each cluster pair that are most similar, until an application-specific criteria is met;

identifying lexically central nodes in each of the clusters, comprising:

determining a lexical centroid of the vectors for the documents of the nodes in the cluster;

selecting those nodes for the documents having vectors in the lexical centroid as the lexically central nodes; and

determining a remainder set comprising the nodes in the cluster that are not one of the lexically central nodes,

identifying auxiliary nodes comprising:

selecting the nodes from the remainder set that are for a document that is a parent of a document for one of the lexically central nodes for the cluster; and

selecting the nodes from the remainder set that are for a document that is a parent of a specified plurality of documents for the nodes in the cluster,

combining the lexically central nodes and the auxiliary nodes to form an extraction node set,

identifying quoting interactions in the document for each node in the extraction node set and selecting an overview string from the document based on the quoting interactions, and

combining the overview strings to form an overview for the hierarchically related information.

2. The method of claim 1 further comprising:

determining whether each node in the extraction set is a root node, which is the node for a root document for plurality of hierarchically related documents; and

selecting an initial string from root document.

3. The method of claim 1 further comprising:

determining whether each node in the extraction set is one of the lexically central nodes, and if the document for the node comprises one or more quoting portions, and

selecting at least one string from the document following a quoting portion.

4. The method of claim 3 wherein an initial string is selected from the document.

5. The method of claim 3 further comprising:

selecting the most lexically central string.

6. The method of claim 1 further comprising:

determining whether the document for the node comprises one or more quoted portions, and

selecting at least one string from the document comprising part of at least one quoted portion.

7. The method of claim 1 further comprising:

determining whether the node is an auxiliary node and the document for the node does not contain any quoted portions, and

selecting at least one string from a final part of the document.

8. The method of claim 1 , further comprising:

determining a lexical centroid of the vectors for selected groups of adjacent nodes; and

selecting the node for the document with the vector closest to the lexical centroid of each such group.

9. The method of 1 , further comprising:

determining whether the nodes of a cluster include conversation root node, which is the node for a root document for a plurality of hierarchically related documents of a tree representing a stored conversation, and

selecting the root node and a proportion of its children nodes in the cluster as the lexically central nodes.

10. The method of claim 1 , further comprising:

determining a vector sum of the vectors for the cluster divided by a count of the vectors; and

finding the vectors for the cluster closest to the lexical centroid, wherein the nodes for the documents for these vectors comprise the lexically central nodes.

11. The method of claim 1 , further comprising:

filling each position in the vector with a weighted frequency for the unique word.

12. The method of claim 11 , further comprising:

stemming the words contained in the document to lemma forms;

ignoring the words contained in the document that occur in a stop list of common words; and

determining the weighted frequency as a tf.idf weighted frequency.

13. The method of claim 1 , further comprising:

filling each position in the vector with a word probability for the unique word.

14. The method of claim 13 , further comprising:

determining the word probability through probabilistic latent semantic indexing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2003
From: NEWMAN, PAULA S.; BLITZER, JOHN C.; BRANTS, THORSTEN H.
To: PALO ALTO RESEARCH CENTER, INCORPORATED
Reel/Frame 013904/0943 →
Continuity (1)
Related Publication 20040117449A1 · Jun 17, 2004