IP Library Granted Patent US 9,317,593
Granted Patent B2
US 9,317,593 · App. 12/243,267 · Granted Apr 19, 2016

Modeling topics using statistical distributions

Inventors: David L. Marvit (San Francisco, CA); Jawahar Jain (Los Altos, CA); Stergios Stergiou (Sunnyvale, CA); Alex Gilman (Fremont, CA); B. Thomas Adler (Santa Cruz, CA); John J. Sidorowich (Santa Cruz, CA); Yannis Labrou (Washington, DC)
Assignee: Fujitsu Limited
G06F17/3071G06F17/30616
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,317,593
App. No.
12/243,267
Granted
Apr 19, 2016
Kind
B2
Abstract

In one embodiment, modeling topics includes accessing a corpus comprising documents that include words. Words of a document are selected as keywords of the document. The documents are clustered according to the keywords to yield clusters, where each cluster corresponds to a topic. A statistical distribution is generated for a cluster from words of the documents of the cluster. A topic is modeled using the statistical distribution generated for the cluster corresponding to the topic.

Claims (48)

1. A computer-implemented method comprising:

accessing a corpus stored in one or more tangible media, the corpus comprising a plurality of documents, a document comprising a plurality of words;

selecting one or more words of each document as one or more keywords of the each document;

clustering the documents according to the keywords to yield a plurality of clusters, each cluster corresponding to a different topic;

generating a statistical distribution for each cluster from a subset of the words of the documents of the each cluster to yield a plurality of statistical distributions, wherein generating the statistical distribution for the each cluster comprises:

determining a co-occurrence value indicating a co-occurrence of the topic of the each cluster with the topics of the other clusters in the plurality of documents; and

generating a co-occurrence distribution from the co-occurrence values;

modeling each topic using the statistical distribution generated for the cluster corresponding to the each topic;

organizing the clusters according to the statistical distributions; and

assigning the topics of the organized clusters to the documents in the organized clusters.

2. The method of claim 1 , the selecting the one or more words of the each document further comprising:

ranking the words of the each document according to a ranking technique; and

selecting one or more highly ranked words as the one or more keywords.

3. The method of claim 1 , the clustering the documents according to the keywords to yield the plurality of clusters further comprising:

removing one or more clusters that fail to satisfy a size threshold.

4. The method of claim 1 , the generating the statistical distribution for the each cluster further comprising:

calculating a term frequency of each word of the subset of the words to yield a plurality of term frequencies; and

generating a term distribution from the term frequencies.

5. The method of claim 1 , the generating the statistical distribution for the each cluster further comprising:

calculating a number of documents that include each word of the subset of the words; and

generating a term distribution from the numbers of documents.

6. The method of claim 1 , further comprising:

identifying at least two clusters with similar statistical distributions; and

consolidating the at least two clusters.

7. One or more non-transitory computer-readable tangible media encoding software operable when executed to:

access a corpus stored in one or more tangible media, the corpus comprising a plurality of documents, a document comprising a plurality of words;

select one or more words of each document as one or more keywords of the each document;

cluster the documents according to the keywords to yield a plurality of clusters, each cluster corresponding to a different topic;

generate a statistical distribution for each cluster from a subset of the words of the documents of the each cluster to yield a plurality of statistical distributions, wherein generating the statistical distribution for the each cluster comprises:

determining a co-occurrence value indicating a co-occurrence of the topic of the each cluster with the topics of the other clusters in the plurality of documents; and

generating a co-occurrence distribution from the co-occurrence values; and

model each topic using the statistical distribution generated for the cluster corresponding to the each topic;

organize the clusters according to the statistical distributions; and

assign the topics of the organized clusters to the documents in the organized clusters.

8. The computer-readable tangible media of claim 7 , further operable to select the one or more words of the each document by:

ranking the words of the each document according to a ranking technique; and

selecting one or more highly ranked words as the one or more keywords.

9. The computer-readable tangible media of claim 7 , further operable to cluster the documents according to the keywords to yield the plurality of clusters by:

removing one or more clusters that fail to satisfy a size threshold.

10. The computer-readable tangible media of claim 7 , further operable to generate the statistical distribution for the each cluster by:

calculating a term frequency of each word of the subset of the words to yield a plurality of term frequencies; and

generating a term distribution from the term frequencies.

11. The computer-readable tangible media of claim 7 , further operable to generate the statistical distribution for the each cluster by:

calculating a number of documents that include each word of the subset of the words; and

generating a term distribution from the numbers of documents.

12. The computer-readable tangible media of claim 7 , further operable to:

identify at least two clusters with similar statistical distributions; and

consolidate the at least two clusters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2008
From: MARVIT, DAVID L.; JAIN, JAWAHAR; STERGIOU, STERGIOS; GILMAN, ALEX; ADLER, B. THOMAS; SIDOROWICH, JOHN J.; LABROU, YANNIS
To: FUJITSU LIMITED
Reel/Frame 021616/0268 →
Continuity (2)
Provisional Application 60977855 · Oct 5, 2007
Related Publication 20090094233A1 · Apr 9, 2009