IP Library Granted Patent US 8,015,188
Granted Patent B2
US 8,015,188 · App. 12/897,710 · Granted Sep 6, 2011

System and method for thematically grouping documents into clusters

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,015,188
App. No.
12/897,710
Granted
Sep 6, 2011
Kind
B2
Abstract

A system and method for thematically grouping documents into clusters is provided. Concepts are extracted from a plurality of documents. The concepts include nouns or noun phrases. A number of occurrences for each concept are determined within each document. A bounded range is applied to the concepts and a subset of the concepts is selected by removing the concepts that fall outside the bounded range. The bounded range includes upper edge conditions and lower edge conditions. Themes are generated from the subset of concepts by identifying two or more concepts with common semantic meaning. Clusters of the documents are generated based on the themes.

Claims (52)

1. A system for thematically grouping documents into clusters, comprising:

an extraction module to extract from a plurality of documents, concepts comprising at least one of nouns and noun phrases;

a frequency determination module to determine a number of occurrences for each concept within each document;

a threshold module to select the documents having the concepts with the occurrences that satisfy a bounded range comprising upper edge conditions and lower edge conditions;

a theme generator module to generate themes for the selected documents from the subset of concepts by identifying two or more concepts with common semantic meaning; and

a cluster module to generate clusters of the selected documents based on the themes; and

a processor to execute the modules.

2. A system according to claim 1 , further comprising:

a document module to remove duplicate documents from the clusters.

3. A system according to claim 2 , further comprising:

a reclustering module to recluster the remaining documents in the clusters.

4. A system according to claim 1 , wherein the clusters each comprise documents with related themes.

5. A system according to claim 1 , further comprising:

a total occurrence module to determine total occurrences for each concept by totaling the occurrences for that concept across all the documents; and

a lexicon to map the concepts against the total occurrences.

6. A system according to claim 1 , further comprising:

a graph module to generate a histogram for each document comprising the concepts extracted from that document and the concept occurrences, wherein the concepts are mapped in order of decreasing frequency to generate a curve representative of semantic content for that document.

7. A system according to claim 1 , further comprising:

a document module to determine a number of the documents that each include a common concept; and

a graph module to generate a corpus graph of the number of documents and common concepts.

8. A system according to claim 1 , further comprising:

a calculation module to determine inner products for at least one document based on the occurrences for one or more of the concepts in that document and a cluster concept weighting for the same concept in at least one of the clusters; and

a document module to identify documents having the smallest inner products as most relevant to the theme represented by the at least one cluster.

9. A system according to claim 1 , further comprising:

a database record for each concept comprising a concept identifier, string, and occurrence frequency.

10. A system according to claim 1 , wherein the documents comprise at least one of electronic messages, word processing documents, and hypertext documents.

11. A method for thematically grouping documents into clusters, comprising the steps of:

extracting from a plurality of documents, concepts comprising at least one of nouns and noun phrases;

determining a number of occurrences for each concept within each document;

selecting the documents having the concepts with the occurrences that satisfy a bounded range comprising upper edge conditions and lower edge conditions;

generating themes for the selected documents from the subset of concepts by identifying two or more concepts with common semantic meaning; and

generating clusters of the selected documents based on the themes,

wherein the steps are performed by a suitably programmed computer.

12. A method according to claim 11 , further comprising:

removing duplicate documents from the clusters.

13. A method according to claim 12 , further comprising:

reclustering the remaining documents in the clusters.

14. A method according to claim 11 , wherein the clusters each comprise documents with related themes.

15. A method according to claim 11 , further comprising:

determining total occurrences for each concept by totaling the occurrences for that concept across all the documents; and

mapping the concepts against the total occurrences via a lexicon.

16. A method according to claim 11 , further comprising:

generating a histogram for each document comprising the concepts extracted from that document and the concept occurrences, wherein the concepts are mapped in order of decreasing frequency to generate a curve representative of semantic content for that document.

17. A method according to claim 11 , further comprising:

determining a number of the documents that each include a common concept; and

generating a corpus graph of the number of documents and common concepts.

18. A method according to claim 11 , further comprising:

determining inner products for at least one document based on the occurrences for one or more of the concepts in that document and a cluster concept weighting for the same concept in at least one of the clusters; and

identifying documents having the smallest inner products as most relevant to the theme represented by the at least one cluster.

19. A method according to claim 11 , further comprising:

generating a database record for each concept comprising a concept identifier, string, and occurrence frequency.

20. A method according to claim 11 , wherein the documents comprise at least one of electronic messages, word processing documents, and hypertext documents.

Assignments (8)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2020
From: ATTENEX CORPORATION
To: FTI TECHNOLOGY LLC
Reel/Frame 051780/0679 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2020
From: GALLIVAN, DAN; KAWAI, KENJI
To: ATTENEX CORPORATION
Reel/Frame 051679/0205 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2018
From: FTI CONSULTING TECHNOLOGY LLC
To: NUIX NORTH AMERICA INC.
Reel/Frame 047237/0019 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS AT REEL/FRAME 036031/0637 Recorded Sep 12, 2018
From: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
To: FTI CONSULTING TECHNOLOGY LLC
Reel/Frame 047060/0107 →
CHANGE OF NAME Recorded Apr 20, 2018
From: FTI TECHNOLOGY LLC
To: FTI CONSULTING TECHNOLOGY LLC
Reel/Frame 045785/0645 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 29, 2015
From: BANK OF AMERICA, N.A.
To: FTI CONSULTING, INC.; FTI CONSULTING TECHNOLOGY LLC
Reel/Frame 036029/0233 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Jun 29, 2015
From: FTI CONSULTING, INC.; FTI CONSULTING TECHNOLOGY LLC; FTI CONSULTING TECHNOLOGY SOFTWARE CORP
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 036031/0637 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Dec 10, 2012
From: FTI CONSULTING, INC.; FTI CONSULTING TECHNOLOGY LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 029434/0087 →