IP Library Granted Patent US 8,626,761
Granted Patent B2
US 8,626,761 · App. 12/606,171 · Granted Jan 7, 2014

System and method for scoring concepts in a document set

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,626,761
App. No.
12/606,171
Granted
Jan 7, 2014
Kind
B2
Abstract

A system and method for scoring concepts in a document set is provided. Concepts including two or more terms extracted from the document set are identified. Each document having one or more of the concepts is designated as a candidate seed document. A score is calculated for each of the concepts identified within each candidate seed document based on a frequency of occurrence, concept weight, structural weight, and corpus weight. A vector is formed for each candidate seed document. The vector is compared with a center of one or more clusters each comprising thematically-related documents. At least one of the candidate seed documents that is sufficiently distinct from the other candidate seed documents is selected as a seed document for a new cluster. Each of the unselected candidate seed documents is placed into one of the clusters having a most similar cluster center.

Claims (95)

1. A system for scoring concepts in a document set, comprising:

a database to maintain a set of documents;

a concept identification module to identify concepts comprising two or more terms extracted from the document set and to designate each document having one or more of the concepts as a candidate seed document;

a value module to determine for each of the concepts identified within each candidate seed document, values for a frequency of occurrence of that concept within that candidate seed document, a concept weight reflecting a specificity of meaning for that concept within that candidate seed document, a structural weight reflecting a degree of significance based on a location of that concept within that candidate seed document, and a corpus weight inversely weighing a reference count of the occurrence for that concept within the document set according to the equation:

r

w

ij

=

{

(

T

-

r

ij

T

)

2

,

r

ij

>

M

1.0

,

r

ij

M

where rw ij comprises the corpus weight, r ij comprises the reference count for the occurrence j for that concept i, T comprises a total number of reference counts of the documents in the document set, and M comprises a maximum reference count of the documents in the document set;

a scoring module to calculate a score for each concept as a function of summation of the values for each of the frequency of occurrence, concept weight, structural weight, and corpus weight;

a vector module to form a vector for each candidate seed document comprising the concepts located in that candidate seed document and the associated concept scores based on the frequency of occurrence, concept weight, structural weight, and corpus weight;

a document comparison module to compare the vector for each candidate seed document with a center of one or more clusters each comprising thematically-related documents and to select at least one of the candidate seed documents that is sufficiently distinct from the other candidate seed documents as a seed document for a new cluster; and

a clustering module to place each of the unselected candidate seed documents into one of the clusters having a most similar cluster center.

2. A system according to claim 1 , further comprising:

a similarity module to determine a similarity between each candidate seed document and each of the one or more cluster centers based on the comparison.

3. A system according to claim 2 , wherein the similarity is determined as an inner product of the candidate seed document and the cluster center.

4. A system according to claim 1 , further comprising:

a preprocessing module to convert each document into a document record and to preprocess the document records to obtain the extracted terms.

5. A system according to claim 1 , wherein the clustering module applies a minimum fit criterion to the placement of the unselected candidate seed documents.

6. A system according to claim 1 , further comprising:

a document relocation module to apply a threshold to each of the clusters, to select those documents within at least one of the clusters that falls outside of the threshold as outlier documents, and to relocate the outlier documents.

7. A system according to claim 6 , wherein each outlier document is placed into the cluster having a best fit based on measures of similarity between that outlier documents and that cluster.

8. A system according to claim 1 , further comprising:

a score compression module to compress the concept scores.

9. A method for scoring concepts in a document set, comprising:

maintaining a set of documents;

identifying concepts comprising two or more terms extracted from the document set and designating each document having one or more of the concepts as a candidate seed document;

determining for each of the concepts identified within each candidate seed document, values for a frequency of occurrence of that concept within that candidate seed document, a concept weight reflecting a specificity of meaning for that concept within that candidate seed document, a structural weight reflecting a degree of significance based on a location of that concept within that candidate seed document, and a corpus weight inversely weighing a reference count of the occurrence for that concept within the document set according to the equation:

r

w

ij

=

{

(

T

-

r

ij

T

)

2

,

r

ij

>

M

1.0

,

r

ij

M

where rw ij comprises the corpus weight, r ij comprises the reference count for the occurrence j for that concept i, T comprises a total number of reference counts of the documents in the document set, and M comprises a maximum reference count of the documents in the document set;

calculating a score for each concept as a function of summation of the values for each of the frequency of occurrence, concept weight, structural weight, and corpus weight;

forming a vector for each candidate seed document comprising the concepts located in that candidate seed document and the associated concept scores based on the frequency of occurrence, concept weight, structural weight, and corpus weight;

comparing the vector for each candidate seed document with a center of one or more clusters each comprising thematically-related documents and selecting at least one of the candidate seed documents that is sufficiently distinct from the other candidate seed documents as a seed document for a new cluster; and

placing each of the unselected candidate seed documents into one of the clusters having a most similar cluster center.

10. A method according to claim 9 , further comprising:

determining a similarity between each candidate seed document and each of the one or more cluster centers based on the comparison.

11. A method according to claim 10 , wherein the similarity is determined as an inner product of the candidate seed document and the cluster center.

12. A method according to claim 9 , further comprising:

converting each document into a document record; and

preprocessing the document records to obtain the extracted terms.

13. A method according to claim 9 , further comprising:

applying a minimum fit criterion to the placement of the unselected candidate seed documents.

14. A method according to claim 9 , further comprising:

applying a threshold to each of the clusters and selecting those documents within at least one of the clusters that falls outside of the threshold as outlier documents; and

relocating the outlier documents.

15. A method according to claim 14 , wherein each outlier document is placed into the cluster having a best fit based on measures of similarity between that outlier documents and that cluster.

16. A method according to claim 9 , further comprising:

compressing the concept scores.

Assignments (11)
SECURITY INTEREST Recorded Apr 4, 2024
From: NUIX NORTH AMERICA INC.
To: THE HONGKONG AND SHANGHAI BANKING CORPORATION LIMITED, SYDNEY BRANCH, AS SECURED PARTY
Reel/Frame 067005/0073 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2019
From: KAWAI, KENJI; EVANS, LYNNE MARIE
To: ATTENEX CORPORATION
Reel/Frame 048749/0445 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2019
From: ATTENEX CORPORATION
To: FTI TECHNOLOGY LLC
Reel/Frame 048749/0452 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2018
From: FTI CONSULTING TECHNOLOGY LLC
To: NUIX NORTH AMERICA INC.
Reel/Frame 047237/0019 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS AT REEL/FRAME 036031/0637 Recorded Sep 12, 2018
From: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
To: FTI CONSULTING TECHNOLOGY LLC
Reel/Frame 047060/0107 →
CHANGE OF NAME Recorded Apr 20, 2018
From: FTI TECHNOLOGY LLC
To: FTI CONSULTING TECHNOLOGY LLC
Reel/Frame 045785/0645 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Jun 29, 2015
From: BANK OF AMERICA, N.A.
To: FTI CONSULTING, INC.; FTI CONSULTING TECHNOLOGY LLC
Reel/Frame 036029/0233 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Jun 29, 2015
From: FTI CONSULTING, INC.; FTI CONSULTING TECHNOLOGY LLC; FTI CONSULTING TECHNOLOGY SOFTWARE CORP
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 036031/0637 →
RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 11, 2012
From: BANK OF AMERICA, N.A.
To: FTI CONSULTING, INC.; FTI TECHNOLOGY LLC; ATTENEX CORPORATION
Reel/Frame 029449/0389 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Dec 10, 2012
From: FTI CONSULTING, INC.; FTI CONSULTING TECHNOLOGY LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 029434/0087 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Mar 14, 2011
From: FTI CONSULTING, INC.; FTI TECHNOLOGY LLC; ATTENEX CORPORATION
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 025943/0038 →