IP Library Granted Patent US 9,619,551
Granted Patent B2
US 9,619,551 · App. 14/949,829 · Granted Apr 11, 2017

Computer-implemented system and method for generating document groupings for display

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,619,551
App. No.
14/949,829
Granted
Apr 11, 2017
Kind
B2
Abstract

A computer-implemented system and method for generating document groupings is provided. A lexicon of terms extracted from a set of documents is generated. The lexicon includes a frequency of each extracted term within each document in the set. Concepts each having two or more of the extracted terms are generated. A subset of the documents in the set is selected based on the term frequencies. The subset of documents is grouped into clusters based on the concepts. A similarity of each document cluster is calculated with one or more documents based on a distance by summing the frequency of each term in that document and a weight of the cluster for each of the terms. The weights are updated until a rate of change for each cluster becomes constant.

Claims (79)

1. A computer-implemented system for generating document groupings, comprising:

a database to store a set of document, a lexicon of terms extracted from the set of documents and comprising a frequency of each extracted term within each document, and concepts each comprising two or more of the extracted terms; and

a server comprising a central processing unit, memory, an input port to receive the documents, lexicon and concepts from the database, and an output port, wherein the central processing unit is configured to:

select a subset of the documents in the set based on the term frequencies;

group the subset of documents into clusters based on the concepts;

calculate a similarity of each document cluster with at least one document based on a distance by summing the frequency of each term in that document and a weight of the cluster for each of the terms; and

update the weights until a rate of change for each cluster becomes constant.

2. A system according to claim 1 , wherein the central processing unit further identifies one or more of the clusters as meaningful by defining a variance, selects those clusters having one or more documents with inner products that fall within the variance as meaningful, and presents those documents identified as meaningful.

3. A system according to claim 1 , wherein the central processing unit further represents semantic content of each document by mapping the terms in that document in order of decreasing frequency.

4. A system according to claim 1 , wherein the central processing unit further represents latent semantics of the document set by mapping each term in the document set against a total frequency occurrence, wherein the total frequency occurrence is calculated as a sum of the term frequencies within each document in the set.

5. A system according to claim 4 , wherein the central processing unit further selects a median value and edge conditions for the total frequency occurrences of the terms and generates the subset of documents from those documents in the set that satisfy the edge conditions.

6. A system according to claim 5 , wherein the central processing unit further re-centers the median value and to generate a different subset of documents for grouping.

7. A system according to claim 5 , wherein the central processing unit further sets the edge conditions based on a size of the documents.

8. A system according to claim 7 , wherein larger documents have tighter edge conditions than shorter documents.

9. A system according to claim 1 , wherein the central processing unit further calculates the distance using the following equation:

d

cluster

=

i

n

doc

term

i

·

cluster

term

i

where doc term represents the frequency of a given term i in one such document and cluster term represents the weight of a given cluster for a given term i.

10. A system according to claim 1 , wherein the central processing unit further calculates the rate of change by determining a first derivative of the inner products over successive iterations.

11. A computer-implemented method for generating document groupings, comprising:

generating a lexicon of terms extracted from a set of documents and comprising a frequency of each extracted term within each document;

generating concepts each comprising two or more of the extracted terms;

selecting a subset of the documents in the set based on the term frequencies;

grouping the subset of documents into clusters based on the concepts;

calculating a similarity of each document cluster with at least one document based on a distance by summing the frequency of each term in that document and a weight of the cluster for each of the terms; and

updating the weights until a rate of change for each cluster becomes constant.

12. A method according to claim 11 , further comprising:

identifying one or more of the clusters as meaningful, comprising:

defining a variance; and

selecting those clusters having one or more documents with inner products that fall within the variance as meaningful; and

displaying those documents identified as meaningful.

13. A method according to claim 11 , further comprising:

representing semantic content of each document by mapping the terms in that document in order of decreasing frequency.

14. A method according to claim 11 , further comprising:

representing latent semantics of the document set by mapping each term in the document set against a total frequency occurrence, wherein the total frequency occurrence is calculated as a sum of the term frequencies within each document in the set.

15. A method according to claim 14 , further comprising:

selecting a median value and edge conditions for the total frequency occurrences of the terms; and

generating the subset of documents from those documents in the set that satisfy the edge conditions.

16. A method according to claim 15 , further comprising:

re-centering the median value; and

generating a different subset of documents for grouping.

17. A method according to claim 15 , further comprising:

setting the edge conditions based on a size of the documents.

18. A method according to claim 17 , wherein larger documents have tighter edge conditions than shorter documents.

19. A method according to claim 11 , further comprising:

calculating the distance using the following equation:

d

cluster

=

i

n

doc

term

i

·

cluster

term

i

where doc term represents the frequency of a given term i in one such document and cluster term represents the weight of a given cluster for a given term i.

20. A method according to claim 11 , further comprising:

calculating the rate of change by determining a first derivative of the inner products over successive iterations.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2020
From: GALLIVAN, DAN; KAWAI, KENJI
To: ATTENEX CORPORATION
Reel/Frame 051679/0205 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2020
From: ATTENEX CORPORATION
To: FTI TECHNOLOGY LLC
Reel/Frame 051679/0232 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2018
From: FTI CONSULTING TECHNOLOGY LLC
To: NUIX NORTH AMERICA INC.
Reel/Frame 047237/0019 →
CHANGE OF NAME Recorded Apr 20, 2018
From: FTI TECHNOLOGY LLC
To: FTI CONSULTING TECHNOLOGY LLC
Reel/Frame 045785/0645 →