IP Library Granted Patent US 9,195,399
Granted Patent B2
US 9,195,399 · App. 14/275,802 · Granted Nov 24, 2015

Computer-implemented system and method for identifying relevant documents for display

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,195,399
App. No.
14/275,802
Granted
Nov 24, 2015
Kind
B2
Abstract

A computer-implemented system and method for identifying relevant documents for display are provided. Themes for a set of documents are generated. The documents are clustered based on the themes. A matrix including an inner product of document frequency occurrences and cluster concept weightings for each theme is generated for the documents. From the matrix, documents most relevant to a particular theme are identified, and the relevant documents are displayed.

Claims (87)

1. A computer-implemented system for identifying relevant documents for display, comprising:

themes for a set of documents;

an extraction module to extract noun phrases from the documents as concepts;

a theme generator to group two or more of the concepts as one such theme;

a frequency table that identifies each of the concepts and a frequency of occurrence of each concept within each of the documents in the set;

a graph generator to generate a graph of the concepts, comprising:

an x-axis of the graph defining the concepts;

a y-axis of the graph defining a number of the documents that reference each concept; and

a mapping module to map the concepts on the graph in order of descending number of referring documents;

a cluster module to cluster the documents based on the themes;

a matrix for the documents comprising an inner product of document frequency occurrences and cluster concept weightings for each theme;

an identification module to identify from the matrix, documents most relevant to a particular theme; and

a display to present the relevant documents.

2. A system according to claim 1 , further comprising:

an assignment module to assign an identifier to each of the concepts, wherein each identifier is a monotonically increasing integer value.

3. A system according to claim 1 , further comprising:

for each document, a histogram comprising a normalized representation of the frequencies of occurrence for those concepts extracted from one such document.

4. A system according to claim 1 , further comprising:

a value determination module to identify a median value of the frequencies of occurrence along the x-axis;

a threshold module to set edge conditions based on the frequencies of occurrence; and

a selection module to select those documents with concepts that fall within the edge conditions as forming a subset of documents that include latent concepts.

5. A system according to claim 1 , further comprising:

a cluster mapping module to map one such document to one of the clusters based on a similarity of the document to the documents in the cluster, wherein the distance is quantified by a distance.

6. A system according to claim 5 , further comprising:

a distance determination module to calculate the distance using the following equation:

d

cluster

=

i

->

n

doc

term

i

·

cluster

term

i

where doc term represents a frequency of occurrence for a given term i in the one such document and cluster term represents the weight of a given cluster for a given term i.

7. A system according to claim 1 , further comprising:

a removal module to remove duplicates of the relevant documents prior to display.

8. A computer-implemented method for identifying relevant documents for display, comprising:

generating themes for a set of documents, comprising:

extracting noun phrases from the documents as concepts; and

grouping two or more of the concepts as one such theme;

generating a frequency table that identifies each of the concepts and a frequency of occurrence of each concept within each of the documents in the set;

generating a graph of the concepts, comprising:

defining an x-axis of the graph as the concepts;

defining a y-axis of the graph as a number of the documents that reference each of the concepts; and

mapping the concepts on the graph in order of descending number of referring documents;

clustering the documents based on the themes;

generating a matrix for the documents comprising an inner product of document frequency occurrences and cluster concept weightings for each theme;

identifying from the matrix, documents most relevant to a particular theme; and

displaying the relevant documents.

9. A method according to claim 8 , further comprising:

assigning an identifier to each of the concepts, wherein each identifier is a monotonically increasing integer value.

10. A method according to claim 8 , further comprising:

generating for each document, a histogram comprising a normalized representation of the frequencies of occurrence for those concepts extracted from one such document.

11. A method according to claim 8 , further comprising:

identifying a median value of the frequencies of occurrence along the x-axis;

setting edge conditions based on the frequencies of occurrence; and

selecting those documents with concepts that fall within the edge conditions as forming a subset of documents that include latent concepts.

12. A method according to claim 8 , further comprising:

mapping one such document to one of the clusters based on a similarity of the document to the documents in the cluster, wherein the distance is quantified by a distance.

13. A method according to claim 12 , further comprising:

calculating the distance using the following equation:

d

cluster

=

i

n

doc

term

i

·

cluster

term

i

where doc term represents a frequency of occurrence for a given term i in the one such document and cluster term represents the weight of a given cluster for a given term i.

14. A method according to claim 8 , further comprising:

removing duplicates of the relevant documents prior to display.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2020
From: GALLIVAN, DAN; KAWAI, KENJI
To: ATTENEX CORPORATION
Reel/Frame 051679/0205 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2020
From: ATTENEX CORPORATION
To: FTI TECHNOLOGY LLC
Reel/Frame 051679/0232 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2018
From: FTI CONSULTING TECHNOLOGY LLC
To: NUIX NORTH AMERICA INC.
Reel/Frame 047237/0019 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS AT REEL/FRAME 036031/0637 Recorded Sep 12, 2018
From: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
To: FTI CONSULTING TECHNOLOGY LLC
Reel/Frame 047060/0107 →
CHANGE OF NAME Recorded Apr 20, 2018
From: FTI TECHNOLOGY LLC
To: FTI CONSULTING TECHNOLOGY LLC
Reel/Frame 045785/0645 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Jun 29, 2015
From: FTI CONSULTING, INC.; FTI CONSULTING TECHNOLOGY LLC; FTI CONSULTING TECHNOLOGY SOFTWARE CORP
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 036031/0637 →