IP Library Granted Patent US 7,844,566
Granted Patent B2
US 7,844,566 · App. 11/431,664 · Granted Nov 30, 2010

Latent semantic clustering

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,844,566
App. No.
11/431,664
Granted
Nov 30, 2010
Kind
B2
Abstract

An embodiment of the present invention provides a computer-based method for automatically identifying clusters of conceptually-related documents in a collection of documents, including the following steps: generating a document-representation of each document in an abstract mathematical space; identifying a plurality of document clusters in the collection of documents based on a conceptual similarity between respective pairs of the document-representations, wherein each document cluster is associated with an exemplary document and a plurality of other documents; and identifying a non-intersecting document cluster from among the plurality of document clusters based on (i) a conceptual similarity between the document-representation of the exemplary document and the document-representation of each document in the non-intersecting cluster and (ii) a conceptual dissimilarity between a cluster-representation of the non-intersecting document cluster and a cluster-representation of each other document cluster. Variants of the method enable creating hierarchy of clusters and conducting incremental updates of preexisting hierarchical structures.

Claims (31)

1. A computer-based method for automatically identifying clusters of conceptually-related documents in a collection of documents, comprising:

(a) generating a document-representation of each document in an abstract mathematical space;

(b) identifying a plurality of document clusters in the collection of documents based on a conceptual similarity between respective pairs of the document-representations, wherein each document cluster is associated with an exemplary document and a plurality of other documents; and

(c) identifying a non-intersecting document cluster from among the plurality of document clusters based on (i) a conceptual similarity between the document-representation of the exemplary document and the document-representation of each document in the non-intersecting cluster and (ii) a conceptual dissimilarity between a cluster-representation of the non-intersecting document cluster and a cluster-representation of each other document cluster, wherein step (c) comprises,

(c1) identifying a non-intersecting document cluster from among the plurality of document clusters if (i) a conceptual similarity between the document-representation of the exemplary document and the document-representation of each document in the non-intersecting cluster is above a predefined similarity threshold and (ii) a conceptual dissimilarity between a cluster-representation of the non-intersecting document cluster and a cluster-representation of each other document cluster is above a predefined dissimilarity threshold; and

(d) iteratively adjusting the predefined similarity threshold from a maximum similarity level to a minimum similarity level via a predefined similarity increment;

(e) iteratively adjusting the predefined dissimilarity threshold from a minimum dissimilarity level to a maximum dissimilarity level via a predefined dissimilarity increment; and

(f) repeating step (c1) for each similarity level and each dissimilarity level.

2. The method of claim 1 , wherein step (b) comprises:

(b1) identifying a plurality of document clusters in the collection of documents based on a conceptual similarity between respective pairs of the document-representations; and

(b2) generating a cluster-representation of each document cluster in the plurality of document clusters, wherein each cluster-representation is associated with an exemplary document and a plurality of other documents.

3. The method of claim 1 , wherein generating a document-representation of each document in an abstract mathematical space comprises:

generating a vector representation of each document in a Latent Semantic Indexing (LSI) space.

4. The method of claim 3 , wherein identifying a plurality of document clusters in the collection of documents based on a conceptual similarity between pairs of the document-representations comprises:

identifying a plurality of document clusters in the collection of documents based on a cosine similarity between pairs of the document-representations.

5. A computer program product for automatically identifying clusters of conceptually-related documents in a collection of documents, comprising:

a computer usable medium having computer readable program code embodied in said medium for causing an application program to execute on an operating system of a computer, said computer readable program code comprising:

computer readable first program code that causes the computer to generate a document-representation of each document in an abstract mathematical space;

computer readable second program code that causes the computer to identify a plurality of document clusters in the collection of documents based on a conceptual similarity between respective pairs of the document-representations, wherein each document cluster includes an exemplary document and a plurality of other documents; and

computer readable third program code that causes the computer to identify a non- intersecting document cluster from among the plurality of document clusters based on (i) a conceptual similarity between the document-representation of the exemplary document and the document-representation of each document in the non-intersecting cluster and (ii) a conceptual dissimilarity between a cluster-representation of the non-intersecting document cluster and a cluster-representation of each other document cluster, wherein the computer readable third program code comprises,

code that causes the computer to identify a non-intersecting document cluster from among the plurality of document clusters if (i) a conceptual similarity between the document-representation of the exemplary document and the document-representation of each document in the non-intersecting cluster is above a predefined similarity threshold and (ii) a conceptual dissimilarity between a cluster-representation of the non-intersecting document cluster and a cluster-representation of each other document cluster is above a predefined dissimilarity threshold; and

computer readable fourth program code that causes the computer to iteratively adjust the predefined similarity threshold from a maximum similarity level to a minimum similarity level via a predefined similarity increment;

computer readable fifth program code that causes the computer to iteratively adjust the predefined dissimilarity threshold from a minimum dissimilarity level to a maximum dissimilarity level via a predefined dissimilarity increment; and

computer readable sixth program code that causes the computer to repeat the third computer readable program code means for each similarity level and each dissimilarity level.

6. The computer program product of claim 5 , wherein the computer readable second program code comprises:

code that causes the computer to identify a plurality of document clusters in the collection of documents based on a conceptual similarity between respective pairs of the document-representations; and

code that causes the computer to generate a cluster-representation of each document cluster in the plurality of document clusters, wherein each cluster-representation is associated with an exemplary document and a plurality of other documents.

7. The computer program product of claim 5 , wherein the computer readable first program code that causes the computer to generate a document-representation of each document in an abstract mathematical space comprises:

code that causes the computer to generate a vector representation of each document in a Latent Semantic Indexing (LSI) space.

8. The computer program product of claim 7 , wherein the computer readable second program code that causes the computer to identify a plurality of document clusters in the collection of documents based on a conceptual similarity between pairs of the document-representations comprises:

code that causes the computer to identify a plurality of document clusters in the collection of documents based on a cosine similarity between pairs of the document-representations.

Assignments (6)
SECURITY INTEREST Recorded Jan 30, 2026
From: RELATIVITY ODA LLC; TEXT IQ, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 074537/0402 →
RELEASE OF SECURITY INTEREST AT REEL/FRAME 056218/0822 Recorded Jan 30, 2026
From: BLUE OWL CAPITAL CORPORATION, AS COLLATERAL AGENT F/K/A OWL ROCK CAPITAL CORPORATION, AS COLLATERAL AGENT
To: RELATIVITY ODA LLC
Reel/Frame 074539/0099 →
SECURITY INTEREST Recorded May 12, 2021
From: RELATIVITY ODA LLC
To: OWL ROCK CAPITAL CORPORATION, AS COLLATERAL AGENT
Reel/Frame 056218/0822 →
CHANGE OF NAME Recorded Aug 28, 2017
From: KCURA LLC
To: RELATIVITY ODA LLC
Reel/Frame 043687/0734 →
MERGER Recorded Mar 6, 2017
From: CONTENT ANALYST COMPANY, LLC
To: KCURA LLC
Reel/Frame 041471/0321 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2006
From: WNEK, JANUSZ
To: CONTENT ANALYST COMPANY, LLC
Reel/Frame 017891/0869 →