IP Library Granted Patent US 7,181,678
Granted Patent B2
US 7,181,678 · App. 10/767,151 · Granted Feb 20, 2007

Document clustering method and system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,181,678
App. No.
10/767,151
Granted
Feb 20, 2007
Kind
B2
Abstract

Document clustering method and system utilizing both the log-based clustering method and the content-based clustering method are disclosed. The method includes the steps of generating log-based document clusters and combining vectors from the log-based document clusters with individual document clusters for content-based clustering analysis. The log-based document clusters are generated by accessing the retrieval session log, clustering the retrieval sessions, and combining the documents opened during each of the sessions of session clusters.

Claims (52)

1. A method for clustering documents, including generating clusters with user perspective comprising:

receiving retrieval session logs;

performing log-based clustering on the session logs to generate session clusters;

representing each of the session clusters as a log-based document suitable for content based clustering;

receiving a plurality of documents that includes a first document that was accessed in one session and a second document that was not accessed in any of the sessions;

replacing the first document with one of the log-based documents, wherein said one of the log-based documents is associated with the session cluster that includes the first documents; and

performing content based clustering on at least one of the log-based document and the second document to generate clusters with user perspective.

2. The method of claim 1 wherein representing each of the session clusters as a log-based document suitable for content based clustering includes modifying each of the log-based documents so that a Euclidean distance between the each of the log-based documents is the same.

3. The method of claim 1 , wherein each of the session logs comprises a query used to retrieve documents.

4. The method of claim 1 , wherein each of the session logs comprises a number of documents found to satisfy a query.

5. The method of claim 1 , wherein each of the session logs comprises a list of documents opened by a user.

6. The method of claim 1 , wherein each of the session logs comprises a length of time that a document was opened.

7. A method for clustering documents comprising:

generating a hybrid matrix of vectors comprising a first vector representing a first document and a second vector representing a log-based document cluster document; and

clustering the documents using the hybrid matrix, wherein the hybrid matrix comprises:

accessing retrieval session logs;

clustering retrieval sessions into session clusters;

generating, a log-based document cluster for each session cluster by combining all documents opened during any retrieval session of the session cluster;

generating a log-based document cluster vector for each of the log-based document clusters;

replacing each document in the log-based document cluster with the log-based document cluster vector;

generating an individual document vector for each document not opened during any retrieval session; and

combining the log-based document cluster vector and the individual document cluster vector.

8. The method of claim 7 wherein the step of clustering retrieval sessions into session clusters comprises the steps of:

generating Boolean session vector for each retrieval session;

forming a matrix of the Boolean session vectors; and

applying a clustering algorithm to the matrix of the Boolean session vectors.

9. A system for clustering documents, the system comprising:

a storage for storing retrieval session logs; and

a processor connected to the storage, configured to cluster the retrieval sessions into session clusters,

generate, for each session cluster, a log-based document cluster,

generate a log-based document cluster vector for each of the log-based document cluster,

generate an individual document vector for each document not opened during any retrieval session,

cluster the documents using the log-based document cluster vectors and individual document vectors.

10. The system of claim 9 wherein the documents are stored in the storage.

11. The system of claim 9 further comprising:

a memory connected to the processor, for storage of a hybrid matrix comprising

the log-based document cluster vectors and the individual document vectors.

12. A data processing system having session logs and documents, the system comprising:

a processor for executing program instructions; and

a media readable by the processor having a document clustering module having a plurality of instructions, that when executed by the processor,

performs log-based clustering on the session logs to generate session clusters,

converts the session clusters into a form suitable for content-based clusters,

performs content-based clustering on the documents and session clusters in a form suitable for content-based clustering to generate document clusters with users' perspective.

13. The system of claim 12 wherein the document clustering module further comprises:

a session vector generation module for receiving the session logs and based thereon for generating a session vector for each session log;

a session cluster generation module coupled to the session vector generation module for receiving the session vectors and based thereon for generating session clusters;

a hybrid matrix builder for receiving the documents, coupled to the session cluster generation module, for receiving the session clusters and based thereon for generating a hybrid matrix having at least one log-based document; and

a topic generation module coupled to the hybrid matrix builder for receiving the hybrid matrix and based thereon for generating document clusters with users' perspective.

14. The system of claim 13 wherein the hybrid matrix builder further comprises:

a session document generation module for receiving session clusters and based thereon generates super documents; and

document modification module coupled to the session document generation module for receiving the super documents, for receiving the documents, and based thereon for generating the hybrid matrix.

15. The system of claim 12 wherein the media is one of a floppy disk, compact disc, a volatile memory, and a non-volatile memory.

Assignments (7)
RELEASE OF SECURITY INTEREST REEL/FRAME 044183/0718 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC (F/K/A ENTIT SOFTWARE LLC); BORLAND SOFTWARE CORPORATION; MICRO FOCUS (US), INC.; SERENA SOFTWARE, INC; ATTACHMATE CORPORATION; MICRO FOCUS SOFTWARE INC. (F/K/A NOVELL, INC.); NETIQ CORPORATION
Reel/Frame 062746/0399 →
RELEASE OF SECURITY INTEREST REEL/FRAME 044183/0577 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC (F/K/A ENTIT SOFTWARE LLC)
Reel/Frame 063560/0001 →
CHANGE OF NAME Recorded Feb 25, 2020
From: ENTIT SOFTWARE LLC
To: MICRO FOCUS LLC
Reel/Frame 052010/0029 →
SECURITY INTEREST Recorded Oct 11, 2017
From: ENTIT SOFTWARE LLC; ARCSIGHT, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 044183/0577 →
SECURITY INTEREST Recorded Oct 11, 2017
From: ATTACHMATE CORPORATION; BORLAND SOFTWARE CORPORATION; NETIQ CORPORATION; MICRO FOCUS (US), INC.; MICRO FOCUS SOFTWARE, INC.; ENTIT SOFTWARE LLC; ARCSIGHT, LLC; SERENA SOFTWARE, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 044183/0718 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2017
From: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
To: ENTIT SOFTWARE LLC
Reel/Frame 042746/0130 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2015
From: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 037079/0001 →