IP Library Granted Patent US 7,996,407
Granted Patent B2
US 7,996,407 · App. 12/018,652 · Granted Aug 9, 2011

System, method and computer executable program for information tracking from heterogeneous sources

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,996,407
App. No.
12/018,652
Granted
Aug 9, 2011
Kind
B2
Abstract

A system for information clustering comprising a data accumulation part for accumulating documents in a document repository, the documents having loosely related attributes, and defining a cluster between the documents being time sliced so as to define chunks of the documents; a vector space generation part for generating document-keyword vectors, the document-keyword vectors consisting of sparse numeral values depending on presence of key words; a dimension reduction part for reducing dimensions of the keywords to create a dimension reduction matrix of the document-keyword matrix; a centroid vector determination part for generating a centroid vector of the cluster, the centroid vectors being defined from keywords and weight of documents within the cluster; and an item repository for storing the centroid vectors together with the keywords and the weights of the centroid vector.

Claims (65)

1. A system for information clustering, said system comprising;

a central processing unit (CPU) for executing parts;

a data accumulation part for accumulating and clustering documents in a document repository, said documents including loosely related clusters between said documents being time sliced so as to define chunks of said documents;

a vector space generation part for generating document-keyword vectors, said document-keyword vectors consisting of sparse numeral values depending on presence of keywords in said documents;

a dimension reduction part for reducing dimensions of said keywords to create a dimension reduction matrix of said document-keyword matrix;

a centroid vector determination part for generating a centroid vector of said cluster, said cluster being retrieved from said document-keyword vector using a principal component in a same line of said dimension reduction matrix, said centroid vectors being defined from keywords and weight of documents within said cluster; and

an item repository for storing said centroid vectors together with said keywords and said weights of said centroid vector.

2. The system of claim 1 , wherein said centroid vector determination part retrieves a principal document in said document using said principal component as a first query vector and subsequently retrieves documents defining said clusters using said principal document as a second query vector.

3. The system of claim 1 , wherein said vector space generation part executes dimension reduction to each of said chunks of said dimension reduction matrix and said centroid vector generation part generates clusters for every chunk of said dimension reduction matrix.

4. The system of claim 1 , wherein said system further comprises an item analyzer part for analyzing evolution of items with respect to said chunk of said document and for information tracking.

5. A computer executable method for information clustering, said method making a computer having a central processing unit (CPU) execute the steps of;

accumulating and clustering documents in a document repository, said documents including loosely related clusters between said documents being time sliced so as to define chunks of said documents;

generating document-keyword vectors, said document-keyword vectors consisting of sparse numeral values depending on presence of keywords in said documents;

reducing dimensions of said keywords to create a dimension reduction matrix of said document-keyword matrix; and

generating a centroid vector of said cluster, said cluster being retrieved from said document-keyword vector using a principal component in a same line of said dimension reduction matrix, said centroid vectors being defined from keywords and weight of documents within said cluster; and

storing said centroid vectors in an item repository together with said keywords and said weights of said centroid vector.

6. The method of claim 5 , said method further comprising the steps of;

retrieving a principal document in said document using said principal component as a first query vector and subsequently retrieving documents defining said clusters using said principal document as a second query vector.

7. The method of claim 5 , said method further comprising the steps of;

executing dimension reduction to each of said chunks of said dimension reduction matrix and

generating clusters for every chunk of said dimension reduction matrix.

8. The method of claim 5 , said method further comprising the steps of;

analyzing evolution of items with respect to said chunk of said document and for information tracking.

9. A system for information tracking, said system comprising;

a central processing unit for executing parts;

a data accumulation part for accumulating and clustering documents in a document repository, said documents including loosely related clusters between said documents being time sliced so as to define chunks of said documents;

a vector space generation part for generating document-keyword vectors, said document-keyword vectors consisting of sparse numeral values depending on presence of keywords in said documents;

a dimension reduction part for reducing dimensions of said keywords to create a dimension reduction matrix of said document-keyword matrix;

a centroid vector determination part for generating a centroid vector of said cluster, said cluster being retrieved from said document-keyword vector using a principal component in a same line of said dimension reduction matrix, said centroid vectors being defined from keywords and weight of documents within said cluster;

an item analyzer part for analyzing evolution of items with respect to said chunk of said document and for information tracking; and

an item repository for storing said centroid vectors together with said keywords and said weights of said centroid vector.

10. The system of claim 9 , wherein said centroid vector determination part retrieves a principal document in said document using said principal component as a first query vector and subsequently retrieves documents defining said clusters using said principal document as a second query vector.

11. A computer executable method for information tracking, said method making a computer having a central processing unit execute the steps of;

accumulating documents in a document repository, said documents including loosely related clusters between said documents being time sliced so as to define chunks of said documents;

generating document-keyword vectors, said document-keyword vectors consisting of sparse numeral values depending on presence of keywords in said documents;

reducing dimensions of said keywords to create a dimension reduction matrix of said document-keyword matrix; and

generating a centroid vector of said cluster, said cluster being retrieved from said document-keyword vector using a principal component in a same line of said dimension reduction matrix, said centroid vectors being defined from keywords and weight of documents within said cluster;

storing said centroid vectors in an item repository together with said keywords and said weights of said centroid vector; and

analyzing evolution of items with respect to said chunk of said document and for information tracking.

12. The method of claim 11 , said method further comprising the steps of;

retrieving a principal document in said document using said principal component as a first query vector and

subsequently retrieving documents defining said clusters using said principal document as a second query vector.

13. A non-transitory computer executable program medium storing a program for making a computer execute a method for information clustering, said method making said computer execute the steps of;

accumulating and clustering documents in a document repository, said documents including loosely related clusters between said documents being time sliced so as to define chunks of said documents;

generating document-keyword vectors, said document-keyword vectors consisting of sparse numeral values depending on presence of keywords in said documents;

reducing dimensions of said keywords to create a dimension reduction matrix of said document-keyword matrix; and

generating a centroid vector of said cluster, said cluster being retrieved from said document-keyword vector using a principal component in a same line of said dimension reduction matrix, said centroid vectors being defined from keywords and weight of documents within said cluster; and

storing said centroid vectors in an item repository together with said keywords and said weights of said centroid vector.

14. The program medium of claim 13 , wherein said method further comprises the steps of;

retrieving a principal document in said document using said principal component as a first query vector and

subsequently retrieving documents defining said clusters using said principal document as a second query vector.

15. The program medium of claim 13 , wherein the method further comprises the steps of;

executing dimension reduction to each of said chunks of said dimension reduction matrix and generating clusters for every chunk of said dimension reduction matrix.

16. The program medium of claim 13 , wherein the method further comprises the steps of;

analyzing evolution of items with respect to said chunk of said document and for information tracking.

17. A non-transitory computer executable program medium storing a program for making a computer execute a method for information tracking, said method making said computer execute the steps of;

accumulating documents in a document repository, said documents including loosely related clusters between said documents being time sliced so as to define chunks of said documents;

generating document-keyword vectors, said document-keyword vectors consisting of sparse numeral values depending on presence of keywords in said documents;

reducing dimensions of said keywords to create a dimension reduction matrix of said document-keyword matrix; and

generating a centroid vector of said cluster, said cluster being retrieved from said document-keyword vector using a principal component in a same line of said dimension reduction matrix, said centroid vectors being defined from keywords and weight of documents within said cluster;

storing said centroid vectors in an item repository together with said keywords and said weights of said centroid vector; and

analyzing evolution of items with respect to said chunk of said document and for information tracking.

18. The program medium of claim 17 , said method further comprising the steps of;

retrieving a principal document in said document using said principal component as a first query vector and

subsequently retrieving documents defining said clusters using said principal document as a second query vector.

Assignments (4)
SECURITY INTEREST Recorded Aug 10, 2023
From: DOMO, INC.
To: OBSIDIAN AGENCY SERVICES, INC.
Reel/Frame 064562/0269 →
SECURITY INTEREST Recorded May 12, 2020
From: DOMO, INC.
To: OBSIDIAN AGENCY SERVICES, INC.
Reel/Frame 052642/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2015
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: DOMO, INC.
Reel/Frame 036087/0456 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2008
From: KOBAYASHI, MEI; YUNG, RAYLENE KAY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 020999/0462 →