IP Library › Granted Patent US 7,937,389
Granted Patent B2
US 7,937,389 · App. 12/072,723 · Granted May 3, 2011

Dynamic reduction of dimensions of a document vector in a document search and retrieval system

Assignee: UT-Battelle, LLC
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,937,389
App. No.
12/072,723
Granted
May 3, 2011
Kind
B2
Abstract

The method and system of the invention involves processing each new document ( 20 ) coming into the system into a document vector ( 16 ), and creating a document vector with reduced dimensionality ( 17 ) for comparison with the data model ( 15 ) without recomputing the data model ( 15 ). These operations are carried out by a first computer ( 11 ) while a second computer ( 12 ) updates the data model ( 18 ), which can be comprised of an initial large group of documents ( 19 ) and is premised on the computing an initial data model ( 13, 14, 15 ) to provide a reference point for determining document vectors from documents processed from the data stream ( 20 ).

Claims (83)

1. A method for reducing dimensions of a document vector used to determine the similarity of a first document to a plurality of other documents in a computer, the method comprising:

receiving a document that is input to the computer for determining the similarity of the document to the plurality of other documents;

preprocessing the document to generate a document vector;

reducing a number of dimensions in the document vector, wherein the dimensionality of a document vector d created by the previous process is m and is reduced to a predefined k with dimensions determined by the following equation:

d

^

⁢

:

⁢

⁢

d

^

=

d

T

⁢

U

k

⁢

∑

k

-

1

,

wherein d is an original document vector, U k and Σ k are the matrices resulted from a truncated singular value decomposition (SVD) process and {circumflex over (d)} is the document vector with reduced dimensionality {circumflex over (d)} having k dimensions;

comparing the document vector with reduced dimensions to at least one document vector for the plurality of documents to determine a similarity of the document to the plurality of other documents; and

displaying a measure of similarity of the document to the other documents to a human observer.

2. The method of claim 1 , further comprising:

preprocessing a plurality of initial documents;

computing a data model representing the plurality of initial documents;

wherein the document vector for the document is compared to at least one document vector for the data model to determine the similarity of the document to the documents forming the data model; and

recomputing the data model upon receiving a predetermined number of new documents.

3. The method of claim 2 , wherein the document vector for the document is compared to a document vector for the data model without updating the data model until a second plurality of documents equaling the predetermined number have been received and processed by the computer.

4. The method of claim 3 , wherein a calculation of the document vector with reduced dimensions is performed by a first computer; and

wherein the updating of the data model is performed by a second computer in communication with the first computer.

5. The method of claim 3 , further comprising updating the data model when a document count reaches the predetermined number of 20,000 documents.

6. The method of claim 2 , wherein the data model is computed from a large corpus of at least 200,000 documents, wherein a global frequency of words in the corpus of documents is expressed in a first table; wherein a term frequency of terms in the corpus of documents is expressed in a second table; and wherein a document matrix of initial dimension, m, is generated; and wherein a truncated singular value decomposition (SVD) is computed according to the expression: SVD k (M)=U k Σ k V k T , where U k is a m×k matrix of initial dimension m; and Σ k is a k×k matrix of reduced dimension k used in the data model.

7. The method of claim 1 , wherein a reduced number of dimensions, k, is preset to a specific number.

8. The method of claim 1 , wherein the preprocessing of the document received by the computer further comprises:

removing punctuation marks and symbols;

removing stop words; and

parsing the document into a list of stemmed words.

9. A computer system for reducing dimensions of a document vector used to determine the similarity of a first document to a plurality of other documents in the computer system, the system comprising:

means for receiving a document that is input to the computer for determining the similarity of the document to the plurality of other documents;

means for preprocessing the document to generate a document vector; and

means for reducing a number of dimensions in the document vector, wherein the dimensionality of a document vector d created by the previous process is m and is reduced to a predefined k with dimensions determined by the following equation:

d

^

⁢

:

⁢

⁢

d

^

=

d

T

⁢

U

k

⁢

∑

k

-

1

,

wherein d is the original document vector, U k and Σ k are the matrices resulted from the truncated singular value decomposition (SVD) process and {circumflex over (d)} is the document vector with reduced dimensionality {circumflex over (d)} having k dimensions;

means for comparing the document vector of reduced dimensions to at least one document vector for the plurality of documents to determine a similarity of the document to the plurality of other documents; and

means for displaying a measure of similarity of the document to the other documents to a human observer.

10. The system of claim 9 , further comprising:

means for preprocessing a plurality of initial documents; and

means for computing a data model representing the plurality of initial documents; and

wherein the document vector for the document is compared to at least document vector for the data model to determine the similarity of the document to the documents forming the data model.

11. The system of claim 10 , wherein the document vector for the document is compared to a document vector for the data model without updating the data model until a second plurality of documents have been received and processed by the computer.

12. The system of claim 11 , wherein means for comparing the document vector to at least one document vector for the plurality of documents is incorporated in a first computer; and

wherein an updating of the data model is performed in a second computer in communication with the first computer.

13. The system of claim 11 , wherein the updating of the data model is performed when a document count reaches 20,000 documents.

14. The system of claim 10 , wherein the data model is computed from a large corpus of at least 200,000 documents, wherein a global frequency of words in the corpus of documents is expressed in a first table; wherein a term frequency of terms in the corpus of documents is expressed in a second table; and wherein a document matrix of initial dimension, m, is generated; and wherein a truncated singular value decomposition (SVD) is computed according to the expression: SVD k (M)=U k Σ k V k T , where U k is a m×k matrix of initial dimension m; and Σ k is a k×k matrix of reduced dimension k used in the data model.

15. The system of claim 9 , wherein a reduced number of dimensions, k, is preset to a specific number.

16. The system of claim 9 , wherein the preprocessing of the document received by the computer further comprises:

removing punctuation marks and symbols;

removing stop words; and

parsing the document into a list of stemmed words.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2010
From: OAK RIDGE ASSOCIATED UNIVERSITIES
To: UT-BATTELLE, LLC
Reel/Frame 024191/0753 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2010
From: JIAO, YU
To: OAK RIDGE ASSOCIATED UNIVERSITIES
Reel/Frame 024172/0582 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2010
From: POTOK, THOMAS E.
To: UT-BATTELLE, LLC
Reel/Frame 024172/0642 →
CONFIRMATORY LICENSE Recorded Jul 23, 2008
From: UT-BATTELLE, LLC
To: ENERGY, U.S. DEPARTMENT OF
Reel/Frame 021284/0266 →
Continuity (2)
Provisional Application 61001437 · Nov 1, 2007
Related Publication 20090119343A1 · May 7, 2009