IP Library Granted Patent US 7,496,561
Granted Patent B2
US 7,496,561 · App. 10/724,170 · Granted Feb 24, 2009

Method and system of ranking and clustering for document indexing and retrieval

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,496,561
App. No.
10/724,170
Granted
Feb 24, 2009
Kind
B2
Abstract

A relevancy ranking and clustering method and system that determines the relevance of a document relative to a user's query using a similarity comparison process. Input queries are parsed into one or more query predicate structures using an ontological parser. The ontological parser parses a set of known documents to generate one or more document predicate structures. A comparison of each query predicate structure with each document predicate structure is performed to determine a matching degree, represented by a real number. A multilevel modifier strategy is implemented to assign different relevance values to the different parts of each predicate structure match to calculate the predicate structure's matching degree. The relevance of a document to a user's query is determined by calculating a similarity coefficient, based on the structures of each pair of query predicates and document predicates. Documents are autonomously clustered using a self-organizing neural network that provides a coordinate system that makes judgments in a non-subjective fashion.

Claims (31)

1. One or more computer readable media storing computer executable instructions to perform a method for vectorizing a set of document predicate structures, the method comprising:

identifying at least one predicate and argument in said set of document predicate structures by a predicate key that is an integer representation;

estimating conceptual nearness of two of said document predicate structures in said set of document predicate structures by subtracting corresponding ones of said predicate keys; and

outputting at least one document based upon the estimated conceptual nearness.

2. The computer readable media of claim 1 , the method further comprising constructing multi-dimensional vectors using said integer representation.

3. The computer readable media of claim 2 , the method further comprising normalizing said multi-dimensional vectors.

4. The computer readable media of claim 3 , the method further comprising identifying at least one query predicate structure by a second predicate key that is a second integer representation, and constructing second multi-dimensional vectors, for said at least one query predicate structure, using said second integer representation.

5. The computer readable media of claim 1 , the method further comprising identifying at least one query predicate structure by a second predicate key that is a second integer representation, and constructing second multi-dimensional vectors, for said at least one query predicate structure, using said second integer representation.

6. The computer readable media of claim 1 , wherein said set of document predicate structures are representations of logical relationships between words in a sentence.

7. The computer readable media of claim 1 , wherein each of said document predicate structures in said set includes a predicate and a set of arguments, wherein the predicate is one of a verb and a preposition.

8. One or more computer readable media storing computer executable instructions to perform a method for vectorizing a set of document predicate structures, the method comprising:

identifying at least one predicate in said set of document predicate structures by a predicate key that is an integer representation;

estimating conceptual nearness of two of said document predicate structures in said set of document predicate structures by subtracting corresponding ones of said predicate keys; and

outputting at least one document based upon the estimated conceptual nearness.

9. The computer readable media of claim 8 , the method further comprising constructing multi-dimensional vectors using said integer representation.

10. The computer readable media of claim 9 , the method further comprising normalizing said multi-dimensional vectors.

11. The computer readable media of claim 10 , the method further comprising identifying at least one query predicate structure by a second predicate key that is a second integer representation, and constructing second multi-dimensional vectors, for said at least one query predicate structure, using said second integer representation.

12. The computer readable media of claim 8 , the method further comprising identifying at least one query predicate structure by a second predicate key that is a second integer representation, and constructing second multi-dimensional vectors, for said at least one query predicate structure, using said second integer representation.

13. The computer readable media of claim 8 , wherein said set of document predicate structures are representations of logical relationships between words in a sentence.

14. One or more computer readable media storing computer executable instructions to perform a method for constructing multi-dimensional vector representations for each document of a set of documents, the method comprising:

determining each predicate structure of one or more predicate structures M in each document of the set of documents, said M predicate structures including a predicate and at least one argument;

identifying the predicate and the at least one argument in each of said M predicate structures by a predicate key that is an integer representation;

determining a fixed number of arguments q for vector construction;

constructing an N-dimensional vector representation of each document based upon the predicate and q arguments; and

outputting at least one document of the set of documents based upon the constructed N-dimensional vector representation of the at least one document,

wherein any predicate structure of said M predicate structures that includes less than q arguments fills unfilled argument positions with a numerical zero.

15. The computer readable media of claim 14 , wherein any predicate structure of said M predicate structures that includes more than q arguments omits remaining arguments after q argument positions are filled.

16. The computer readable media of claim 15 , wherein conceptual nearness of two of said N-dimensional vector representations is estimated by subtracting corresponding ones of said predicate keys.

17. The computer readable media of claim 15 , the method further comprising normalizing said N-dimensional vector representations.

18. The computer readable media of claim 14 , wherein conceptual nearness of two of said N-dimensional vector representations is estimated by subtracting corresponding ones of said predicate keys.

19. The computer readable media of claim 14 , the method further comprising normalizing said N-dimensional vector representations.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Jan 17, 2020
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: LEIDOS, INC.
Reel/Frame 051632/0742 →
RELEASE OF SECURITY INTEREST Recorded Jan 17, 2020
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: LEIDOS, INC.
Reel/Frame 051632/0819 →
SECURITY INTEREST Recorded Aug 25, 2016
From: LEIDOS, INC.
To: CITIBANK, N.A.
Reel/Frame 039809/0801 →
SECURITY INTEREST Recorded Aug 25, 2016
From: LEIDOS, INC.
To: CITIBANK, N.A.
Reel/Frame 039818/0272 →
CHANGE OF NAME Recorded Apr 14, 2014
From: SCIENCE APPLICATIONS INTERNATIONAL CORPORATION
To: LEIDOS, INC.
Reel/Frame 032674/0817 →