IP Library Granted Patent US 7,197,451
Granted Patent B1
US 7,197,451 · App. 09/615,726 · Granted Mar 27, 2007

Method and mechanism for the creation, maintenance, and comparison of semantic abstracts

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,197,451
App. No.
09/615,726
Granted
Mar 27, 2007
Kind
B1
Abstract

Codifying the “most prominent measurement points” of a document can be used to measure semantic distances given an area of study (e.g., white papers on some subject area). A semantic abstract is created for each document. The semantic abstract is a semantic measure of the subject or theme of the document providing a new and unique mechanism for characterizing content. The semantic abstract includes state vectors in the topological vector space, each state vector representing one lexeme or lexeme phrase about the document. The state vectors can be dominant phrase vectors in the topological vector space mapped from dominant phrases extracted from the document. The state vectors can also correspond to words in the document that are most significant to the document's meaning (the state vectors are called dominant vectors in this case). One semantic abstract can be directly compared with another semantic abstract, resulting in a numeric semantic distance between the semantic abstracts being compared.

Claims (116)

1. A method for determining dominant phrase vectors in a topological vector space for a semantic content of a document on a computer system, the method comprising:

identifying a directed set of concepts as a dictionary, the directed set including a maximal element and at least one concept, and at least one chain from the maximal element to every concept;

selecting a subset of the chains to form a basis for the dictionary;

accessing dominant phrases for the document, the dominant phrases representing a condensed content for the document;

measuring how concretely each dominant phrase is represented in each chain in the basis and the dictionary;

constructing at least one state vector in the topological vector space for each dominant phrase using the measures of how concretely each dominant phrase is represented in each chain in the dictionary and the basis; and

collecting the state vectors into the dominant phrase vectors for the document.

2. A method according to claim 1 , wherein accessing dominant phrases includes extracting the dominant phrases from the document using a phrase extractor.

3. A method according to claim 1 , wherein accessing dominant phrases includes storing the dominant phrases in computer memory accessible by the computer system.

4. A method according to claim 1 , the method further comprising forming a semantic abstract comprising the dominant phrase vectors.

5. A method according to claim 1 , wherein:

measuring how concretely each dominant phrase is represented in each chain in the basis and the dictionary includes:

identifying at least one lexeme in each dominant phrase; and

measuring how concretely each lexeme in each dominant phrase is represented in each chain in the basis and the dictionary; and

constructing at least one state vector in the topological vector space for each dominant phrase includes constructing at least one state vector in the topological vector space for each lexeme in each dominant phrase using the measures of how concretely each lexeme in each dominant phrase is represented in each chain in the basis and the dictionary.

6. A method according to claim 1 , wherein constructing at least one state vector includes constructing the state vectors in the topological vector space for each dominant phrase using the measures of how concretely each dominant phrase is represented in each chain in the dictionary and the basis, the state vectors independent of the document.

7. A method for determining dominant vectors in a topological vector space for a semantic content of a document on a computer system, the method comprising:

identifying a directed set of concepts as a dictionary, the directed set including a maximal element and at least one concept, and at least one chain from the maximal element to every concept;

selecting a subset of the chains to form a basis for the dictionary;

storing the document in computer memory accessible by the computer system;

extracting words from at least a portion of the document;

measuring how concretely each word is represented in each chain in the basis and the dictionary;

constructing a state vector in the topological vector space for each word using the measures of how concretely each word is represented in each chain in the dictionary and the basis;

filtering the state vectors; and

collecting the filtered state vectors into the dominant vectors for the document.

8. A method according to claim 7 , wherein extracting words includes extracting words from the entire document.

9. A method according to claim 7 , wherein filtering the state vectors includes selecting the state vectors that occur with highest frequencies.

10. A method according to claim 7 , wherein filtering the state vectors includes:

calculating a centroid in the topological vector space for the state vectors; and

selecting the state vectors nearest the centroid.

11. A method according to claim 7 , the method further comprising forming a semantic abstract comprising the dominant vectors.

12. A computer-readable medium containing a program to determine dominant vectors in a topological vector space for a semantic content of a document on a computer system, the program being executable on the computer system to implement the method of claim 7 .

13. A method according to claim 7 , wherein constructing a state vector includes constructing the state vector in the topological vector space for each word using the measures of how concretely each word is represented in each chain in the dictionary and the basis, the state vectors independent of the document.

14. A method for determining a semantic abstract in a topological vector space for a semantic content of a document on a computer system, the method comprising:

identifying a directed set of concepts as a dictionary, the directed set including a maximal element and at least one concept, and at least one chain from the maximal element to every concept;

selecting a subset of the chains to form a basis for the dictionary;

storing the document in computer memory accessible by the computer system;

determining dominant phrases for the document;

measuring how concretely each dominant phrase is represented in each chain in the basis and the dictionary;

constructing dominant phrase vectors in the topological vector space for the dominant phrases using the measures of how concretely each dominant phrase is represented in each chain in the dictionary and the basis;

selecting words for the document;

measuring how concretely each word is represented in each chain in the basis and the dictionary;

constructing dominant vectors in the topological vector space for the words using the measures of how concretely each word is represented in each chain in the dictionary and the basis; and

generating the semantic abstract using the dominant phrase vectors and the dominant vectors.

15. A method according to claim 14 , wherein generating the semantic abstract includes reducing the dominant phrase vectors based on the dominant vectors.

16. A method according to claim 15 , wherein constructing at least one state vector includes constructing at least one state vector for the first document in the topological vector space for each dominant phrase for the first document using the measures of how concretely each dominant phrase for the first document is represented in each chain in the dictionary and the basis, the state vectors independent of the first document and the second document.

17. A method according to claim 14 , wherein generating the semantic abstract includes reducing the dominant vectors based on the dominant phrase vectors.

18. A method according to claim 14 , wherein generating the semantic abstract includes obtaining a probability distribution function for a reduced set of the dominant phrase vectors similar to a probability distribution function for the dominant phrase vectors.

19. A method according to claim 14 , the method further comprising identifying the lexemes or lexeme phrases corresponding to state vectors in the semantic abstract.

20. A computer-readable medium containing a program to determine a semantic abstract in a topological vector space for a semantic content of a document on a computer system, the program being executable on the computer system to implement the method of claim 14 .

21. A method according to claim 14 , wherein constructing dominant vectors includes constructing dominant vectors in the topological vector space for the words using the measures of how concretely each word is represented in each chain in the dictionary and the basis, the dominant vectors independent of the document.

22. A method for comparing the semantic content of first and second documents on a computer system, the method comprising:

identifying a directed set of concepts as a dictionary, the directed set including a maximal element and at least one concept, and at least one chain from the maximal element to every concept;

selecting a subset of the chains to form a basis for the dictionary;

accessing dominant phrases for the first document, the dominant phrases representing a condensed content for the first document;

measuring how concretely each dominant phrase for the first document is represented in each chain in the basis and the dictionary;

constructing at least one state vector for the first document in a topological vector space for each dominant phrase for the first document using the measures of how concretely each dominant phrase for the first document is represented in each chain in the dictionary and the basis;

collecting the state vectors for the first document into the semantic abstract for the first document;

determining a semantic abstract for the second document;

measuring a distance between the semantic abstracts; and

classifying how closely related the first and second documents are using the distance.

23. A method according to claim 22 , wherein measuring a distance includes measuring a Hausdorff distance between the semantic abstracts.

24. A method according to claim 22 , wherein measuring a distance includes determining a centroid vector in the topological vector space for each semantic abstract.

25. A method according to claim 24 , wherein measuring a distance further includes measuring an angle between the centroid vectors.

26. A method according to claim 24 , wherein measuring a distance further includes measuring a Euclidean distance between the centroid vectors.

27. A computer-readable medium containing a program to compare the semantic content of first and second documents on a computer system, the program being executable on the computer system to implement the method of claim 22 .

28. A method according to claim 22 , wherein determining a semantic abstract for the second document includes:

accessing dominant phrases for the second document, the dominant phrases for the second document representing a condensed content for the second document;

measuring how concretely each dominant phrase for the second document is represented in each chain in the basis and the dictionary;

constructing at least one state vector for the second document in the topological vector space for each dominant phrase for the second document using the measures of how concretely each dominant phrase for the second document is represented in each chain in the dictionary and the basis; and

collecting the state vectors for the second document into the semantic abstract for the second document.

29. A method according to claim 22 , wherein constructing at least one state vector includes constructing the state vectors for the first document in the topological vector space for each dominant phrase for the first document using the measures of how concretely each dominant phrase for the first document is represented in each chain in the dictionary and the basis, the state vectors independent of the first document and the second document.

30. A method for locating a second document on a computer with a semantic content similar to a first document, the method comprising:

identifying a directed set of concepts as a dictionary, the directed set including a maximal element and at least one concept, and at least one chain from the maximal element to every concept;

selecting a subset of the chains to form a basis for the dictionary;

accessing dominant phrases for the first document, the dominant phrases representing a condensed content for the first document;

measuring how concretely each dominant phrase for the first document is represented in each chain in the basis and the dictionary;

constructing at least one state vector for the first document in a topological vector space for each dominant phrase for the first document using the measures of how concretely each dominant phrase for the first document is represented in each chain in a dictionary and the basis;

collecting the state vectors for the first document into the semantic abstract for the first document;

locating a second document;

determining a semantic abstract for the second document;

measuring a distance between the semantic abstracts for the first and second documents;

classifying how closely related the first and second documents are using the distance; and

if the second document is classified as having a semantic content similar to the semantic content of the first document, selecting the second document.

31. A method according to claim 30 , the method further comprising, if the second document is classified as not having a semantic content similar to the semantic content of the first document, rejecting the second document.

32. A method according to claim 30 , wherein determining a semantic abstract for the second document includes:

accessing dominant phrases for the second document, the dominant phrases for the second document representing a condensed content for the second document;

measuring how concretely each dominant phrase for the second document is represented in each chain in the basis and the dictionary;

constructing at least one state vector for the second document in the topological vector space for each dominant phrase for the second document using the measures of how concretely each dominant phrase for the second document is represented in each chain in the dictionary and the basis; and

collecting the state vectors for the second document into the semantic abstract for the second document.

33. An apparatus on a computer system to determine a semantic abstract in a topological vector space for a semantic content of a document stored on the computer system, the apparatus comprising:

a phrase extractor adapted to extract phrases from the document;

a state vector constructor adapted to construct state vectors in the topological vector space for each phrase extracted by the phrase extractor, the state vectors measuring how concretely each phrase extracted by the phrase extractor is represented in each chain in a basis and a dictionary, the dictionary including a directed set of concepts including a maximal element and at least one chain from the maximal element to every concept in the directed set, the basis including a subset of chains in the directed set; and

collection means for collecting the state vectors into the semantic abstract for the document.

34. An apparatus according to claim 33 , the apparatus further comprising filter means for filtering the state vectors to reduce the size of the semantic abstract.

35. An apparatus according to claim 33 , wherein the state vector constructor is further adapted to construct a state vector for each word in the document.

36. An apparatus according to claim 33 , wherein the state vector constructor is operative to construct the state vectors independent of the document.

37. A method for determining a semantic abstract in a topological vector space for a semantic content of a document on a computer system, the method comprising:

extracting dominant phrases from the document using a phrase extractor, the dominant phrases representing a condensed content for the document;

identifying a directed set of concepts as a dictionary, the directed set including a maximal element and at least one concept, and at least one chain from the maximal element to every concept;

selecting a subset of the chains to form a basis for the dictionary;

measuring how concretely each dominant phrase is represented in each chain in the basis and the dictionary;

constructing at least one first state vector in the topological vector space for each dominant phrase using the measures of how concretely each dominant phrase is represented in each chain in the dictionary and the basis;

collecting the first state vectors into dominant phrase vectors for the document;

extracting words from at least a portion of the document;

constructing a second state vector in the topological vector space for each word using the dictionary and the basis;

filtering the second state vectors;

collecting the filtered second state vectors into dominant vectors for the document; and

generating the semantic abstract using the dominant phrase vectors and the dominant vectors.

38. A method according to claim 37 , the method further comprising comparing the semantic abstract with a second semantic abstract for a second document to determine how closely related the contents of the documents are.

39. A method according to claim 37 , wherein constructing a second state vector in the topological vector space for each word using the dictionary and the basis includes:

measuring how concretely each word is represented in each chain in the basis and the dictionary; and

constructing the second state vectors in the topological vector space for each word using the measures of how concretely each word is represented in each chain in the dictionary and the basis.

40. A method according to claim 37 , wherein:

constructing at least one first state vector includes constructing the first state vector in the topological vector space for each dominant phrase using the measures of how concretely each dominant phrase is represented in each chain in the dictionary and the basis, the first state vector independent of the document; and

constructing a second state vector includes constructing the second state vector in the topological vector space for each word using the dictionary and the basis, the second state vector independent of the document.

Assignments (10)
RELEASE OF SECURITY INTEREST Recorded Oct 26, 2020
From: JEFFERIES FINANCE LLC
To: RPX CORPORATION
Reel/Frame 054486/0422 →
SECURITY INTEREST Recorded Jun 29, 2018
From: RPX CORPORATION
To: JEFFERIES FINANCE LLC
Reel/Frame 046486/0433 →
RELEASE (REEL 038041 / FRAME 0001) Recorded Jan 2, 2018
From: JPMORGAN CHASE BANK, N.A.
To: RPX CORPORATION; RPX CLEARINGHOUSE LLC
Reel/Frame 044970/0030 →
SECURITY AGREEMENT Recorded Mar 9, 2016
From: RPX CORPORATION; RPX CLEARINGHOUSE LLC
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 038041/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2016
From: NOVELL INTELLECTUAL PROPERTY HOLDINGS, INC.
To: RPX CORPORATION
Reel/Frame 037809/0057 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2011
From: CPTN HOLDINGS LLC
To: NOVELL INTELLECTUAL PROPERTY HOLDING, INC.
Reel/Frame 027325/0131 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2011
From: CPTN HOLDINGS LLC
To: NOVELL INTELLECTUAL PROPERTY HOLDINGS, INC.
Reel/Frame 027465/0206 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2011
From: CPTN HOLDINGS LLC
To: NOVELL INTELLECTUAL PROPERTY HOLDINGS INC.
Reel/Frame 027162/0342 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2011
From: NOVELL, INC.
To: CPTN HOLDINGS LLC
Reel/Frame 027157/0583 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2000
From: CARTER, STEPHEN R.; JENSEN, DELOS C.; MILLETT, RONALD P.
To: NOVELL, INC.
Reel/Frame 010951/0428 →