IP Library Granted Patent US 8,600,922
Granted Patent B2
US 8,600,922 · App. 12/294,589 · Granted Dec 3, 2013

Methods and systems for knowledge discovery

Inventors: Edwin Adriaansen (Culemborg, NL); Bob Schijvenaars (Zoeterwoude, NL)
Assignee: Elsevier Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,600,922
App. No.
12/294,589
Granted
Dec 3, 2013
Kind
B2
Abstract

Provided are methods and systems for knowledge discovery utilizing knowledge profiles.

Claims (65)

1. A computer-implemented method for textual analysis comprising:

a. determining, by a computer processor, a co-occurrence of a long form and an associated short form of a term in a document;

b. locating, by a computer processor, a plurality of occurrences of the associated short form; and

c. expanding, by a computer processor, the plurality of occurrences of the associated short form with the long form wherein the document has a more accurate representation of frequency of occurrence of the term;

d. receiving a context fingerprint for each of a plurality of concepts;

e. determining an overlap of context fingerprints among the plurality of the concepts;

f. determining a similarity score between the context fingerprints; and

g. predicting that two or more of the plurality of concepts have a relationship, wherein the overlap is above a first threshold and the similarity score is above a second threshold.

2. The method of claim 1 , wherein the long form of the term comprises at least one word.

3. The method of claim 2 , wherein the associated short form is an abbreviation of the at least one word.

4. The method of claim 1 , wherein the term represents a concept.

5. The method of claim 1 , further comprising determining a frequency of occurrence of the term in the document.

6. The method of claim 1 , further comprising generating a fingerprint of the document.

7. The method of claim 1 , further comprising performing steps a-g for a plurality of documents.

8. A system for textual analysis comprising:

a memory configured for storing text data; and

a processor, coupled to the memory, configured for performing steps comprising,

a. determining a co-occurrence of a long form and an associated short form of a term in a document,

b. locating a plurality of occurrences of the associated short form,

c. expanding the plurality of occurrences of the associated short form with the long form wherein the document has a more accurate representation of frequency of occurrence of the term;

d. receiving a context fingerprint for each of a plurality of concepts;

e. determining an overlap of context fingerprints among the plurality of the concepts;

f. determining a similarity score between the context fingerprints; and

g. predicting that two or more of the plurality of concepts have a relationship, wherein the overlap is above a first threshold and the similarity score is above a second threshold.

9. The system of claim 8 , wherein the long form of the term comprises at least one word.

10. The system of claim 9 , wherein the associated short form is an abbreviation of the at least one word.

11. The system of claim 8 , wherein the term represents a concept.

12. The system of claim 8 , wherein the processor is further configured for determining a frequency of occurrence of the term in the document.

13. The system of claim 8 , wherein the processor is further configured for generating a fingerprint of the document.

14. The system of claim 8 , wherein the processor is further configured for performing steps a-g for a plurality of documents.

15. A non-transitory computer-readable storage medium with computer executable instructions embodied thereon for textual analysis, that when executed by a computer processor, causes said computer processor to perform steps comprising:

a. determining a co-occurrence of a long form and an associated short form of a term in a document;

b. locating a plurality of occurrences of the associated short form; and

c. expanding the plurality of occurrences of the associated short form with the long form wherein the document has a more accurate representation of frequency of occurrence of the term;

d. receiving a context fingerprint for each of a plurality of concepts;

e. determining an overlap of context fingerprints among the plurality of the concepts;

f. determining a similarity score between the context fingerprints; and

g. predicting that two or more of the plurality of concepts have a relationship, wherein the overlap is above a first threshold and the similarity score is above a second threshold.

16. The computer-readable storage medium of claim 15 , wherein the long form of the term comprises at least one word.

17. The computer-readable storage medium of claim 16 , wherein the associated short form is an abbreviation of the at least one word.

18. The computer-readable storage medium of claim 15 , wherein the term represents a concept.

19. The computer-readable storage medium of claim 15 , further comprising computer executable instructions for determining a frequency of occurrence of the term in the document.

20. The computer-readable storage medium of claim 15 , further comprising computer executable instructions for generating a fingerprint of the document.

21. The computer-readable storage medium of claim 15 , further comprising computer executable instructions for performing steps a-g for a plurality of documents.

22. The computer-implemented method of claim 7 , wherein at least two of the plurality of documents have an associated fingerprint, further comprising the step of combining said associated fingerprints having a relationship.

23. The computer-implemented method of claim 22 wherein combining said associated fingerprints having a relationship comprises averaging the fingerprints.

24. The computer-implemented method of claim 22 wherein combining said associated fingerprints having a relationship comprises:

taking a square of the respective relevance weights;

averaging the squares of the respective relevance weights; and

taking the root of the averages.

25. The system of claim 8 , wherein the plurality of concepts do not co-occur in a plurality of documents.

26. The system of claim 8 , wherein the plurality of concepts do not co-occur within the same sentence of a single document.

27. The system of claim 8 , wherein the plurality of concepts do not co-occur within the same paragraph of a single document.

28. The system of claim 8 , wherein a context fingerprint is a list of concepts and their associated relevance weights which are constructed based on co-occurrence of concepts in documents with the concept the context fingerprint is created for.

29. The system of claim 8 , wherein determining an overlap of context fingerprints among the plurality of concepts comprises determining a number of concepts the two context fingerprints have in common.

30. The system of claim 8 , wherein determining a similarity score between the context fingerprints comprises performing a matching algorithm.

31. The system of claim 30 , wherein performing a matching algorithm comprises:

storing each context fingerprint as a vector; and

performing a vector matching algorithm.

32. The computer-readable storage medium of claim 21 , wherein at least two of the plurality of documents have an associated fingerprint, further comprising computer executable instructions for combining said associated fingerprints having a relationship.

33. The computer-readable storage medium of claim 32 wherein combining said associated fingerprints having a relationship comprises averaging the fingerprints.

34. The computer-readable storage medium of claim 32 wherein combining said associated fingerprints having a relationship comprises:

taking a square of the respective relevance weights;

averaging the squares of the respective relevance weights; and

taking the root of the averages.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2013
From: ADRIAANSEN, EDWIN; SCHIJVENAARS, BOB
To: COLLEXIS HOLDINGS, INC.
Reel/Frame 031045/0136 →
MERGER Recorded Jun 1, 2011
From: SCIENCE INFORMATION SOLUTIONS LLC
To: ELSEVIER INC.
Reel/Frame 026372/0095 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2010
From: COLLEXIS HOLDINGS, INC.
To: SCIENCE INFORMATION SOLUTIONS LLC
Reel/Frame 025088/0390 →
Continuity (2)
Provisional Application 60829424 · Oct 13, 2006
Related Publication 20100049684A1 · Feb 25, 2010