IP Library Granted Patent US 8,166,051
Granted Patent B1
US 8,166,051 · App. 12/364,753 · Granted Apr 24, 2012

Computation of term dominance in text documents

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,166,051
App. No.
12/364,753
Granted
Apr 24, 2012
Kind
B1
Abstract

An improved entropy-based term dominance metric useful for characterizing a corpus of text documents, and is useful for comparing the term dominance metrics of a first corpus of documents to a second corpus having a different number of documents.

Claims (74)

1. A computer-implemented method comprising:

generating, with a processor of a hardware computing system, a set of dominance metrics for a first corpus A having n documents, wherein the set of dominance metrics includes for each term i in a collection C a respective dominance metric D A (i) which is based on a quotient of:

an exponentiation of two by an entropy value for the respective term i; and

the number n of documents in the first corpus A;

wherein the entropy value for the respective term i is based on a respective sum of product values for each document j of first corpus A, wherein each of the product values is based on a respective product of:

a respective value p ij ; and

a logarithm of the respective value p ij ;

wherein the respective value p ij is based on a quotient of:

a number of times term i occurs in document j; and

a total number of times term i occurs in first corpus A; and

storing information characterizing the first corpus A to a data storage unit communicatively coupled to the processor, the information based on the generated set of dominance metrics.

2. The method of claim 1 , further comprising:

sorting the set of dominance metrics; and

based on the sorted set of dominance metrics, identifying a sub-set of terms of the collection C; and

wherein the stored information characterizes the first corpus A based on the identified sub-set of terms.

3. The method of claim 1 , further comprising:

using the set of dominance metrics to characterize a person who generated or accessed the first corpus A.

4. The method of claim 1 , further comprising:

comparing a second set of dominance values from a second corpus B to the set of dominance values from the first corpus A.

5. The method of claim 1 , further comprising:

storing the first corpus A within a computer accessible database; and

displaying the set of dominance values to a display screen.

6. The method of claim 4 , further comprising:

storing the first corpus A within a computer accessible database; and

displaying a result of the comparing to a display screen.

7. The method of claim 4 , wherein the comparing step comprises:

creating a collection of dominant terms common to both first corpus A and second corpus B;

based on the collection of dominant terms, constructing two vectors of paired term-by-term similarities, one vector each for first corpus A and second corpus B; and

computing a scalar similarity measure between first corpus A and second corpus B including determining a cosine similarity of the two vectors of paired term-by-term similarities.

8. The method of claim 7 , further comprising:

calculating a first term-similarity measure for two terms in the first corpus A; and

calculating a second term-similarity measure for the same two terms in the second.

9. The method of claim 8 , further comprising:

finding an intersection of the most dominant terms from first corpus A and second corpus B; and

selecting the top N most dominant terms from both first corpus A and second corpus B.

10. The method of claim 9 , wherein N is less than or equal to 50.

11. The method of claim 9 , further comprising:

based on the intersection, calculating a first paired-term dominance vector, M A , for first corpus A;

wherein the first paired-term dominance vector, M A , is computed using all possible non-reflexive pairs of terms from intersection; and

based on the intersection, calculating a second paired-term dominance vector, M B , for second corpus B.

12. The method of claim 11 , further comprising generating a model-to-model comparison of first corpus A to second corpus B, including computing a cosine similarity, s AB , between vector M A for first corpus A, and vector M B for second corpus B.

13. The method of claim 1 , further comprising using the set of dominance metrics to characterize a group of people who have generated or accessed the first corpus A.

14. The method of claim 1 , further comprising using the set of dominance metrics to characterize a news source or official body that has generated or accessed the first corpus A.

15. An article of manufacture comprising a computer readable storage medium having instructions stored thereon which, when executed by one or more processors, cause the one or more processors to perform a method comprising:

generating a set of dominance metrics for a first corpus A having n documents, wherein the set of dominance metrics includes for each term i in a collection C a respective dominance metric D A (i) which is based on a quotient of:

an exponentiation of two by an entropy value for the respective term i; and

the number n of documents in the first corpus A;

wherein the entropy value for the respective term i is based on a respective sum of product values for each document j of first corpus A, wherein each of the product values is based on a respective product of:

a respective value p ij ; and

a logarithm of the respective value p ij ;

wherein the respective value p ij is based on a quotient of

a number of times term i occurs in document j; and

a total number of times term i occurs in first corpus A; and

storing information characterizing the first corpus A, the information based on the generated set of dominance metrics.

16. The article of manufacture of claim 15 , wherein the method steps further comprise:

sorting the set of dominance metrics; and

based on the sorted set of dominance metrics, identifying a sub-set of terms of the collection C; and

wherein the stored information characterizes the first corpus A based on the identified sub-set of terms.

17. The article of manufacture of claim 15 , wherein the method steps further comprise:

using the set of dominance metrics to characterize a person who generated or accessed the first corpus A.

18. The article of manufacture of claim 15 , wherein the method steps further comprise:

comparing a second set of dominance values from a second corpus B to the set of dominance values from the first corpus A.

19. The article of manufacture of claim 15 , wherein the method steps further comprise:

storing the first corpus A within a computer accessible database; and

displaying the set of dominance values to a display screen.

20. The article of manufacture of claim 18 , wherein the method steps further comprise:

storing the first corpus A within a computer accessible database; and

displaying a result of the comparing to a display screen.

21. The article of manufacture of claim 18 , wherein the comparing step further comprises:

creating a collection of dominant terms common to both first corpus A and second corpus B;

based on the collection of dominant terms, constructing two vectors of paired term-by-term similarities, one vector each for first corpus A and second corpus B; and

computing a final similarity measure between first corpus A and second corpus B including determining a cosine similarity of the two vectors of paired term-by-term similarities.

22. The method of claim 15 , further comprising using the set of dominance metrics to characterize a group of people who have generated or accessed the first corpus A.

23. The method of claim 15 , further comprising using the set of dominance metrics to characterize a news source or official body that has generated or accessed the first corpus A.

Assignments (3)
CHANGE OF NAME Recorded Jan 19, 2018
From: SANDIA CORPORATION
To: NATIONAL TECHNOLOGY & ENGINEERING SOLUTIONS OF SANDIA, LLC
Reel/Frame 045102/0144 →
CONFIRMATORY LICENSE Recorded Apr 23, 2009
From: SANDIA CORPORATION
To: U.S. DEPARTMENT OF ENERGY
Reel/Frame 022586/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2009
From: BAUER, TRAVIS L.; BENZ, ZACHARY O.; VERZI, STEPHEN J.
To: SANDIA CORPORATION, OPERATOR OF SANDIA NATIONAL LABORATORIES
Reel/Frame 022201/0590 →