IP Library Granted Patent US 8,626,491
Granted Patent B2
US 8,626,491 · App. 13/655,337 · Granted Jan 7, 2014

Selecting terms in a document

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,626,491
App. No.
13/655,337
Granted
Jan 7, 2014
Kind
B2
Abstract

Determining a mapping between a textual representation in a document and a concept is disclosed. A document is received. A set of candidate textual representations in the document is identified. For at least one candidate textual representation included in the set, an associated concept included in a taxonomy of concepts is determined. The candidate textual representation and the associated concept are provided as output.

Claims (56)

1. A method for document characterization, the method comprising:

identifying, by a computer system, a plurality of candidate textual representations in a document;

identifying, by a computer system, unambiguous textual representations from among the plurality of candidate textual representations, each unambiguous textual representation having a single concept associated therewith;

identifying, by a computer system, ambiguous textual representations from among the plurality of candidate textual representations, each ambiguous textual representation having a plurality of concepts associated therewith;

for each ambiguous textual representation:

determining, by the computer system, that one of the plurality of concepts associated with the ambiguous textual representation is also associated with an unambiguous textual representation; and

in response to determining that one of the plurality of concepts associated with the ambiguous representation is also associated with the unambiguous textual representation, pruning, by the computer system, concepts of the plurality of concepts other than the concept to which an unambiguous textual representation is mapped; and

outputting a modified version of the document having a portion of the candidate textual representations corresponding to at least a portion of the list of concepts following pruning hyperlinked to supplemental materials relating to the at least a portion of the list of concepts.

2. The method of claim 1 , further comprising pruning the list of concepts according to a blacklist.

3. The method of claim 1 , further comprising for remaining concepts of the plurality of concepts associated with an ambiguous textual representation:

identifying if a concept of the remaining concepts is listed in a whitelist; and

if so, pruning concepts of the remaining concepts other than the concept listed in the whitelist.

4. The method of claim 1 , further comprising:

calculating, for each concept of at least a portion of the concepts in the list of concepts, a score; and

pruning the list of concepts according to the scores for the at least a portion of the concepts.

5. The method of claim 4 , wherein calculating the score for each concept of at least a portion of the concepts in the list of concepts further comprises:

calculating the score at least in part based on an inverse document frequency for the concept according to a corpus of documents.

6. The method of claim 4 , wherein calculating the score for each concept of at least a portion of the concepts in the list of concepts further comprises:

calculating the score at least in part based on a number of homonyms associated with the concept.

7. The method of claim 4 , wherein calculating the score for each concept of at least a portion of the concepts in the list of concepts further comprises:

calculating the score at least in part based on linkworthiness of the concept according to a corpus of documents.

8. The method of claim 7 , wherein the linkworthiness of a concept is calculated as a number of times the concept is linked in the corpus of documents divided by the total number of times the concept is included in the corpus of documents.

9. The method of claim 4 , wherein calculating the score for each concept of at least a portion of the concepts in the list of concepts further comprises:

calculating the score at least in part based on current references to the concept in social media content.

10. The method of claim 4 , wherein calculating the score for each concept of at least a portion of the concepts in the list of concepts further comprises:

calculating the score at least in part based on one or more locations in the document of one or more textual representations associated with the concept.

11. The method of claim 4 , wherein calculating the score for each concept of at least a portion of the concepts in the list of concepts further comprises:

calculating the score at least in part based on a comparison of a category vector for the concept to a document vector, the category vector including concepts adjacent the concept in a taxonomy and the document vector including a combination of category vectors for at least a portion of the concepts of the list of concepts.

12. A system for document characterization, the system comprising one or more processors and one or more memory devices operably coupled to the one or more processors, the one or more memory devices storing executable data effective to cause the one or more processors to:

identify a plurality of candidate textual representations in a document;

identify unambiguous textual representations from among the plurality of candidate textual representations, each unambiguous textual representation having a single concept associated therewith;

identify ambiguous textual representations from among the plurality of candidate textual representations, each ambiguous textual representation having a plurality of concepts associated therewith;

for each ambiguous textual representation:

determine that one of the plurality of concepts associated with the ambiguous textual representation is also associated with an unambiguous textual representation; and

in response to determining that one of the plurality of concepts associated with the ambiguous representation is also associated with the unambiguous textual representation, prune concepts of the plurality of concepts other than the concept to which an unambiguous textual representation is mapped; and

output a modified version of the document having a portion of the candidate textual representations corresponding to at least a portion of the list of concepts following pruning hyperlinked to supplemental materials relating to the at least a portion of the list of concept and without any additional hyperlinks to the one of the pruned concepts of the plurality of concepts.

13. The system of claim 12 , wherein the executable data is further effective to cause the one or more processors to prune the list of concepts according to a blacklist.

14. The system of claim 12 , wherein the executable data is further effective to cause the one or more processors to, for remaining concepts of the plurality of concepts associated with an ambiguous textual representation:

identify if a concept of the remaining concepts is listed in a whitelist; and

if so, prune concepts of the remaining concepts other than the concept listed in the whitelist.

15. The system of claim 12 , wherein the executable data is further effective to cause the one or more processors to:

calculate, for each concept of at least a portion of the concepts in the list of concepts, a score; and

prune the list of concepts according to the scores for the at least a portion of the concepts.

16. The system of claim 15 , wherein the executable data is further effective to cause the one or more processors to calculate the score for each concept of at least a portion of the concepts in the list of concepts by:

calculating the score at least in part based on an inverse document frequency for the concept according to a corpus of documents.

17. The system of claim 15 , wherein the executable data is further effective to cause the one or more processors to calculate the score for each concept of at least a portion of the concepts in the list of concepts by:

calculating the score at least in part based on a number of homonyms associated with the concept.

18. The system of claim 15 , wherein the executable data is further effective to cause the one or more processors to calculate the score for each concept of at least a portion of the concepts in the list of concepts by:

calculating the score at least in part based on linkworthiness of the concept according to a corpus of documents.

19. The system of claim 18 , wherein the linkworthiness of a concept is calculated as a number of times the concept is linked in the corpus of documents divided by the total number of times the concept is included in the corpus of documents.

20. The system of claim 15 , wherein the executable data is further effective to cause the one or more processors to calculate the score for each concept of at least a portion of the concepts in the list of concepts by:

calculating the score at least in part based on current references to the concept in social media content.

21. The system of claim 15 , wherein the executable data is further effective to cause the one or more processors to calculate the score for each concept of at least a portion of the concepts in the list of concepts by:

calculating the score at least in part based on one or more locations in the document of one or more textual representations associated with the concept.

22. The system of claim 15 , wherein the executable data is further effective to cause the one or more processors to calculate the score for each concept of at least a portion of the concepts in the list of concepts by:

calculating the score at least in part based on a comparison of a category vector for the concept to a document vector, the category vector including concepts adjacent the concept in a taxonomy and the document vector including a combination of category vectors for at least a portion of the concepts of the list of concepts.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2018
From: WAL-MART STORES, INC.
To: WALMART APOLLO, LLC
Reel/Frame 045817/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2013
From: GATTANI, ABHISHEK; RAJARAMAN, ANAND
To: WAL-MART STORES, INC.
Reel/Frame 029589/0069 →