IP Library Granted Patent US 8,315,849
Granted Patent B1
US 8,315,849 · App. 12/757,899 · Granted Nov 20, 2012

Selecting terms in a document

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,315,849
App. No.
12/757,899
Granted
Nov 20, 2012
Kind
B1
Abstract

Determining a mapping between a textual representation in a document and a concept is disclosed. A document is received. A set of candidate textual representations in the document is identified. For at least one candidate textual representation included in the set, an associated concept included in a taxonomy of concepts is determined. The candidate textual representation and the associated concept are provided as output.

Claims (51)

1. A system for determining a mapping between a textual representation in a document and a concept, comprising:

a communications interface configured to receive a document; and

a processor configured to:

identify a set of candidate textual representations in the document;

determine, the set of candidate textual representation included in the set, a set of associated concepts included in a taxonomy of concepts; and

sum a plurality of category vectors to generate a document vector, each category vector associated with an associated concept of the set of associated concepts and indicating correspondence of related concepts to the associated concept;

compute a set of document similarity scores for the set of associated concepts according to a correspondence of the category vectors corresponding thereto and the document vector;

select at least one representative concept of the associated concepts according to the set of document similarity scores;

provide as output the representative concept and a candidate textual representation of the set of candidate textual representations corresponding thereto; and

a memory coupled to the processor and configured to provide the processor with instructions.

2. The system of claim 1 wherein the processor is configured to identify the set of candidate textual representations at least in part by performing a greedy match against a list of entries included in the taxonomy.

3. The system of claim 2 wherein performing a greedy match includes detecting a preposition.

4. The system of claim 1 wherein the processor is further configured to prune the set of candidate textual representations based at least in part on a blacklist.

5. The system of claim 1 wherein the processor is further configured to prune the set of candidate textual representations based at least in part on a regular expression.

6. The system of claim 1 wherein the processor is further configured to prune the set of candidate textual representations based at least in part on a detection of a proper noun sequence.

7. The system of claim 1 wherein the processor is further configured to determine a part of speech of the at least one candidate textual representation.

8. The system of claim 1 wherein the processor is further configured to populate a feature vector associated with the candidate textual representation.

9. The system of claim 8 wherein the processor is configured to populate the feature vector at least in part by computing a part-of-speech score.

10. The system of claim 1 wherein the processor is further configured to determine an inverse document frequency score.

11. The system of claim 1 wherein determining the associated concept includes resolving a synonym.

12. The system of claim 1 wherein determining the associated concept includes performing a disambiguation.

13. The system of claim 12 wherein performing a disambiguation includes determining that a second textual representation is synonymous with the first textual representation.

14. The system of claim 1 wherein the processor is further configured to determine a threshold number of pairs of candidate textual representations and associated concepts to be provided as output.

15. The system of claim 1 , wherein the processor is further configured to:

determine a set of linkworthiness scores for the set of associated concepts; and

further select the at lest one representative concept according to the linkworthiness scores of the set of associated concepts.

16. The system of claim 1 , wherein the processor is further configured to:

determine a set of freshness scores for the set of associated concepts; and

further select the at lest one representative concept according to the freshness scores of the set of associated concepts.

17. A method for determining a mapping between a textual representation in a document and a concept, comprising:

receiving a document;

identifying a set of candidate textual representations in the document;

determining, for the set of candidate textual representations, a set of associated concepts included in a taxonomy of concepts;

summing a plurality of category vectors to generate a document vector, each category vector associated with an associated concept of the set of associated concepts and indicating correspondence of related concepts to the associated concept;

computing a set of document similarity scores for the set of associated concepts according to a correspondence of the category vectors corresponding thereto and the document vector;

selecting at least one representative concept of the associated concepts according to the set of document similarity scores; and

providing as output the at least one representative concept and a candidate textual representation of the set of candidate textual representations corresponding thereto.

18. The method of claim 17 , further comprising:

determining a set of linkworthiness scores for the set of associated concepts; and

further selecting the at lest one representative concept according to the linkworthiness scores of the set of associated concepts.

19. The method of claim 17 , further comprising:

determining a set of freshness scores for the set of associated concepts; and

further selecting the at lest one representative concept according to the freshness scores of the set of associated concepts.

20. A computer program product for determining a mapping between a textual representation in a document and a concept, the computer program product being embodied in a computer readable storage medium and comprising computer instructions for:

receiving a document;

identifying a set of candidate textual representations in the document;

determining, for the set of candidate textual representation included in the set, a set of associated concept included in a taxonomy of concepts;

summing a plurality of category vectors to generate a document vector, each category vector associated with an associated concept of the set of associated concepts and indicating correspondence of related concepts to the associated concept;

computing a set of document similarity scores for the set of associated concepts according to a correspondence of the category vectors corresponding thereto and the document vector;

selecting at least one representative concept of the associated concepts according to the set of document similarity scores; and

providing as output the at least one representative concept and a candidate textual representation of the set of candidate textual representations corresponding thereto.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2018
From: WAL-MART STORES, INC.
To: WALMART APOLLO, LLC
Reel/Frame 045817/0115 →
MERGER Recorded Apr 19, 2012
From: KOSMIX CORPORATION
To: WAL-MART STORES, INC.
Reel/Frame 028074/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2010
From: GATTANI, ABHISHEK; RAJARAMAN, ANAND
To: KOSMIX CORPORATION
Reel/Frame 024637/0259 →