IP Library › Granted Patent US 8,239,349
Granted Patent B2
US 8,239,349 · App. 12/900,133 · Granted Aug 7, 2012

Extracting data

Assignee: Hewlett-Packard Development Company, L.P.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,239,349
App. No.
12/900,133
Filed
Oct 7, 2010
Granted
Aug 7, 2012
Kind
B2
Examiner
WOO, ISAAC M
Art Unit
2155
USPC
707/749
Abstract

Information can be extracted from unstructured documents using embodiments described herein. An entity recognition may be performed on an unstructured document and found entities may be annotated. Annotating includes inserting tags around the found entities to generate marked entities. A rule is applied to each of the marked entities in the unstructured document to generate a confidence value for every marked entity, wherein the rule comprises a plurality of prefixes for a target entity and a plurality of suffixes for the target entity. A marked entity with the highest confidence value is selected as an extraction target.

Claims (73)

1. A system for extracting data, comprising:

a processor;

a storage system, comprising:

an unstructured document;

a rule comprising a plurality of prefixes for a target entity and a plurality of suffixes for the target entity, and

code configured to direct the processor to:

perform an entity recognition on the unstructured document to mark entities by type;

apply the rule to marked entities; and

generate a confidence value for each of the marked entities, where the confidence value indicates a degree of match to the rule.

2. The system of claim 1 , comprising code configured to direct the processor to select the entity with a highest confidence value as an extraction target.

3. The system of claim 1 , comprising code configured to direct the processor to annotate each token of the unstructured document with a parts-of-speech tag.

4. The system of claim 3 , wherein the memory comprises code configured to direct the processor to direct the processor to annotate each token of the unstructured document with a parts-of-speech tag.

5. The system of claim 1 , wherein the memory comprises code configured to direct the processor to:

obtain a plurality of training documents, wherein an entity of interest is marked in each of the plurality of training documents;

generate a plurality of prefixes for the entity of interest;

generate a plurality of suffixes for the entity of interest;

create a plurality of individuals, wherein each of the plurality of individuals comprises a sequence of prefix words and a sequence of suffix words;

apply each of the plurality of individuals to each of the plurality of training documents to generate a fitness score associated with each of the plurality of individuals; and

convert an individual having the highest fitness score into a new rule.

6. The system of claim 5 , comprising code configured to direct the processor to annotate each token of each of the plurality of training documents with a parts-of-speech tag.

7. The system of claim 5 , wherein the storage system comprises code configured to direct the processor to:

evolve the plurality of individuals for a preselected number of generations; and

convert an individual rule having the highest fitness score into a rule.

8. The system of claim 7 , wherein the memory comprises code configured to direct the processor to:

remove suffixes associated with the rule from the plurality of suffixes;

remove prefixes associated with the rule from the plurality of prefixes;

generate a new plurality of individuals;

apply each of the new plurality of individuals to each of the plurality of documents to generate a fitness score associated with each of the new plurality of individuals;

evolve the new plurality of individuals for a preselected number of generations; and

convert an individual in the new plurality of individuals having the highest fitness score into a new rule.

9. The system of claim 5 , wherein evolving the plurality of individuals comprises:

calculating a fitness score for each individual in a population;

identifying a subset of individuals having a fitness score above a threshold;

discarding all of the individuals having fitness scores less than the threshold; and

recombining the individuals to form a new generation of individuals.

10. The system of claim 9 , wherein the threshold is selected to eliminate a portion of individuals from the population.

11. The system of claim 9 , wherein recombining the individuals comprises randomly selecting prefixes and suffixes from at least two individuals to create a new individual.

12. The system of claim 1 , comprising a network interface, wherein the storage system comprises code configured to direct the processor to obtain an unstructured document through the network interface.

13. The system of claim 1 , comprising a display, wherein the storage system comprises code configured to direct the processor to display an unstructured document, wherein the entities and extraction targets of the unstructured document are marked.

14. A method of extracting a target entity from a document, comprising:

performing an entity recognition on an unstructured document in a database to create found entities;

annotating the found entities in the unstructured document, wherein annotating comprises inserting tags around the found entities to generate marked entities;

applying a rule to each of the marked entities in the unstructured document to generate a confidence value for every entity, wherein the rule comprises a plurality of prefixes for a target entity and a plurality of suffixes for the target entity; and

selecting a marked entity with the highest confidence value as an extraction target.

15. The method of claim 14 , comprising generating the rule through the application of a genetic algorithm to a plurality of training documents.

16. The method of claim 15 , comprising annotating each token of the plurality of training documents and the unstructured document with a parts-of-speech tag.

17. The method of claim 15 , comprising generating additional rules by:

removing suffixes associated with the rule from a plurality of suffixes;

removing prefixes associated with the rule from a plurality of prefixes;

generating a new plurality of individuals;

applying each of the new plurality of individuals to each of the plurality of training documents to generate a fitness score associated with each of the new plurality of individuals;

evolve the new plurality of individuals for a preselected number of generations; and

convert an individual in the new plurality of individuals having the highest fitness score into a new rule.

18. A non-transitory, computer readable medium comprising code configured to direct a processor to:

load an unstructured document into a memory;

perform an entity recognition on the unstructured document;

annotate entities recognized in the unstructured document to generate marked entities;

apply a rule to the marked entities to generate a confidence value for every marked entity, wherein the rule comprises a plurality of prefixes for a target entity and a plurality of suffixes for the target entity; and

select the marked entity with a highest confidence value as an extraction target.

19. The non-transitory computer readable medium of claim 18 , comprising code configured to direct the processor to:

obtain a plurality of training documents, wherein an entity of interest is marked in each of the plurality of training documents;

generate a bag of words of prefixes for the entity of interest;

generate a bag of words of suffixes for the entity of interest;

create a plurality of individuals, wherein each of the plurality of individuals comprises a plurality of prefix words selected the bag of words of prefixes and a plurality of suffix words selected from the bag of words of suffixes;

apply each of the plurality of individuals to each of the plurality of training documents to generate a fitness score associated with each of the plurality of individuals; and

convert an individual having the highest fitness score into a new rule.

20. The non-transitory computer readable medium of claim 19 , comprising code configured to direct the processor to:

remove suffixes associated with the rule from the bag of words of prefixes;

remove prefixes associated with the rule from the bag of words of suffixes;

generate a new plurality of individuals by selecting a plurality of prefix words from the bag of words of prefixes and selecting a plurality of suffix words from the bag of words of suffixes;

apply each of the new plurality of individuals to each of the plurality of documents to generate a fitness score for each of the new plurality of individuals;

evolve the new plurality of individuals for a preselected number of generations; and

convert an individual in the new plurality of individuals having the highest fitness score into a new rule.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2015
From: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 037079/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 7, 2010
From: CASTELLANOS, MARIA G.; DURAZO, MIGUEL; DAYAL, UMESHWAR
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P. COMPANY
Reel/Frame 025110/0357 →
Continuity (1)
Related Publication 20120089620A1 · Apr 12, 2012