IP Library › Granted Patent US 8,595,245
Granted Patent B2
US 8,595,245 · App. 11/493,085 · Granted Nov 26, 2013

Reference resolution for text enrichment and normalization in mining mixed data

Inventors: Bruno Cavestro (Grenoble, FR); Jean-Michel Renders (Saint-Nazaire-les-Eymes, FR)
Assignee: Xerox Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,595,245
App. No.
11/493,085
Granted
Nov 26, 2013
Kind
B2
Abstract

A method for enrichment of text which enables mixed data mining includes generating a model for structured data found in tables of a database. In the model, semantically-linked terms are associated with referents, such as field names or cell content of the fields, of the structured data. The referents may be a business object or refer to a business object. A plurality of candidate referring entities in textual data in the database, such as chunks of free text, is identified. For each candidate referring entity, a similarity measure between the candidate referring entity in the textual data and the model is computed to identify referring entities of the candidate referring entities and corresponding business objects/referents to which the referring entities refer. The textual data is enriched with information derived from the business objects.

Claims (45)

1. A method for enrichment of text comprising:

providing a database in which defined hierarchical relationships exist between different parts of the data, the database having a structured part and an unstructured part, the database including a set of fields and a set of records, each of the fields being distinguished as comprising either structured data or unstructured data, the structured data fields having a predefined relationship to the structured data of each field, the structured part of the database including the structured data in structured data fields of records, and the unstructured part including unstructured data comprising textual data for unstructured data fields of the records, whereby some of records include both structured data in structured data fields and textual data for unstructured data fields;

after providing the database, generating a model for the structured data in the structured data fields of the structured part of the database, the generating comprising associating referents in the database with designating terms which each describe a business object, the referents each comprising or referring to one of the business objects;

identifying a plurality of candidate referring entities in the textual data of the unstructured data fields of the unstructured part of the provided database;

for each candidate referring entity, computing a similarity measure which includes:

comparing the candidate referring entity in the textual data with the model to identify referring entities of the candidate referring entities and corresponding objects to which the referring entities refer, and

comparing the candidate referring entity in the textual data with the structured data of the same record; and

based on the computed similarity measure, enriching the textual data for the unstructured data fields with information derived from the business objects for that record, the enrichment including annotating a free text entry in the database with information relating to a business object or referent.

2. The method of claim 1 , wherein the business objects comprise physical or logical object of significance to a business.

3. The method of claim 1 , wherein the referents comprise contents of fields for the structured data.

4. The method of claim 1 , wherein the generation of the model includes for each of the referents, identifying designating terms which are in a semantic relationship with the referent.

5. The method of claim 4 , wherein the identification of designating terms which are in a semantic relationship with the referent comprises accessing a lexical resource.

6. The method of claim 1 , wherein one of the designating terms associated with each of the referents comprises a normalized unique identifier of the object.

7. The method of claim 1 , wherein the identifying of candidate referring entities in the textual data comprises identifying noun phrases.

8. The method of claim 1 , wherein the identifying of candidate referring entities comprises identifying a normalized form of the candidate referring entity.

9. The method of claim 1 , wherein the identifying of candidate referring entities comprises at least one of:

expanding the textual data with external information relating to the textual data; and

co-reference resolution within the textual data or expanded textual data.

10. The method of claim 1 , wherein the associating of designating terms with the referents comprises, for each referent, associating a plurality of designating terms with the referent, each of the designating terms having a weight.

11. The method of claim 10 , wherein the computing of the similarity measure comprises computing the similarity measure between a candidate referring entity and a referent as a function of the weight of each of the designating terms.

12. The method of claim 1 , wherein the computing of the similarity measure comprises computing at least one of a string kernel value and a minimum edit distance between a candidate referring entity and a designating term of the model.

13. The method of claim 1 , wherein the comparing of the candidate referring entity in the textual data with structured data of the same record of a table or with structured data of another record linked to that record.

14. The method of claim 1 , wherein the enrichment comprises enriching the textual data with a normalized identifier of the business object to which the identified referent refers.

15. The method of claim 14 , wherein the candidate referring entity in the textual data is enriched with a normalized identifier of the business object, which comprises at least one of:

a name of a person, product, or service, and

an associated role or function of the person, product, or service.

16. The method of claim 1 , further comprising enriching the structured data with information derived from the object to which the referent refers.

17. A system comprising:

a database comprising records stored in memory which include structured data arranged in records comprising fields of structured data and textual data in fields of textual data, the textual data comprising annotations which identify business objects referred to by the structured data, developed by the method of claim 1 ; and

a processor which executes instructions in memory for querying the database to analyze text from the textual data and structured data.

18. A method of retrieving text responsive to a query comprising:

inputting a query;

retrieving information responsive to the query from stored structured and textual data, the textual data having been enriched according to the method of claim 1 .

19. The method of claim 1 , wherein each of the records includes both a structured part and an unstructured part.

20. The method of claim 1 , wherein the referents comprise field names for the structured data fields.

21. The method of claim 1 , wherein the database comprises a table in the form of cells, each of a plurality of the records having structured data in structured data cells of a respective row of cells and unstructured data in or linked to an unstructured data cell in the row of cells, each field comprising a column of cells which includes cells of the plurality of rows.

22. A system for enrichment of text comprising:

a database in which defined hierarchical relationships exist between different parts of the data, the database having a structured part and an unstructured part, the database including a set of fields and a set of records, each of the fields being distinguished as comprising either structured data or unstructured data, the structured data fields having a predefined relationship to the structured data of each field, the structured part of the database including the structured data in structured data fields of records, and the unstructured part including unstructured data comprising textual data for unstructured data fields of the records, whereby some of records include both structured data in structured data fields and textual data for unstructured data fields;

a model for structured data in structured data fields of the structured part of the database, the model associating referents in the database with designating terms which each describe a business object, the referents each comprising or referring to one of the business objects, the model having been generated after providing the database; and

a processor which:

identifies a plurality of candidate referring entities in the textual data of the unstructured data fields of the unstructured part of the database,

for each candidate referring entity, computes a similarity measure which includes:

comparing the candidate referring entity in the textual data with the model to identify referring entities of the candidate referring entities and corresponding objects to which the referring entities refer, and

comparing the candidate referring entity in the textual data with the structured data of the same record; and

based on the computed similarity measure, enriches the textual data for the unstructured data fields with information derived from the business objects for that record, the enrichment including annotating a free text entry in the database with information relating to a business object or referent.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2006
From: CAVESTRO, BRUNO; RENDERS, JEAN-MICHEL
To: XEROX CORPORATION
Reel/Frame 018133/0949 →
Continuity (1)
Related Publication 20080027893A1 · Jan 31, 2008