IP Library Granted Patent US 7,426,520
Granted Patent B2
US 7,426,520 · App. 10/938,205 · Granted Sep 16, 2008

Method and apparatus for semantic discovery and mapping between data sources

Assignee: Exeros, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,426,520
App. No.
10/938,205
Filed
Sep 9, 2004
Granted
Sep 16, 2008
Kind
B2
Examiner
WOO, ISAAC M
Art Unit
2166
USPC
707/101
Abstract

An apparatus and method are described for the discovery of semantics, relationships and mappings between data in different software applications, databases, files, reports, messages, or systems. In one aspect, semantics and relationships and mappings are identified between a first and a second data source. A binding condition is discovered between portions of data in the first and the second data source. The binding condition is used to discover correlations between portions of data in the first and the second data source. The binding condition and the correlations are used to discover a transformation function between portions of data in the first and the second data source.

Claims (79)

1. A method for generating a mapping suitable for analyzing data in a first data source and a second data source, for identifying semantics and relationships between the first data source and the second data source, the method comprising:

determining a binding condition between portions of data in the first data source and the second data source, wherein determining the binding condition comprises:

creating a first column index table for the first data source;

creating a second column index table for the second data source;

creating an overlap column table from the first and second column index tables to identify potential binding conditions;

constructing binding conditions from the potential binding conditions; and

validating the binding conditions using the correlations;

using the binding condition to determine correlations between portions of data in the first data source and the second data source; and

using the binding condition and the correlations to generate the mapping between portions of data in the first data source and the second data source.

2. The method of claim 1 , wherein constructing the binding conditions comprises:

building a first row index table for the first data source;

building a second row index table for the second data source;

using the first and second row index tables to create a row match table to identify the potential binding conditions having a high co-occurrence;

combining the potential binding conditions having a high co-occurrence; and

generating binding condition strings from the combined potential binding conditions.

3. The method of claim 1 , wherein the correlations are determined by a correlation determining process comprising:

for each combination of a first and second data source column, identifying a maximum count of correlated rows between the first and second data source columns.

4. The method of claim 1 , further comprising determining a filter for the binding condition.

5. The method of claim 1 , further comprising determining a filter for the mapping.

6. The method of claim 1 , further comprising:

determining a join condition; and

using the join condition to generate a first data object for the first data source.

7. The method of claim 1 , further comprising performing schema matching between the first data object and a second data object of the second data source.

8. The method of claim 7 , wherein performing schema matching comprises constructing a metadata index including relevance scores.

9. The method of claim 8 , wherein the relevance scores are determined using a multiplier.

10. The method of claim 1 , wherein determining the mapping comprises:

for each second data source column, reading correlated first data source columns;

identifying constants;

for each column in the second data source, finding a best match; and

generating the mapping based on the best matches.

11. A system comprising:

a processing unit coupled to a memory through a bus; and

a process for generating a mapping suitable for analyzing data in a first data source and a second data source, for identifying semantics and relationships between the first data source and the second data source, the method comprising:

determining a binding condition between portions of data in the first data source and the second data source, wherein determining the binding condition comprises:

creating a first column index table for the first data source;

creating a second column index table for the second data source;

creating an overlap column table from the first and second column index tables to identify potential binding conditions;

constructing binding conditions from the potential binding conditions; and

validating the binding conditions using the correlations;

using the binding condition to determine correlations between portions of data in the first data source and the second data source; and

using the binding condition and the correlations to generate the mapping between portions of data in the first data source and the second data source.

12. The system of claim 11 , wherein constructing the binding conditions comprises:

building a first row index table for the first data source;

building a second row index table for the second data source;

using the first and second row index tables to create a row match table to identify the potential binding conditions having a high co-occurrence;

combining the potential binding conditions having a high co-occurrence; and

generating binding condition strings from the combined potential binding conditions.

13. The system of claim 11 , wherein the correlations are determined by a correlation determining process comprising:

for each combination of a first and second data source column, identifying a maximum count of correlated rows between the first and second data source columns.

14. The system of claim 11 , further comprising determining a filter for the binding condition.

15. The system of claim 11 , further comprising determining a filter for the mapping.

16. The system of claim 11 , further comprising:

determining a join condition; and

using the join condition to generate a first data object for the first data source.

17. The system of claim 11 , further comprising performing schema matching between the first data object and a second data object of the second data source.

18. The system of claim 17 , wherein performing schema matching comprises constructing a metadata index including relevance scores.

19. The system of claim 18 , wherein the relevance scores are determined using a multiplier.

20. The system of claim 11 , wherein determining the mapping comprises:

for each second data source column, reading correlated first data source columns;

identifying constants;

for each column in the second data source, finding a best match; and

generating the mapping based on the best matches.

21. A machine-readable medium having instructions to cause a machine to perform a machine-implemented method for generating a mapping suitable for analyzing data in a first data source and a second data source, for identifying semantics and relationships between the first data source and the second data source, the method comprising:

determining a binding condition between portions of data in the first data source and the second data source, wherein determining the binding condition comprises:

creating a first column index table for the first data source;

creating a second column index table for the second data source;

creating an overlap column table from the first and second column index tables to identify potential binding conditions;

constructing binding conditions from the potential binding conditions; and

validating the binding conditions using the correlations;

using the binding condition to determine correlations between portions of data in the first data source and the second data source; and

using the binding condition and the correlations to generate the mapping between portions of data in the first data source and the second data source.

22. The machine-readable medium of claim 21 , wherein constructing the binding conditions comprises:

building a first row index table for the first data source;

building a second row index table for the second data source;

using the first and second row index tables to create a row match table to identify the potential binding conditions having a high co-occurrence;

combining the potential binding conditions having a high co-occurrence; and

generating binding condition strings from the combined potential binding conditions.

23. The machine-readable medium of claim 21 , wherein the correlations are determined by a correlation determining process comprising:

for each combination of a first and second data source column, identifying a maximum count of correlated rows between the first and second data source columns.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2018
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: WORKDAY, INC.
Reel/Frame 046311/0942 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2009
From: EXEROS, INC.; MDM UNIVERSITY, LLC
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 022846/0309 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2004
From: GORELIK, ALEXANDER; YAN, LINGLING
To: EXEROS, INC.
Reel/Frame 015784/0945 →
Continuity (2)
Provisional Application 6050204300 · Sep 10, 2003
Related Publication 20050055369A1 · Mar 10, 2005