IP Library › Granted Patent US 8,874,613
Granted Patent B2
US 8,874,613 · App. 13/891,130 · Granted Oct 28, 2014

Semantic discovery and mapping between data sources

Inventors: Alexander Gorelik (Palo Alto, CA); Lingling Yan (Morgan Hill, CA)
Assignee: International Business Machines Corporation
G06F17/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,874,613
App. No.
13/891,130
Filed
May 9, 2013
Granted
Oct 28, 2014
Kind
B2
Examiner
WOO, ISAAC M
Art Unit
2155
USPC
707/791
Abstract

An apparatus and method are described for the discovery of semantics, relationships and mappings between data in different software applications, databases, files, reports, messages, or systems. In one aspect, semantics and relationships and mappings are identified between a first and a second data source. A binding condition is discovered between portions of data in the first and the second data source. The binding condition is used to discover correlations between portions of data in the first and the second data source. The binding condition and the correlations are used to discover a transformation function between portions of data in the first and the second data source.

Claims (24)

1. A method for determining a mapping between a plurality of columns of a first data source and a second data source by using a fuzzy match, the method comprising:

identifying a join condition between the plurality of columns of the first data source and the second data source;

constructing source data objects based on a set of relationships between the plurality of columns of the first data source and the second data source and the identified join condition;

determining a fuzzy match between the plurality of columns of the first data source and the second data;

identifying the mapping corresponding to the fuzzy match between the plurality of columns of the first data source and the second data source; and

generating a window providing a display of success for the mapping between the columns of the first data source and the second data source.

2. The method as recited in claim 1 , wherein the determining the fuzzy match comprises:

matching the columns of the first data source with the second data source;

identifying a score based on degree of matching of the columns of the first data source with the columns of the second data source; and

identifying a binding condition based on the score of matching, when the score of matching is more than a threshold.

3. The method as recited in claim 1 , wherein identifying the join condition between the plurality of columns of the first data source and the second data source, comprises:

discovering key columns, wherein the key columns are based on a unique index of the columns in the first data source and the second data source; and

discovering a foreign key, wherein the foreign key is based on value match between the discovered key columns.

4. A method for determining a mapping between a plurality of columns of a first data source and a second data source by using a fuzzy match, the method comprising:

identifying a join condition between the plurality of columns of the first data source and the second data source;

constructing source data objects based on a set of relationships between the plurality of columns of the first data source and the second data source and the identified join condition;

determining a fuzzy match between the plurality of columns of the first data source and the second data source by performing:

identifying a list of pre-loaded synonyms corresponding to the columns in the first data source;

comparing the list of pre-loaded synonyms to the columns in the second data source;

identifying a relevancy score based on the comparison between the list of the pre-loaded synonyms and the columns in the second data source; and

determining the mapping based on the relevancy score; and

identifying the mapping corresponding to the fuzzy match between the plurality of columns of the first data source and the second data source.

5. The method as recited in claim 4 , wherein the mapping is used to identify a binding condition between the first data source and the second data source.

6. The method as recited in claim 4 , wherein the list of synonyms is discovered.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2018
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: WORKDAY, INC.
Reel/Frame 046311/0942 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2013
From: GORELIK, ALEXANDER; YAN, LINGLING
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 030432/0359 →
Continuity (5)
Division 13267292 · Oct 6, 2011
Division 12283477 · Sep 12, 2008
Continuation 10938205 · Sep 9, 2004
Provisional Application 60502043 · Sep 10, 2003
Related Publication 20130254183A1 · Sep 26, 2013