IP Library Granted Patent US 10,528,532
Granted Patent B2
US 10,528,532 · App. 14/905,240 · Granted Jan 7, 2020

Systems and methods for data integration

Inventors: Paolo Papotti (Doha, QA); Felix Naumann (Doha, QA); Sebastian Kruse (Doha, QA); El Kindi Rezig (Doha, QA)
Assignees: Qatar Foundation; Hasso-Plattner-Institut Für Softwaresystemtechnik GmbH
G06F16/215G06F16/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,528,532
App. No.
14/905,240
Granted
Jan 7, 2020
Kind
B2
Abstract

A computer implemented method for integrating data into a target database may include: providing a plurality of source databases which each may include a relational schema and data for integration into the target database; generating at least one complexity model based on the relational schema and data of each source database, each complexity model indicating at least one inconsistency between two or more of the data sources which may be require to be resolved to integrate the data from the data sources into the target database; and generating an effort model that may include an effort value for each inconsistency indicated by each complexity model, each effort value indicating at least one of a time period and a financial cost to resolve the inconsistency to integrate data from the data sources into the target database.

Claims (43)

1. A computer implemented method for integrating data into a target database, the method comprising:

providing a plurality of source databases which each comprise a relational schema and the data for integration into the target database;

identifying a correspondence between data for a source element and a target element of the each source database and the target database;

identifying at least one integration challenge and, providing an extensible automatic effort estimation framework (Efes) which utilizes a plurality of plug-ins submodels to resolve the at least one integration challenge, the plurality of plug-ins submodels are configured to:

generate at least one complexity model based on the relational schema and the identified correspondence between the source element and target element data of each source database, each complexity model indicating at least one inconsistency between two or more of the data sources which must be resolved to integrate the data from the data sources into the target database;

generate an effort model comprising an effort value for each inconsistency indicated by each complexity model using integration tools to integrate the source element into the target element, each effort value indicating cost required to resolve the inconsistency to integrate data from the data sources into the target database; and

integrate data from at least one of the source databases into the target database.

2. The method of claim 1 , wherein the method further comprises:

comparing the effort values for a plurality of data sources; and

selecting at least one data source that produce a minimum effort value for integration into the target database.

3. The method of claim 1 , wherein the method further comprises:

outputting at least one effort value to a user for the user to determine from the effort value the effort required to integrate data into the target database.

4. The method of claim 1 , further comprising generating the effort model using a goal profile wherein the goal profile is based on task models, wherein the task model is a specified set of operations on the data.

5. The method of claim 1 , further comprising generating the effort model using an effort-calculation function for each task type, wherein the task type is an operation on the data.

6. The method of claim 1 , wherein the method further comprises generating a graph using the at least one complexity model to indicate the parts of the schema of the target database that are more complex than other parts of the schema of the target database.

7. The method of claim 1 , wherein generating the complexity model is independent of the language used to express the data transformation, the expressiveness of the data model and the constraint language.

8. The method of claim 1 , further comprising generating the complexity model using integration tools to inspect a plurality of single attributes of each data source.

9. The method of claim 1 , wherein the method comprises generating a plurality of complexity models based on the relational schema and data of each source database.

10. The method of claim 1 , wherein the method further comprises:

receiving at least one sub-model for use with each effort model and each complexity model to exchange, improve or extend each effort model or each complexity model.

11. The method of claim 10 , wherein the method comprises receiving a plurality of data complexity sub-models, each data complexity sub-model representing the data integration scenario with respect to one certain complexity aspect.

12. The method of claim 1 , wherein the method further comprises:

comparing the effort values for a plurality of data sources;

selecting at least one data source that produce a minimum effort value for integration into the target database; and

outputting at least one effort value to a user for the user to determine from the effort value the effort required to integrate data into the target database.

13. The method of claim 1 , wherein the method further comprises:

comparing the effort values for a plurality of data sources; and

selecting at least one data source that produce a minimum effort value for integration into the target database; and

wherein generating the complexity model is independent of the language used to express the data transformation, the expressiveness of the data model and the constraint language.

14. A non-transitory computer readable storage medium storing instructions which, when executed by a computer, cause the computer to:

provide a plurality of source databases which each comprise a relational schema and the data for integration into the target database;

identify a correspondence between data for a source element and a target element of the each source database and the target database;

identify at least one integration challenge and, provide an extensible automatic effort estimation framework (Efes) which utilizes a plurality of plug-ins submodels to resolve the at least one integration challenge, the plurality of plug-ins submodels are configured to:

generate at least one complexity model based on the relational schema and the identified correspondence between the source element and target element data of each source database, each complexity model indicating at least one inconsistency between two or more of the data sources which must be resolved to integrate the data from the data sources into the target database;

generate an effort model comprising an effort value for each inconsistency indicated by each complexity model using integration tools to integrate the source element into the target element, each effort value indicating cost required to resolve the inconsistency to integrate data from the data sources into the target database; and

integrate data from at least one of the source databases into the target database.

15. A system for integrating data into a target database, the system comprising computing device which incorporates a processor and a memory, the memory storing instructions which, when executed by the processor, cause the processor to:

provide a plurality of source databases which each comprise a relational schema and the data for integration into the target database;

identify a correspondence between data for a source element and a target element of the each source database and the target database;

identify at least one integration challenge and, provide an extensible automatic effort estimation framework (Efes) which utilizes a plurality of plug-ins submodels to resolve the at least one integration challenge, the plurality of plug-ins submodels are configured to:

generate at least one complexity model based on the relational schema and the identified correspondence between the source element and target element data of each source database, each complexity model indicating at least one inconsistency between two or more of the data sources which must be resolved to integrate the data from the data sources into the target database;

generate an effort model comprising an effort value for each inconsistency indicated by each complexity model using integration tools to integrate the source element into the target element, each effort value indicating cost required to resolve the inconsistency to integrate data from the data sources into the target database; and

integrate data from at least one of the source databases into the target database.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2025
From: QATAR FOUNDATION FOR EDUCATION, SCIENCE & COMMUNITY DEVELOPMENT
To: HAMAD BIN KHALIFA UNIVERSITY
Reel/Frame 069936/0656 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2018
From: PAPOTTI, PAOLO; REZIG, EL KINDI
To: QATAR FOUNDATION
Reel/Frame 046041/0409 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2018
From: NAUMANN, FELIX; KRUSE, SEBASTIAN
To: HASSO-PLATTNER-INSTITUT FUR SOFTWARESYSTEMTECHNIK GMBH
Reel/Frame 046041/0501 →
Priority Claims (1)
GB 1312776.6 · Jul 17, 2013 · national
Continuity (1)
Related Publication 20160154830A1 · Jun 2, 2016
Cited By (3)
US 12,242,460 US 12,265,519 US 12,326,839