IP Library Granted Patent US 9,501,567
Granted Patent B2
US 9,501,567 · App. 13/368,845 · Granted Nov 22, 2016

User-guided multi-schema integration

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,501,567
App. No.
13/368,845
Granted
Nov 22, 2016
Kind
B2
Abstract

Methods, systems, and computer-readable storage media for user-guided multi-schema integration and include actions of receiving a plurality of schemas, each schema defining a data structure and including a plurality of intermediate elements and a plurality of leaf elements, receiving leaf correspondences that match leaf elements between schemas of the plurality of schemas, processing the plurality of schemas and the leaf correspondences using closed frequent itemset mining to define a first plurality of redundancy groups, each redundancy group including a proposed correspondence between intermediate elements of schemas, displaying the first plurality of redundancy groups to a user, receiving user input, the user input including one or more actions to one or more redundancy groups in the first plurality of redundancy groups, processing the plurality of schemas, the leaf correspondences and the one or more actions to define a second plurality of redundancy groups, and displaying the second plurality of redundancy groups.

Claims (46)

1. A computer-implemented method of providing a user-guided multi-schema integration, the method being executed using one or more processors and comprising:

receiving a plurality of schemas from computer-readable memory, each schema of the plurality of schemas defining a data structure and comprising a plurality of intermediate elements and a plurality of leaf elements;

receiving leaf correspondences from a computer-readable memory, the leaf correspondences matching leaf elements between schemas of the plurality of schemas and being associated with a maximum of a confidence level,

processing the plurality of schemas and the leaf correspondences using closed frequent itemset mining (CFIM) to determine intermediate correspondences, the intermediate correspondences matching intermediate elements between schemas of the plurality of schemas and being associated to the confidence level that depends on a position of the intermediate elements in each schema of the plurality of schemas and to define a first plurality of redundancy groups, each redundancy group in the first plurality of redundancy groups comprising a proposed correspondence between intermediate elements of schemas of the plurality of schemas;

displaying, using a display device, the first plurality of redundancy groups to a user;

receiving user input, the user input comprising one or more actions to one or more redundancy groups in the first plurality of redundancy groups;

processing the plurality of schemas and the leaf correspondences to hide at least one of the plurality of schemas based on the one or more actions to define a second plurality of redundancy groups; and

displaying, using a display device, the second plurality of redundancy groups to the user.

2. The method of claim 1 , wherein the one or more actions comprise approving a subset of redundancy groups of the first plurality of redundancy groups, the subset comprising at least one redundancy group, and, in response to approving the subset of redundancy groups, defining one or more respective correspondences between intermediate elements of schemas of the plurality of schemas.

3. The method of claim 2 , wherein the at least one redundancy group is absent from the second plurality of redundancy groups.

4. The method of claim 2 , further comprising identifying one or more sub-correspondences based on the subset of redundancy groups, wherein redundancy groups associated with the one or more sub-correspondences are absent from the second plurality of redundancy groups.

5. The method of claim 2 , further comprising identifying one or more conflicting correspondences based on the subset of redundancy groups, wherein redundancy groups associated with the one or more conflicting correspondences are absent from the second plurality of redundancy groups.

6. The method of claim 2 , wherein processing the plurality of schemas, the leaf correspondences and the one or more actions to define the second plurality of redundancy groups comprises processing the plurality of schemas, the leaf correspondences and the one or more respective correspondences between intermediate elements of schemas of the plurality of schemas.

7. The method of claim 1 , wherein the one or more actions comprise disapproving a subset of redundancy groups of the first plurality of redundancy groups, the subset comprising at least one redundancy group.

8. The method of claim 7 , wherein the at least one redundancy group is absent from the second plurality of redundancy groups.

9. The method of claim 1 , further comprising:

determining, for each redundancy group in the first plurality of redundancy groups, a rank to provide a plurality of ranks; and

determining a rank order based on the plurality of ranks,

wherein displaying the first plurality of redundancy groups comprises displaying redundancy groups of the first plurality of redundancy groups based on the rank order.

10. The method of claim 1 , further comprising:

determining, for each redundancy group in the second plurality of redundancy groups, a rank to provide a plurality of ranks; and

determining a rank order based on the plurality of ranks,

wherein displaying the second plurality of redundancy groups comprises displaying redundancy groups of the second plurality of redundancy groups based on the rank order.

11. The method of claim 1 , wherein processing the plurality of schemas and the leaf correspondences using closed frequent itemset mining (CFIM) to define the first plurality of redundancy groups comprises transforming schemas of the plurality of schemas into respective linear inputs.

12. The method of claim 1 , wherein the second plurality of redundancy groups comprises redundancy groups of the first plurality of redundancy groups.

13. The method of claim 1 , further comprising:

defining one or more respective correspondences between intermediate elements of schemas of the plurality of schemas based on the one or more actions; and

providing a unified data model based on the leaf correspondences and the one or more respective correspondences.

14. A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for improving keyword searches, the operations comprising:

receiving a plurality of schemas, each schema of the plurality of schemas defining a data structure and comprising a plurality of intermediate elements and a plurality of leaf elements;

receiving leaf correspondences, the leaf correspondences matching leaf elements between schemas of the plurality of schemas and being associated with a maximum of a confidence level;

processing the plurality of schemas and the leaf correspondences using closed frequent itemset mining (CFIM) to determine intermediate correspondences, the intermediate correspondences matching intermediate elements between schemas of the plurality of schemas and being associated to the confidence level that depends on a position of the intermediate elements in each schema of the plurality of schemas and to define a first plurality of redundancy groups, each redundancy group in the first plurality of redundancy groups comprising a proposed correspondence between intermediate elements of schemas of the plurality of schemas;

providing the first plurality of redundancy groups for display to a user;

receiving user input, the user input comprising one or more actions to one or more redundancy groups in the first plurality of redundancy groups;

processing the plurality of schemas and the leaf correspondences to hide at least one of the plurality of schemas based on the one or more actions to define a second plurality of redundancy groups; and

providing the second plurality of redundancy groups for display to the user.

15. A system, comprising:

a computing device; and

a computer-readable storage device coupled to the computing device and having instructions stored thereon which, when executed by the computing device, cause the computing device to perform operations for improving keyword searches for enterprise services, the operations comprising:

receiving a plurality of schemas, each schema of the plurality of schemas defining a data structure and comprising a plurality of intermediate elements and a plurality of leaf elements;

receiving leaf correspondences, the leaf correspondences matching leaf elements between schemas of the plurality of schemas and being associated with a maximum of a confidence level;

processing the plurality of schemas and the leaf correspondences using closed frequent itemset mining (CFIM) to determine intermediate correspondences, the intermediate correspondences matching intermediate elements between schemas of the plurality of schemas and being associated to the confidence level that depends on a position of the intermediate elements in each schema of the plurality of schemas and to define a first plurality of redundancy groups, each redundancy group in the first plurality of redundancy groups comprising a proposed correspondence between intermediate elements of schemas of the plurality of schemas;

providing the first plurality of redundancy groups for display to a user;

receiving user input, the user input comprising one or more actions to one or more redundancy groups in the first plurality of redundancy groups;

processing the plurality of schemas and the leaf correspondences to hide at least one of the plurality of schemas based on the one or more actions to define a second plurality of redundancy groups; and

providing the second plurality of redundancy groups for display to the user.

Assignments (2)
CHANGE OF NAME Recorded Aug 26, 2014
From: SAP AG
To: SAP SE
Reel/Frame 033625/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2012
From: LEMCKE, JENS; KHAN, MUHAMMAD WASIMULLAH; STUHEC, GUNTHER
To: SAP AG
Reel/Frame 027885/0703 →