IP Library › Granted Patent US 11,334,268
Granted Patent B2
US 11,334,268 · App. 16/740,212 · Granted May 17, 2022

Data lineage and data provenance enhancement

Inventors: Paul R. Bastide (Ashland, MA); Aris Gkoulalas-Divanis (Waltham, MA); Rohit Ranchal (Austin, TX)
Assignee: International Business Machines Corporation
G06F3/0644G06F3/064G06F3/0622G06F3/0673G06F16/219G06F16/244G06F16/24578
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,334,268
App. No.
16/740,212
Granted
May 17, 2022
Kind
B2
Abstract

One embodiment of the invention provides a method for data lineage and data provenance enhancement. The method comprises arranging a data set into a logical ordering, and partitioning the data set into at least one set of partitions based on the logical ordering. The method further comprises, for each partition of the at least one set of partitions, determining a corresponding score for the partition, and determining a data similarity between the partition and each other partition of each other data set based on the corresponding score for the partition and another score corresponding to the other partition. The method further comprises determining data lineage of the data set based on each data similarity determined.

Claims (54)

1. A method for data lineage and data provenance enhancement, comprising:

determining a logical ordering in which to arrange a dimension of a data set based on application context and behavioral metadata;

arranging the data set into the logical ordering;

partitioning the data set into at least one set of partitions based on the logical ordering;

for each partition of the at least one set of partitions:

determining a corresponding score for the partition; and

determining a data similarity between the partition and each other partition of each other data set based on the corresponding score for the partition and another score corresponding to the other partition;

determining data lineage of the data set based on each data similarity determined;

determining whether the data set has a risk based on the data lineage; and

triggering a review cycle for the data set in response to determining the data set has a risk.

2. The method of claim 1 , wherein the logical ordering comprises a structured list of columns and rows.

3. The method of claim 1 , further comprising:

managing the data lineage, data pedigree, and data provenance of the data set based on each score determined for each partition of the data set.

4. The method of claim 1 , wherein the determining the corresponding score for the partition comprises determining a hash value of the partition.

5. The method of claim 1 , wherein the partitioning the data set into the at least one set of partitions based on the logical ordering comprises:

for each data block of the data set, progressively partitioning the data block into larger partitions until there is no data similarity between a partition that the data block is partitioned into and one or more other partitions of one or more other data sets.

6. The method of claim 1 , wherein the logical ordering is further based on change related data and shape related data, and wherein the application context, the behavioral metadata, the change related data, and the shape related data includes at least one of the following factors: location, time, labels, entry type, licensing, authorization, and statistical.

7. The method of claim 3 , wherein the managing comprises blocking propagation of the data set into a trusted data catalog.

8. The method of claim 3 , wherein the managing comprises labeling the data set as having a risk.

9. The method of claim 3 , wherein the managing comprises blocking compilation of the data set and enforcing secondary confirmation before authorizing the compilation of the data set.

10. A system for data lineage and data provenance enhancement, comprising:

at least one processor; and

a non-transitory processor-readable memory device storing instructions that when executed by the at least one processor causes the at least one processor to perform operations including:

receiving a data set;

determining a logical ordering in which to arrange a dimension of the data set based on application context and behavioral metadata;

arranging the data set into the logical ordering;

partitioning the data set into at least one set of partitions based on the logical ordering;

for each partition of the at least one set of partitions:

determining a corresponding score for the partition; and

determining a data similarity between the partition and each other partition of each other data set based on the corresponding score for the partition and another score corresponding to the other partition;

determining data lineage of the data set based on each data similarity determined;

determining whether the data set has a risk based on the data lineage; and

triggering a review cycle for the data set in response to determining the data set has a risk.

11. The system of claim 10 , wherein the logical ordering comprises a structured list of columns and rows.

12. The system of claim 10 , wherein the operations further include:

managing the data lineage, data pedigree, and data provenance of the data set based on each score determined for each partition of the data set.

13. The system of claim 10 , wherein the determining the corresponding score for the partition comprises determining a hash value of the partition.

14. The system of claim 10 , wherein the partitioning the data set into the at least one set of partitions based on the logical ordering comprises:

for each data block of the data set, progressively partitioning the data block into larger partitions until there is no data similarity between a partition that the data block is partitioned into and one or more other partitions of one or more other data sets.

15. The system of claim 10 , wherein the logical ordering is further based on change related data and shape related data, and wherein the application context, the behavioral metadata, the change related data, and the shape related data includes at least one of the following factors: location, time, labels, entry type, licensing, authorization, and statistical.

16. The system of claim 12 , wherein the managing comprises blocking propagation of the data set into a trusted data catalog.

17. The system of claim 12 , wherein the managing comprises labeling the data set as having a risk.

18. The system of claim 12 , wherein the managing comprises blocking compilation of the data set and enforcing secondary confirmation before authorizing the compilation of the data set.

19. A computer program product for data lineage and data provenance enhancement, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

determine a logical ordering in which to arrange a dimension of a data set based on application context and behavioral metadata;

arrange the data set into the logical ordering;

partition the data set into at least one set of partitions based on the logical ordering;

for each partition of the at least one set of partitions:

determine a corresponding score for the partition; and

determine a data similarity between the partition and each other partition of each other data set based on the corresponding score for the partition and another score corresponding to the other partition;

determine data lineage of the data set based on each data similarity determined;

determine whether the data set has a risk based on the data lineage; and

trigger a review cycle for the data set in response to determining the data set has a risk.

20. The computer program product of claim 19 , wherein the logical ordering comprises a structured list of columns and rows.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2020
From: BASTIDE, PAUL R.; GKOULALAS-DIVANIS, ARIS; RANCHAL, ROHIT
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 051483/0422 →
Continuity (1)
Related Publication 20210216229A1 · Jul 15, 2021