IP Library Granted Patent US 9,811,573
Granted Patent B1
US 9,811,573 · App. 14/039,537 · Granted Nov 7, 2017

Lineage information management in data analytics

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,811,573
App. No.
14/039,537
Granted
Nov 7, 2017
Kind
B1
Abstract

A data analytics workload is obtained, wherein the data analytics workload includes one or more execution parameters and an input data set. An identifier specific to the data analytics workload is generated. The data analytics workload is at least partially executed based on the one or more execution parameters and the input data set to generate an output data set. Meta data associated with the output data set generated by execution of the data analytics workload is obtained, wherein the meta data includes lineage information corresponding to the output data set generated by execution of the input data set. The meta data is registered in a meta data store.

Claims (52)

1. A method comprising steps of:

obtaining an input data set and a first data analytics workload, wherein the first data analytics workload comprises one or more execution parameters for executing the first data analytics workload based on the input data set, and wherein the execution parameters comprise a location of the input data set;

obtaining an identifier specific to the first data analytics workload according to one or more predefined rules, wherein obtaining the identifier comprises creating at least one content address according to and uniquely identifying both of the input data set and the one or more execution parameters to generate the identifier;

commencing execution of the first data analytics workload based on the one or more execution parameters and the input data set, wherein the first data analytics workload writes output data generated during the execution of the first data analytics workload;

registering meta data associated with the output data generated during the execution of the first data analytics workload at a meta data store using said at least one identifier, wherein the meta data comprises lineage information for tracing and verifying at least one of creation, movement, use and alteration of the output data generated during execution of the first data analytics workload based at least in part on said identifier;

storing the meta data associated with the output data temporarily in a temporary store prior to completing execution of the first data analytics workload, wherein the lineage information is stored with both parent meta data associated with a parent file and child meta data associated with one or more child data files associated with the parent file;

merging the temporarily stored meta data from the temporary store into the meta data store upon completing execution of the first data analytics workload; and

performing one or more functions utilizing the lineage information from the merged meta data, the one or more functions comprising one or more of a data provenance function, a data analytics work scheduling function, and a data de-duplication function;

wherein performing the data provenance function comprises replaying the first data analytics workload by tracking a footprint derived from the lineage information to validate the output data generated during the execution of the first data analytics workload;

wherein performing the data analytics workload scheduling function comprises querying the meta data server to locate the lineage information responsive to re-obtaining the first data analytics workload, and returning data from an enterprise data store without re-executing the first data analytics workload responsive to locating the lineage information;

wherein performing the data de-duplication function comprises, responsive to obtaining the first data analytics workload and a second data analytics workload, comparing the first identifier with a second identifier specific to the second data analytics workload and, responsive to the first identifier matching the second identifier, registering the first and second data analytics workloads to point to common data; and

wherein one or more of the above steps are performed via at least one processing device.

2. The method of claim 1 , wherein the at least one processing device is part of a distributed computing platform.

3. The method of claim 2 , wherein the distributed computing platform is a massively distributed computing platform.

4. An article of manufacture comprising a processor-readable storage medium having encoded therein executable code of one or more software programs, wherein the one or more software programs when executed by the at least one processing device implement the steps of the method of claim 1 .

5. The method of claim 1 , wherein the lineage information is stored in a data analytics result hierarchy, wherein the data analytics hierarchy comprises first meta data for a first input data file and at least second meta data extending from the first meta data for at least a second input data file, the second input data file being part of the input data set of the data analytics workload.

6. The method of claim 5 , further comprising updating the output data generated during execution of the data analytics workload responsive to identifying a change to the first input data file.

7. An apparatus, comprising:

a memory; and

a processor operatively coupled to the memory and configured to:

obtain an input data set and a first data analytics workload, wherein the first data analytics workload comprises one or more execution parameters for executing the first data analytics workload based on the input data set, and wherein the execution parameters comprise a location of the input data set;

obtain an identifier specific to the first data analytics workload according to one or more predefined rules, wherein, in obtaining the identifier, the processor is configured to create at least one content address according to and uniquely identifying both of the input data set and the one or more execution parameters to generate the identifier;

commence execution of the first data analytics workload based on the one or more execution parameters and the input data set, wherein the data analytics workload writes output data generated during the execution of the first data analytics workload;

register meta data associated with the output data generated during the execution of the first data analytics workload at a meta data store using said at least one identifier, wherein the meta data comprises lineage information for tracing and verifying at least one of creation, movement, use and alteration of the output data generated during execution of the first data analytics workload based at least in part on said identifier;

store the meta data associated with the output data temporarily in a temporary store prior to completing execution of the first data analytics workload, wherein the lineage information is stored with both parent meta data associated with a parent file and child meta data associated with one or more child data files associated with the parent file;

merge the temporarily stored meta data from the temporary store into the meta data store upon completing execution of the data analytics workload; and

perform one or more functions utilizing the lineage information from the merged meta data, the one or more functions comprising one or more of a data provenance function, a data analytics work scheduling function, and a data de-duplication function;

wherein, in performing the data provenance function, the processor is configured to replay the first data analytics workload by tracking a footprint derived from the lineage information to validate the output data generated during the execution of the first data analytics workload;

wherein, in performing the data analytics workload scheduling function, the processor is configured to query the meta data server to locate the lineage information responsive to re-obtaining the first data analytics workload, and return data from an enterprise data store without re-executing the first data analytics workload responsive to locating the lineage information; and

wherein, in performing the data de-duplication function, the processor is configured to, responsive to obtaining the first data analytics workload and a second data analytics workload, compare the first identifier with a second identifier specific to the second data analytics workload and, responsive to the first identifier matching the second identifier, register the first and second data analytics workloads to point to common data.

8. The apparatus of claim 7 , wherein the processor is part of a distributed computing platform.

9. The apparatus of claim 8 , wherein the distributed computing platform is a massively distributed computing platform.

10. The apparatus of claim 7 , wherein the lineage information is stored in a data analytics result hierarchy, wherein the data analytics hierarchy comprises first meta data for a first input data file and at least second meta data extending from the first meta data for at least a second input data file, the second input data file being part of the input data set of the data analytics workload.

11. The apparatus of claim 10 , further comprising updating the output data generated during execution of the data analytics workload responsive to identifying a change to the first input data file.

12. A system, comprising:

a meta data server; and

an execution master node operatively coupled to the meta data server and configured to:

obtain an input data set and a first data analytics workload, wherein the first data analytics workload comprises one or more execution parameters for executing the first data analytics workload based on the input data set, and wherein the execution parameters comprise a location of the input data set;

obtain an identifier specific to the first data analytics workload according to one or more predefined rules, wherein, in obtaining the identifier, the execution master node is configured to create at least one content address according to and uniquely identifying both of the input data set and the one or more execution parameters to generate the identifier;

commence execution of the first data analytics workload based on the one or more execution parameters and the input data set, wherein the data analytics workload writes output data generated during the execution of the first data analytics workload;

register meta data associated with the output data generated during the execution of the first data analytics workload at a meta data store using said at least one identifier, wherein the meta data comprises lineage information for tracing and verifying at least one of creation, movement, use and alteration of the output data generated during execution of the first data analytics workload based at least in part on said identifier;

store the meta data associated with the output data temporarily in a temporary store prior to completing execution of the first data analytics workload, wherein the lineage information is stored with both parent meta data associated with a parent file and child meta data associated with one or more child data files associated with the parent file;

merge the temporarily stored meta data from the temporary store into the meta data store upon completing execution of the data analytics workload; and

perform one or more functions utilizing the lineage information from the merged meta data, the one or more functions comprising one or more of a data provenance function, a data analytics work scheduling function, and a data de-duplication function;

wherein, in performing the data provenance function, the execution master node is configured to replay the first data analytics workload by tracking a footprint derived from the lineage information to validate the output data generated during the execution of the first data analytics workload;

wherein, in performing the data analytics workload scheduling function, the execution master node is configured to query the meta data server to locate the lineage information responsive to re-obtaining the first data analytics workload, and return data from an enterprise data store without re-executing the first data analytics workload responsive to locating the lineage information; and

wherein, in performing the data de-duplication function, the execution master node is configured to, responsive to obtaining the first data analytics workload and a second data analytics workload, compare the first identifier with a second identifier specific to the second data analytics workload and, responsive to the first identifier matching the second identifier, register the first and second data analytics workloads to point to common data.

13. The system of claim 12 , further comprising one or more slave nodes operatively coupled to the meta data server and the execution master node.

14. The system of claim 12 , wherein the execution master node is part of a distributed computing platform.

15. The system of claim 14 , wherein the distributed computing platform is a massively distributed computing platform.

16. The system of claim 12 , wherein the lineage information is stored in a data analytics result hierarchy, wherein the data analytics hierarchy comprises first meta data for a first input data file and at least second meta data extending from the first meta data for at least a second input data file, the second input data file being part of the input data set of the data analytics workload.

17. The system of claim 16 , further comprising updating the output data generated during execution of the data analytics workload responsive to identifying a change to the first input data file.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (045482/0131) Recorded May 20, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO WYSE TECHNOLOGY L.L.C.)
Reel/Frame 061749/0924 →
RELEASE OF SECURITY INTEREST AT REEL 045482 FRAME 0395 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
Reel/Frame 058298/0314 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
PATENT SECURITY AGREEMENT (CREDIT) Recorded Mar 1, 2018
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 045482/0395 →
PATENT SECURITY AGREEMENT (NOTES) Recorded Mar 1, 2018
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 045482/0131 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2017
From: EMC CORPORATION
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 043732/0276 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2014
From: XIANG, DONG; TODD, STEPHEN; CHEN, QIYAN
To: EMC CORPORATION
Reel/Frame 031931/0113 →