Systems and methods for spark lineage data capture
Systems and methods for SPARK lineage data capture are disclosed. In one embodiment, in an information processing apparatus comprising at least one computer processor, a method for lineage data capture may include: (1) receiving, at a lineage engine and from a listener service, a decisive logical plan for a job; (2) extracting, using a plan parser, lineage data from the decisive logical plan; (3) producing, by a job lineage builder, job lineage data and job attribute data from the lineage data; (4) extracting, by the job lineage builder and from the job lineage data and the job attribute data, attribute information, transformation information, and estimate information for the job; and (5) storing, in a database, the attribute information, the transformation information, and the estimate information.
1. A method for lineage data capture, comprising:
receiving, at a lineage engine and from a listener service, a parsed or indecisive logical plan for a job;
converting, by a query manager, the parsed or indecisive logical plan to a decisive logical plan;
extracting, using a plan parser, lineage data from the decisive logical plan;
receiving, from a metadata repository, existing metadata associated with the job;
integrating the existing metadata into the extracted lineage data to generate supplemented lineage data;
producing, by a job lineage builder, job lineage data and job attribute data from the supplemented lineage data;
extracting, by the job lineage builder and from the job lineage data and the job attribute data, attribute information, transformation information, and estimate information for the job;
storing, in a database, the attribute information, the transformation information, and the estimate information;
determining, based on the attribute information and using an attribute traversing engine, at least one other job associated with the job;
stitching the at least one other job to the job; and
presenting the stitched at least one other job, the attribute information, the transformation information, and the estimate information on a graphical user interface (GUI).
2. The method of claim 1 , further comprising:
receiving, at the GUI, a job query for data lineage for the job;
identifying, by a lineage relationship engine, a base job for the job and identifying at least one dependency for the base job;
executing, by the lineage relationship engine, a recursive search on the database until an origin and a destination for the base job are identified; and
outputting, at the GUI, the origin and the destination for the base job.
3. The method of claim 2 , wherein the GUI comprises a web service, a command line interface, or a database interface.
4. The method of claim 1 , wherein the associated attributes comprise one or more of an attribute name, an attribute type, an attribute classification, and an attribute complexity.
5. The method of claim 1 , wherein the decisive logical plan comprises a plurality of stages.
6. The method of claim 5 , wherein each stage comprises a direct acyclic graph.
7. A system for lineage data capture, comprising:
a job lineage builder executed by a computer processor;
a graphical user interface (GUI);
a lineage engine executed by a computer processor and comprising a plan parser; and
an attribute database;
wherein:
the lineage engine is configured to receive a decisive logical plan for a job and from a listener service, the decisive logical plan converted from a parsed or indecisive plan by a query manager;
the plan parser is configured to extract lineage data from the decisive logical plan, to receive existing metadata associated with the job from a metadata repository, and to integrate the existing metadata into the extracted lineage data to generate supplemented lineage data;
the job lineage builder is configured to produce job lineage data and job attribute data from the supplemented lineage data;
the job lineage builder is configured to extract attribute information, transformation information, and estimate information for the job from the job lineage data and the job attribute data;
the job lineage builder is configured to store the attribute information, the transformation information, and the estimate information in the attribute database
the job lineage builder is configured to determine, based on the attribute information and using an attribute traversing engine, at least one other job associated with the job;
the job lineage builder is configured to stitch the at least one other job to the job; and
the job lineage builder is configured to present the stitched at least one other job, the attribute information, the transformation information, and the estimate information on the GUI.
8. The system of claim 7 , further comprising:
a lineage relationship engine;
wherein:
the GUI is configured to receive a job query for data lineage for the job;
the lineage relationship engine is configured to identify a base job for the job and at least one dependency for the base job;
the lineage relationship engine is configured to execute a recursive search on the attribute database until an origin and a destination for the base job are identified; and
the GUI is configured to output the origin and the destination for the base job.
9. The system of claim 8 , wherein the GUI comprises a web service, a command line interface, or a database interface.
10. The system of claim 7 , wherein the associated attributes comprise one or more of an attribute name, an attribute type, an attribute classification, and an attribute complexity.
11. The system of claim 7 , wherein the decisive logical plan comprises a plurality of stages.
12. The system of claim 11 , wherein each stage comprises a direct acyclic graph.
13. A non-transitory computer readable medium having stored thereon software instructions that, when executed by a processor, cause the processor to perform the following:
receive a parsed or indecisive logical plan for a job and from a listener service;
convert, by a query manager, the parsed or indecisive logical plan to a decisive logical plan;
extract lineage data from the decisive logical plan;
receive existing metadata associated with the job from a metadata repository;
integrate the existing metadata into the extracted lineage data to generate supplemented lineage data;
produce job lineage data and job attribute data from the supplemented lineage data;
extract attribute information, transformation information, and estimate information for the job from the job lineage data and the job attribute data;
store the attribute information, the transformation information, and the estimate information in an attribute database;
determining, based on the attribute information and using an attribute traversing engine, at least one other job associated with the job;
stitching the at least one other job to the job; and
presenting the stitched at least one other job, the attribute information, the transformation information, and the estimate information on a graphical user interface (GUI).
14. The non-transitory computer readable medium of claim 13 , further comprising software instructions that, when executed by a processor, cause the processor to:
receive a job query for data lineage for the job from the GUI;
identify a base job for the job and at least one dependency for the base job;
execute a recursive search on the attribute database until an origin and a destination for the base job are identified; and
output the origin and the destination for the base job to the GUI.
15. The non-transitory computer readable medium of claim 14 , wherein the GUI comprises a web service, a command line interface, or a database interface.
16. The non-transitory computer readable medium of claim 14 , further comprising software instructions that, when executed by a processor, cause the processor to output the attributes for the base job.