IP Library Granted Patent US 10,915,544
Granted Patent B2
US 10,915,544 · App. 15/227,265 · Granted Feb 9, 2021

Transforming and loading data utilizing in-memory processing

Inventors: Lawrence A. Greene (Plainville, MA); Yong Li (Newton, MA); Xiaoyan Pu (Chelmsford, MA); Yeh-Heng Sheng (Cupertino, CA)
Assignee: International Business Machines Corporation
G06F16/254
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,915,544
App. No.
15/227,265
Granted
Feb 9, 2021
Kind
B2
Abstract

A system includes at least one processor and processes an ETL job. The system analyzes a specification of the ETL job including one or more functional expressions to load data from one or more source data stores, process the data in memory, and store the processed data to one or more target data stores. One or more data flows are produced from the specification based on the one or more functional expressions. The one or more data flows utilize in-memory distributed data sets generated to accommodate parallel processing for loading and processing the data. The one or more data flows are optimized to assign operations to be performed on the one or more source data stores. The optimized data flows are executed to load the data to the one or more target data stores in accordance with the specification. Present invention embodiments further include methods and computer program products.

Claims (20)

1. A method of processing an Extract, Transform, Load (ETL) job comprising:

analyzing a specification of the ETL job including one or more functional expressions to load data from one or more source data stores, process the data in memory, and store the processed data to one or more target data stores;

producing one or more data flows from the specification based on the one or more functional expressions, wherein the one or more data flows utilize in-memory distributed data sets generated to accommodate parallel processing for loading and processing the data, wherein producing the one or more data flows comprises transforming the ETL job into an in-memory computational model comprising a plurality of query language statements by:

converting each source table identified in the ETL job into a read function executable on the in-memory distributed data sets,

converting each target table identified in the ETL job into a write function executable on the in-memory distributed data sets, and

converting each shaping or transformation operation identified in the ETL job into an operational statement executable on the in-memory distributed data sets;

optimizing the one or more data flows to assign operations to be performed on the one or more source data stores, wherein optimizing the one or more data flows comprises consolidating two or more query language statements of the plurality of query language statements; and

transmitting the in-memory computational model to a cluster comprising a plurality of nodes to execute, in parallel by the plurality of nodes, the optimized data flows to load the data to the one or more target data stores in accordance with the specification.

2. The method of claim 1 , further comprising:

storing results of one or more designated operations on an in-memory distributed data set of a data flow.

3. The method of claim 2 , further comprising:

re-starting the ETL job from a previously executed designated operation based on corresponding stored results.

4. The method of claim 2 , further comprising:

re-using the stored results of a designated operation in response to a subsequent execution of that operation.

5. The method of claim 1 , further comprising:

maintaining a filtered status of data within an in-memory distributed data set to accommodate filtering conditions.

6. The method of claim 1 , wherein executing the optimized data flows comprises:

generating query language constructs for functions of the optimized data flows;

generating a graph of objects of the in-memory distributed data sets corresponding to the query language constructs; and

transforming the optimized data flows to the query language based on the generated graph.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2016
From: GREENE, LAWRENCE A.; LI, YONG; PU, XIAOYAN; SHENG, YEH-HENG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 039332/0144 →
Continuity (2)
Continuation 14851061 · Sep 11, 2015
Related Publication 20170075966A1 · Mar 16, 2017
Cited By (2)
US 12,405,963 US 12,443,373