IP Library › Granted Patent US 10,884,793
Granted Patent B2
US 10,884,793 · App. 15/496,137 · Granted Jan 5, 2021

Parallelization of data processing

Inventors: Ning Duan (Beijing, CN); Wei Huang (Beijing, CN); Peng Ji (Beijing, CN); Yi Qi (Beijing, CN); Qi Zhang (Beijing, CN); Jun Zhu (Beijing, CN)
Assignee: International Business Machines Corporation
G06F9/4881G06F16/23G06F16/245G06F16/24554
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,884,793
App. No.
15/496,137
Granted
Jan 5, 2021
Kind
B2
Abstract

A method and apparatus for parallelization of data processing. The method including: parsing a data processing flow to split a write table sequence for the data processing flow; generating a plurality of instances of the data processing flow based at least in part on the split write table sequence; and scheduling the plurality of instances for parallelization of data processing.

Claims (49)

1. A method for parallelization of data processing, the method comprising:

parsing a data processing flow to split a write table sequence for the data processing flow, wherein the write table sequence is split into a plurality of segments, and each neighboring segment of the plurality of segments respectively indicate different database tables;

generating a plurality of instances of the data processing flow based on a frequency of a database table appearing in the write table sequence; and

scheduling the plurality of instances for parallelization of data processing with pipeline technology, wherein the plurality of instances perform write operations on a different database tables at the same time.

2. The method according to claim 1 , wherein the data processing comprises data extraction, transformation, and loading.

3. The method according to claim 1 , wherein the data processing flow comprises any one of a plurality of data processing subtasks executed in parallel.

4. The method according to claim 3 , further comprising:

scanning database partitions; and

dispatching the plurality of data processing subtasks of a data processing task to the database partitions based at least in part on the scanning result.

5. The method according to claim 4 , wherein said scanning the database partitions comprises:

scanning a database partition key table to obtain database partition keys; and

mapping the database partitions and the database partition keys to learn a number of the database partitions.

6. The method according to claim 5 , wherein said dispatching the plurality of data processing subtasks to the database partitions comprises:

parallelizing the data processing task into the plurality of data processing subtasks based at least in part on the number of the database partitions;

dispatching the plurality of data processing subtasks to corresponding database partitions; and

executing the plurality of data processing subtasks in parallel.

7. An apparatus for parallelization of data processing, the apparatus comprising:

a memory;

a processor device communicatively coupled to the memory; and

a module configured for parallelization of data processing coupled to the memory and the processor device to carry out the steps of a method comprising:

parsing a data processing flow to split a write table sequence for the data processing flow, wherein the write table sequence is split into a plurality of segments, and each neighboring segment of the plurality of segments respectively indicate different database tables;

generating a plurality of instances of the data processing flow based on a frequency of a database table appearing in the write table sequence; and

scheduling the plurality of instances for parallelization of data processing with pipeline technology, wherein the plurality of instances perform write operations on different database tables at the same time.

8. The apparatus according to claim 7 , wherein the data processing comprises data extraction, transformation, and loading.

9. The apparatus according to claim 8 , wherein the data processing flow comprises any one of a plurality of data processing subtasks executed in parallel.

10. The apparatus according to claim 9 , further comprising:

scanning database partitions; and

dispatching the plurality of data processing subtasks of a data processing task to the database partitions based at least in part on the scanning result.

11. The apparatus according to claim 10 , wherein said scanning the database partitions comprises:

scanning a database partition key table to obtain database partition keys; and

mapping the database partitions and the database partition keys to learn a number of the database partitions.

12. The apparatus according to claim 11 , wherein said dispatching the plurality of data processing subtasks to the database partitions comprises:

parallelizing the data processing task into the plurality of data processing subtasks based at least in part on the number of the database partitions;

dispatching the plurality of data processing subtasks to corresponding database partitions; and

executing the plurality of data processing subtasks in parallel.

13. A computer program product comprising:

a computer readable storage medium readable by one or more processing circuit and storing instructions for execution by one or more processor for performing a method comprising:

parsing a data processing flow to split a write table sequence for the data processing flow, wherein the write table sequence is split into a plurality of segments, and each neighboring segment of the plurality of segments respectively indicate different database tables:

generating a plurality of instances of the data processing flow based on a frequency of a database table appearing in the write table sequence; and

scheduling the plurality of instances for parallelization of data processing with pipeline technology, wherein the plurality of instances perform write operations on different database tables at the same time.

14. The method according to claim 1 , wherein the write table sequence is split according to an assemble structure of the write table sequence.

15. The method according to claim 1 , wherein the plurality of instances perform write operations on different database tables at the same time.

16. The method according to claim 1 , further comprising:

in response to a conflict in database access during a data extraction transformation and loading (ETL) processing, increasing efficiency of data processing by combining parallelization of data processing with a database partition feature, wherein combining parallelization of data processing comprises:

scanning database partitions; and

dispatching a plurality of data processing subtasks of a data processing task to database partitions based, at least in part, on the scanning result, wherein dispatching a plurality of data processing subtasks comprises:

parallelizing the data processing task into the plurality of data processing subtasks based, at least in part, on the number of the database partitions;

dispatching the plurality of data processing subtasks to corresponding database partitions; and

executing the plurality of data processing subtasks in parallel.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2017
From: DUAN, NING; HUANG, WEI; JI, PENG; QI, YI; ZHANG, QI; ZHU, JUN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 042134/0844 →
Priority Claims (1)
CN 2013 1 0261903 · Jun 27, 2013 · national
Continuity (2)
Continuation 14299042 · Jun 9, 2014
Related Publication 20170228255A1 · Aug 10, 2017