IP Library › Granted Patent US 9,665,608
Granted Patent B2
US 9,665,608 · App. 14/299,042 · Granted May 30, 2017

Parallelization of data processing

Inventors: Ning Duan (Beijing, CN); Wei Huang (Beijing, CN); Peng Ji (Beijing, CN); Yi Qi (Beijing, CN); Qi Zhang (Beijing, CN); Jun Zhu (Beijing, CN)
Assignee: International Business Machines Corporation
G06F17/30345
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,665,608
App. No.
14/299,042
Granted
May 30, 2017
Kind
B2
Abstract

A method and apparatus for parallelization of data processing. The method including: parsing a data processing flow to split a write table sequence for the data processing flow; generating a plurality of instances of the data processing flow based at least in part on the split write table sequence; and scheduling the plurality of instances for parallelization of data processing.

Claims (28)

1. A method for parallelization of data processing, the method comprising:

parsing a data processing flow to split a write table sequence for the data processing flow, wherein the write table sequence is split into a plurality of segments, and neighboring segments indicate different database tables;

generating a plurality of instances of the data processing flow based at least in part on the split write table sequence; and

scheduling the plurality of instances for parallelization of data processing with pipeline technology.

2. The method according to claim 1 , wherein the write table sequence is split according to an assemble structure of the write table sequence.

3. The method according to claim 1 , wherein the plurality of instances perform write operations on different database tables at the same time.

4. The method according to claim 1 , wherein the data processing comprises data extraction, transformation, and loading.

5. An apparatus for parallelization of data processing, the apparatus comprising:

a memory;

a processor device communicatively coupled to the memory; and

a module configured for parallelization of data processing coupled to the memory and the processor device to carry out the steps of a method comprising:

parsing a data processing flow to split a write table sequence for the data processing flow, wherein the write table sequence is split into a plurality of segments, and neighboring segments indicate different database tables;

generating a plurality of instances of the data processing flow based at least in part on the split write table sequence; and

scheduling the plurality of instances for parallelization of data processing with pipeline technology.

6. The apparatus according to claim 5 , wherein the write table sequence is split according to an assemble structure of the write table sequence.

7. The apparatus according to claim 5 , wherein the plurality of instances perform write operations on different database tables at the same time.

8. The apparatus according to claim 5 , wherein the data processing comprises data extraction, transformation, and loading.

9. The apparatus according to claim 5 , wherein the data processing flow comprises any one of a plurality of data processing subtasks executed in parallel.

10. The apparatus according to claim 9 , further comprising:

scanning database partitions; and

dispatching the plurality of data processing subtasks of a data processing task to the database partitions based at least in part on the scanning result.

11. The apparatus according to claim 10 , wherein said scanning the database partitions comprises:

scanning a database partition key table to obtain database partition keys; and

mapping the database partitions and the database partition keys to learn a number of the database partitions.

12. The apparatus according to claim 11 , wherein said dispatching the plurality of data processing subtasks to the database partitions comprises:

parallelizing the data processing task into the plurality of data processing subtasks based at least in part on the number of the database partitions;

dispatching the plurality of data processing subtasks to corresponding database partitions; and

executing the plurality of data processing subtasks in parallel.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2014
From: DUAN, NING; HUANG, WEI; JI, PENG; QI, YI; ZHANG, QI; ZHU, JUN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 033054/0568 →
Priority Claims (1)
CN 2013 1 0261903 · Jun 27, 2013 · national
Continuity (1)
Related Publication 20150006468A1 · Jan 1, 2015