Parallelization of data processing
A method and apparatus for parallelization of data processing. The method including: parsing a data processing flow to split a write table sequence for the data processing flow; generating a plurality of instances of the data processing flow based at least in part on the split write table sequence; and scheduling the plurality of instances for parallelization of data processing.
1. A method for parallelization of data processing, the method comprising:
parsing a data processing flow to split a write table sequence for the data processing flow, wherein the write table sequence is split into a plurality of segments, and neighboring segments indicate different database tables;
generating a plurality of instances of the data processing flow based at least in part on the split write table sequence; and
scheduling the plurality of instances for parallelization of data processing with pipeline technology.
2. The method according to claim 1 , wherein the write table sequence is split according to an assemble structure of the write table sequence.
3. The method according to claim 1 , wherein the plurality of instances perform write operations on different database tables at the same time.
4. The method according to claim 1 , wherein the data processing comprises data extraction, transformation, and loading.
5. An apparatus for parallelization of data processing, the apparatus comprising:
a memory;
a processor device communicatively coupled to the memory; and
a module configured for parallelization of data processing coupled to the memory and the processor device to carry out the steps of a method comprising:
parsing a data processing flow to split a write table sequence for the data processing flow, wherein the write table sequence is split into a plurality of segments, and neighboring segments indicate different database tables;
generating a plurality of instances of the data processing flow based at least in part on the split write table sequence; and
scheduling the plurality of instances for parallelization of data processing with pipeline technology.
6. The apparatus according to claim 5 , wherein the write table sequence is split according to an assemble structure of the write table sequence.
7. The apparatus according to claim 5 , wherein the plurality of instances perform write operations on different database tables at the same time.
8. The apparatus according to claim 5 , wherein the data processing comprises data extraction, transformation, and loading.
9. The apparatus according to claim 5 , wherein the data processing flow comprises any one of a plurality of data processing subtasks executed in parallel.
10. The apparatus according to claim 9 , further comprising:
scanning database partitions; and
dispatching the plurality of data processing subtasks of a data processing task to the database partitions based at least in part on the scanning result.
11. The apparatus according to claim 10 , wherein said scanning the database partitions comprises:
scanning a database partition key table to obtain database partition keys; and
mapping the database partitions and the database partition keys to learn a number of the database partitions.
12. The apparatus according to claim 11 , wherein said dispatching the plurality of data processing subtasks to the database partitions comprises:
parallelizing the data processing task into the plurality of data processing subtasks based at least in part on the number of the database partitions;
dispatching the plurality of data processing subtasks to corresponding database partitions; and
executing the plurality of data processing subtasks in parallel.