IP Library › Granted Patent US 11,461,135
Granted Patent B2
US 11,461,135 · App. 16/663,428 · Granted Oct 4, 2022

Dynamically modifying the parallelism of a task in a pipeline

Inventors: Yannick Saillet (Stuttgart, DE); Namit Kabra (Hyderabad, IN); Ritesh Kumar Gupta (Hyderabad, IN)
Assignee: International Business Machines Corporation
G06F9/4887G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,461,135
App. No.
16/663,428
Granted
Oct 4, 2022
Kind
B2
Abstract

In an approach to dynamically identifying and modifying the parallelism of a particular task in a pipeline, the optimal execution time of each stage in a dynamic pipeline is calculated. The actual execution time of each stage in the dynamic pipeline is measured. Whether the actual time of completion of the data processing job will exceed a threshold is determined. If it is determined that the actual time of completion of the data processing job will exceed the threshold, then additional instances of the stages are created.

Claims (75)

1. A computer-implemented method for dynamically modifying the parallelism of a task in a pipeline, the computer-implemented method comprising:

calculating, by one or more computer processors, an optimal execution time of each stage of a plurality of stages of a dynamic pipeline for a data processing job;

using a machine learning decision tree model, determining, by one or more computer processors, an actual execution time of each stage of the plurality of stages of the dynamic pipeline for the data processing job;

determining, by one or more computer processors, whether an actual time of completion of the data processing job will exceed a threshold, based on the actual execution time of each stage of the plurality of stages of the dynamic pipeline for the data processing job;

responsive to determining that the actual time of completion of the data processing job will exceed the threshold, creating, by one or more computer processors, one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job, wherein the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job reduce the actual time of completion of the data processing job;

training, by one or more computer processors, the machine learning decision tree model to predict a behavior of the dynamic pipeline for the data processing job; and

using the machine learning decision tree model, predicting, by one or more computer processors, an optimum configuration of the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job.

2. The computer-implemented method of claim 1 , wherein determining, by one or more computer processors, whether the actual time of completion of the data processing job will exceed the threshold, based on the actual execution time of each stage of the plurality of stages of the dynamic pipeline for the data processing job, further comprises using a machine learning model to predict the actual time of completion of the data processing job.

3. The computer-implemented method of claim 1 further comprising:

determining, by one or more computer processors, one or more outlier data points, wherein the one or more outlier data points identify pipeline stages of the plurality of stages of the dynamic pipeline for the data processing job where the actual execution time differs from the optimal execution time; and

determining, by one or more computer processors, if a predicted time of completion of the data processing job will exceed the threshold, based on the one or more outlier data points.

4. The computer-implemented method of claim 1 , wherein creating, by one or more computer processors, the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job comprises:

determining, by one or more computer processors, whether a sufficient resources are available to spawn a first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

responsive to determining that the sufficient resources are available to spawn the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job, spawning, by one or more computer processors, the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job; and

partitioning, by one or more computer processors, a pipeline data to the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job.

5. The computer-implemented method of claim 4 , wherein creating, by one or more computer processors, the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job further comprises:

measuring, by one or more computer processors, a throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

determining, by one or more computer processors, whether the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has increased; and

responsive to determining that the throughput of the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has increased, spawning, by one or more computer processors, a second one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job.

6. The computer-implemented method of claim 4 , wherein creating, by one or more computer processors, the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job further comprises:

measuring, by one or more computer processors, a throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

determining, by one or more computer processors, whether the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has increased;

responsive to determining that the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has not increased, removing, by one or more computer processors, the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job; and

storing, by one or more computer processors, a current configuration of the dynamic pipeline for the data processing job.

7. A computer program product for dynamically modifying the parallelism of a task in a pipeline, the computer program product comprising:

one or more computer-readable storage devices and program instructions stored on the one or more computer readable storage devices, the stored program instructions comprising:

program instructions to calculate an optimal execution time of each stage of a plurality of stages of a dynamic pipeline for a data processing job;

program instructions to determine, using a machine learning decision tree model, an actual execution time of each stage of the plurality of stages of the dynamic pipeline for the data processing job;

program instructions to determine whether the actual time of completion of the data processing job will exceed a threshold, based on the actual execution time of each stage of the plurality of stages of the dynamic pipeline for the data processing job;

responsive to determining that the actual time of completion of the data processing job will exceed the threshold, program instructions to create one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job, wherein the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job reduce the actual time of completion of the data processing job;

program instructions to train, by one or more computer processors, the machine learning decision tree model to predict a behavior of the dynamic pipeline for the data processing job; and

using the machine learning decision tree model, program instructions to predict, by one or more computer processors, an optimum configuration of the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job.

8. The computer program product of claim 7 , wherein program instructions to determine whether the actual time of completion of the data processing job will exceed the threshold, based on the actual execution time of each stage of the plurality of stages of the dynamic pipeline for the data processing job, further comprises using a machine learning model to predict the actual time of completion of the data processing job.

9. The computer program product of claim 7 further comprising:

program instructions to determine one or more outlier data points, wherein the one or more outlier data points identify pipeline stages of the plurality of stages of the dynamic pipeline for the data processing job where the actual execution time differs from the optimal execution time; and

program instructions to determine if a predicted time of completion of the data processing job will exceed the threshold, based on the one or more outlier data points.

10. The computer program product of claim 7 , wherein program instructions to create the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job comprises:

program instructions to determine whether a sufficient resources are available to spawn a first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

responsive to determining that the sufficient resources are available to spawn the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job, program instructions to spawn the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job; and

program instructions to partition a pipeline data to the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job.

11. The computer program product of claim 10 , wherein program instructions to create the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job further comprises:

program instructions to measure a throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

program instructions to determine whether the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has increased; and

responsive to determining that the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has increased, program instructions to spawn a second one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job.

12. The computer program product of claim 10 , wherein program instructions to create the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job further comprises:

program instructions to measure a throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

program instructions to determine whether the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

responsive to determining that the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has not increased, program instructions to remove the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job; and

program instructions to store a current configuration of the dynamic pipeline for the data processing job.

13. A computer system for dynamically modifying the parallelism of a task in a pipeline, the computer program product comprising:

one or more computer processors;

one or more computer-readable storage media; and

program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the stored program instructions comprising:

program instructions to calculate an optimal execution time of each stage of a plurality of stages of a dynamic pipeline for a data processing job;

using a machine learning decision tree model, program instructions to determine an actual execution time of each stage of the plurality of stages of the dynamic pipeline for the data processing job;

program instructions to determine whether the actual time of completion of the data processing job will exceed a threshold, based on the actual execution time of each stage of the plurality of stages of the dynamic pipeline for the data processing job;

responsive to determining that the actual time of completion of the data processing job will exceed the threshold, program instructions to create one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job, wherein the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job reduce the actual time of completion of the data processing job;

program instructions to train, by one or more computer processors, the machine learning decision tree model to predict a behavior of the dynamic pipeline for the data processing job; and

using the machine learning decision tree model, program instructions to predict, by one or more computer processors, an optimum configuration of the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job.

14. The computer system of claim 13 , wherein program instructions to determine whether the actual time of completion of the data processing job will exceed the threshold, based on the actual execution time of each stage of the plurality of stages of the dynamic pipeline for the data processing job, further comprises using a machine learning model to predict the actual time of completion of the data processing job.

15. The computer system of claim 13 further comprising,

program instructions to determine one or more outlier data points, wherein the one or more outlier data points identify pipeline stages of the plurality of stages of the dynamic pipeline for the data processing job where the actual execution time differs from the optimal execution time; and

program instructions to determine if a predicted time of completion of the data processing job will exceed the threshold, based on the one or more outlier data points.

16. The computer system of claim 13 , wherein program instructions to create the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job comprises:

program instructions to determine whether a sufficient resources are available to spawn a first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

responsive to determining that the sufficient resources are available to spawn the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job, program instructions to spawn the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job; and

program instructions to partition a pipeline data to the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job.

17. The computer system of claim 16 , wherein program instructions to create the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job further comprises:

program instructions to measure a throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

program instructions to determine whether the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has increased; and

responsive to determining that the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has increased, program instructions to spawn a second one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job.

18. The computer system of claim 16 , wherein program instructions to create the one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job further comprises:

program instructions to measure a throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job;

program instructions to determine whether the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job; responsive to determining that the throughput of the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job has not increased, program instructions to remove the first one or more additional instances of the one or more of the plurality of stages of the dynamic pipeline for the data processing job; and

program instructions to store a current configuration of the dynamic pipeline for the data processing job.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2019
From: SAILLET, YANNICK; KABRA, NAMIT; GUPTA, RITESH KUMAR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 050822/0923 →
Continuity (1)
Related Publication 20210124611A1 · Apr 29, 2021