System and method to dynamically add nodes to a container management system cluster for AI workloads
A computer-implemented method for labeling and managing cloud computing resources includes receiving one or more computing jobs in a job queue and obtaining resource requirements for a first one of the one or more computing jobs. Nodes are placed into a cluster for the resource requirements from one or more cloud providers and the nodes are labelled to correspond to the first one of the one or more computing jobs. The first one of the one or more computing jobs from the job queue and is executed after the labelled aggregated resources are ready.
1 . A computer-implemented method for labeling and managing cloud computing resources, the computer-implemented method comprising:
receiving a plurality of computing jobs in a job queue;
obtaining resource requirements for a first computing job of the plurality of computing jobs;
placing nodes from one or more cloud providers into a cluster for the resource requirements;
labeling the placed nodes to correspond to the first computing job;
removing, based on the labeling of the placed nodes, the first computing job from the job queue;
aggregating the labeled nodes for the first computing job;
initiating the first computing job after the labeled nodes are aggregated for the first computing job and after the removing of the first computing job from the job queue;
monitoring, with a cluster monitor, a computing job status of the initiated first computing job; and
after the cluster monitor provides an indication that the first computing job is complete, determining whether one or more nodes of the placed nodes used by the first computing job are usable in a second computing job of remaining computing jobs of the plurality of computing jobs, wherein the remaining computing jobs are remaining in the job queue after the removing of the first computing job from the job queue;
relabeling the one or more labeled nodes to correspond to the second computing job that requires resources used by the first computing job.
2 . The computer-implemented method of claim 1 , wherein
the one or more cloud providers are multiple cloud providers, and
the nodes are obtained from the multiple cloud providers.
3 . The computer-implemented method of claim 1 , further comprising requesting heterogenous resources from the one or more cloud providers.
4 . The computer-implemented method of claim 1 , further comprising relabeling the one or more nodes of the placed nodes to correspond to the second computing job when the relabeled one or more nodes of the placed nodes are usable in the second computing job.
5 . The computer-implemented method of claim 1 , further comprising deleting nodes from the placed nodes of the cluster, that are not required for the plurality of computing jobs in the job queue.
6 . The computer-implemented method of claim 1 , further comprising relabeling multiple nodes of the placed nodes to correspond to multiple computing jobs of the plurality of computing jobs when the relabeled multiple nodes are usable in the multiple computing jobs.
7 . The computer-implemented method of claim 1 , further comprising:
adding additional nodes to the cluster to permit the second computing job to have all resources required for execution; and
labelling the additional nodes to correspond to the second computing job.
8 . A computer-implemented method for labeling and managing cloud computing resources, the computer-implemented method comprising:
receiving multiple computing jobs in a job queue;
obtaining resource requirements for each computing job of the multiple computing jobs;
placing nodes from one or more cloud providers into a cluster for the resource requirements for a first computing job of the multiple computing jobs;
labeling the placed nodes to correspond to the first computing job;
removing, based on the labeling of the placed nodes, the first computing job from the job queue;
aggregating the labeled nodes for the first computing job;
initiating the first computing job after the labeled nodes are aggregated for the first computing job and after the removing of the first computing job from the job queue;
monitoring a computing job status of the initiated first computing job;
determining, based on the monitoring of the computing job status, that the first computing job is completed;
determining, based on the determining that the first computing job is completed, that one or more labeled nodes among the labeled nodes used by the first computing job are usable in a second computing job of remaining computing jobs of the multiple computing jobs, wherein the remaining computing jobs are remaining in the job queue after the removing of the first computing job from the job queue; and
relabeling the one or more labeled nodes to correspond to the second computing job that requires resources used by the first computing job.
9 . The computer-implemented method of claim 8 , further comprising deleting, from the cluster, a portion of the placed nodes that are not required for the remaining computing jobs.
10 . The computer-implemented method of claim 8 , wherein
the one or more cloud providers are multiple cloud providers, and
the nodes are obtained from the multiple cloud providers.
11 . The computer-implemented method of claim 8 , further comprising requesting heterogenous resources from the one or more cloud providers.
12 . The computer-implemented method of claim 8 , further comprising:
adding additional nodes to the cluster to permit one or more computing jobs of the remaining computing jobs to have all resources required for execution; and
labelling the additional nodes to correspond to each computing job of the one or more computing jobs of the remaining computing jobs.
13 . A non-transitory computer readable storage medium tangibly embodying a computer readable program code having computer readable instructions that, when executed, causes a computer device to carry out a method of computing resource labeling and management, the method comprising:
receiving a plurality of computing jobs in a job queue;
obtaining resource requirements for a first computing job of the plurality of computing jobs;
placing nodes from one or more cloud providers into a cluster for the resource requirements;
labeling the placed nodes to correspond to the first computing job;
removing, based on the labeling of the placed nodes, the first computing job from the job queue;
aggregating the labeled nodes for the first computing job;
initiating the first computing job after the labeled nodes are aggregated for the first computing job and after the removing of the first computing job from the job queue;
monitoring, with a cluster monitor, a computing job status of the initiated first computing job; and
after the cluster monitor provides an indication that the first computing job is complete, determining whether one or more nodes of the placed nodes used by the first computing job are usable in a second computing job of remaining computing jobs of the plurality of computing jobs, wherein
the remaining computing jobs are remaining in the job queue after the removing of the first computing job from the job queue;
relabeling the one or more labeled nodes to correspond to the second computing job that requires resources used by the first computing job.
14 . The non-transitory computer readable storage medium of claim 13 , the method further comprising obtaining the nodes from the one or more cloud providers.
15 . The non-transitory computer readable storage medium of claim 13 , the method further comprising requesting heterogenous resources from the one or more cloud providers.
16 . The non-transitory computer readable storage medium of claim 13 , the method further comprising: monitoring, with a cluster monitor, a computing job status of the initiated first computing job; determining, after the cluster monitor provides an indication that the first computing job is complete, whether one or more nodes of the placed nodes used by the first computing job are usable in a second computing job of remaining computing jobs of the plurality of computing jobs, wherein the remaining computing jobs are remaining in the job queue after the removing of the first computing job from the job queue; and relabeling the one or more nodes of the placed nodes to correspond to the second computing job when the relabeled one or more nodes of the placed nodes are usable in the second computing job.
17 . The non-transitory computer readable storage medium of claim 13 , the method further comprising deleting nodes from the placed nodes of the cluster, that are not required for the remaining computing jobs in the job queue.
18 . The non-transitory computer readable storage medium of claim 13 , the method further comprising:
adding additional nodes to the cluster to permit the second computing job of the plurality of computing jobs to have all resources required for execution; and
labelling the additional nodes to correspond to the second computing job.