IP Library Granted Patent US 12,664,032
Granted Patent B2
US 12,664,032 · App. 18/329,556 · Granted Jun 23, 2026

System and method to dynamically add nodes to a container management system cluster for AI workloads

Inventors: Abhishek Malvankar (White Plains, NY); Alaa S. Youssef (Valhalla, NY); Diana Jeanne Arroyo (Austin, TX)
Assignee: International Business Machines Corporation
G06F9/54G06F9/4881
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,032
App. No.
18/329,556
Filed
Jun 5, 2023
Granted
Jun 23, 2026
Kind
B2
Art Unit
2194
USPC
718/104
Abstract

A computer-implemented method for labeling and managing cloud computing resources includes receiving one or more computing jobs in a job queue and obtaining resource requirements for a first one of the one or more computing jobs. Nodes are placed into a cluster for the resource requirements from one or more cloud providers and the nodes are labelled to correspond to the first one of the one or more computing jobs. The first one of the one or more computing jobs from the job queue and is executed after the labelled aggregated resources are ready.

Claims (60)

1 . A computer-implemented method for labeling and managing cloud computing resources, the computer-implemented method comprising:

receiving a plurality of computing jobs in a job queue;

obtaining resource requirements for a first computing job of the plurality of computing jobs;

placing nodes from one or more cloud providers into a cluster for the resource requirements;

labeling the placed nodes to correspond to the first computing job;

removing, based on the labeling of the placed nodes, the first computing job from the job queue;

aggregating the labeled nodes for the first computing job;

initiating the first computing job after the labeled nodes are aggregated for the first computing job and after the removing of the first computing job from the job queue;

monitoring, with a cluster monitor, a computing job status of the initiated first computing job; and

after the cluster monitor provides an indication that the first computing job is complete, determining whether one or more nodes of the placed nodes used by the first computing job are usable in a second computing job of remaining computing jobs of the plurality of computing jobs, wherein the remaining computing jobs are remaining in the job queue after the removing of the first computing job from the job queue;

relabeling the one or more labeled nodes to correspond to the second computing job that requires resources used by the first computing job.

2 . The computer-implemented method of claim 1 , wherein

the one or more cloud providers are multiple cloud providers, and

the nodes are obtained from the multiple cloud providers.

3 . The computer-implemented method of claim 1 , further comprising requesting heterogenous resources from the one or more cloud providers.

4 . The computer-implemented method of claim 1 , further comprising relabeling the one or more nodes of the placed nodes to correspond to the second computing job when the relabeled one or more nodes of the placed nodes are usable in the second computing job.

5 . The computer-implemented method of claim 1 , further comprising deleting nodes from the placed nodes of the cluster, that are not required for the plurality of computing jobs in the job queue.

6 . The computer-implemented method of claim 1 , further comprising relabeling multiple nodes of the placed nodes to correspond to multiple computing jobs of the plurality of computing jobs when the relabeled multiple nodes are usable in the multiple computing jobs.

7 . The computer-implemented method of claim 1 , further comprising:

adding additional nodes to the cluster to permit the second computing job to have all resources required for execution; and

labelling the additional nodes to correspond to the second computing job.

8 . A computer-implemented method for labeling and managing cloud computing resources, the computer-implemented method comprising:

receiving multiple computing jobs in a job queue;

obtaining resource requirements for each computing job of the multiple computing jobs;

placing nodes from one or more cloud providers into a cluster for the resource requirements for a first computing job of the multiple computing jobs;

labeling the placed nodes to correspond to the first computing job;

removing, based on the labeling of the placed nodes, the first computing job from the job queue;

aggregating the labeled nodes for the first computing job;

initiating the first computing job after the labeled nodes are aggregated for the first computing job and after the removing of the first computing job from the job queue;

monitoring a computing job status of the initiated first computing job;

determining, based on the monitoring of the computing job status, that the first computing job is completed;

determining, based on the determining that the first computing job is completed, that one or more labeled nodes among the labeled nodes used by the first computing job are usable in a second computing job of remaining computing jobs of the multiple computing jobs, wherein the remaining computing jobs are remaining in the job queue after the removing of the first computing job from the job queue; and

relabeling the one or more labeled nodes to correspond to the second computing job that requires resources used by the first computing job.

9 . The computer-implemented method of claim 8 , further comprising deleting, from the cluster, a portion of the placed nodes that are not required for the remaining computing jobs.

10 . The computer-implemented method of claim 8 , wherein

the one or more cloud providers are multiple cloud providers, and

the nodes are obtained from the multiple cloud providers.

11 . The computer-implemented method of claim 8 , further comprising requesting heterogenous resources from the one or more cloud providers.

12 . The computer-implemented method of claim 8 , further comprising:

adding additional nodes to the cluster to permit one or more computing jobs of the remaining computing jobs to have all resources required for execution; and

labelling the additional nodes to correspond to each computing job of the one or more computing jobs of the remaining computing jobs.

13 . A non-transitory computer readable storage medium tangibly embodying a computer readable program code having computer readable instructions that, when executed, causes a computer device to carry out a method of computing resource labeling and management, the method comprising:

receiving a plurality of computing jobs in a job queue;

obtaining resource requirements for a first computing job of the plurality of computing jobs;

placing nodes from one or more cloud providers into a cluster for the resource requirements;

labeling the placed nodes to correspond to the first computing job;

removing, based on the labeling of the placed nodes, the first computing job from the job queue;

aggregating the labeled nodes for the first computing job;

initiating the first computing job after the labeled nodes are aggregated for the first computing job and after the removing of the first computing job from the job queue;

monitoring, with a cluster monitor, a computing job status of the initiated first computing job; and

after the cluster monitor provides an indication that the first computing job is complete, determining whether one or more nodes of the placed nodes used by the first computing job are usable in a second computing job of remaining computing jobs of the plurality of computing jobs, wherein

the remaining computing jobs are remaining in the job queue after the removing of the first computing job from the job queue;

relabeling the one or more labeled nodes to correspond to the second computing job that requires resources used by the first computing job.

14 . The non-transitory computer readable storage medium of claim 13 , the method further comprising obtaining the nodes from the one or more cloud providers.

15 . The non-transitory computer readable storage medium of claim 13 , the method further comprising requesting heterogenous resources from the one or more cloud providers.

16 . The non-transitory computer readable storage medium of claim 13 , the method further comprising: monitoring, with a cluster monitor, a computing job status of the initiated first computing job; determining, after the cluster monitor provides an indication that the first computing job is complete, whether one or more nodes of the placed nodes used by the first computing job are usable in a second computing job of remaining computing jobs of the plurality of computing jobs, wherein the remaining computing jobs are remaining in the job queue after the removing of the first computing job from the job queue; and relabeling the one or more nodes of the placed nodes to correspond to the second computing job when the relabeled one or more nodes of the placed nodes are usable in the second computing job.

17 . The non-transitory computer readable storage medium of claim 13 , the method further comprising deleting nodes from the placed nodes of the cluster, that are not required for the remaining computing jobs in the job queue.

18 . The non-transitory computer readable storage medium of claim 13 , the method further comprising:

adding additional nodes to the cluster to permit the second computing job of the plurality of computing jobs to have all resources required for execution; and

labelling the additional nodes to correspond to the second computing job.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2023
From: MALVANKAR, ABHISHEK; YOUSSEF, ALAA S.; ARROYO, DIANA JEANNE
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063859/0737 →
Continuity (1)
Related Publication 20240403143A1 · Dec 5, 2024
References Cited (34)
US 7870234B2 · Arendt · 2011 [cited by applicant]
US 9280390B2 · Sirota · 2016 [cited by applicant]
US 9384571B1 · Covell · 2016 [cited by examiner]
US 10554574B2 · Bhandaru · 2020 [cited by examiner]
US 10721141B1 · Verma · 2020 [cited by examiner]
US 10997538B1 · Chandrachood · 2021 [cited by examiner]
US 20060123428A1 · Burns · 2006 [cited by examiner]
US 20100077300A1 · Dugan · 2010 [cited by examiner]
US 20120233623A1 · van Riel · 2012 [cited by examiner]
US 20180067919A1 · Ning · 2018 [cited by examiner]
US 20210004163A1 · Xu · 2021 [cited by examiner]
US 20210019179A1 · Yadav · 2021 [cited by examiner]
US 20210058934A1 · Jiang · 2021 [cited by examiner]
US 20210081217A1 · Xiao · 2021 [cited by examiner]
US 20210263667A1 · Whitlock · 2021 [cited by applicant]
US 20210271521A1 · Kang · 2021 [cited by applicant]
US 20220100573A1 · Allen · 2022 [cited by examiner]
US 20220113993A1 · Vogt · 2022 [cited by examiner]
US 20220138168A1 · Veselova · 2022 [cited by examiner]
US 20220329651A1 · Kim · 2022 [cited by applicant]
US 20230037783A1 · Huang · 2023 [cited by examiner]
US 20230273830A1 · Shi · 2023 [cited by examiner]
US 20230305905A1 · Chen · 2023 [cited by examiner]
US 20240069964A1 · Chatterjee · 2024 [cited by examiner]
US 20240069998A1 · Chatterjee · 2024 [cited by examiner]
US 20240403143A1 · Malvankar · 2024 [cited by examiner]
US 20250238261A1 · Tu · 2025 [cited by examiner]
CN 111124765A · 2020 [cited by examiner]
CN 106170782B · 2020 [cited by applicant]
Lei Li, Resource Allocation and Task Offloading for Heterogeneous Real-Time Tasks With Uncertain Duration Time in a Fog Queueing System. (Year: 2015). [cited by examiner]
Masoud Mansoury, FairMatch: A Graph-based Approach for Improving Aggregate Diversity in Recommender Systems. (Year: 2020). [cited by examiner]
Caballer, M. et al., “Deployment of Elastic Virtual Hybrid Clusters Across Cloud Sites”, Journal of Grid Computing 19, No. 1, 2021, 33 pages. [cited by applicant]
Karpenter with AWS Node Termination Handler, downloaded Mar. 22, 2023 from https://dev.to/aws-builders/karpenter-with-aws-node-termination-handler-149d, 6 pgs. [cited by applicant]
IBM Spectrum LSF Resource Connector Overview, downloaded Mar. 22, 2023 from ibm.com/docs/en/spectrum-lsf/10.1.0?topic=connector-lsf-resource-connector-overview, 4 pgs. [cited by applicant]