IP Library › Granted Patent US 12,554,553
Granted Patent B2
US 12,554,553 · App. 17/498,011 · Granted Feb 17, 2026

Dynamic scaling for workload execution

Inventors: Jin Chi JC He (Xian, CN); Guang Han Sui (Beijing, CN); Peng Li (Xian, CN); Gang Pu (Xian, CN); Gang Wang (Xi'an, CN); Liang Wang (Xian, CN)
Assignee: International Business Machines Corporation
G06F9/5077G06F9/5016G06F9/5044G06F9/5083G06F9/546
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,553
App. No.
17/498,011
Granted
Feb 17, 2026
Kind
B2
Abstract

Aspects of the invention include receiving, by a controller, a workload comprising one or more tasks, generating a first pod comprising a first sidecar container, generating one or more ephemeral containers for the first pod based on the workload and one or more resource allocation metrics for the pod, executing the one or more tasks in the one or more ephemeral containers, monitoring the one or more resource allocation metrics for the pod, and generating at least one new ephemeral container in the first pod based on the one or more resource allocation metrics for the pod and the workload.

Claims (58)

1 . A computer-implemented method comprising:

receiving, by a controller, a plurality of workloads each comprising one or more tasks, wherein the one or more tasks of the plurality of workloads are stored in a workload queue;

generating a first pod comprising a main container that is configured as a first sidecar container, wherein the first sidecar container executes a loop that is configured to keep the first pod from exiting;

generating within the first pod, by a pod manager, one or more initial ephemeral containers for the first pod;

executing the one or more tasks in the one or more ephemeral containers, wherein executing the one or more tasks comprises:

executing the one or more tasks in the one or more initial ephemeral containers;

monitoring one or more resource allocation metrics for the pod based on the execution of the initial ephemeral containers; and

based on the workload queue and the monitoring, scaling the first pod using a pod manager and a horizontal pod autoscaler (HPA) to execute the one or more tasks in the workload queue, wherein the scaling the first pod includes the steps:

dynamically generating within the first pod, by a pod manager, one or more additional ephemeral containers for the first pod based on the workloads and one or more resource allocation metrics for the pod, wherein all ephemeral containers are created with resource limits to avoid exceeding pod resource limits;

scaling the first pod using the HPA based on the workload queue and resource metrics, wherein scaling the first pod creates additional pods having one or more ephemeral containers in a Kubernetes cluster when the workload queue exceeds a predefined threshold; and

executing the one or more tasks in the plurality ephemeral containers.

2 . The computer-implemented method of claim 1 , further comprising:

terminating at least one ephemeral container in the one or more ephemeral containers in the first pod based on the one or more resource allocations metrics.

3 . The computer-implemented method of claim 1 , further comprising:

determining a maximum ephemeral containers for the first pod based on the one or more resource allocation metrics;

generating a second pod comprising a second one or more ephemeral containers based on the workload requiring a number of ephemeral containers exceeding the maximum ephemeral containers for the first pod.

4 . The computer-implemented method of claim 1 , wherein the queue comprises a message queuing telemetry transport queue.

5 . The computer-implemented method of claim 1 , wherein the first pod comprises a Kubernetes pod.

6 . The computer-implemented method of claim 1 , wherein the first sidecar container transmits results of the workload execution by executing a lightweight loop.

7 . The computer-implemented method of claim 1 , wherein dynamically generating the one or more ephemeral containers comprises allocating resources to the ephemeral containers based on a predefined threshold for resource utilization metrics.

8 . A system comprising:

a memory having computer readable instructions; and

one or more processors for executing the computer readable instructions, the computer readable instructions controlling the one or more processors to perform operations comprising:

receiving, a plurality of workloads each comprising one or more tasks, wherein the one or more tasks of the plurality of workloads are stored in a workload queue;

generating a first pod comprising a main container that is configured as a first sidecar container, wherein the first sidecar container executes a loop that is configured to keep the first pod from exiting;

generating within the first pod, by a pod manager, one or more initial ephemeral containers for the first pod;

executing the one or more tasks in the one or more ephemeral containers, wherein executing the one or more tasks comprises:

executing the one or more tasks in the one or more initial ephemeral containers;

monitoring one or more resource allocation metrics for the pod based on the execution of the initial ephemeral containers; and

based on the workload queue and the monitoring, scaling the first pod using a pod manager and a horizontal pod autoscaler (HPA) to execute the one or more tasks in the workload queue, wherein the scaling the first pod includes the steps:

dynamically generating within the first pod, by a pod manager, one or more additional ephemeral containers for the first pod based on the workloads and one or more resource allocation metrics for the pod, wherein all ephemeral containers are created with resource limits to avoid exceeding pod resource limits;

scaling the first pod using the HPA based on the workload queue and resource metrics, wherein scaling the first pod creates additional pods having one or more ephemeral containers in a Kubernetes cluster when the workload queue exceeds a predefined threshold; and

executing the one or more tasks in the plurality ephemeral containers.

9 . The system of claim 8 , wherein the operations further comprise:

terminating at least one ephemeral container in the one or more ephemeral containers in the first pod based on the one or more resource allocations metrics.

10 . The system of claim 8 , wherein the operations further comprise:

determining a maximum ephemeral containers for the first pod based on the one or more resource allocation metrics;

generating a second pod comprising a second one or more ephemeral containers based on the workload requiring a number of ephemeral containers exceeding the maximum ephemeral containers for the first pod.

11 . The system of claim 8 , wherein the queue comprises a message queuing telemetry transport queue.

12 . The system of claim 8 , wherein the first pod comprises a Kubernetes pod.

13 . A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by one or more processors to cause the one or more processors to perform operations comprising:

receiving, by a controller, a plurality of workloads each comprising one or more tasks, wherein the one or more tasks of the plurality of workloads are stored in a workload queue;

generating a first pod comprising a main container that is configured as a first sidecar container, wherein the first sidecar container executes a loop that is configured to keep the first pod from exiting;

generating within the first pod, by a pod manager, one or more initial ephemeral containers for the first pod;

executing the one or more tasks in the one or more ephemeral containers, wherein executing the one or more tasks comprises:

executing the one or more tasks in the one or more initial ephemeral containers;

monitoring one or more resource allocation metrics for the pod based on the execution of the initial ephemeral containers; and

based on the workload queue and the monitoring, scaling the first pod using a pod manager and a horizontal pod autoscaler (HPA) to execute the one or more tasks in the workload queue, wherein the scaling the first pod includes the steps:

dynamically generating within the first pod, by a pod manager, one or more additional ephemeral containers for the first pod based on the workloads and one or more resource allocation metrics for the pod, wherein all ephemeral containers are created with resource limits to avoid exceeding pod resource limits;

scaling the first pod using the HPA based on the workload queue and resource metrics, wherein scaling the first pod creates additional pods having one or more ephemeral containers in a Kubernetes cluster when the workload queue exceeds a predefined threshold; and

executing the one or more tasks in the plurality ephemeral containers.

14 . The computer program product of claim 13 , further comprising:

terminating at least one ephemeral container in the one or more ephemeral containers in the first pod based on the one or more resource allocations metrics.

15 . The computer program product of claim 13 , further comprising:

determining a maximum ephemeral containers for the first pod based on the one or more resource allocation metrics;

generating a second pod comprising a second one or more ephemeral containers based on the workload requiring a number of ephemeral containers exceeding the maximum ephemeral containers for the first pod.

16 . The computer program product of claim 13 , wherein the queue comprises a message queuing telemetry transport queue.

17 . The computer program product of claim 13 , wherein the first pod comprises a Kubernetes pod.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2021
From: HE, JIN CHI JC; SUI, GUANG HAN; LI, PENG; PU, GANG; WANG, GANG; WANG, LIANG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057749/0052 →
Continuity (1)
Related Publication 20230114504A1 · Apr 13, 2023
References Cited (14)
US 8621052B2 · Keohane et al. · 2013 [cited by applicant]
US 10303492B1 · Wagner · 2019 [cited by examiner]
US 11481243B1 · Wang et al. · 2022 [cited by applicant]
US 20200225990A1 · Vaddi · 2020 [cited by examiner]
US 20210064433A1 · Nakfour · 2021 [cited by applicant]
US 20210294651A1 · Misca · 2021 [cited by examiner]
US 20230109368A1 · Ni · 2023 [cited by examiner]
US 20230325258A1 · Wang · 2023 [cited by examiner]
CN 102929725A · 2013 [cited by applicant]
CN 111880816A · 2020 [cited by applicant]
WO 2020151306A1 · 2020 [cited by applicant]
“Diego Najar, Keep a container running for troubleshooting on Kubernetes, https://varlogdiego.com/keep-a-container-running-on-kubernetes” (Year: 2019). [cited by examiner]
Anonymous, “Kubernetes Cluster Autoscaling Based on Cost Optimization, Including Licensing Costs,” IP.com Prior Art Database Technical Disclosure, IP.com No. IPCOM000263720D, Sep. 29, 2020, 4 pages. [cited by applicant]
Fu et al., “Fast and Efficient Container Startup at the Edge via Dependency Scheduling,” 3rd USENIX Workshop on Hot Topics in Edge Computing (HotEdge 20), 2020, 7 pages. [cited by applicant]