IP Library › Granted Patent US 12,632,314
Granted Patent B2
US 12,632,314 · App. 18/142,041 · Granted May 19, 2026

Elastic provisioning of container-based graphics processing unit (GPU) nodes

Inventors: Yisan Zhao (Beijing, CN); Xiaoyu Hu (Austin, TX); Robert Riemer (Siegburg, DE); Aidan Cully (Saint Augustine, FL)
Assignee: VMware LLC
G06F9/505G06F11/3442
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,314
App. No.
18/142,041
Granted
May 19, 2026
Kind
B2
Abstract

Example methods and systems for elastic provisioning of container-based graphics processing unit (GPU) nodes are described. In one example, a computer system may monitor usage information associated with a pool of multiple container-based GPU nodes. Based on the usage information, the computer system may apply rule(s) to determine whether capacity adjustment is required. In response to determination that capacity expansion is required, the computer system may configure the pool to expand by adding (a) at least one container-based GPU node to the pool, or (b) at least one container pod to one of the multiple container-based GPU nodes. Otherwise, in response to determination that capacity shrinkage is required, the computer system may configure the pool to shrink by removing (a) at least one container-based GPU node, or (b) at least one container pod from the pool.

Claims (47)

1 . A method for a computer system to perform elastic scaling of a graphics processing unit (GPU) node pool that comprises a plurality of container-based GPU nodes hosting container pods that are capable of accessing GPUs, wherein the method comprises:

obtaining usage information of the GPU node pool, the usage information indicating a number of the container pods that are each available to use the GPUs to execute computational tasks initiated outside the GPU node pool, the container pods that are available each being determined based on either a user not being logged in to the container pod or the container pod not currently processing computational tasks using the GPUs for acceleration;

based on the usage information, applying one or more rules to determine whether the number of available container pods satisfies a minimum or maximum number of available container pods;

in response to determination at a first time that the number of available container pods does not satisfy the minimum number of available container pods, adding (a) at least one container-based GPU node, or (b) at least one container pod to one of the container-based GPU nodes; and

in response to determination at a second time that the number of available container pods does not satisfy the maximum number of available container pods, removing (a) at least one of the multiple-container-based GPU nodes, or (b) at least one container pod from one of the container-based GPU nodes.

2 . The method of claim 1 , wherein the usage information is obtained from a plurality of GPU manager pods executing respectively in the container-based GPU nodes, the GPU manager pods each managing one or more GPU worker pods executing in the same one of the container-based GPU nodes as the GPU manager pod, and the usage information being based on availability of GPU worker pods.

3 . The method of claim 1 ,

wherein adding the at least one container-based GPU node or at least one container pod includes requesting a container orchestration layer to add the at least one container-based GPU node or at least one container pod, the container orchestration layer being configured to create and manage the container pods hosted by the container-based GPU nodes, and

wherein removing the at least one of the container-based GPU nodes or at least one container pod includes requesting the container orchestration layer to remove the at least one of the container-based GPU nodes or at least one container pod.

4 . The method of claim 1 , wherein

adding the at least one container-based GPU node or at least one container pod is triggered automatically by a container orchestration layer that is configured to create and manage the container pods hosted by the container-based GPU nodes, and

removing the at least one of the container-based GPU nodes or at least one container pod is triggered automatically by the container orchestration layer.

5 . The method of claim 1 , wherein

the usage information is obtained by a control plane entity of the computer system, and

the GPU node pool is deployed across multiple cloud providers.

6 . One or more non-transitory computer-readable storage media including instructions which, in response to execution by one or more processors of a computer system, cause the one or more processors to perform a method for elastic scaling of a graphics processing unit (GPU) node pool that comprises a plurality of container-based GPU nodes hosting container pods that are capable of accessing GPUs, wherein the method comprises:

obtaining usage information of the GPU node pool, the usage information indicating a number of the container pods that are each available to use the GPUs to execute computational tasks initiated outside the GPU node pool, the container pods that are available each being determined based on either a user not being logged in to the container pod or the container pod not currently processing computational tasks using the GPUs for acceleration;

based on the usage information, applying one or more rules to determine whether the number of available container pods satisfies a minimum or maximum number of available container pods;

in response to determination at a first time that the number of available container pods does not satisfy the minimum number of available container pods, adding (a) at least one container-based GPU node, or (b) at least one container pod to one of the container-based GPU nodes; and

in response to determination at a second time that the number of available container pods does not satisfy the maximum number of available container pods, removing (a) at least one of the container-based GPU nodes, or (b) at least one container pod from one of the container-based GPU nodes.

7 . The one or more non-transitory computer-readable storage media of claim 6 , wherein the usage information is obtained from a plurality of GPU manager pods executing respectively in the container-based GPU nodes, the GPU manager pods each managing of one or more GPU worker pods executing in the same one of the container-based GPU nodes as the GPU manager pod, and the usage information being based on availability of GPU worker pods.

8 . The one or more non-transitory computer-readable storage media of claim 6 ,

wherein adding the at least one container-based GPU node or at least one container pod includes requesting a container orchestration layer to add the at least one container-based GPU node or at least one container pod, the container orchestration layer being configured to create and manage the container pods hosted by the container-based GPU nodes, and

wherein removing the at least one of the container-based GPU nodes or at least one container pod includes requesting the container orchestration layer to remove the at least one of the container-based GPU nodes or at least one container pod.

9 . The one or more non-transitory computer-readable storage media of claim 6 , wherein

adding the at least one container-based GPU node or at least one container pod is triggered automatically by a container orchestration layer that is configured to create and manage the container pods hosted by the container-based GPU nodes, and

removing the at least one of the container-based GPU nodes or at least one container pod is triggered automatically by the container orchestration layer.

10 . The one or more non-transitory computer-readable storage media of claim 6 , wherein

the usage information is obtained by a control plane entity of the computer system, and

the GPU node pool is deployed across multiple cloud providers.

11 . A computer system, comprising:

one or more processors; and

one or more non-transitory computer-readable media having stored thereon instructions that, when executed by the one or more processors, cause the one or more processors to perform a method for elastic scaling of a graphics processing unit (GPU) node pool that comprises a plurality of container-based GPU nodes hosting container pods that are capable of accessing GPUs, wherein the method comprises:

obtaining usage information of the GPU node pool, the usage information indicating a number of the container pods that are each available to use the GPUs to execute computational tasks initiated outside the GPU node pool, the container pods that are available each being determined based on either a user not being logged in to the container pod or the container pod not currently processing computational tasks using the GPUs for acceleration;

based on the usage information, applying one or more rules to determine whether the number of available container pods satisfies a minimum or maximum number of available container pods;

in response to determination at a first time that the number of available container pods does not satisfy the minimum number of available container pods, adding (a) at least one container-based GPU node, or (b) at least one container pod to one of the container-based GPU nodes; and

in response to determination at a second time that the number of available container pods does not satisfy the maximum number of available container pods, removing (a) at least one of the container-based GPU nodes, or (b) at least one container pod from one of the container-based GPU nodes.

12 . The computer system of claim 11 , wherein the usage information is obtained from a plurality of GPU manager pods executing respectively in the container-based GPU nodes, the GPU manager pods each managing one or more GPU worker pods executing in the same one of the container-based GPU nodes as the GPU manager pod, and the usage information being based on availability of GPU worker pods.

13 . The computer system of claim 11 ,

wherein adding the at least one container-based GPU node or at least one container pod includes requesting a container orchestration layer to add the at least one container-based GPU node or at least one container pod, the container orchestration layer being configured to create and manage the container pods hosted by the container-based GPU nodes, and

wherein removing the at least one of the container-based GPU nodes or at least one container pod includes requesting the container orchestration layer to remove the at least one of the container-based GPU nodes or at least one container pod.

14 . The computer system of claim 11 , wherein

adding the at least one container-based GPU node or at least one container pod is triggered automatically by a container orchestration layer that is configured to create and manage the container pods hosted by the container-based GPU nodes, and

removing the at least one of the container-based GPU nodes or at least one container pod is triggered automatically by the container orchestration layer.

15 . The computer system of claim 11 , wherein

the usage information is obtained by a control plane entity of the computer system, and

the GPU node pool is deployed across multiple cloud providers.

Assignments (2)
CHANGE OF NAME Recorded May 16, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 067456/0166 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2023
From: ZHAO, YISAN; HU, XIAOYU; RIEMER, ROBERT; CULLY, AIDAN
To: VMWARE, INC.
Reel/Frame 063501/0709 →
Priority Claims (1)
WO PCT/CN2023/000007 · Jan 12, 2023 · international
Continuity (1)
Related Publication 20240241760A1 · Jul 18, 2024
References Cited (32)
US 8260840B1 · Sirota · 2012 [cited by examiner]
US 8719415B1 · Sirota · 2014 [cited by examiner]
US 8966030B1 · Sirota · 2015 [cited by examiner]
US 10262390B1 · Sun · 2019 [cited by examiner]
US 10275851B1 · Zhao · 2019 [cited by examiner]
US 10325343B1 · Zhao · 2019 [cited by examiner]
US 12154025B1 · Savic · 2024 [cited by examiner]
US 20150089034A1 · Stickle · 2015 [cited by examiner]
US 20150135185A1 · Sirota · 2015 [cited by examiner]
US 20160277488A1 · Fallon · 2016 [cited by examiner]
US 20180131581A1 · DeLuca · 2018 [cited by examiner]
US 20210081265A1 · Mariyappa · 2021 [cited by examiner]
US 20210117249A1 · Doshi · 2021 [cited by examiner]
US 20210149745A1 · An · 2021 [cited by examiner]
US 20210263780A1 · Ou · 2021 [cited by examiner]
US 20220094743A1 · Asveren · 2022 [cited by examiner]
US 20220318060A1 · Choochotkaew · 2022 [cited by examiner]
US 20220329651A1 · Kim · 2022 [cited by examiner]
US 20220405134A1 · Guo · 2022 [cited by examiner]
US 20230050796A1 · Ghergu · 2023 [cited by examiner]
US 20230080046A1 · Paul · 2023 [cited by examiner]
US 20230089925A1 · Cho · 2023 [cited by examiner]
US 20230114504A1 · He · 2023 [cited by examiner]
US 20230153162A1 · Baillargeon · 2023 [cited by examiner]
US 20230289206A1 · Kim · 2023 [cited by examiner]
US 20230403198A1 · Cai · 2023 [cited by examiner]
US 20230409368A1 · Siddiqui · 2023 [cited by examiner]
US 20240069964A1 · Chatterjee · 2024 [cited by examiner]
US 20240103926A1 · Saxena · 2024 [cited by examiner]
US 20240176662A1 · Liu · 2024 [cited by examiner]
US 20240241760A1 · Zhao · 2024 [cited by examiner]
US 20250156213A1 · Ai · 2025 [cited by examiner]