Elastic provisioning of container-based graphics processing unit (GPU) nodes
Example methods and systems for elastic provisioning of container-based graphics processing unit (GPU) nodes are described. In one example, a computer system may monitor usage information associated with a pool of multiple container-based GPU nodes. Based on the usage information, the computer system may apply rule(s) to determine whether capacity adjustment is required. In response to determination that capacity expansion is required, the computer system may configure the pool to expand by adding (a) at least one container-based GPU node to the pool, or (b) at least one container pod to one of the multiple container-based GPU nodes. Otherwise, in response to determination that capacity shrinkage is required, the computer system may configure the pool to shrink by removing (a) at least one container-based GPU node, or (b) at least one container pod from the pool.
1 . A method for a computer system to perform elastic scaling of a graphics processing unit (GPU) node pool that comprises a plurality of container-based GPU nodes hosting container pods that are capable of accessing GPUs, wherein the method comprises:
obtaining usage information of the GPU node pool, the usage information indicating a number of the container pods that are each available to use the GPUs to execute computational tasks initiated outside the GPU node pool, the container pods that are available each being determined based on either a user not being logged in to the container pod or the container pod not currently processing computational tasks using the GPUs for acceleration;
based on the usage information, applying one or more rules to determine whether the number of available container pods satisfies a minimum or maximum number of available container pods;
in response to determination at a first time that the number of available container pods does not satisfy the minimum number of available container pods, adding (a) at least one container-based GPU node, or (b) at least one container pod to one of the container-based GPU nodes; and
in response to determination at a second time that the number of available container pods does not satisfy the maximum number of available container pods, removing (a) at least one of the multiple-container-based GPU nodes, or (b) at least one container pod from one of the container-based GPU nodes.
2 . The method of claim 1 , wherein the usage information is obtained from a plurality of GPU manager pods executing respectively in the container-based GPU nodes, the GPU manager pods each managing one or more GPU worker pods executing in the same one of the container-based GPU nodes as the GPU manager pod, and the usage information being based on availability of GPU worker pods.
3 . The method of claim 1 ,
wherein adding the at least one container-based GPU node or at least one container pod includes requesting a container orchestration layer to add the at least one container-based GPU node or at least one container pod, the container orchestration layer being configured to create and manage the container pods hosted by the container-based GPU nodes, and
wherein removing the at least one of the container-based GPU nodes or at least one container pod includes requesting the container orchestration layer to remove the at least one of the container-based GPU nodes or at least one container pod.
4 . The method of claim 1 , wherein
adding the at least one container-based GPU node or at least one container pod is triggered automatically by a container orchestration layer that is configured to create and manage the container pods hosted by the container-based GPU nodes, and
removing the at least one of the container-based GPU nodes or at least one container pod is triggered automatically by the container orchestration layer.
5 . The method of claim 1 , wherein
the usage information is obtained by a control plane entity of the computer system, and
the GPU node pool is deployed across multiple cloud providers.
6 . One or more non-transitory computer-readable storage media including instructions which, in response to execution by one or more processors of a computer system, cause the one or more processors to perform a method for elastic scaling of a graphics processing unit (GPU) node pool that comprises a plurality of container-based GPU nodes hosting container pods that are capable of accessing GPUs, wherein the method comprises:
obtaining usage information of the GPU node pool, the usage information indicating a number of the container pods that are each available to use the GPUs to execute computational tasks initiated outside the GPU node pool, the container pods that are available each being determined based on either a user not being logged in to the container pod or the container pod not currently processing computational tasks using the GPUs for acceleration;
based on the usage information, applying one or more rules to determine whether the number of available container pods satisfies a minimum or maximum number of available container pods;
in response to determination at a first time that the number of available container pods does not satisfy the minimum number of available container pods, adding (a) at least one container-based GPU node, or (b) at least one container pod to one of the container-based GPU nodes; and
in response to determination at a second time that the number of available container pods does not satisfy the maximum number of available container pods, removing (a) at least one of the container-based GPU nodes, or (b) at least one container pod from one of the container-based GPU nodes.
7 . The one or more non-transitory computer-readable storage media of claim 6 , wherein the usage information is obtained from a plurality of GPU manager pods executing respectively in the container-based GPU nodes, the GPU manager pods each managing of one or more GPU worker pods executing in the same one of the container-based GPU nodes as the GPU manager pod, and the usage information being based on availability of GPU worker pods.
8 . The one or more non-transitory computer-readable storage media of claim 6 ,
wherein adding the at least one container-based GPU node or at least one container pod includes requesting a container orchestration layer to add the at least one container-based GPU node or at least one container pod, the container orchestration layer being configured to create and manage the container pods hosted by the container-based GPU nodes, and
wherein removing the at least one of the container-based GPU nodes or at least one container pod includes requesting the container orchestration layer to remove the at least one of the container-based GPU nodes or at least one container pod.
9 . The one or more non-transitory computer-readable storage media of claim 6 , wherein
adding the at least one container-based GPU node or at least one container pod is triggered automatically by a container orchestration layer that is configured to create and manage the container pods hosted by the container-based GPU nodes, and
removing the at least one of the container-based GPU nodes or at least one container pod is triggered automatically by the container orchestration layer.
10 . The one or more non-transitory computer-readable storage media of claim 6 , wherein
the usage information is obtained by a control plane entity of the computer system, and
the GPU node pool is deployed across multiple cloud providers.
11 . A computer system, comprising:
one or more processors; and
one or more non-transitory computer-readable media having stored thereon instructions that, when executed by the one or more processors, cause the one or more processors to perform a method for elastic scaling of a graphics processing unit (GPU) node pool that comprises a plurality of container-based GPU nodes hosting container pods that are capable of accessing GPUs, wherein the method comprises:
obtaining usage information of the GPU node pool, the usage information indicating a number of the container pods that are each available to use the GPUs to execute computational tasks initiated outside the GPU node pool, the container pods that are available each being determined based on either a user not being logged in to the container pod or the container pod not currently processing computational tasks using the GPUs for acceleration;
based on the usage information, applying one or more rules to determine whether the number of available container pods satisfies a minimum or maximum number of available container pods;
in response to determination at a first time that the number of available container pods does not satisfy the minimum number of available container pods, adding (a) at least one container-based GPU node, or (b) at least one container pod to one of the container-based GPU nodes; and
in response to determination at a second time that the number of available container pods does not satisfy the maximum number of available container pods, removing (a) at least one of the container-based GPU nodes, or (b) at least one container pod from one of the container-based GPU nodes.
12 . The computer system of claim 11 , wherein the usage information is obtained from a plurality of GPU manager pods executing respectively in the container-based GPU nodes, the GPU manager pods each managing one or more GPU worker pods executing in the same one of the container-based GPU nodes as the GPU manager pod, and the usage information being based on availability of GPU worker pods.
13 . The computer system of claim 11 ,
wherein adding the at least one container-based GPU node or at least one container pod includes requesting a container orchestration layer to add the at least one container-based GPU node or at least one container pod, the container orchestration layer being configured to create and manage the container pods hosted by the container-based GPU nodes, and
wherein removing the at least one of the container-based GPU nodes or at least one container pod includes requesting the container orchestration layer to remove the at least one of the container-based GPU nodes or at least one container pod.
14 . The computer system of claim 11 , wherein
adding the at least one container-based GPU node or at least one container pod is triggered automatically by a container orchestration layer that is configured to create and manage the container pods hosted by the container-based GPU nodes, and
removing the at least one of the container-based GPU nodes or at least one container pod is triggered automatically by the container orchestration layer.
15 . The computer system of claim 11 , wherein
the usage information is obtained by a control plane entity of the computer system, and
the GPU node pool is deployed across multiple cloud providers.