Virtualized computing resource management for machine learning model-based processing in computing environment
Techniques are disclosed for virtualized computing resource management for machine learning model-based processing in a computing environment. For example, a method maintains one or more virtualized computing resources, wherein each of the one or more virtualized computing resources is created and one or more initializations are caused to be performed. After creation and performance of the one or more initializations, each of the one or more virtualized computing resources is placed in an idle state. The method then receives a machine learning model-based request, and removes at least one of the one or more virtualized computing resources from the idle state to process the machine learning model-based request.
1 . A method, comprising:
determining a given number of virtualized computing resources to maintain in an idle state in a system, wherein the given number of virtualized computing resources to maintain in an idle state is determined based on available computational resources and traffic volume of the system;
maintaining the given number of the virtualized computing resources in an idle state in a standby pool of the system, wherein maintaining comprises creating the given number of virtualized computing resources, and performing initialization operations for each of the virtualized computing resources created, wherein the initialization operations comprise:
initializing a given virtualized computing resource on an operating system of a compute node of the system;
initializing a machine learning framework for use by the given virtualized computing resource;
initializing a hardware accelerator processing device for use by the given virtualized computing resource, wherein initializing the hardware accelerator processing device comprises creating a placeholder session to register the given virtualized computing resource on a memory space of the hardware accelerator processing device; and
initializing the given virtualized computing resource into an idle state;
receiving a machine learning model-based request; and
in response to receiving the machine learning model-based request:
waking up and removing from the standby pool one or more of the virtualized computing resources from the idle state to process the machine learning model-based request; and
creating and initializing one or more new virtualized computing resources into an idle state to replace the virtualized computing resources removed from the standby pool;
wherein the method is performed by at least one processor accessing and executing program instructions stored in at least one memory.
2 . The method of claim 1 , wherein the machine learning model-based request comprises an inference serving request.
3 . The method of claim 2 , further comprising processing the inference serving request by:
loading a trained machine learning model;
processing input associated with the inference serving request using the trained machine learning model; and
returning a result of the input processing by the trained machine learning model.
4 . The method of claim 1 , wherein creating the given number of virtualized computing resources comprises:
creating a given virtualized computing resource; and
mounting a model registry to the given virtualized computing resource, the model registry comprising a directory of one or more machine learning models.
5 . The method of claim 1 , wherein initializing the machine learning framework for use by the given virtualized computing resource comprises loading a plurality of machine learning libraries associated with the machine learning framework.
6 . The method of claim 1 , wherein the virtualized computing resources comprise containers.
7 . The method of claim 6 , wherein the at least one processor and the at least one memory comprises a worker node in a container orchestration framework.
8 . The method of claim 7 , wherein the worker node is part of an edge computing platform.
9 . An apparatus, comprising:
at least one processor and at least one memory storing computer program instructions wherein, when the at least one processor executes the computer program instructions, the apparatus is configured to:
determine a given number of virtualized computing resources to maintain in an idle state in a system, wherein the given number of virtualized computing resources to maintain in an idle state is determined based on available computational resources and traffic volume of the system;
maintaining the given number of the virtualized computing resources in an idle state in a standby pool of the system, wherein maintaining comprises creating the given number of virtualized computing resources, and performing initialization operations for each of the virtualized computing resources created, wherein the initialization operations comprise:
initializing a given virtualized computing resource on an operating system of a compute node of the system;
initializing a machine learning framework for use by the given virtualized computing resource;
initializing a hardware accelerator processing device for use by the given virtualized computing resource, wherein initializing the hardware accelerator processing device comprises creating a placeholder session to register the given virtualized computing resource on a memory space of the hardware accelerator processing device; and
initializing the given virtualized computing resource into an idle state;
receive a machine learning model-based request; and
in response to receiving the machine learning model-based request:
wake up and remove from the standby pool one or more of the virtualized computing resources from the idle state to process the machine learning model-based request; and
create and initialize one or more new virtualized computing resources into an idle state to replace the virtualized computing resources removed from the standby pool.
10 . The apparatus of claim 9 , wherein the machine learning model-based request comprises an inference serving request.
11 . The apparatus of claim 10 , wherein in processing the inference serving request the apparatus is configured to:
load a trained machine learning model;
process input associated with the inference serving request using the trained machine learning model; and
return a result of the input processing by the trained machine learning model.
12 . The apparatus of claim 9 , wherein:
creating the given number of virtualized computing resources comprises creating a given virtualized computing resource, and mounting a model registry to the given virtualized computing resource, the model registry comprising a directory of one or more machine learning models; and
initializing the machine learning framework for use by the given virtualized computing resource comprises loading a plurality of machine learning libraries associated with the machine learning framework.
13 . The apparatus of claim 9 , wherein:
the virtualized computing resources comprise containers;
the at least one processor and the at least one memory comprises a worker node in a container orchestration framework; and
the worker node is part of an edge computing platform.
14 . The apparatus of claim 9 , wherein the virtualized computing resources comprise containers.
15 . A computer program product stored on a non-transitory computer-readable medium and comprising machine executable instructions, the machine executable instructions, when executed, causing a processing device to perform steps of:
determining a given number of virtualized computing resources to maintain in an idle state in a system, wherein the given number of virtualized computing resources to maintain in an idle state is determined based on available computational resources and traffic volume of the system;
maintaining the given number of the virtualized computing resources in an idle state in a standby pool of the system, wherein maintaining comprises creating the given number of virtualized computing resources, and performing initialization operations for each of the virtualized computing resources created, wherein the initialization operations comprise:
initializing a given virtualized computing resource on an operating system of a compute node of the system;
initializing a machine learning framework for use by the given virtualized computing resource;
initializing a hardware accelerator processing device for use by the given virtualized computing resource, wherein initializing the hardware accelerator processing device comprises creating a placeholder session to register the given virtualized computing resource on a memory space of the hardware accelerator processing device; and
initializing the given virtualized computing resource into an idle state;
receiving a machine learning model-based request; and
in response the receiving the machine learning model-based request:
waking up and removing from the standby pool one or more of the virtualized computing resources from the idle state to process the machine learning model-based request; and
creating and initializing one or more new virtualized computing resources into an idle state to replace the virtualized computing resources removed from the standby pool to process the machine learning model-based request.
16 . The computer program product of claim 15 , wherein the machine learning model-based request comprises an inference serving request.
17 . The computer program product of claim 16 , further comprising processing the inference serving request by:
loading a trained machine learning model;
processing input associated with the inference serving request using the trained machine learning model; and
returning a result of the input processing by the trained machine learning model.
18 . The computer program product of claim 15 , wherein:
creating the given number of virtualized computing resources comprises creating a given virtualized computing resource, and mounting a model registry to the given virtualized computing resource, the model registry comprising a directory of one or more machine learning models; and
initializing the machine learning framework for use by the given virtualized computing resource comprises loading a plurality of machine learning libraries associated with the machine learning framework.
19 . The computer program product of claim 15 , wherein the virtualized computing resources comprise containers.
20 . The computer program product of claim 19 , wherein the processing device comprises a worker node in a container orchestration framework.