IP Library › Granted Patent US 12,724,645
Granted Patent B2
US 12,724,645 · App. 17/681,288 · Granted Sep 1, 2026

Virtualized computing resource management for machine learning model-based processing in computing environment

Inventor: Victor Fong (Melrose, MA)
Assignee: Dell Products L.P.
G06F9/5077G06F9/45558G06N5/027G06F2009/4557
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,645
App. No.
17/681,288
Granted
Sep 1, 2026
Kind
B2
Abstract

Techniques are disclosed for virtualized computing resource management for machine learning model-based processing in a computing environment. For example, a method maintains one or more virtualized computing resources, wherein each of the one or more virtualized computing resources is created and one or more initializations are caused to be performed. After creation and performance of the one or more initializations, each of the one or more virtualized computing resources is placed in an idle state. The method then receives a machine learning model-based request, and removes at least one of the one or more virtualized computing resources from the idle state to process the machine learning model-based request.

Claims (70)

1 . A method, comprising:

determining a given number of virtualized computing resources to maintain in an idle state in a system, wherein the given number of virtualized computing resources to maintain in an idle state is determined based on available computational resources and traffic volume of the system;

maintaining the given number of the virtualized computing resources in an idle state in a standby pool of the system, wherein maintaining comprises creating the given number of virtualized computing resources, and performing initialization operations for each of the virtualized computing resources created, wherein the initialization operations comprise:

initializing a given virtualized computing resource on an operating system of a compute node of the system;

initializing a machine learning framework for use by the given virtualized computing resource;

initializing a hardware accelerator processing device for use by the given virtualized computing resource, wherein initializing the hardware accelerator processing device comprises creating a placeholder session to register the given virtualized computing resource on a memory space of the hardware accelerator processing device; and

initializing the given virtualized computing resource into an idle state;

receiving a machine learning model-based request; and

in response to receiving the machine learning model-based request:

waking up and removing from the standby pool one or more of the virtualized computing resources from the idle state to process the machine learning model-based request; and

creating and initializing one or more new virtualized computing resources into an idle state to replace the virtualized computing resources removed from the standby pool;

wherein the method is performed by at least one processor accessing and executing program instructions stored in at least one memory.

2 . The method of claim 1 , wherein the machine learning model-based request comprises an inference serving request.

3 . The method of claim 2 , further comprising processing the inference serving request by:

loading a trained machine learning model;

processing input associated with the inference serving request using the trained machine learning model; and

returning a result of the input processing by the trained machine learning model.

4 . The method of claim 1 , wherein creating the given number of virtualized computing resources comprises:

creating a given virtualized computing resource; and

mounting a model registry to the given virtualized computing resource, the model registry comprising a directory of one or more machine learning models.

5 . The method of claim 1 , wherein initializing the machine learning framework for use by the given virtualized computing resource comprises loading a plurality of machine learning libraries associated with the machine learning framework.

6 . The method of claim 1 , wherein the virtualized computing resources comprise containers.

7 . The method of claim 6 , wherein the at least one processor and the at least one memory comprises a worker node in a container orchestration framework.

8 . The method of claim 7 , wherein the worker node is part of an edge computing platform.

9 . An apparatus, comprising:

at least one processor and at least one memory storing computer program instructions wherein, when the at least one processor executes the computer program instructions, the apparatus is configured to:

determine a given number of virtualized computing resources to maintain in an idle state in a system, wherein the given number of virtualized computing resources to maintain in an idle state is determined based on available computational resources and traffic volume of the system;

maintaining the given number of the virtualized computing resources in an idle state in a standby pool of the system, wherein maintaining comprises creating the given number of virtualized computing resources, and performing initialization operations for each of the virtualized computing resources created, wherein the initialization operations comprise:

initializing a given virtualized computing resource on an operating system of a compute node of the system;

initializing a machine learning framework for use by the given virtualized computing resource;

initializing a hardware accelerator processing device for use by the given virtualized computing resource, wherein initializing the hardware accelerator processing device comprises creating a placeholder session to register the given virtualized computing resource on a memory space of the hardware accelerator processing device; and

initializing the given virtualized computing resource into an idle state;

receive a machine learning model-based request; and

in response to receiving the machine learning model-based request:

wake up and remove from the standby pool one or more of the virtualized computing resources from the idle state to process the machine learning model-based request; and

create and initialize one or more new virtualized computing resources into an idle state to replace the virtualized computing resources removed from the standby pool.

10 . The apparatus of claim 9 , wherein the machine learning model-based request comprises an inference serving request.

11 . The apparatus of claim 10 , wherein in processing the inference serving request the apparatus is configured to:

load a trained machine learning model;

process input associated with the inference serving request using the trained machine learning model; and

return a result of the input processing by the trained machine learning model.

12 . The apparatus of claim 9 , wherein:

creating the given number of virtualized computing resources comprises creating a given virtualized computing resource, and mounting a model registry to the given virtualized computing resource, the model registry comprising a directory of one or more machine learning models; and

initializing the machine learning framework for use by the given virtualized computing resource comprises loading a plurality of machine learning libraries associated with the machine learning framework.

13 . The apparatus of claim 9 , wherein:

the virtualized computing resources comprise containers;

the at least one processor and the at least one memory comprises a worker node in a container orchestration framework; and

the worker node is part of an edge computing platform.

14 . The apparatus of claim 9 , wherein the virtualized computing resources comprise containers.

15 . A computer program product stored on a non-transitory computer-readable medium and comprising machine executable instructions, the machine executable instructions, when executed, causing a processing device to perform steps of:

determining a given number of virtualized computing resources to maintain in an idle state in a system, wherein the given number of virtualized computing resources to maintain in an idle state is determined based on available computational resources and traffic volume of the system;

maintaining the given number of the virtualized computing resources in an idle state in a standby pool of the system, wherein maintaining comprises creating the given number of virtualized computing resources, and performing initialization operations for each of the virtualized computing resources created, wherein the initialization operations comprise:

initializing a given virtualized computing resource on an operating system of a compute node of the system;

initializing a machine learning framework for use by the given virtualized computing resource;

initializing a hardware accelerator processing device for use by the given virtualized computing resource, wherein initializing the hardware accelerator processing device comprises creating a placeholder session to register the given virtualized computing resource on a memory space of the hardware accelerator processing device; and

initializing the given virtualized computing resource into an idle state;

receiving a machine learning model-based request; and

in response the receiving the machine learning model-based request:

waking up and removing from the standby pool one or more of the virtualized computing resources from the idle state to process the machine learning model-based request; and

creating and initializing one or more new virtualized computing resources into an idle state to replace the virtualized computing resources removed from the standby pool to process the machine learning model-based request.

16 . The computer program product of claim 15 , wherein the machine learning model-based request comprises an inference serving request.

17 . The computer program product of claim 16 , further comprising processing the inference serving request by:

loading a trained machine learning model;

processing input associated with the inference serving request using the trained machine learning model; and

returning a result of the input processing by the trained machine learning model.

18 . The computer program product of claim 15 , wherein:

creating the given number of virtualized computing resources comprises creating a given virtualized computing resource, and mounting a model registry to the given virtualized computing resource, the model registry comprising a directory of one or more machine learning models; and

initializing the machine learning framework for use by the given virtualized computing resource comprises loading a plurality of machine learning libraries associated with the machine learning framework.

19 . The computer program product of claim 15 , wherein the virtualized computing resources comprise containers.

20 . The computer program product of claim 19 , wherein the processing device comprises a worker node in a container orchestration framework.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: FONG, VICTOR
To: DELL PRODUCTS L.P.
Reel/Frame 059106/0621 →
Continuity (1)
Related Publication 20230273837A1 · Aug 31, 2023
References Cited (40)
US 9891954B2 · Antony · 2018 [cited by examiner]
US 9971621B1 · Berg · 2018 [cited by examiner]
US 10628195B2 · Miller · 2020 [cited by examiner]
US 11373119B1 · Doshi · 2022 [cited by examiner]
US 11900174B2 · Hou · 2024 [cited by examiner]
US 12219057B2 · Rogers · 2025 [cited by examiner]
US 12443425B1 · Raslan · 2025 [cited by examiner]
US 20080177424A1 · Wheeler · 2008 [cited by examiner]
US 20080201711A1 · Amir Husain · 2008 [cited by examiner]
US 20120011509A1 · Husain · 2012 [cited by examiner]
US 20130179895A1 · Calder · 2013 [cited by examiner]
US 20140096132A1 · Wang · 2014 [cited by examiner]
US 20160092250A1 · Wagner · 2016 [cited by examiner]
US 20160103698A1 · Yang · 2016 [cited by examiner]
US 20160150053A1 · Janczuk · 2016 [cited by examiner]
US 20170322834A1 · de Sene · 2017 [cited by examiner]
US 20190155633A1 · Faulhaber, Jr. · 2019 [cited by examiner]
US 20190156244A1 · Faulhaber, Jr. · 2019 [cited by examiner]
US 20200026576A1 · Kaplan · 2020 [cited by examiner]
US 20200027210A1 · Haemel · 2020 [cited by examiner]
US 20200183723A1 · Sabev · 2020 [cited by examiner]
US 20200311617A1 · Swan · 2020 [cited by examiner]
US 20210125104A1 · Christiansen · 2021 [cited by examiner]
US 20220188138A1 · Akkur Rajamannar · 2022 [cited by examiner]
US 20220206873A1 · He · 2022 [cited by examiner]
US 20220237505A1 · Feldman · 2022 [cited by examiner]
US 20220261631A1 · Cohen · 2022 [cited by examiner]
US 20220311594A1 · Kadam · 2022 [cited by examiner]
US 20220318647A1 · Ashrafzadeh · 2022 [cited by examiner]
US 20220342649A1 · Cao · 2022 [cited by examiner]
US 20220342714A1 · Lincourt, Jr. · 2022 [cited by examiner]
US 20220382601A1 · Feldman · 2022 [cited by examiner]
US 20220414503A1 · Park · 2022 [cited by examiner]
US 20230273837A1 · Fong · 2023 [cited by examiner]
US 20240171657A1 · Sharma Banjade · 2024 [cited by examiner]
US 20240256313A1 · Sanapo · 2024 [cited by examiner]
L. Wang et al., “Peeking Behind the Curtains of Serverless Platforms,” USENIX Annual Technical Conference, Jul. 2018, 13 pages. [cited by applicant]
F. Romero et al., “INFaaS: Automated Model-less Inference Serving,” USENIX Annual Technical Conference, Jul. 2021, 15, pages. [cited by applicant]
Wikipedia, “Edge Computing,” https://en.wikipedia.org/w/index.php?title=Edge_computing&oldid=1073476548, Feb. 22, 2022, 7 pages. [cited by applicant]
The Kubernetes Authors, “Cluster Architecture,” https://kubernetes.io/docs/concepts/architecture/_print/, Accessed Feb. 25, 2022, 23 pages. [cited by applicant]