IP Library › Granted Patent US 12,639,114
Granted Patent B2
US 12,639,114 · App. 17/811,693 · Granted May 26, 2026

Swarm multi-agent reinforcement learning-based pipeline for workload placement

Inventors: Eduardo Vera Sousa (Niterói/RJ, BR); Hugo de Oliveira Barbalho (Rio de Janeiro, BR)
Assignee: Dell Products L.P.
G06F9/5011G06F9/5088G06F2209/501G06F2209/506
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,114
App. No.
17/811,693
Filed
Jul 11, 2022
Granted
May 26, 2026
Kind
B2
Art Unit
2196
USPC
718/104
Abstract

Multi-agent reinforcement learning-based workload placement is disclosed. A placement engine is configured to use the state of a system and actual rewards to generate expected rewards that correspond to actions. Agents can take actions for corresponding workloads based on the expected rewards output by the placement engine. This allows workloads to be placed in a manner that conservers power relative to load placement policies while helping avoid service level agreement violations.

Claims (128)

1 . A method comprising:

receiving input into a placement engine implemented on a cloud infrastructure, the input including an actual reward of a workload operating on a first virtual machine and a state, wherein the actual reward corresponds to a service level agreement (SLA) response time metric for the workload on the first virtual machine, and the state comprises a one-hot encoding vector representing, for each of a plurality of virtual machines and workloads executing thereon, resource usage, workload state, and time-to-completion values using floating-point entries;

inputting the state and actual reward into a trained neural network configured to map the input to expected rewards;

generating, by the trained neural network, the expected rewards for a plurality of candidate actions, including a first expected reward associated with maintaining the workload on the first virtual machine and a second expected reward associated with migrating the workloads to a second virtual machine; and

performing, by an agent associated with the workload, a first action when the first expected reward exceeds the second expected reward, or a second action when the second expected reward exceeds the first expected reward,

wherein the trained neural network is trained using a process that includes:

random migration of workloads among the plurality of virtual machines during training to explore workload placement conditions; and

use of a reward function having a curve that is asymmetric and nonlinear, defined by a difference between an SLA response-time metric and an actual response time for the workload, wherein a portion of the reward curve corresponding to SLA violations decays faster than a portion corresponding to SLA compliance, such that the training of the placement engine is biased toward maintaining SLA targets in workload placement decisions.

2 . The method of claim 1 , wherein the cloud infrastructure comprises the plurality of virtual machines, and wherein the actual reward corresponds to a service level agreement metric of the workload operating on the first virtual machine.

3 . The method of claim 2 , wherein the state includes a one hot encoding style of the plurality of virtual machines.

4 . The method of claim 3 , wherein the one hot encoding style includes a resource usage per virtual machine, a resource usage per workload, a state of each workload, and a time to completion for each workload using floating point values.

5 . The method of claim 2 , wherein the first action is to keep the workload at the first virtual machine and wherein the second action is to migrate the workload to the second virtual machine.

6 . The method of claim 1 , wherein the placement engine comprises a neural network configured to map the input to the expected rewards.

7 . The method of claim 1 , wherein the placement engine outputs an expected reward for performing an action relative to each virtual machine in the cloud infrastructure.

8 . The method of claim 1 , further comprising adjusting the reward function when an SLA violation is detected.

9 . The method of claim 8 , wherein the reward function is:

f

⁡

(

Δ

,

σ

L

,

σ

R

)

=

-

(

Δ

)

2

e

2

⁢

σ

L

2

⁢

if

⁢

Δ

>

1

,

otherwise

⁢

-

(

Δ

)

2

e

2

⁢

σ

L

2

-

1

,

wherein Δ is a difference between an SLA response time metric and an actual response time for the workload in the environment, wherein σ L and σ R define, respectively, how fast a left and a right portion of the reward function decay.

10 . The method of claim 1 , wherein the placement engine is trained by randomly migrating workloads amongst virtual machines in the cloud infrastructure.

11 . The method of claim 1 , wherein the placement engine is configured to place the workload in a manner that includes both minimum virtual machine placement usage and balancing the workload across the plurality of virtual machines.

12 . A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising:

receiving input into a placement engine implemented on a cloud infrastructure, the input including an actual reward of a workload operating on a first virtual machine and a state, wherein the actual reward corresponds to a service level agreement (SLA) response time metric for the workload on the first virtual machine, and the state comprises a one-hot encoding vector representing, for each of a plurality of virtual machines and workloads executing thereon, resource usage, workload state, and time-to-completion values using floating-point entries;

inputting the state and actual reward into a trained neural network configured to map the input to expected rewards;

generating, by the trained neural network, the expected rewards for a plurality of candidate actions, including a first expected reward associated with maintaining the workload on the first virtual machine and a second expected reward associated with migrating the workloads to a second virtual machine; and

performing, by an agent associated with the workload, a first action when the first expected reward exceeds the second expected reward, or a second action when the second expected reward exceeds the first expected reward,

wherein the trained neural network is trained using a process that includes:

random migration of workloads among the plurality of virtual machines during training to explore workload placement conditions; and

use of a reward function having a curve that is asymmetric and nonlinear, defined by a difference between an SLA response-time metric and an actual response time for the workload, wherein a portion of the reward curve corresponding to SLA violations decays faster than a portion corresponding to SLA compliance, such that the training of the placement engine is biased toward maintaining SLA targets in workload placement decisions.

13 . The non-transitory storage medium of claim 12 , wherein the cloud infrastructure comprises the plurality of virtual machines and wherein the actual reward corresponds to a service level agreement metric of the workload operating on the first virtual machine.

14 . The non-transitory storage medium of claim 13 , wherein the state includes a one hot encoding style of the plurality of virtual machines.

15 . The non-transitory storage medium of claim 14 , wherein the one hot encoding style includes a resource usage per virtual machine, a resource usage per workload, a state of each workload, and a time to completion for each workload using floating point values.

16 . The non-transitory storage medium of claim 13 , wherein the first action is to keep the workload at the first virtual machine and wherein the second action is to migrate the workload to the second virtual machine.

17 . The non-transitory storage medium of claim 12 , wherein the placement engine comprises a neural network configured to map the input to the expected rewards, wherein the placement engine is configured to place the workload in a manner that includes both minimizing virtual machine usage and balancing the workload across the plurality of virtual machines.

18 . The non-transitory storage medium of claim 12 , wherein the placement engine outputs an expected reward for performing an action relative to each virtual machine in the cloud infrastructure.

19 . The non-transitory storage medium of claim 12 , further comprising adjusting a reward function when an SLA violation is detected.

20 . The non-transitory storage medium of claim 19 , wherein the reward function is:

f

⁡

(

Δ

,

σ

L

,

σ

R

)

=

-

(

Δ

)

2

e

2

⁢

σ

L

2

⁢

if

⁢

Δ

>

1

,

otherwise

⁢

-

(

Δ

)

2

e

2

⁢

σ

L

2

-

1

,

wherein Δ is a difference between an SLA response time metric and an actual response time for the workload in the environment, wherein σ L and σ R define, respectively, how fast a left and a right portion of the reward function decay.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2022
From: SOUSA, EDUARDO VERA; BARBALHO, HUGO DE OLIVEIRA
To: DELL PRODUCTS L.P.
Reel/Frame 060474/0045 →
Continuity (1)
Related Publication 20240012685A1 · Jan 11, 2024
References Cited (27)
US 7467145B1 · Castellanos et al. · 2008 [cited by applicant]
US 10963810B2 · Dirac et al. · 2021 [cited by applicant]
US 11574243B1 · Mallya Kasaragod · 2023 [cited by examiner]
US 20100242045A1 · Swamy · 2010 [cited by examiner]
US 20180260746A1 · Xiong et al. · 2018 [cited by applicant]
US 20190266015A1 · Chandra · 2019 [cited by examiner]
US 20200090819A1 · Durduran et al. · 2020 [cited by applicant]
US 20200241921A1 · Calmon · 2020 [cited by examiner]
US 20200257968A1 · Mitra · 2020 [cited by examiner]
US 20210058455A1 · Kozhaya et al. · 2021 [cited by applicant]
US 20210092071A1 · Tortosa et al. · 2021 [cited by applicant]
US 20220180275A1 · Parizi et al. · 2022 [cited by applicant]
US 20230086563A1 · Min et al. · 2023 [cited by applicant]
US 20230171340A1 · Saravanan et al. · 2023 [cited by applicant]
US 20230334320A1 · Zhang et al. · 2023 [cited by applicant]
US 20240015080A1 · Illikkal et al. · 2024 [cited by applicant]
US 20240386054A1 · Wouhaybi et al. · 2024 [cited by applicant]
Zhang, X., Shae, Z. Y., Zheng, S., & Jamjoom, H. (Apr. 2012). Virtual machine migration in an over-committed cloud. In 2012 IEEE Network Operations and Management Symposium (pp. 196-203). IEEE. (Year: 2012). [cited by examiner]
Mignon, Alexandre dos Santos; Luis de Azevedo da Rocha, Ricardo, An adaptive implementation of e-greedy in reinforcement learning, Procedia Computer Science, 2017. [cited by applicant]
Crawford, Daniel, et al. Reinforcement Learning using Quantum Boltzmann Machines. arXiv preprint arXiv:1612.05695, 2016. [cited by applicant]
Ghojogh, Benyamin, et al. Restricted Boltzmann Machine and Deep Belief Network: Tutorial and Survey. arXiv preprint arXiv:2107.12521 (2021). [cited by applicant]
Kullback, Solomon, and Richard A. Leibler. “On information and sufficiency.” The annals of mathematical statistics 22.1 (1951): 79-86. [cited by applicant]
Gronauer, Sven; Diepold, Klaus. Multi-agent deep reinforcement learning: a survey. Artificial Intelligence Review, 2022, vol. 55, No. 2, p. 895-943. [cited by applicant]
Kalyan et al. (“Unsupervised Synaptic Pruning Strategies for Restricted Boltzmann Machines”; 2018 IEEE Biomedical Circuits and Systems Conference (BioCAS); Oct. 17-19, 2018). [cited by applicant]
Berman H.B., “Probability Rules”, archived at https://web.archive.org/web/20220625141414/https://stattrek.com/probability/probability-rules, Jun. 25, 2022. Accessed Dec. 1, 2025. (Year: 2022). [cited by applicant]
Mondal et al, “Scheduling of Time-Varying Workloads Using Reinforcement Learning”, 2021, The Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI-21), pp. 9000-9008. (Year: 2021). [cited by applicant]
Tong et al, “Proactive scheduling in distributed computing—A reinforcement learning approach”, 2014, J. Parallel Distrib. Comput. 74, pp. 2662-2672. (Year: 2014). [cited by applicant]