Swarm multi-agent reinforcement learning-based pipeline for workload placement
Multi-agent reinforcement learning-based workload placement is disclosed. A placement engine is configured to use the state of a system and actual rewards to generate expected rewards that correspond to actions. Agents can take actions for corresponding workloads based on the expected rewards output by the placement engine. This allows workloads to be placed in a manner that conservers power relative to load placement policies while helping avoid service level agreement violations.
1 . A method comprising:
receiving input into a placement engine implemented on a cloud infrastructure, the input including an actual reward of a workload operating on a first virtual machine and a state, wherein the actual reward corresponds to a service level agreement (SLA) response time metric for the workload on the first virtual machine, and the state comprises a one-hot encoding vector representing, for each of a plurality of virtual machines and workloads executing thereon, resource usage, workload state, and time-to-completion values using floating-point entries;
inputting the state and actual reward into a trained neural network configured to map the input to expected rewards;
generating, by the trained neural network, the expected rewards for a plurality of candidate actions, including a first expected reward associated with maintaining the workload on the first virtual machine and a second expected reward associated with migrating the workloads to a second virtual machine; and
performing, by an agent associated with the workload, a first action when the first expected reward exceeds the second expected reward, or a second action when the second expected reward exceeds the first expected reward,
wherein the trained neural network is trained using a process that includes:
random migration of workloads among the plurality of virtual machines during training to explore workload placement conditions; and
use of a reward function having a curve that is asymmetric and nonlinear, defined by a difference between an SLA response-time metric and an actual response time for the workload, wherein a portion of the reward curve corresponding to SLA violations decays faster than a portion corresponding to SLA compliance, such that the training of the placement engine is biased toward maintaining SLA targets in workload placement decisions.
2 . The method of claim 1 , wherein the cloud infrastructure comprises the plurality of virtual machines, and wherein the actual reward corresponds to a service level agreement metric of the workload operating on the first virtual machine.
3 . The method of claim 2 , wherein the state includes a one hot encoding style of the plurality of virtual machines.
4 . The method of claim 3 , wherein the one hot encoding style includes a resource usage per virtual machine, a resource usage per workload, a state of each workload, and a time to completion for each workload using floating point values.
5 . The method of claim 2 , wherein the first action is to keep the workload at the first virtual machine and wherein the second action is to migrate the workload to the second virtual machine.
6 . The method of claim 1 , wherein the placement engine comprises a neural network configured to map the input to the expected rewards.
7 . The method of claim 1 , wherein the placement engine outputs an expected reward for performing an action relative to each virtual machine in the cloud infrastructure.
8 . The method of claim 1 , further comprising adjusting the reward function when an SLA violation is detected.
9 . The method of claim 8 , wherein the reward function is:
f
(
Δ
,
σ
L
,
σ
R
)
=
-
(
Δ
)
2
e
2
σ
L
2
if
Δ
>
1
,
otherwise
-
(
Δ
)
2
e
2
σ
L
2
-
1
,
wherein Δ is a difference between an SLA response time metric and an actual response time for the workload in the environment, wherein σ L and σ R define, respectively, how fast a left and a right portion of the reward function decay.
10 . The method of claim 1 , wherein the placement engine is trained by randomly migrating workloads amongst virtual machines in the cloud infrastructure.
11 . The method of claim 1 , wherein the placement engine is configured to place the workload in a manner that includes both minimum virtual machine placement usage and balancing the workload across the plurality of virtual machines.
12 . A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising:
receiving input into a placement engine implemented on a cloud infrastructure, the input including an actual reward of a workload operating on a first virtual machine and a state, wherein the actual reward corresponds to a service level agreement (SLA) response time metric for the workload on the first virtual machine, and the state comprises a one-hot encoding vector representing, for each of a plurality of virtual machines and workloads executing thereon, resource usage, workload state, and time-to-completion values using floating-point entries;
inputting the state and actual reward into a trained neural network configured to map the input to expected rewards;
generating, by the trained neural network, the expected rewards for a plurality of candidate actions, including a first expected reward associated with maintaining the workload on the first virtual machine and a second expected reward associated with migrating the workloads to a second virtual machine; and
performing, by an agent associated with the workload, a first action when the first expected reward exceeds the second expected reward, or a second action when the second expected reward exceeds the first expected reward,
wherein the trained neural network is trained using a process that includes:
random migration of workloads among the plurality of virtual machines during training to explore workload placement conditions; and
use of a reward function having a curve that is asymmetric and nonlinear, defined by a difference between an SLA response-time metric and an actual response time for the workload, wherein a portion of the reward curve corresponding to SLA violations decays faster than a portion corresponding to SLA compliance, such that the training of the placement engine is biased toward maintaining SLA targets in workload placement decisions.
13 . The non-transitory storage medium of claim 12 , wherein the cloud infrastructure comprises the plurality of virtual machines and wherein the actual reward corresponds to a service level agreement metric of the workload operating on the first virtual machine.
14 . The non-transitory storage medium of claim 13 , wherein the state includes a one hot encoding style of the plurality of virtual machines.
15 . The non-transitory storage medium of claim 14 , wherein the one hot encoding style includes a resource usage per virtual machine, a resource usage per workload, a state of each workload, and a time to completion for each workload using floating point values.
16 . The non-transitory storage medium of claim 13 , wherein the first action is to keep the workload at the first virtual machine and wherein the second action is to migrate the workload to the second virtual machine.
17 . The non-transitory storage medium of claim 12 , wherein the placement engine comprises a neural network configured to map the input to the expected rewards, wherein the placement engine is configured to place the workload in a manner that includes both minimizing virtual machine usage and balancing the workload across the plurality of virtual machines.
18 . The non-transitory storage medium of claim 12 , wherein the placement engine outputs an expected reward for performing an action relative to each virtual machine in the cloud infrastructure.
19 . The non-transitory storage medium of claim 12 , further comprising adjusting a reward function when an SLA violation is detected.
20 . The non-transitory storage medium of claim 19 , wherein the reward function is:
f
(
Δ
,
σ
L
,
σ
R
)
=
-
(
Δ
)
2
e
2
σ
L
2
if
Δ
>
1
,
otherwise
-
(
Δ
)
2
e
2
σ
L
2
-
1
,
wherein Δ is a difference between an SLA response time metric and an actual response time for the workload in the environment, wherein σ L and σ R define, respectively, how fast a left and a right portion of the reward function decay.