IP Library Granted Patent US 11,562,223
Granted Patent B2
US 11,562,223 · App. 15/961,033 · Granted Jan 24, 2023

Deep reinforcement learning for workflow optimization

Inventors: Vinícius Michel Gottin (Rio de Janeiro, BR); Jonas F. Dias (Rio de Janeiro, BR); Daniel Sadoc Menasché (Rio de Janeiro, BR); Alex Laier Bordignon (Niterói, BR); Angelo Ernani Maia Ciarlini (Rio de Janeiro, BR)
Assignee: EMC IP Holding Company LLC
G06N3/08G06F9/5027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,223
App. No.
15/961,033
Granted
Jan 24, 2023
Kind
B2
Abstract

Deep reinforcement learning techniques are provided for resource allocation in a shared computing environment. An exemplary method comprises: obtaining a specification of a workflow of a plurality of concurrent workflows in a shared computing environment, wherein the specification comprises a plurality of workflow states and one or more control variables for the workflow in the shared computing environment; evaluating values of the control variables for an execution of the concurrent workflows using a reinforcement learning agent by (i) observing the states, including a current state, and (ii) obtaining an expected utility score for combinations of the control variables for the execution of the concurrent workflows given an allocation of resources of the shared computing environment corresponding to the combination of control variables in the current state; and providing an allocation of the resources of the shared computing environment reflecting the combination having the expected utility score that satisfies a predefined score criteria.

Claims (31)

1. A method, comprising:

obtaining a specification of a plurality of concurrent workflows in a shared computing environment, wherein the specification of a given one of the plurality of concurrent workflows comprises a plurality of states of the given workflow and one or more control variables indicating an allocation of one or more resources for the given workflow in the shared computing environment;

evaluating, using at least one processing device, a plurality of values of the one or more control variables indicating the allocation of the one or more resources, for an execution of said plurality of concurrent workflows, using at least one reinforcement learning agent, wherein said evaluating comprises observing said plurality of states, including a current state comprising a current configuration of said plurality of concurrent workflows and said shared computing environment, and obtaining, from the at least one reinforcement learning agent, an expected utility score for a plurality of combinations of said control variables for the execution of said plurality of concurrent workflows given an allocation of the one or more resources of the shared computing environment corresponding to said combination of said control variables in said current state, wherein the at least one reinforcement learning agent traverses said plurality of states and is trained to select a particular action for a given state, wherein the particular action corresponds to the allocation of the one or more resources of the shared computing environment for the given state; and

initiating an adjustment of the allocation of the one or more resources of the shared computing environment reflecting the combination of the control variables having the expected utility score, from the at least one reinforcement learning agent, that satisfies one or more predefined score criteria.

2. The method of claim 1 , further comprising the step of applying the allocation of the one or more resources of the shared environment.

3. The method of claim 1 , further comprising the step of updating said at least one reinforcement learning agent by further training a model with the states that result from said allocation as new training samples.

4. The method of claim 1 , wherein said expected utility score further comprises an expected cost depending on one or more of an execution time of the given workflow and a consumption of resources in said shared computing environment.

5. The method of claim 1 , wherein said states further comprise provenance data of said plurality of workflows.

6. The method of claim 1 , wherein said states further comprise telemetry data of said shared computing environment.

7. The method of claim 1 , wherein said at least one reinforcement learning agent comprises a Deep Q-Learning agent using a Q-Deep Neural Network (QDNN) as a representation of a Q-Function, and wherein said obtaining the expected utility score for the plurality of combinations of said control variables comprises selecting an action at random and computing a cost-to-go from the expected utility score of the selected action updated by an observation of the current state, and wherein an updating of the at least one reinforcement learning agent comprises a training of the QDNN given new samples in iterative epochs.

8. The method of claim 7 , wherein the values of the expected utility score are given by predictions from a Deep Neural Network for a predefined number of training epochs.

9. The method of claim 7 , wherein the obtaining the expected utility score for the plurality of combinations of said control variables further comprises the step of updating the cost-to-go from a previous state with a future value given by a pretrained neural network.

10. The method of claim 7 , wherein said computation of the cost-to-go from the expected utility score of the selected action updated by the observation of the current state additionally comprises the storage of input/output pairs in a database of samples and wherein a training batch for said training of the QDNN is comprised of new samples from the database processed in iterative epochs.

11. The method of claim 7 , wherein estimates of substantially optimal future values in said computation of the cost-to-go from the expected utility score of the selected action updated by the observation of the current state are given by the outputs of the QDNN with an architecture that yields multiple outputs configuring the expected utility scores for substantially all actions.

12. The method of claim 1 , wherein the one or more control variables comprise one or more of a number of processing cores allocated to a given workflow and an amount of memory allocated to the given workflow.

13. A system, comprising:

a memory; and

at least one processing device, coupled to the memory, operative to implement the following steps:

obtaining a specification of a plurality of concurrent workflows in a shared computing environment, wherein the specification of a given one of the plurality of concurrent workflows comprises a plurality of states of the given workflow and one or more control variables indicating an allocation of one or more resources for the given workflow in the shared computing environment;

evaluating, using at least one processing device, a plurality of values of the one or more control variables indicating the allocation of the one or more resources, for an execution of said plurality of concurrent workflows, using at least one reinforcement learning agent, wherein said evaluating comprises observing said plurality of states, including a current state comprising a current configuration of said plurality of concurrent workflows and said shared computing environment, and obtaining, from the at least one reinforcement learning agent, an expected utility score for a plurality of combinations of said control variables for the execution of said plurality of concurrent workflows given an allocation of the one or more resources of the shared computing environment corresponding to said combination of said control variables in said current state, wherein the at least one reinforcement learning agent traverses said plurality of states and is trained to select a particular action for a given state, wherein the particular action corresponds to the allocation of the one or more resources of the shared computing environment for the given state; and

initiating an adjustment of the allocation of the one or more resources of the shared computing environment reflecting the combination of the control variables having the expected utility score, from the at least one reinforcement learning agent, that satisfies one or more predefined score criteria.

14. The system of claim 13 , further comprising the step of updating said at least one reinforcement learning agent by further training a model with the states that result from said allocation as new training samples.

15. The system of claim 13 , wherein said expected utility score further comprises an expected cost depending on one or more of an execution time of the given workflow and a consumption of resources in said shared computing environment.

16. The system of claim 13 , wherein said at least one reinforcement learning agent comprises a Deep Q-Learning agent using a Q-Deep Neural Network (QDNN) as a representation of a Q-Function, and wherein said obtaining the expected utility score for the plurality of combinations of said control variables comprises selecting an action at random and computing a cost-to-go from the expected utility score of the selected action updated by an observation of the current state, and wherein an updating of the at least one reinforcement learning agent comprises a training of the QDNN given new samples in iterative epochs.

17. The system of claim 16 , wherein the values of the expected utility score are given by predictions from a Deep Neural Network for a predefined number of training epochs.

18. The system of claim 16 , wherein the obtaining the expected utility score for the plurality of combinations of said control variables further comprises the step of updating the cost-to-go from a previous state with a future value given by a pretrained neural network.

19. The system of claim 16 , wherein said computation of the cost-to-go from the expected utility score of the selected action updated by the observation of the current state additionally comprises the storage of input/output pairs in a database of samples and wherein a training batch for said training of the QDNN is comprised of new samples from the database processed in iterative epochs.

20. A computer program product, comprising a tangible machine-readable storage medium having encoded therein executable code of one or more software programs, wherein the one or more software programs when executed by at least one processing device perform the following steps:

obtaining a specification of a plurality of concurrent workflows in a shared computing environment, wherein the specification of a given one of the plurality of concurrent workflows comprises a plurality of states of the given workflow and one or more control variables indicating an allocation of one or more resources for the given workflow in the shared computing environment;

evaluating, using at least one processing device, a plurality of values of the one or more control variables indicating the allocation of the one or more resources, for an execution of said plurality of concurrent workflows, using at least one reinforcement learning agent, wherein said evaluating comprises observing said plurality of states, including a current state comprising a current configuration of said plurality of concurrent workflows and said shared computing environment, and obtaining, from the at least one reinforcement learning agent, an expected utility score for a plurality of combinations of said control variables for the execution of said plurality of concurrent workflows given an allocation of the one or more resources of the shared computing environment corresponding to said combination of said control variables in said current state, wherein the at least one reinforcement learning agent traverses said plurality of states and is trained to select a particular action for a given state, wherein the particular action corresponds to the allocation of the one or more resources of the shared computing environment for the given state; and

initiating an adjustment of the allocation of the one or more resources of the shared computing environment reflecting the combination of the control variables having the expected utility score, from the at least one reinforcement learning agent, that satisfies one or more predefined score criteria.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (046366/0014) Recorded May 20, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
Reel/Frame 060450/0306 →
RELEASE OF SECURITY INTEREST AT REEL 046286 FRAME 0653 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
Reel/Frame 058298/0093 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
PATENT SECURITY AGREEMENT (CREDIT) Recorded Jun 1, 2018
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 046286/0653 →
PATENT SECURITY AGREEMENT (NOTES) Recorded Jun 1, 2018
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 046366/0014 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2018
From: GOTTIN, VINÍCIUS MICHEL; DIAS, JONAS F.; MENASCHÉ, DANIEL SADOC; BORDIGNON, ALEX LAIER; CIARLINI, ANGELO E. M.
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 045622/0675 →
Continuity (1)
Related Publication 20190325304A1 · Oct 24, 2019
Cited By (1)
US 12,645,985