IP Library › Granted Patent US 12,596,594
Granted Patent B2
US 12,596,594 · App. 18/345,827 · Granted Apr 7, 2026

Reinforcement learning policy serving and training framework in production cloud systems

Inventors: Haoran Qiu (Champaign, IL); Chen Wang (Chappaqua, NY); Alaa S. Youssef (Valhalla, NY); Hubertus Franke (Cortlandt Manor, NY); Ravishankar K. Iyer (Champaign, IL); Zbigniew Tomasz Kalbarczyk (Urbana, IL)
Assignee: International Business Machines Corporation
G06F9/5077G06F11/3442
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,594
App. No.
18/345,827
Filed
Jun 30, 2023
Granted
Apr 7, 2026
Kind
B2
Examiner
KIM, DONG U
Art Unit
2197
USPC
718/1
Abstract

Method and systems for online training management of reinforcement learning policy serving for cloud computing systems are discloses. An example method includes controlling a cloud computing system using a first reinforcement learning (RL) model; training the first RL model to generate a second RL model in response to one or more first criteria being satisfied; and controlling the cloud computing system using the second RL model in response to one or more second criteria being satisfied.

Claims (73)

1 . A method, comprising:

controlling resources of a cloud computing system using a first reinforcement learning (RL) model, the first RL model performs tasks to automate control of the resources while refraining from performing online training of the first RL model, wherein the first RL model, for each of the tasks, receives state information from the cloud computing system, returns an action to the cloud computing system, receives reward information from the cloud computing system, and records a first trajectory comprising the state, action, and reward for each of the tasks;

monitoring at least one first trajectory generated while using the first RL model to identify a first trigger to start online training of the first RL model;

performing online training of the first RL model to generate a second RL model in response to the first trigger;

monitoring one or more second trajectories generated while performing the online training of the first RL model, to identify a second trigger to stop the online training of the first RL model; and

controlling the resources of the cloud computing system using the second RL model in response to the second trigger while refraining from performing online training of the second RL model.

2 . The method of claim 1 , wherein monitoring the at least one first trajectory further comprises evaluating, based on the at least one first trajectory, one or more first properties of the reward information received from the cloud computing system for each of the tasks, the one or more first properties of the reward information comprises one or more performance metrics or utilization metrics associated with the action for each task; and

wherein monitoring the one or more second trajectories further comprises evaluating, based on the one or more second trajectories, one or more second properties of the reward information received from the cloud computing system for each of the tasks, the one or more second properties of the reward information comprises one or more performance metrics or utilization metrics associated with the action for each task.

3 . The method of claim 2 , wherein the first trigger to start online training of the first RL model is identified based on the one or more performance metrics or utilization metrics of the one or more first properties of the reward information being less than a first threshold; and

wherein the second trigger to stop online training of the first RL model is identified based on the one or more performance metrics or utilization metrics of the one or more second properties of the reward information being less than or equal to a second threshold or being greater than or equal to a third threshold.

4 . The method of claim 3 , wherein:

the one or more performance metrics or utilization metrics of the one or more first properties of the reward information include an average reward value among a number of past trajectories generated using the first RL model and wherein the one or more performance metrics or utilization metrics of the one or more second properties of the reward information include a variance of reward values among a number of past trajectories generated using the first RL model during the online training or an average reward value among the number of past trajectories generated using the first RL model during the online training.

5 . The method of claim 2 , further comprising:

storing the at least one first trajectory and the one or more second trajectories in a database;

wherein evaluating the one or more first properties comprises evaluating the one or more first properties via the database; and

wherein evaluating the one or more second properties comprises evaluating the one or more second properties via the database.

6 . The method of claim 1 , further comprising:

outputting a first indication to start the online training the first RL model in response to the first trigger; and

outputting a second indication to stop the online training the first RL model in response to the second trigger.

7 . The method of claim 1 , wherein controlling resources of the cloud computing system comprises:

controlling resource autoscaling associated with the resources of the cloud computing system;

controlling power management associated with the resources of the cloud computing system;

controlling load balancing associated with the resources of the cloud computing system;

controlling congestion control associated with the resources of the cloud computing system;

controlling database query performance associated with the resources of the cloud computing system; or

a combination thereof.

8 . The method of claim 1 , wherein performing the online training of the first RL model comprises performing online training of the first RL model while using the first RL model to perform the tasks to automate control of the resources.

9 . A system, comprising:

a memory; and

one or more processors coupled to the memory, the one or more processors being configured to:

control resources of a cloud computing system using a first reinforcement learning (RL) model, the first RL model performs tasks to automate control of the resources while refraining from performing online training of the first RL model, wherein the first RL model, for each of the tasks, receives state information from the cloud computing system, returns an action to the cloud computing system, receives reward information from the cloud computing system, and records a first trajectory comprising the state, action, and reward for each of the tasks;

monitor at least one first trajectory generated while using the first RL model, to identify a first trigger to start online training of the first RL model;

perform online training of the first RL model to generate a second RL model in response to the first trigger;

monitor one or more second trajectories generated while performing the online training of the first RL model, to identify a second trigger to stop the online training of the first RL model; and

control the resources of the cloud computing system using the second RL model in response to the second trigger, while refraining from performing online training of the second RL model.

10 . The system of claim 9 , wherein the one or more processors are further configured to:

evaluate, based on the at least one first trajectory, one or more first properties of the reward information received from the cloud computing system for each of the tasks, the one or more first properties of the reward information comprises one or more performance metrics or utilization metrics associated with the action for each task; and

evaluate, based on the one or more second trajectories, one or more second properties of the reward information received from the cloud computing system for each of the tasks, the one or more second properties of the reward information comprises one or more performance metrics or utilization metrics associated with the action for each task.

11 . The system of claim 10 , wherein the first trigger to start online training of the first RL model is identified based on the one or more performance metrics or utilization metrics of the one or more first properties of the reward information

being less than a first threshold; and

wherein the second trigger to stop online training of the first RL model is identified based on the one or more performance metrics or utilization metrics of the one or more second properties of the reward information being less than or equal to a second threshold, or being greater than or equal to a third threshold.

12 . The system of claim 11 , wherein:

the one or more performance metrics or utilization metrics of the one or more first properties of the reward information include an average reward value among a number of past trajectories generated using the first RL model; wherein the one or more performance metrics or utilization metrics of the one or more second properties of the reward information include a variance of reward values among a number of past trajectories generated using the first RL model during the online training; or an average reward value among the number of past trajectories generated using the first RL model during the online training.

13 . The system of claim 10 , wherein:

the one or more processors are further configured to store the at least one first trajectory and the one or more second trajectories in a database;

the one or more processors are further configured to evaluate the one or more first properties via the database; and

the one or more processors are further configured to evaluate the one or more second properties via the database.

14 . The system of claim 9 , wherein the one or more processors are further configured to:

output a first indication to start the online training the first RL model in response to the first trigger; and

output a second indication to stop the online training the first RL model in response to the second trigger.

15 . The system of claim 9 , wherein to control resources of the cloud computing system, the one or more processors are further configured to:

control resource autoscaling associated with the resources of the cloud computing system;

control power management associated with the resources of the cloud computing system;

control load balancing associated with the resources of the cloud computing system;

control congestion control associated with the resources of the cloud computing system;

control database query performance associated with the resources of the cloud computing system; or

a combination thereof.

16 . The system of claim 9 , wherein to perform online training of the first RL model, the one or more processors are further configured to perform the online training of the first RL model while using the first RL model to perform the tasks to automate control of the resources.

17 . A computer program product for online training management, the computer program product comprising:

a computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code executable by one or more computer processors to:

control resources of a cloud computing system using a first reinforcement learning (RL) model, the first RL model performs tasks to automate control of the resources while refraining from performing online training of the first RL model, wherein the first RL model, for each of the tasks, receives state information from the cloud computing system, returns an action to the cloud computing system, receives reward information from the cloud computing system, and records a first trajectory comprising the state, action, and reward for each of the tasks;

monitor at least one first trajectory generated while using the first RL model, to identify a first trigger to start online training of the first RL model;

perform online training of the first RL model to generate a second RL model in response to the first trigger;

monitor one or more second trajectories generated while performing the online training of the first RL model, to identify a second trigger to stop the online training of the first RL model; and

control the resources of the cloud computing system using the second RL model in response to the second trigger, while refraining from performing online training of the second RL model.

18 . The computer program product of claim 17 , wherein the computer-readable program code being further executable by the one or more computer processors to:

evaluate, based on the at least one first trajectory, one or more first properties of the reward information received from the cloud computing system for each of the tasks, the one or more first properties of the reward information comprises one or more performance metrics or utilization metrics associated with the action for each task; and

evaluate, based on the one or more second trajectories, one or more second properties of the reward information received from the cloud computing system for each of the tasks, the one or more second properties of the reward information comprises one or more performance metrics or utilization metrics associated with the action for each task.

19 . The computer program product of claim 18 , wherein the first trigger to start online training of the first RL model is identified based on the one or more performance metrics or utilization metrics of the one or more first properties of the reward information

being less than a first threshold; and

wherein the second trigger to stop online training of the first RL model is identified based on the one or more performance metrics or utilization metrics of the one or more second properties of the reward information being less than or equal to a second threshold, or being greater than or equal to a third threshold.

20 . The computer program product of claim 19 , wherein:

the one or more performance metrics or utilization metrics of the one or more first properties of the reward information include an average reward value among a number of past trajectories generated using the first RL model; wherein the one or more performance metrics or utilization metrics of the one or more second properties of the reward information include a variance of reward values among a number of past trajectories generated using the first RL model during the online training, or an average reward value among the number of past trajectories generated using the first RL model during the online training.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: QIU, HAORAN; IYER, RAVISHANKAR K.; KALBARCZYK, ZBIGNIEW TOMASZ
To: THE BOARD OF TRUSTEES OF THE UNIVERSITY OF ILLINOIS
Reel/Frame 064916/0848 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2023
From: WANG, CHEN; YOUSSEF, ALAA S.; FRANKE, HUBERTUS
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064132/0773 →
Continuity (1)
Related Publication 20250004858A1 · Jan 2, 2025
References Cited (34)
US 11250313B2 · Dasgupta et al. · 2022 [cited by applicant]
US 11468334B2 · Chaudhury et al. · 2022 [cited by applicant]
US 12192820B2 · Yeh · 2025 [cited by examiner]
US 20150178638A1 · Deshpande et al. · 2015 [cited by applicant]
US 20170032279A1 · Miserendino et al. · 2017 [cited by applicant]
US 20190385061A1 · Chaudhury et al. · 2019 [cited by applicant]
US 20200242449A1 · Dasgupta et al. · 2020 [cited by applicant]
US 20200285503A1 · Dou · 2020 [cited by examiner]
US 20220188690A1 · Rawat · 2022 [cited by examiner]
US 20220343117A1 · Jeong · 2022 [cited by examiner]
US 20220374704A1 · Feng · 2022 [cited by examiner]
US 20220390909A1 · Kubota et al. · 2022 [cited by applicant]
US 20230164817A1 · Bhamri · 2023 [cited by examiner]
US 20230388856A1 · Yan · 2023 [cited by examiner]
US 20230409387A1 · Gupta · 2023 [cited by examiner]
US 20230409393A1 · Misra · 2023 [cited by examiner]
US 20240007414A1 · Jain · 2024 [cited by examiner]
US 20240112065A1 · Rezaeian · 2024 [cited by examiner]
US 20240355318A1 · Beaver · 2024 [cited by examiner]
US 20240422650A1 · Hashmi · 2024 [cited by examiner]
WO 2017201107A1 · 2017 [cited by applicant]
Siqiao Xue et al., “A Meta Reinforcement Learning Approach for Predictive Autoscaling in the Cloud,” arXiv.org, Dated: May 31, 2022, pp. 1-10. [cited by applicant]
Hamid Arabnejad et al., “A Comparison of Reinforcement Learning Techniques for Fuzzy Cloud Auto-Scaling,” arXiv.org, Dated: May 19, 2017, pp. 1-11. [cited by applicant]
Yisel Gari et al., “Reinforcement Learning-based Application Autoscaling in the Cloud: a Survey,” arXiv.org, Dated: Nov. 17, 2020, pp. 1-40. [cited by applicant]
Jason Gauci et al., “Horizon: Facebook's Open Source Applied Reinforcement Learning Platform,” arXiv.org, Dated: Sep. 4, 2019, pp. 1-10. [cited by applicant]
Shumpei Kubosawa et al., (Sep. 2021). Non-steady-state Control under Disturbances: Navigating Plant Operation via Simulation-Based Reinforcement Learning. In 2021 60th Annual Conference of the Society of Instrument and … [cited by applicant]
Peizheng Li et al., “RLOps: Development Life-cycle of Reinforcement Learning Aided Open RAN,” arXiv.org, Dated: Nov. 25, 2022, pp. 1-17. [cited by applicant]
Feng Liu et al., (Jul. 2005). Neural network model for time series prediction by reinforcement learning. In Proceedings. 2005 IEEE International Joint Conference on Neural Networks, 2005. (vol. 2, pp. 809-814). IEEE. [cited by applicant]
Jose Antonio Martin H., “Reinforcement Learning in System Identification,” arXiv.org, Dated: Dec. 14, 2022, pp. 1-21. [cited by applicant]
Marcel Panzer & Benedict Bender (2022) Deep reinforcement learning in production systems: a systematic literature review, International Journal of Production Research, 60:13, 4316-4341, DOI: 10.1080/00207543.2021.197313… [cited by applicant]
Lucia Schuler et al., “AI-based Resource Allocation: Reinforcement Learning for Adaptive Auto-scaling in Serverless Environments,” arXiv.org, Dated: May 29, 2020, pp. 1-8. [cited by applicant]
Spathis, Dimitrios. Machine learning to model health with multimodal mobile sensor data. Diss. University of Cambridge, Year: 2022, pp. 1-172. [cited by applicant]
Zhengjie Sun et al., “Cloud-Edge Collaboration in Industrial Internet of Things: a Joint Offloading Scheme Based on Resource Prediction,” IEEE Internet of Things Journal, 9(18), Year: 2021, pp. 17014-17025 (Abstract Onl… [cited by applicant]
Wei, Yi, et al. “A reinforcement learning based auto-scaling approach for SaaS providers in dynamic cloud environment.” Mathematical Problems in Engineering 2019, Year: 2019, pp. 1-2. [cited by applicant]