IP Library › Granted Patent US 12,580,824
Granted Patent B2
US 12,580,824 · App. 18/541,996 · Granted Mar 17, 2026

Intelligent management of machine learning inference in edge-cloud systems

Inventors: Sharath Gopal (Fremont, CA); Baolin Li (Saugus, MA); Marcus Gualtieri (Fremont, CA); Xiaowei Zhou (Fremont, CA); Liu Ren (Saratoga, CA)
Assignee: Robert Bosch GmbH
H04L41/16H04L41/0826
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,580,824
App. No.
18/541,996
Granted
Mar 17, 2026
Kind
B2
Abstract

A system and method relate to managing a cloud computing system. Queries are received from one or more edge devices of a set of edge devices. Each query includes sensor data from the respective edge device. Prediction data is generated via one or more cloud machine learning models, using the sensor data, during a current time period. System state data is generated and indicates a current state of an environment during the current time period. The environment is defined by the cloud computing system and the set of edge devices. A machine learning system generates policy data by optimizing an expected return of a reward with respect to taking a particular action given the system state data. The machine learning system is employed by the cloud computing system. The policy data indicates a recommended action from a set of actions. The cloud computing system performs the recommended action.

Claims (85)

1 . A computer-implemented method for managing a cloud computing system, the computer-implemented method comprising:

receiving queries from one or more edge devices of a set of edge devices that are connected to the cloud computing system, each query including sensor data from the respective edge device;

generating prediction data, via one or more machine learning models, using the sensor data, the one or more machine learning models being employed by the cloud computing system during a current time period;

generating system state data indicating a current state of an environment during the current time period, the environment being defined by the cloud computing system and the set of edge devices;

generating, via a reinforcement learning (RL) agent, policy data by optimizing an expected return of a reward with respect to taking a particular action given the system state data, the policy data indicating a recommended action from a set of actions; and

performing the recommended action,

wherein the cloud computing system employs the RL agent as a machine learning system to control an offload tendency of transmitting a respective query of a machine learning inference task from one or more of the edge devices to the cloud computing system via the policy data.

2 . The computer-implemented method of claim 1 , wherein:

the machine learning system comprises a deep neural network (DNN) that is configured to optimize the expected return by approximating a Q-value for each action of the set of actions; and

the recommended action returns a particular action having the best Q-value from among the set of actions.

3 . The computer-implemented method of claim 1 , wherein the system state data is generated for the current time period based on

a number of the edge devices being serviced by the cloud computing system;

a change in the number of edge devices between a previous time period and the current time period,

an amount of computer resources that are being used to process the queries;

a number of the queries being processed by the cloud computing system;

latency data regarding the queries being processed by the cloud computing system;

cost data of the cloud computing system incurred by processing the queries from the previous time period to the current time period; and

query threshold data being used by each edge device to determine whether or not to generate a respective query for transmission to the cloud computing system.

4 . The computer-implemented method of claim 1 , wherein the set of actions comprises

allocating one or more computer resources of the cloud computing system,

deallocating the one or more computer resources,

maintaining the one or more computer resources, and

modifying a query threshold of the cloud computing system.

5 . The computer-implemented method of claim 1 , further comprising:

generating query threshold data of the cloud computing system for the current time period, the query threshold data being used by each edge device to determine whether or not to generate a respective query for transmission to the cloud computing system; and

transmitting the query threshold data to each edge device of the set of edge devices.

6 . The computer-implemented method of claim 1 , wherein the reward is computed such that a predetermined cost budget is not exceeded by cost data associated with processing the queries on the cloud computing system and additional cost data associated with performing the particular action.

7 . The computer-implemented method of claim 1 , wherein the reward is computed such that a predetermined latency target is not exceeded by latency data associated with processing the queries.

8 . A system comprising:

one or more processors; and

one or more memory in data communication with the one or more processors, the one or more memory including computer readable data stored thereon that, when executed by the one or more processors, causes the one or more processors to perform a method for managing a cloud computing system, the method including

receiving queries from one or more edge devices of a set of edge devices that are connected to the cloud computing system, each query including sensor data from the respective edge device;

generating prediction data, via one or more machine learning models, using the sensor data, the one or more machine learning models being employed by the cloud computing system during a current time period;

generating system state data indicating a current state of an environment during the current time period, the environment being defined by the cloud computing system and the set of edge devices;

generating, via a reinforcement learning (RL) agent, policy data by optimizing an expected return of a reward with respect to taking a particular action given the system state data, the policy data indicating a recommended action from a set of actions; and

performing the recommended action,

wherein the cloud computing system employs the RL agent as a machine learning system to control an offload tendency of transmitting a respective query of a machine learning inference task from one or more of the edge devices to the cloud computing system via the policy data.

9 . The system of claim 8 , wherein:

the machine learning system comprises a deep neural network (DNN) that is configured to optimize the expected return by approximating a Q-value for each action of the set of actions; and

the recommended action returns a particular action having the best Q-value from among the set of actions.

10 . The system of claim 8 , wherein the system state data is generated for the current time period based on

a number of the edge devices being serviced by the cloud computing system;

a change in the number of edge devices between a previous time period and the current time period,

an amount of computer resources that are being used to process the queries;

a number of the queries being processed by the cloud computing system;

latency data regarding the queries being processed by the cloud computing system;

cost data of the cloud computing system incurred by processing the queries from the previous time period to the current time period; and

query threshold data being used by each edge device to determine whether or not to generate a respective query for transmission to the cloud computing system.

11 . The system of claim 8 , wherein the set of actions comprises

allocating one or more computer resources of the cloud computing system,

deallocating the one or more computer resources,

maintaining the one or more computer resources, and

modifying a query threshold of the cloud computing system.

12 . The system of claim 8 , wherein the method further comprises:

generating query threshold data of the cloud computing system for the current time period, the query threshold data being used by each edge device to determine whether or not to generate a respective query for transmission to the cloud computing system; and

transmitting the query threshold data to each edge device of the set of edge devices.

13 . The system of claim 8 , wherein the reward is computed such that a predetermined cost budget is not exceeded by cost data associated with processing the queries on the cloud computing system and additional cost data associated with performing the particular action.

14 . The system of claim 8 , wherein the reward is computed such that a predetermined latency target is not exceeded by latency data associated with processing the queries.

15 . One or more non-transitory computer-readable media that store instructions that, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising:

receiving queries from one or more edge devices of a set of edge devices that are connected to the cloud computing system, each query including sensor data from the respective edge device;

generating prediction data, via one or more machine learning models, using the sensor data, the one or more machine learning models being employed by the cloud computing system during a current time period;

generating system state data indicating a current state of an environment during the current time period, the environment being defined by the cloud computing system and the set of edge devices;

generating, via a reinforcement learning (RL) agent, policy data by optimizing an expected return of a reward with respect to taking a particular action given the system state data, the policy data indicating a recommended action from a set of actions; and

performing the recommended action,

wherein the cloud computing system employs the RL agent as a machine learning system to control an offload tendency of transmitting a respective query of a machine learning inference task from one or more of the edge devices to the cloud computing system via the policy data.

16 . The one or more non-transitory computer-readable media of claim 15 , wherein

the machine learning system comprises a deep neural network (DNN) that is configured to optimize the expected return by approximating a Q-value for each action of the set of actions; and

the recommended action returns a particular action having the best Q-value from among the set of actions.

17 . The one or more non-transitory computer-readable media of claim 15 , wherein the system state data is generated for the current time period based on

a number of the edge devices being serviced by the cloud computing system;

a change in the number of edge devices between a previous time period and the current time period,

an amount of computer resources that are being used to process the queries;

a number of the queries being processed by the cloud computing system;

latency data regarding the queries being processed by the cloud computing system;

cost data of the cloud computing system incurred by processing the queries from the previous time period to the current time period; and

query threshold data being used by each edge device to determine whether or not to generate a respective query for transmission to the cloud computing system.

18 . The one or more non-transitory computer-readable media of claim 15 , wherein the set of actions comprises

allocating one or more computer resources of the cloud computing system,

deallocating the one or more computer resources,

maintaining the one or more computer resources, and

modifying a query threshold of the cloud computing system.

19 . The one or more non-transitory computer-readable media of claim 15 , wherein the method further comprises:

generating query threshold data of the cloud computing system for the current time period, the query threshold data being used by each edge device to determine whether or not to generate a respective query for transmission to the cloud computing system; and

transmitting the query threshold data to each edge device of the set of edge devices.

20 . The one or more non-transitory computer-readable media of claim 15 , wherein the reward is computed such that a predetermined cost budget is not exceeded by cost data associated with processing the queries on the cloud computing system and additional cost data associated with performing the particular action.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2025
From: GOPAL, SHARATH; LI, BAOLIN; GUALTIERI, MARCUS; ZHOU, XIAOWEI; REN, LIU
To: ROBERT BOSCH GMBH
Reel/Frame 070808/0295 →
Continuity (1)
Related Publication 20250202778A1 · Jun 19, 2025
References Cited (12)
US 9876851B2 · Chandramouli et al. · 2018 [cited by applicant]
US 11153177B1 · Hermoni · 2021 [cited by examiner]
US 11503113B2 · Dong · 2022 [cited by applicant]
US 20210350213A1 · Sethi · 2021 [cited by examiner]
US 20240007414A1 · Jain · 2024 [cited by examiner]
David et al., “Tensorflow Lite Micro: Embedded Machine Learning on TINYML Systems,” Proceedings of the 4th MLSys Conference, San Jose, CA, USA, 2021, pp. 1-12. [cited by applicant]
Verhelst et al., “Chapter 18: Machine Learning at the Edge, ”NANO-CHIPS 2030, 2020, pp. 293-322. [cited by applicant]
Wang et al., “Enabling Edge-Cloud Video Analytics for Robotics Applications,” DOI 10.1109/TCC.2022.3142066, IEEE Transactions on Cloud Computing, pp. 1-14. [cited by applicant]
Banitalebi-Dehkordi et al., “Auto-Split: A General Framework of Collaborative Edge-Cloud AI,” ADS Track Paper, KDD '21, Aug. 14-18, 2021, Virtual Event, Singapore, pp. 2543-2553. [cited by applicant]
Chinchali et al., “Network offloading policies for cloud robotics: a learning-based approach,” Autonomous Robots (2021) 45:997-1012, https://doi.org/10.1007/s10514-021-09987-4, pp. 997-1012. [cited by applicant]
Redmon et al., “You Only Look Once: Unified, Real-Time Object Detection,” Computer Vision Foundation, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 779-788. [cited by applicant]
Mnih et al., “Human-level control through deep reinforcement learning,” Letter, doi:10.1038/nature 14236, Nature, Feb. 26, 2015, vol. 518, 13 pages. [cited by applicant]