IP Library Granted Patent US 12,632,282
Granted Patent B2
US 12,632,282 · App. 18/459,884 · Granted May 19, 2026

Method and system for latency optimization in a distributed compute network

Inventor: Anton Zvonko Gazvoda (Ljubljana, SI)
Assignee: BunnyWay Informacijske storitve d.o.o
G06F9/45558G06N20/00H04L41/0895H04L43/0852G06F2009/4557
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,282
App. No.
18/459,884
Granted
May 19, 2026
Kind
B2
Abstract

There is provided a computer-implemented method of provisioning resources in a distributed compute network comprising one or more routing nodes and one or more compute nodes configured to host one or more virtual application instances of an application thereon, the method being performed by at least one hardware processor and comprising: receiving, by a system manager, operational parameter data for each available routing node for a current state of the distributed compute network, wherein the operational parameter data comprises measured and/or predicted values of the local latency between a respective routing node and any available compute nodes accessible by the respective routing node; determining, by the system manager, an optimized global latency value for the current state of the distributed compute network based on the received operational parameter data; defining a latency threshold for a global latency for the application based on the optimized global latency value; generating, utilizing a trained machine learning model, a proposed new state of the distributed compute network having a global latency for the application which meets or exceeds the global latency threshold; and implementing the proposed new state on the distributed compute network by selecting and/or deselecting one or more compute nodes for provisioning of virtual application instances of the application.

Claims (41)

1 . A computer-implemented method of provisioning resources in a distributed compute network comprising one or more routing nodes and one or more compute nodes configured to host one or more virtual application instances of an application thereon, the method being performed by at least one hardware processor and comprising:

a) receiving, by a system manager, operational parameter data for each available routing node for a current state of the distributed compute network, wherein the operational parameter data comprises measured and/or predicted values of a local latency between a respective routing node and any available compute nodes accessible by the respective routing node;

b) determining, by the system manager, an optimized global latency value for the current state of the distributed compute network based on the received operational parameter data;

c) defining a latency threshold for a global latency for the application based on the optimized global latency value;

d) generating, utilizing a trained machine learning model, a proposed new state of the distributed compute network having a global latency for the application which meets or exceeds the latency threshold; and

e) implementing the proposed new state on the distributed compute network by selecting and/or deselecting one or more compute nodes for provisioning of virtual application instances of the application.

2 . A computer-implemented method according to claim 1 , wherein the global latency of the application is a function of the local latencies of any provisioned virtual application instances.

3 . A computer-implemented method according to claim 1 , wherein the optimized global latency value is a function of the local latency values for each available routing node between a respective routing node and an available compute node accessible by the respective routing node having a lowest local latency value.

4 . A computer-implemented method according to claim 3 , wherein the optimized global latency value comprises an averaged sum of the local latency values for each available routing node between a respective routing node and the available compute node accessible by the respective routing node having the lowest latency.

5 . A computer-implemented method according to claim 1 , wherein step c) further comprises defining one or more further latency thresholds for the global latency for the application based on the optimized global latency value.

6 . A computer-implemented method according to claim 1 , wherein the trained machine learning model is trained using reinforcement-learning.

7 . A computer-implemented method according to claim 1 , wherein step c) further comprises:

f) proposing one or more actions to the current state to generate a proposed new state, the one or more actions comprising selecting and/or deselecting one or more compute nodes for provisioning of virtual application instances of the application;

g) determining whether the global latency of the proposed new state meets or exceeds the latency threshold and, if so determined, implementing the proposed new state at step e).

8 . A computer-implemented method according to claim 7 , wherein if, at step g) the global latency of the proposed new state does not meet or exceed the latency threshold, the method further comprises:

h) Iteratively repeating steps f) and g) until the latency threshold is met.

9 . A system for provisioning resources in a distributed compute network comprising one or more routing nodes and one or more compute nodes configured to host one or more virtual application instances of an application thereon, the system comprising:

at least one hardware processor operable to perform the steps of:

a) receiving, by a system manager, operational parameter data for each available routing node for a current state of the distributed compute network, wherein the operational parameter data comprises measured and/or predicted values of a local latency between a respective routing node and any available compute nodes accessible by the respective routing node;

b) determining, by the system manager, an optimized global latency value for the current state of the distributed compute network based on the received operational parameter data;

c) defining a latency threshold for a global latency for the application based on the optimized global latency value;

d) generating, utilizing a trained machine learning model, a proposed new state of the distributed compute network having a global latency for the application which meets or exceeds the latency threshold; and

e) implementing the proposed new state on the distributed compute network by selecting and/or deselecting one or more compute nodes for provisioning of virtual application instances of the application.

10 . A system according to claim 9 , wherein the global latency of the application is a function of the local latencies of any provisioned virtual application instances.

11 . A system according to claim 9 , wherein the optimized global latency value is a function of the local latency values for each available routing node between a respective routing node and an available compute node accessible by the respective routing node having a lowest local latency value.

12 . A system according to claim 11 , wherein the optimized global latency value comprises an averaged sum of the local latency values for each available routing node between a respective routing node and the available compute node accessible by the respective routing node having the lowest latency.

13 . A system according to claim 9 , wherein step c) further comprises defining one or more further latency thresholds for the global latency for the application based on the optimized global latency value.

14 . A system according to claim 9 , wherein the trained machine learning model is trained using reinforcement-learning.

15 . A system according to claim 9 , wherein step c) further comprises:

f) proposing one or more actions to the current state to generate a proposed new state, the one or more actions comprising selecting and/or deselecting one or more compute nodes for provisioning of virtual application instances of the application;

g) determining whether the global latency of the proposed new state meets or exceeds the latency threshold and, if so determined, implementing the proposed new state at step e).

16 . A computer-implemented method according to claim 7 , wherein if, at step g) the global latency of the proposed new state does not meet or exceed the latency threshold, the method further comprises:

h) Iteratively repeating steps f) and g) until the latency threshold is met.

17 . A non-transitory computer readable storage medium storing a program of instructions executable by a machine to perform a computer-implemented method of provisioning resources in a distributed compute network comprising one or more routing nodes and one or more compute nodes configured to host one or more virtual application instances of an application thereon, the method comprising:

a) receiving, by a system manager, operational parameter data for each available routing node for a current state of the distributed compute network, wherein the operational parameter data comprises measured and/or predicted values of a local latency between a respective routing node and any available compute nodes accessible by the respective routing node;

b) determining, by the system manager, an optimized global latency value for the current state of the distributed compute network based on the received operational parameter data;

c) defining a latency threshold for a global latency for the application based on the optimized global latency value;

d) generating, utilizing a trained machine learning model, a proposed new state of the distributed compute network having a global latency for the application which meets or exceeds the latency threshold; and

e) implementing the proposed new state on the distributed compute network by selecting and/or deselecting one or more compute nodes for provisioning of virtual application instances of the application.

18 . A non-transitory computer readable storage medium according to claim 17 , wherein the optimized global latency value is a function of the local latency values for each available routing node between a respective routing node and an available compute node accessible by the respective routing node having a lowest local latency value.

19 . A non-transitory computer readable storage medium according to claim 18 , wherein the optimized global latency value comprises an averaged sum of the local latency values for each available routing node between a respective routing node and the available compute node accessible by the respective routing node having the lowest latency.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2023
From: GAZVODA, ANTON ZVONKO
To: BUNNYWAY INFORMACIJSKE STORITVE D.O.O
Reel/Frame 064776/0436 →
Continuity (1)
Related Publication 20250077257A1 · Mar 6, 2025
References Cited (24)
US 10320610B2 · Matni · 2019 [cited by examiner]
US 10979534B1 · Parulkar · 2021 [cited by examiner]
US 11132226B2 · Jadhav · 2021 [cited by examiner]
US 11368517B2 · Beveridge · 2022 [cited by examiner]
US 20130297802A1 · Laribi · 2013 [cited by examiner]
US 20150163162A1 · DeCusatis · 2015 [cited by examiner]
US 20220114033A1 · Arvinte · 2022 [cited by examiner]
US 20230297433A1 · Mishra · 2023 [cited by examiner]
US 20240251376A1 · Cheung · 2024 [cited by examiner]
Zhang et al., “A-SARSA: A Predictive Container Auto-Scaling Algorithm Based on Reinforcement Learning,” 2020 IEEE International Conference on Web Services (2020). [cited by applicant]
Rossi et al., “Geo-distributed efficient deployment of containers with Kubernetes,” Computer Communications 159: 161-174 (2020). [cited by applicant]
Lara Lorna Jiménez, “Decentralized Location-aware Orchestration of Containerized Microservice Applications,” Doctoral Thesis, Lulea University of Technology, Department of Computer Science and Electrical Engineering, Re… [cited by applicant]
Ma et al., “Location-Aware Cloud Service Brokering in Multi-cloud Environment” In: Troya, J., et al. Service-Oriented Computing—ICSOC 2022 Workshops, Lecture Notes in Computer Science, vol. 13821, Retrieved from https./… [cited by applicant]
Knob et al., “Improving Container Deployment in Edge Computing Using the Infrastructure Aware Scheduling Algorithm,” IEEE Symposium on Computers and Communications, Retrieved from https://ieeexplore.ieee org/document/96… [cited by applicant]
Kaur et al., “Live migration of containerized microservices between remote Kubernetes Clusters,” Infocom Workshops, IEEE International Conference on Computer Communications, Retrieved from https://hal.science/hal-034667… [cited by applicant]
Santos et al., “gym-hpa: Efficient Auto-Scaling via Reinforcement Learning for Complex Microservice-based Applications in Kubernetes,” NOMS 2023-2023 IEEE/IFIP Network Operations and Management Symposium, Retrieved from… [cited by applicant]
Pelle et al., “Cost and Latency Optimized Edge Computing Platform,” Electronics 11(561):1-24 (2022). [cited by applicant]
Wang et al., “Container Scaling Strategy Based on Reinforcement Learning,” Security and Communication Networks, 2023(7400235):1-10 (2023). [cited by applicant]
Daniel Edsinger, “Auto-scaling cloud infrastructure with Reinforcement Learning,” Masters Thesis, Chalmers University of Technology, Department of Computer Science and Engineering, Retrieved from https://odr.chalmers.se… [cited by applicant]
Schuler et al., “AI-based Resource Allocation: Reinforcement Learning for Adaptive Auto-scaling in Serverless Environments,” IEEE/ACM 21st International Symposium on Cluster, Cloud and Internet Computing (CCGrid), pp. 8… [cited by applicant]
Gari et al., “Reinforcement Learning-based Application Autoscaling in the Cloud: A Survey,” arXiv:2001.09957, Retrieved from https://doi.org/10.48550/arXiv.2001.09957 (2021). [cited by applicant]
Xue et al., “A Meta Reinforcement Learning Approach for Predictive Autoscaling in the Cloud,” arXiv:2205.15795v1, Retrieved from https://doi.org/10.48560/arXiv.2205 15795 (2022). [cited by applicant]
Cheng et al., “DRL-cloud: Deep reinforcement learning-based resource provisioning and task scheduling for cloud service providers,” ASPDAC '18: Proceedings of the 23rd Asia and South Pacific Design Automation Conference… [cited by applicant]
Xiao et al., “DScaler: A Horizontal Autoscaler of Microservice Based on Deep Reinforcement Learning,” 23rd Asia-Pacific Network Operations and Management Symposium (APNOMS), pp. 1-6, Retrieved from https://ieeexplore.ie… [cited by applicant]