IP Library › Granted Patent US 12,748,615
Granted Patent B2
US 12,748,615 · App. 18/365,922 · Granted Sep 29, 2026

Predicting worker instance count for cloud-based computing platforms

Inventors: Neha Keshari (Bothell, WA); Abhisek Pan (Bellevue, WA); David Allen Dion (Bothell, WA); Brendon Machado (Seattle, WA); Karthik Subramaniam Hariharan (Seattle, WA); Karthikeyan Subramanian (Redmond, WA); Thomas Moscibroda (Bellevue, WA); Karel Trueba Nobregas (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F9/45558G06F2009/4557
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,615
App. No.
18/365,922
Granted
Sep 29, 2026
Kind
B2
Abstract

Described are examples for recommending increase in worker instance count for an availability zone in a cloud-based computing platform. A machine learning (ML) model can be used to predict a time series forecast of a workload for the availability zone in a future time period. A predicted number of worker instances to handle the predicted workload can be computed, and if the number of worker instances in the availability zone is less than the predicted number of worker instances, a recommendation to increase the number of worker instances in the availability zone can be generated.

Claims (34)

1 . A device for increasing worker instance count for an availability zone in a cloud-based computing platform, comprising:

one or more memories storing instructions; and

one or more processors coupled to the one or more memories and configured to execute the instructions to:

provide resource allocation information as input to a machine learning (ML) model to receive an output of a time series forecast of a workload for the availability zone in a future time period, wherein the resource allocation information comprises a history of resource allocation requests for the availability zone including original allocation requests, retry requests that occur when an original allocation request fails, and requests related to underlying platform issues, wherein the ML model is trained on a timespan of workload history for the availability zone such that the time series forecast weighs workload increases due to customer demand over workload increases due to system disruption;

compute a throughput threshold for the worker instances in the availability zone by creating throughput bins from a history of workload data and identifying a throughput bin value at which a number of throttling or timeout exceptions is at least a percentage;

compute a predicted number of worker instances in the availability zone for handling the workload in the future time period based at least in part on the outputted time series forecast and the computed throughput threshold;

when a number of worker instances in the availability zone is less than the predicted number of worker instances, provision a recommended number of servers for the number of worker instances in the availability zone, and process the workload with the provisioned servers at the future time period.

2 . The device of claim 1 , wherein the one or more processors are further configured to execute the instructions to predict a peak workload over the future time period based on the predicted time series forecast of the workload and a daily to minute peak workload ratio, wherein the predicted number of worker instances is based on the peak workload.

3 . The device of claim 2 , wherein the one or more processors are further configured to execute the instructions to compute a throughput threshold for the number of worker instances, wherein the predicted number of worker instances is based on a comparison of the throughput threshold with the peak workload.

4 . The device of claim 3 , wherein the one or more processors are further configured to execute the instructions to compute the throughput threshold based on one or more of a throttling metric, a total allocation time, or a number of timeout exceptions observed from a history of workload data for the availability zone.

5 . The device of claim 1 , wherein the time series forecast of the workload is based on a history of workload data, statistical trend of the workload data, and one or more properties of a time period related to the workload data.

6 . The device of claim 1 , wherein the time series forecast is based on a time series prediction model of historical workload data for the availability zone.

7 . The device of claim 1 , wherein the time series forecast is based on fitting an empirical statistical distribution of a history of resource allocation requests of the availability zone.

8 . The device of claim 1 , wherein the time series forecast for the availability zone is based on the one or more processors execute the instructions to determine the availability zone is one of multiple availability zones having a threshold confidence for accuracy of predicting the time series forecast.

9 . The device of claim 1 , wherein the time series forecast is based on the history of resource allocation requests.

10 . The device of claim 1 , wherein the one or more processors are further configured to execute the instructions to provide, to the ML model, recent workload data for the availability zone and associated performance metrics for use in predicting subsequent time series forecasts of the workload for the availability zone.

11 . The device of claim 1 , wherein the one or more processors are further configured to execute the instructions to generate the recommendation to increase the number of worker instances further based on one or more of a customer priority of a customer corresponding to the workload, a scalability of the cloud-based computing platform, or a number of throttling failures during deployment of virtual machines in the availability zone.

12 . A method for increasing worker instance count for an availability zone in a cloud-based computing platform, comprising:

predicting, using a machine learning (ML) model and by providing resource allocation information as input to the ML model, a time series forecast of a workload for the availability zone in a future time period, wherein the resource allocation information comprises a history of resource allocation requests for the availability zone including original allocation requests, retry requests that occur when an original allocation request fails, and requests related to underlying platform issues, wherein the ML model is trained on a timespan of workload history for the availability zone such that the time series forecast weighs workload increases due to customer demand over workload increases due to system disruption;

computing a throughput threshold for the worker instances in the availability zone by creating throughput bins from a history of workload data and identifying a throughput bin value at which a number of throttling or timeout exceptions is at least a percentage;

determining, based on the throughput threshold, whether a number of worker instances in the availability zone is sufficient for satisfying the workload in the future time period; and

based on determining that the number of worker instances is not sufficient, provisioning a recommended number of servers for the number of worker instances in the availability zone, and processing the workload with the provisioned servers at the future time period.

13 . The method of claim 12 , further comprising predicting a peak workload over the future time period based on the predicted time series forecast of the workload and a daily to minute peak workload ratio, wherein determining whether the number of worker instances is sufficient is based on the peak workload.

14 . The method of claim 13 , further comprising computing a throughput threshold for the number of worker instances, wherein determining whether the number of worker instances is sufficient based on a comparison of the throughput threshold with the peak workload.

15 . The method of claim 14 , wherein computing the throughput threshold is based on one or more of a throttling metric, a total allocation time, or a number of timeout exceptions observed from a history of workload data for the availability zone.

16 . The method of claim 12 , wherein predicting the time series forecast of the workload is based on one or more of a history of workload data, statistical trend of the workload data, and one or more properties of a time period related to the workload data.

17 . The method of claim 12 , wherein predicting the time series forecast is based on a history of resource allocation requests including initial requests, retries, and requests due to throttled workload.

18 . The method of claim 12 , wherein generating the recommendation to increase the number of worker instances is further based on one or more of a customer priority of a customer corresponding to the workload, a scalability of the cloud-based computing platform, or a number of throttling failures during deployment of virtual machines in the availability zone.

19 . A non-transitory computer-readable device storing instructions thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations for increasing worker instance count for an availability zone in a cloud-based computing platform, comprising:

predicting, using a machine learning (ML) model and by providing resource allocation information as input to the ML model, a time series forecast of a workload for the availability zone in a future time period, wherein the resource allocation information comprises a history of resource allocation requests for the availability zone including original allocation requests, retry requests that occur when an original allocation request fails, and requests related to underlying platform issues, wherein the ML model is trained on a timespan of workload history for the availability zone such that the time series forecast weighs workload increases due to customer demand over workload increases due to system disruption;

computing a throughput threshold for the worker instances in the availability zone by creating throughput bins from a history of workload data and identifying a throughput bin value at which a number of throttling or timeout exceptions is at least a percentage;

determining, based on the throughput threshold, whether a number of worker instances in the availability zone is sufficient for satisfying the workload in the future time period; and

based on determining that the number of worker instances is not sufficient, provisioning a recommended number of servers for the number of worker instances in the availability zone, and processing the workload with the provisioned servers at the future time period.

20 . The non-transitory computer-readable device of claim 19 , the operations further comprising predicting a peak workload over the future time period based on the predicted time series forecast of the workload and a daily to minute peak workload ratio, wherein determining whether the number of worker instances is sufficient is based on the peak workload.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2023
From: KESHARI, NEHA; PAN, ABHISEK; DION, DAVID ALLEN; MACHADO, BRENDON; HARIHARAN, KARTHIK SUBRAMANIAM; SUBRAMANIAN, KARTHIKEYAN; MOSCIBRODA, THOMAS; NOBREGAS, KAREL TRUEBA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 064541/0913 →
Continuity (1)
Related Publication 20250045088A1 · Feb 6, 2025
References Cited (26)
US 11113120B1 · Greenfield et al. · 2021 [cited by applicant]
US 20110029970A1 · Arasaratnam · 2011 [cited by examiner]
US 20150207851A1 · Nampally · 2015 [cited by examiner]
US 20170199564A1 · Saxena et al. · 2017 [cited by applicant]
US 20180302340A1 · Alvarez Callau et al. · 2018 [cited by applicant]
US 20190243686A1 · LaBute · 2019 [cited by examiner]
US 20190342184A1 · May · 2019 [cited by applicant]
US 20200074323A1 · Martin · 2020 [cited by examiner]
US 20220253790A1 · Volkov et al. · 2022 [cited by applicant]
US 20230071347A1 · Hudis et al. · 2023 [cited by applicant]
US 20230153165A1 · Higginson · 2023 [cited by examiner]
US 20230214308A1 · Chen · 2023 [cited by applicant]
US 20230342658A1 · Tripathi · 2023 [cited by examiner]
US 20230401089A1 · Armangau · 2023 [cited by examiner]
US 20240289111A1 · Kaveri Poompatnam Chandrasekaran · 2024 [cited by examiner]
US 20250021388A1 · Agrahari · 2025 [cited by examiner]
CN 108270805B · 2021 [cited by applicant]
KR 20210149576A · 2021 [cited by applicant]
WO 2023287969A1 · 2023 [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/038258, Nov. 4, 2024, 18 pages. [cited by applicant]
“Applying Machine Type Recommendations for VM Instances”, Retrieved from: https://cloud.google.com/compute/docs/instances/apply-machine-type-recommendations-for-instances, Retrieved On: May 2, 2023, 10 Pages. [cited by applicant]
“Horizontal Autoscaling”, Retrieved from: https://cloud.google.com/dataflow/docs/horizontal-autoscaling, Retrieved On: Jun. 15, 2023, 7 Pages. [cited by applicant]
“Viewing Resource Recommendations”, Retrieved From: https://docs.aws.amazon.com/compute-optimizer/latest/ug/viewing-recommendations.html, Retrieved On: May 2, 2023, 1 Page. [cited by applicant]
Liu, et al., “Worker Recommendation for Crowdsourced Q&A Services: A Triple-Factor Aware Approach”, In Proceedings of the VLDB Endowment, vol. 11, Issue 3, Nov. 1, 2017, pp. 380-392. [cited by applicant]
Yadav, et al., “Resource Provisioning Through Machine Learning in Cloud Services”, In Arabian Journal for Science and Engineering, vol. 47, Jul. 24, 2021, pp. 1483-1505. [cited by applicant]
International Preliminary Report On Patentability received for PCT Application No. PCT/US2024/038258, mailed on Feb. 19, 2026, 13 pages. [cited by applicant]