Predicting worker instance count for cloud-based computing platforms
Described are examples for recommending increase in worker instance count for an availability zone in a cloud-based computing platform. A machine learning (ML) model can be used to predict a time series forecast of a workload for the availability zone in a future time period. A predicted number of worker instances to handle the predicted workload can be computed, and if the number of worker instances in the availability zone is less than the predicted number of worker instances, a recommendation to increase the number of worker instances in the availability zone can be generated.
1 . A device for increasing worker instance count for an availability zone in a cloud-based computing platform, comprising:
one or more memories storing instructions; and
one or more processors coupled to the one or more memories and configured to execute the instructions to:
provide resource allocation information as input to a machine learning (ML) model to receive an output of a time series forecast of a workload for the availability zone in a future time period, wherein the resource allocation information comprises a history of resource allocation requests for the availability zone including original allocation requests, retry requests that occur when an original allocation request fails, and requests related to underlying platform issues, wherein the ML model is trained on a timespan of workload history for the availability zone such that the time series forecast weighs workload increases due to customer demand over workload increases due to system disruption;
compute a throughput threshold for the worker instances in the availability zone by creating throughput bins from a history of workload data and identifying a throughput bin value at which a number of throttling or timeout exceptions is at least a percentage;
compute a predicted number of worker instances in the availability zone for handling the workload in the future time period based at least in part on the outputted time series forecast and the computed throughput threshold;
when a number of worker instances in the availability zone is less than the predicted number of worker instances, provision a recommended number of servers for the number of worker instances in the availability zone, and process the workload with the provisioned servers at the future time period.
2 . The device of claim 1 , wherein the one or more processors are further configured to execute the instructions to predict a peak workload over the future time period based on the predicted time series forecast of the workload and a daily to minute peak workload ratio, wherein the predicted number of worker instances is based on the peak workload.
3 . The device of claim 2 , wherein the one or more processors are further configured to execute the instructions to compute a throughput threshold for the number of worker instances, wherein the predicted number of worker instances is based on a comparison of the throughput threshold with the peak workload.
4 . The device of claim 3 , wherein the one or more processors are further configured to execute the instructions to compute the throughput threshold based on one or more of a throttling metric, a total allocation time, or a number of timeout exceptions observed from a history of workload data for the availability zone.
5 . The device of claim 1 , wherein the time series forecast of the workload is based on a history of workload data, statistical trend of the workload data, and one or more properties of a time period related to the workload data.
6 . The device of claim 1 , wherein the time series forecast is based on a time series prediction model of historical workload data for the availability zone.
7 . The device of claim 1 , wherein the time series forecast is based on fitting an empirical statistical distribution of a history of resource allocation requests of the availability zone.
8 . The device of claim 1 , wherein the time series forecast for the availability zone is based on the one or more processors execute the instructions to determine the availability zone is one of multiple availability zones having a threshold confidence for accuracy of predicting the time series forecast.
9 . The device of claim 1 , wherein the time series forecast is based on the history of resource allocation requests.
10 . The device of claim 1 , wherein the one or more processors are further configured to execute the instructions to provide, to the ML model, recent workload data for the availability zone and associated performance metrics for use in predicting subsequent time series forecasts of the workload for the availability zone.
11 . The device of claim 1 , wherein the one or more processors are further configured to execute the instructions to generate the recommendation to increase the number of worker instances further based on one or more of a customer priority of a customer corresponding to the workload, a scalability of the cloud-based computing platform, or a number of throttling failures during deployment of virtual machines in the availability zone.
12 . A method for increasing worker instance count for an availability zone in a cloud-based computing platform, comprising:
predicting, using a machine learning (ML) model and by providing resource allocation information as input to the ML model, a time series forecast of a workload for the availability zone in a future time period, wherein the resource allocation information comprises a history of resource allocation requests for the availability zone including original allocation requests, retry requests that occur when an original allocation request fails, and requests related to underlying platform issues, wherein the ML model is trained on a timespan of workload history for the availability zone such that the time series forecast weighs workload increases due to customer demand over workload increases due to system disruption;
computing a throughput threshold for the worker instances in the availability zone by creating throughput bins from a history of workload data and identifying a throughput bin value at which a number of throttling or timeout exceptions is at least a percentage;
determining, based on the throughput threshold, whether a number of worker instances in the availability zone is sufficient for satisfying the workload in the future time period; and
based on determining that the number of worker instances is not sufficient, provisioning a recommended number of servers for the number of worker instances in the availability zone, and processing the workload with the provisioned servers at the future time period.
13 . The method of claim 12 , further comprising predicting a peak workload over the future time period based on the predicted time series forecast of the workload and a daily to minute peak workload ratio, wherein determining whether the number of worker instances is sufficient is based on the peak workload.
14 . The method of claim 13 , further comprising computing a throughput threshold for the number of worker instances, wherein determining whether the number of worker instances is sufficient based on a comparison of the throughput threshold with the peak workload.
15 . The method of claim 14 , wherein computing the throughput threshold is based on one or more of a throttling metric, a total allocation time, or a number of timeout exceptions observed from a history of workload data for the availability zone.
16 . The method of claim 12 , wherein predicting the time series forecast of the workload is based on one or more of a history of workload data, statistical trend of the workload data, and one or more properties of a time period related to the workload data.
17 . The method of claim 12 , wherein predicting the time series forecast is based on a history of resource allocation requests including initial requests, retries, and requests due to throttled workload.
18 . The method of claim 12 , wherein generating the recommendation to increase the number of worker instances is further based on one or more of a customer priority of a customer corresponding to the workload, a scalability of the cloud-based computing platform, or a number of throttling failures during deployment of virtual machines in the availability zone.
19 . A non-transitory computer-readable device storing instructions thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations for increasing worker instance count for an availability zone in a cloud-based computing platform, comprising:
predicting, using a machine learning (ML) model and by providing resource allocation information as input to the ML model, a time series forecast of a workload for the availability zone in a future time period, wherein the resource allocation information comprises a history of resource allocation requests for the availability zone including original allocation requests, retry requests that occur when an original allocation request fails, and requests related to underlying platform issues, wherein the ML model is trained on a timespan of workload history for the availability zone such that the time series forecast weighs workload increases due to customer demand over workload increases due to system disruption;
computing a throughput threshold for the worker instances in the availability zone by creating throughput bins from a history of workload data and identifying a throughput bin value at which a number of throttling or timeout exceptions is at least a percentage;
determining, based on the throughput threshold, whether a number of worker instances in the availability zone is sufficient for satisfying the workload in the future time period; and
based on determining that the number of worker instances is not sufficient, provisioning a recommended number of servers for the number of worker instances in the availability zone, and processing the workload with the provisioned servers at the future time period.
20 . The non-transitory computer-readable device of claim 19 , the operations further comprising predicting a peak workload over the future time period based on the predicted time series forecast of the workload and a daily to minute peak workload ratio, wherein determining whether the number of worker instances is sufficient is based on the peak workload.