Techniques for automating resilience testing of clusters in a distributed computing environment
One embodiment sets forth a technique for automatically identifying clusters within a network to be tested for resilience. The technique includes the steps of determining that a first condition precedent for testing a first cluster within the network for resilience has been satisfied; determining that the first cluster should be tested for resilience based on a first property associated with the first cluster; in response to determining that the first condition precedent has been satisfied and determining that the first cluster should be tested for resilience, generating a cluster resilience test package for the first cluster; automatically causing the first cluster to be tested under a cluster resilience test in accordance with the cluster resilience test package. Another embodiment sets forth techniques for automatically carrying out the cluster resilience test in accordance with the cluster resilience test package.
1 . A computer-implemented method for automatically testing clusters within a network for resilience, the method comprising:
receiving a cluster resilience test package associated with a cluster resilience test for a cluster;
establishing one or more configuration settings for the cluster based on the cluster resilience test package;
causing the cluster to implement the one or more configuration settings;
causing network traffic to be routed to the cluster based on one or more traffic shaping rules included in the cluster resilience test package;
analyzing telemetry information associated with the cluster to determine whether one or more performance goals included in the cluster resilience test package are satisfied;
responsive to determining that the one or more performance goals are satisfied, causing at least one cluster to implement the one or more configuration settings; and
responsive to determining that the one or more performance goals are not satisfied, iteratively adjusting the one or more configuration settings based on the telemetry information, and analyzing updated telemetry information associated with the cluster, until the one or more performance goals are satisfied or a predefined number of adjustments are made to the one or more configuration settings.
2 . The computer-implemented method of claim 1 , wherein the one or more configuration settings include at least one of a set of autoscaling rules to be implemented by the cluster or a set of load shedding rules to be implemented by the cluster, and the one or more performance goals are associated with at least one of a success buffer, a failure buffer, or a recovery time constant associated with the cluster.
3 . The computer-implemented method of claim 2 , wherein the success buffer is associated with a first amount of additional load that can be handled by the cluster before service degradation occurs, the failure buffer is associated with a second amount of additional load that can be handled by the cluster before service failure occurs, and the recovery time constant represents an amount of time required for the success buffer to recover relative to an occurrence of a load spike experienced by the cluster.
4 . The computer-implemented method of claim 2 , wherein the cluster implements at least one of one or more load shedding operations in accordance with the set of load shedding rules or one or more autoscaling operations in accordance with the set of autoscaling rules.
5 . The computer-implemented method of claim 2 , wherein the one or more performance goals are satisfied when the telemetry information indicates that a recovery time exhibited by the cluster satisfies a threshold level of alignment with the recovery time constant.
6 . The computer-implemented method of claim 1 , wherein the cluster comprises a clone of a different cluster that is operating within a service and that is generated in conjunction with carrying out the cluster resilience test.
7 . The computer-implemented method of claim 1 , wherein an amount of the network traffic routed to the cluster is multiple times larger than an average amount of network traffic managed by the cluster.
8 . The computer-implemented method of claim 1 , wherein causing the at least one cluster to implement the one or more configuration settings comprises scheduling the at least one cluster to implement the one or more configuration settings during a time at which the at least one cluster experiences a threshold amount of network traffic.
9 . The computer-implemented method of claim 1 , wherein the at least one cluster includes the cluster and at least one different cluster.
10 . The computer-implemented method of claim 9 , further comprising, prior to causing the at least one cluster to implement the one or more configuration settings:
determining that the at least one different cluster and the cluster satisfy a threshold level of similarity.
11 . One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to automatically test clusters within a network for resilience, by performing the operations of:
receiving a cluster resilience test package associated with a cluster resilience test for a cluster;
establishing one or more configuration settings for the cluster based on the cluster resilience test package;
causing the cluster to implement the one or more configuration settings;
causing network traffic to be routed to the cluster based on one or more traffic shaping rules included in the cluster resilience test package;
analyzing telemetry information associated with the cluster to determine whether one or more performance goals included in the cluster resilience test package are satisfied;
responsive to determining that the one or more performance goals are satisfied, causing at least one cluster to implement the one or more configuration settings; and
responsive to determining that the one or more performance goals are not satisfied, iteratively adjusting the one or more configuration settings based on the telemetry information, and analyzing updated telemetry information associated with the cluster, until the one or more performance goals are satisfied or a predefined number of adjustments are made to the one or more configuration settings.
12 . The one or more non-transitory computer readable media of claim 11 , wherein the one or more configuration settings include at least one of a set of autoscaling rules to be implemented by the cluster or a set of load shedding rules to be implemented by the cluster, and the one or more performance goals are associated with at least one of a success buffer, a failure buffer, or a recovery time constant associated with the cluster.
13 . The one or more non-transitory computer readable media of claim 12 , wherein the success buffer is associated with a first amount of additional load that can be handled by the cluster before service degradation occurs, the failure buffer is associated with a second amount of additional load that can be handled by the cluster before service failure occurs, and the recovery time constant represents an amount of time required for the success buffer to recover relative to an occurrence of a load spike experienced by the cluster.
14 . The one or more non-transitory computer readable media of claim 12 , wherein the cluster implements at least one of one or more load shedding operations in accordance with the set of load shedding rules or one or more autoscaling operations in accordance with the set of autoscaling rules.
15 . The one or more non-transitory computer readable media of claim 12 , wherein the one or more performance goals are satisfied when the telemetry information indicates that a recovery time exhibited by the cluster satisfies a threshold level of alignment with the recovery time constant.
16 . The one or more non-transitory computer readable media of claim 11 , wherein the cluster comprises a clone of a different cluster that is operating within a service and that is generated in conjunction with carrying out the cluster resilience test.
17 . The one or more non-transitory computer readable media of claim 11 , wherein a traffic management server causes the network traffic to be routed to the cluster based on the one or more traffic shaping rules.
18 . The one or more non-transitory computer readable media of claim 11 , wherein a set of load shedding rules included in the one or more configuration settings causes the cluster to divert the network traffic to at least one other cluster in accordance with the set of load shedding rules.
19 . The one or more non-transitory computer readable media of claim 11 , wherein a set of autoscaling rules included in the one or more configuration settings causes the cluster to increase or decrease resources utilized by the cluster.
20 . A system, comprising:
one or more memories that include instructions; and
one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the operations of:
receiving a cluster resilience test package associated with a cluster resilience test for a cluster;
establishing one or more configuration settings for the cluster based on the cluster resilience test package;
causing the cluster to implement the one or more configuration settings;
causing network traffic to be routed to the cluster based on one or more traffic shaping rules included in the cluster resilience test package;
analyzing telemetry information associated with the cluster to determine whether one or more performance goals included in the cluster resilience test package are satisfied;
responsive to determining that the one or more performance goals are satisfied, causing at least one cluster to implement the one or more configuration settings; and
responsive to determining that the one or more performance goals are not satisfied, iteratively adjusting the one or more configuration settings based on the telemetry information, and analyzing updated telemetry information associated with the cluster, until the one or more performance goals are satisfied or a predefined number of adjustments are made to the one or more configuration settings.