IP Library › Granted Patent US 12,750,421
Granted Patent B2
US 12,750,421 · App. 19/041,809 · Granted Sep 29, 2026

Techniques for automating resilience testing of clusters in a distributed computing environment

Inventors: Sasha Elia Joseph (Campbell, CA); Annies Abduljaffar (San Jose, CA); Qasim Khawaja (Los Angeles, CA)
Assignee: NETFLIX, INC.
H04L67/1008H04L41/5012H04L47/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,750,421
App. No.
19/041,809
Granted
Sep 29, 2026
Kind
B2
Abstract

One embodiment sets forth a technique for automatically identifying clusters within a network to be tested for resilience. The technique includes the steps of determining that a first condition precedent for testing a first cluster within the network for resilience has been satisfied; determining that the first cluster should be tested for resilience based on a first property associated with the first cluster; in response to determining that the first condition precedent has been satisfied and determining that the first cluster should be tested for resilience, generating a cluster resilience test package for the first cluster; automatically causing the first cluster to be tested under a cluster resilience test in accordance with the cluster resilience test package. Another embodiment sets forth techniques for automatically carrying out the cluster resilience test in accordance with the cluster resilience test package.

Claims (44)

1 . A computer-implemented method for automatically testing clusters within a network for resilience, the method comprising:

receiving a cluster resilience test package associated with a cluster resilience test for a cluster;

establishing one or more configuration settings for the cluster based on the cluster resilience test package;

causing the cluster to implement the one or more configuration settings;

causing network traffic to be routed to the cluster based on one or more traffic shaping rules included in the cluster resilience test package;

analyzing telemetry information associated with the cluster to determine whether one or more performance goals included in the cluster resilience test package are satisfied;

responsive to determining that the one or more performance goals are satisfied, causing at least one cluster to implement the one or more configuration settings; and

responsive to determining that the one or more performance goals are not satisfied, iteratively adjusting the one or more configuration settings based on the telemetry information, and analyzing updated telemetry information associated with the cluster, until the one or more performance goals are satisfied or a predefined number of adjustments are made to the one or more configuration settings.

2 . The computer-implemented method of claim 1 , wherein the one or more configuration settings include at least one of a set of autoscaling rules to be implemented by the cluster or a set of load shedding rules to be implemented by the cluster, and the one or more performance goals are associated with at least one of a success buffer, a failure buffer, or a recovery time constant associated with the cluster.

3 . The computer-implemented method of claim 2 , wherein the success buffer is associated with a first amount of additional load that can be handled by the cluster before service degradation occurs, the failure buffer is associated with a second amount of additional load that can be handled by the cluster before service failure occurs, and the recovery time constant represents an amount of time required for the success buffer to recover relative to an occurrence of a load spike experienced by the cluster.

4 . The computer-implemented method of claim 2 , wherein the cluster implements at least one of one or more load shedding operations in accordance with the set of load shedding rules or one or more autoscaling operations in accordance with the set of autoscaling rules.

5 . The computer-implemented method of claim 2 , wherein the one or more performance goals are satisfied when the telemetry information indicates that a recovery time exhibited by the cluster satisfies a threshold level of alignment with the recovery time constant.

6 . The computer-implemented method of claim 1 , wherein the cluster comprises a clone of a different cluster that is operating within a service and that is generated in conjunction with carrying out the cluster resilience test.

7 . The computer-implemented method of claim 1 , wherein an amount of the network traffic routed to the cluster is multiple times larger than an average amount of network traffic managed by the cluster.

8 . The computer-implemented method of claim 1 , wherein causing the at least one cluster to implement the one or more configuration settings comprises scheduling the at least one cluster to implement the one or more configuration settings during a time at which the at least one cluster experiences a threshold amount of network traffic.

9 . The computer-implemented method of claim 1 , wherein the at least one cluster includes the cluster and at least one different cluster.

10 . The computer-implemented method of claim 9 , further comprising, prior to causing the at least one cluster to implement the one or more configuration settings:

determining that the at least one different cluster and the cluster satisfy a threshold level of similarity.

11 . One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to automatically test clusters within a network for resilience, by performing the operations of:

receiving a cluster resilience test package associated with a cluster resilience test for a cluster;

establishing one or more configuration settings for the cluster based on the cluster resilience test package;

causing the cluster to implement the one or more configuration settings;

causing network traffic to be routed to the cluster based on one or more traffic shaping rules included in the cluster resilience test package;

analyzing telemetry information associated with the cluster to determine whether one or more performance goals included in the cluster resilience test package are satisfied;

responsive to determining that the one or more performance goals are satisfied, causing at least one cluster to implement the one or more configuration settings; and

responsive to determining that the one or more performance goals are not satisfied, iteratively adjusting the one or more configuration settings based on the telemetry information, and analyzing updated telemetry information associated with the cluster, until the one or more performance goals are satisfied or a predefined number of adjustments are made to the one or more configuration settings.

12 . The one or more non-transitory computer readable media of claim 11 , wherein the one or more configuration settings include at least one of a set of autoscaling rules to be implemented by the cluster or a set of load shedding rules to be implemented by the cluster, and the one or more performance goals are associated with at least one of a success buffer, a failure buffer, or a recovery time constant associated with the cluster.

13 . The one or more non-transitory computer readable media of claim 12 , wherein the success buffer is associated with a first amount of additional load that can be handled by the cluster before service degradation occurs, the failure buffer is associated with a second amount of additional load that can be handled by the cluster before service failure occurs, and the recovery time constant represents an amount of time required for the success buffer to recover relative to an occurrence of a load spike experienced by the cluster.

14 . The one or more non-transitory computer readable media of claim 12 , wherein the cluster implements at least one of one or more load shedding operations in accordance with the set of load shedding rules or one or more autoscaling operations in accordance with the set of autoscaling rules.

15 . The one or more non-transitory computer readable media of claim 12 , wherein the one or more performance goals are satisfied when the telemetry information indicates that a recovery time exhibited by the cluster satisfies a threshold level of alignment with the recovery time constant.

16 . The one or more non-transitory computer readable media of claim 11 , wherein the cluster comprises a clone of a different cluster that is operating within a service and that is generated in conjunction with carrying out the cluster resilience test.

17 . The one or more non-transitory computer readable media of claim 11 , wherein a traffic management server causes the network traffic to be routed to the cluster based on the one or more traffic shaping rules.

18 . The one or more non-transitory computer readable media of claim 11 , wherein a set of load shedding rules included in the one or more configuration settings causes the cluster to divert the network traffic to at least one other cluster in accordance with the set of load shedding rules.

19 . The one or more non-transitory computer readable media of claim 11 , wherein a set of autoscaling rules included in the one or more configuration settings causes the cluster to increase or decrease resources utilized by the cluster.

20 . A system, comprising:

one or more memories that include instructions; and

one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the operations of:

receiving a cluster resilience test package associated with a cluster resilience test for a cluster;

establishing one or more configuration settings for the cluster based on the cluster resilience test package;

causing the cluster to implement the one or more configuration settings;

causing network traffic to be routed to the cluster based on one or more traffic shaping rules included in the cluster resilience test package;

analyzing telemetry information associated with the cluster to determine whether one or more performance goals included in the cluster resilience test package are satisfied;

responsive to determining that the one or more performance goals are satisfied, causing at least one cluster to implement the one or more configuration settings; and

responsive to determining that the one or more performance goals are not satisfied, iteratively adjusting the one or more configuration settings based on the telemetry information, and analyzing updated telemetry information associated with the cluster, until the one or more performance goals are satisfied or a predefined number of adjustments are made to the one or more configuration settings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2025
From: JOSEPH, SASHA ELIA; ABDULJAFFAR, ANNIES; KHAWAJA, QASIM
To: NETFLIX, INC.
Reel/Frame 070080/0672 →
Continuity (1)
Related Publication 20260222469A1 · Jul 30, 2026
References Cited (33)
US 7370101B1 · Lakkapragada et al. · 2008 [cited by applicant]
US 7434104B1 · Skeoch et al. · 2008 [cited by applicant]
US 9065783B2 · Ding · 2015 [cited by examiner]
US 10073763B1 · Raman et al. · 2018 [cited by applicant]
US 10346367B1 · Luszcz · 2019 [cited by examiner]
US 10866851B2 · Basiri et al. · 2020 [cited by applicant]
US 11361846B1 · Jain et al. · 2022 [cited by applicant]
US 11467945B1 · Cao et al. · 2022 [cited by applicant]
US 11757982B2 · Nair · 2023 [cited by examiner]
US 20060224725A1 · Bali · 2006 [cited by examiner]
US 20080209044A1 · Forrester · 2008 [cited by applicant]
US 20110145643A1 · Kumar et al. · 2011 [cited by applicant]
US 20150134825A1 · Alshinnawi et al. · 2015 [cited by applicant]
US 20170366604A1 · McDuff · 2017 [cited by examiner]
US 20190250954A1 · Sethi · 2019 [cited by examiner]
US 20210073110A1 · Vidal et al. · 2021 [cited by applicant]
US 20210203550A1 · Thakkar · 2021 [cited by examiner]
US 20220114041A1 · Tiwari et al. · 2022 [cited by applicant]
US 20230409418A1 · Sood · 2023 [cited by applicant]
US 20240286624A1 · Kaveri Poompatnam Chandrasekaran · 2024 [cited by examiner]
US 20240291738A1 · Patronas et al. · 2024 [cited by applicant]
US 20250147889A1 · Dayanand et al. · 2025 [cited by applicant]
US 20250321869A1 · Bartram · 2025 [cited by applicant]
US 20260050418A1 · Banerjee · 2026 [cited by examiner]
CN 101309167B · 2011 [cited by examiner]
CN 112671029A · 2021 [cited by examiner]
Non Final Office received for U.S. Appl. No. 19/041,806, dated Apr. 30, 2026, 10 pages. [cited by applicant]
International Search Report for Application No. PCT/US2026/012730 dated Apr. 22, 2026. [cited by applicant]
International Search Report for Application No. PCT/US2026/012735 dated Apr. 23, 2026. [cited by applicant]
Cotroneo et al., “ThorFI: A Novel Approach for Network Fault Injection as a Service”, arXiv:2201.07521, Jan. 19, 2022, pp. 1-21. [cited by applicant]
Burns et al., “Continuous Delivery with Spinnaker”, Fast, Safe, Repeatable Multi-Cloud Deployments, May 11, 2018, 81 pages. [cited by applicant]
Rosenthal et al., “Chaos Engineering”, System Resiliency in Practice, retrieved from https://www.oreilly.com/library/view/chaos-engineering/9781492043850/, Apr. 3, 2020, 372 pages. [cited by applicant]
Arundel et al., “Cloud Native DevOps with Kubernetes”, Cloud Native DevOps with Kubernetes: Building, Deploying, and Scaling Modern Applications in the Cloud, Jan. 24, 2019, 344 pages. [cited by applicant]