IP Library › Granted Patent US 11,841,780
Granted Patent B1
US 11,841,780 · App. 17/548,264 · Granted Dec 12, 2023

Simulated network outages to manage backup network scaling

Inventors: Rishi Baldawa (Vancouver, CA); Shawn Patrick Jones (Fernie, CA)
Assignee: Amazon Technologies, Inc.
G06F11/2025G06F11/0709G06F11/263G06F11/3006H04L67/1074G06F11/203G06F2201/85
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,841,780
App. No.
17/548,264
Granted
Dec 12, 2023
Kind
B1
Abstract

A system is configured to simulate outages of network resources. The system is configured to provide a control plane for computing resources of a provider network. The control plane is configured to cause simulated outages of a primary region of the plurality of regions selected to host the plurality of different computing resources. During the simulated outages, the control plane moves respective workloads of the plurality of different computing resources to be performed in the one or more secondary networks and tracks a performance of the one or more secondary regions hosting the moved respective workloads of the plurality of computing resources. After completing individual ones of the simulated outages of the first network, the control plane moves the respective workloads of the plurality of different computing resources back to the primary region.

Claims (68)

1. A system, comprising:

one or more processors; and

a memory storing instructions that, when executed by on or across the one or more processors, cause the one or more processors to implement a control plane for a plurality of computing resources of a provider network, wherein the provider network is implemented across a plurality of regions, wherein the control plane is configured to:

cause simulated outages of a primary region of the plurality of regions selected to host the plurality of different computing resources, wherein the simulated outages are initiated at different periods of time and when a performance of one or more secondary regions of the plurality of regions satisfies one or more performance thresholds, and wherein the one or more secondary regions of the plurality of regions are selected to host the plurality of computing resources for the simulated outages of the primary region;

during the simulated outages of the primary region:

move respective workloads of the plurality of different computing resources to be performed in the one or more secondary regions; and

track a performance of the one or more secondary regions hosting the moved respective workloads of the plurality of computing resources; and

after completing individual ones of the simulated outages of the primary region, move, by the control plane, the respective workloads of the plurality of different computing resources back to the primary region.

2. The system of claim 1 , wherein the different periods of time are separated by a simulated outage interval, and wherein individual ones of the simulated outages are completed after a simulated outage duration elapses.

3. The system of claim 1 , wherein the control plane is further configured to:

monitor the performance of the one or more secondary regions during one or more time periods while the primary region is hosting the plurality of computing resources;

determine whether the performance of the one or more secondary regions does not satisfy one or more performance thresholds; and

based on a determination that the performance of the one or more secondary regions does not satisfy the one or more performance thresholds, disable the simulated outages until the performance of the one or more secondary regions satisfies the one or more performance thresholds.

4. The system of claim 1 , wherein the control plane is further configured to:

monitoring the performance of the primary region;

based on the performance of the primary region, detect a region change event occurred at the primary region; and

in response to detection of the region change event, moving the respective workloads of the plurality of computing resources to be performed in the one or more secondary regions.

5. The system of claim 4 , wherein to cause the simulated outages, the control plane is further configured to:

generate a simulated indication of the region change event; and

detect the region change event according to the simulated indication.

6. A method, comprising:

causing, by a control plane for a plurality of computing resources, simulated outages of a primary network of a plurality of networks selected to host the plurality of different computing resources, wherein the simulated outages are initiated at different periods of time and when a performance of one or more secondary networks of the plurality of networks satisfies one or more performance thresholds, and wherein the one or more secondary networks of the plurality of networks are selected to host the plurality of computing resources for the simulated outages of the primary network;

during the simulated outages of the primary network, moving, by the control plane, respective workloads of the plurality of computing resources to be performed in the one or more secondary networks; and

after completing individual ones of the simulated outages of the primary network, moving, by the control plane, the respective workloads of the plurality of computing resources back to the primary network.

7. The method of claim 6 , wherein the different periods of time are separated by a simulated outage interval, and wherein individual ones of the simulated outages are completed after a simulated outage duration elapses.

8. The method of claim 6 , further comprising:

tracking the performance of the one or more secondary networks during one or more time periods while the primary network is hosting the plurality of computing resources;

determining whether the performance of the one or more secondary networks does not satisfy the one or more performance thresholds; and

based on a determination that the performance of the one or more secondary networks does not satisfy the one or more performance thresholds, disabling the simulated outages until the performance of the one or more secondary networks satisfies the one or more performance thresholds.

9. The method of claim 6 , further comprising:

monitoring performance of the primary network;

based on the performance of the primary network, detecting a network change event occurred at the primary network; and

in response to detection of the network change event, moving the respective workloads of the plurality of computing resources to be performed in the one or more secondary networks.

10. The method of claim 9 , wherein causing the simulated outages comprises:

generating a simulated indication of the network change event; and

detecting the network change event according to the simulated indication.

11. The method of claim 10 , wherein completing the individual ones of the simulated outages comprises:

removing the simulated indication of the network change event; and

continuing the monitoring of the performance of the primary network.

12. The method of claim 6 , further comprising:

receiving indications of the respective workloads while the simulated outages are not active; and

direct the respective workloads to the primary network selected to host the plurality of different computing resources when the simulated outages are not active.

13. The method of claim 6 , further comprising:

receiving a request to temporarily disable the simulated outages for an opt-out duration; and

disabling the simulated outages until the opt-out duration elapses.

14. One or more computer-readable storage media storing instructions that, when executed on or across one or more processors, cause the one or more processors to:

cause simulated outages of a primary network of a plurality of networks selected to host a plurality of different computing resources, wherein the simulated outages are initiated at different periods of time and when a performance of one or more secondary networks of the plurality of networks satisfies one or more performance thresholds, and wherein one or more secondary networks of the plurality of networks are selected to host the plurality of computing resources for the simulated outages of the primary network;

during the simulated outages of the primary network, move respective workloads of the plurality of computing resources to be performed in the one or more secondary networks; and

after completing individual ones of the simulated outages of the primary network, move the respective workloads of the plurality of computing resources back to the primary network.

15. The one or more computer-readable storage media of claim 14 , wherein the different periods of time are separated by a simulated outage interval, and wherein individual ones of the simulated outages are completed after a simulated outage duration elapses.

16. The one or more computer-readable storage media of claim 14 , further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to:

monitor performance of the one or more secondary networks during one or more time periods while the primary network is hosting the plurality of computing resources;

determine whether the performance of the one or more secondary networks does not satisfy the one or more performance thresholds; and

based on a determination that the performance of the one or more secondary networks does not satisfy the one or more performance thresholds, disable the simulated outages until the performance of the one or more secondary networks satisfies the one or more performance thresholds.

17. The one or more computer-readable storage media of claim 14 , further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to:

monitor performance of the primary network;

based on the performance of the primary network, detect a network change event occurred at the primary network; and

in response to detection of the network change event, move the respective workloads of the plurality of computing resources to be performed in the one or more secondary networks.

18. The one or more computer-readable storage media of claim 14 , further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to:

to cause the simulated outages:

generate a simulated indication of the network change event; and

detect the network change event according to the simulated indication.

19. The one or more computer-readable storage media of claim 14 , further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to:

receive indications of the respective workloads while the simulated outages are not active; and

direct the respective workloads to the primary network selected to host the plurality of different computing resources when the simulated outages are not active.

20. The one or more computer-readable storage media of claim 14 , further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to:

receive a request to temporarily disable the simulated outages for an opt-out duration; and

disable the simulated outages until the opt-out duration elapses.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2023
From: BALDAWA, RISHI; JONES, SHAWN PATRICK
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064968/0679 →
Cited By (1)
US 12,500,887