IP Library Granted Patent US 11,886,932
Granted Patent B1
US 11,886,932 · App. 17/023,602 · Granted Jan 30, 2024

Managing resource instances

Inventors: Suvojit Dasgupta (Atlanta, GA); Anand Kumar Sivasamy Kaliaperumal (Plano, TX); Gregory Harrison Fina (Petaluma, CA)
Assignee: Amazon Technologies, Inc.
G06F9/5083G06F9/4812G06F9/4881G06F9/5027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,886,932
App. No.
17/023,602
Granted
Jan 30, 2024
Kind
B1
Abstract

Reliability monitoring can be performed for compute instances in a cluster with auto-scaling capability. Such monitoring can analyze state information for various instances, such as spot instances, to determine when an interruption or termination is to occur. An impact assessor can determine the impact on performance due to any such interruption or termination, and if necessary to maintain at least a minimum level of performance then an action performer can obtain additional or alternate instances, which may be of a different type, to make up for lost capacity. Any tasks being performed can be migrated to the newly-allocated instances without any failures or significant impact on performance, and the previously-utilized instances can be released corresponding to the termination or interruption.

Claims (48)

1. A computer-implemented method, comprising:

allocating a cluster of compute instances for performance of a task, wherein the compute instances are reclaimable instances of a first instance type in a first instance category;

receiving, in a state machine comprising memory to store state information related to the compute instances and comprising at least one processor to monitor the state information, a notification that a compute instance of the cluster will be terminated prior to completion of the task;

determining that the performance of the task will fall outside an acceptable range of performance;

enabling, by the state machine, an identification of a second compute instance to replace capacity of the compute instance that will be terminated, the second compute instance being of a different instance type or different instance category relative to the first instance type or the first instance category; and

transitioning at least a portion of the task to the second compute instance for completion of the performance.

2. The computer-implemented method of claim 1 , further comprising:

tagging the cluster for reliability monitoring, wherein the capacity of the compute instances is monitored for potential replacement.

3. The computer-implemented method of claim 1 , wherein the compute instances in the cluster are of one of a plurality of instance types inside an auto-scaling group of resources.

4. The computer-implemented method of claim 1 , wherein the task is part of a long-running job requiring multiple concurrent compute instances for successful performance.

5. The computer-implemented method of claim 1 , further comprising:

capturing a snapshot of the state information before transitioning at least a portion of the performance of the task to the second compute instance; and

performing a rollback of the performance to an instance of the first instance type and the first instance category using the state information in the snapshot.

6. A computer-implemented method, comprising:

allocating a cluster of first compute instances to perform a task;

determining, based in part on a notification received, that at least a subset of the first compute instances will be unable to perform at least a portion of the task, the notification received in a state machine comprising memory to store state information related to the first compute instances and comprising at least one processor to monitor the state information;

enabling, by the state machine, a determination that additional resource capacity will be required to perform the task with at least a minimum level of performance; and

allocating one or more second compute instances to provide the additional resource capacity, the one or more second compute instances capable of being of a different instance type or instance category relative to that of the first compute instances, the one or more second compute instances corresponding to reclaimable capacity or interruptible capacity.

7. The computer-implemented method of claim 6 , wherein the first compute instances are spot instances providing capacity subject to potential interruption, and wherein determining that at least a subset of the first compute instances will be unable to perform at least a portion of the task includes receiving the notification of an interruption event corresponding to a potential interruption.

8. The computer-implemented method of claim 6 , further comprising:

launching an impact assessor to determine that additional resource capacity will be required, the impact assessor further utilizing instance state to determine the one or more second compute instances.

9. The computer-implemented method of claim 8 , further comprising:

launching an action performer to allocate the one or more second compute instances and transition at least a portion of the task to be performed by the one or more second compute instances.

10. The computer-implemented method of claim 9 , further comprising:

releasing the subset of the first compute instances after at least a portion of the task is transitioned to the one or more second compute instances.

11. The computer-implemented method of claim 6 , further comprising:

tagging the cluster for reliability monitoring, wherein the first compute instances will be monitored to determine whether the first compute instances will be able to perform the task.

12. The computer-implemented method of claim 6 , wherein the first compute instances are of one of a plurality of instance types inside an auto-scaling group of resources.

13. The computer-implemented method of claim 6 , wherein the task is part of a long-running job requiring multiple concurrent compute instances for successful performance.

14. The computer-implemented method of claim 6 , further comprising:

capturing a snapshot of the state information before transitioning at least a portion of the performance of the task to the one or more second compute instances; and

performing a rollback of the performance to instances of an instance type and an instance category of the first compute instances using the state information in the snapshot.

15. The computer-implemented method of claim 6 , further comprising:

enabling a user to specify a preferred type of the first compute instances.

16. A system, comprising:

a processor; and

memory including instructions that, when executed by the processor, cause the system to:

allocate a cluster of first compute instances to perform a task, the compute instances corresponding to a first instance type of reclaimable capacity or interruptible capacity;

determine, based in part on a notification received, that at least a subset of the first compute instances will be unable to perform at least a portion of the task, the notification received in a state machine of the system that stores state information related to the first compute instances and that monitors the state information;

enable, by the state machine, a determination in the system that additional resource capacity will be required to perform the task with at least a minimum level of performance; and

allocate one or more second compute instances to provide the additional resource capacity, the one or more second compute instances capable of being of a different instance type relative to the first compute instances.

17. The system of claim 16 , wherein the first compute instances are spot instances providing capacity subject to potential interruption, and wherein determining that at least a subset of the first compute instances will be unable to perform at least a portion of the task includes receiving an interruption event corresponding to a potential interruption.

18. The system of claim 16 , wherein the instructions when executed further cause the system to:

launch an impact assessor to determine that additional resource capacity will be required, the impact assessor further utilizing instance state to determine the one or more second compute instances.

19. The system of claim 16 , wherein the instructions when executed further cause the system to:

launch an action performer to allocate the one or more second compute instances and transition at least a portion of the task to be performed by the one or more second compute instances.

20. The system of claim 16 , wherein the instructions when executed further cause the system to:

release the subset of the first compute instances after at least a portion of the task is transitioned to the one or more second compute instances.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2020
From: DASGUPTA, SUVOJIT; SIVASAMY KALIAPERUMAL, ANAND KUMAR; FINA, GREGORY HARRISON
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 053800/0589 →
Cited By (7)
US 12,210,895 US 12,321,771 US 12,386,636 US 12,474,945 US 12,664,028 US 12,681,754 US 12,682,392