IP Library Granted Patent US 11,789,800
Granted Patent B2
US 11,789,800 · App. 17/548,635 · Granted Oct 17, 2023

Degraded availability zone remediation for multi-availability zone clusters of host computers

Inventors: Piyush Parmar (Bangalore, IN); Pawan Saxena (Palo Alto, CA); Gabriel Tarasuk-Levin (Sunnyvale, CA); Dhaval Shah (Bangalore, IN); Umesha Margi (Bangalore, IN)
Assignee: VMWARE, INC.
G06F11/0772G06F11/2041G06F11/3006G06F11/3072
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,789,800
App. No.
17/548,635
Granted
Oct 17, 2023
Kind
B2
Abstract

System and computer-implemented method for managing multi-availability zone (AZ) clusters of host computers in a cloud computing environment automatically detects a degraded state of a first AZ in the cloud computing environment based on host failure events for host computers in a first cluster section of a multi-AZ cluster of host computers located in the first AZ and a recovered state of the first AZ based a successful scale-in operation of another multi-AZ cluster located partially in the first AZ. In response to the detection of the degraded state of the first AZ, a second cluster section of the multi-AZ cluster of host computers located in a second AZ is scaled out. In response to the detection of the recovered state of the first AZ, the second cluster section of the multi-AZ cluster of host computers located in the second AZ is scaled in.

Claims (34)

1. A computer-implemented method for managing multi-availability zone (AZ) clusters of host computers in a cloud computing environment, the method comprising:

automatically detecting a degraded state of a first AZ in the cloud computing environment based on host failure events for host computers in a first cluster section of a multi-AZ cluster of host computers located in the first AZ;

scaling out a second cluster section of the multi-AZ cluster of host computers located in a second AZ in response to the detecting of the degraded state of the first AZ;

automatically detecting a recovered state of the first AZ based a successful scale-in operation of another multi-AZ cluster located partially in the first AZ; and

scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ in response to the detecting of the recovered state of the first AZ.

2. The computer-implemented method of claim 1 , wherein automatically detecting the degraded state of the first AZ includes receiving auto-remediation failure notifications of replacement host computers being unable to be provisioned in the first cluster section of the multi-AZ cluster of host computers located in the first AZ, wherein the auto-remediation failure notifications indicate that the first AZ is potentially a degraded AZ.

3. The computer-implemented method of claim 2 , wherein automatically detecting the degraded state of the first AZ includes determining that the first AZ is a degraded AZ when a virtualization manager for the multi-AZ cluster of host computers has lost connection to all the hosts in the first cluster section of the multi-AZ cluster of host computers located in the first AZ.

4. The computer-implemented method of claim 1 , wherein scaling out the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes adding a number of new host computers to the second cluster section of the multi-AZ cluster of host computers, wherein the number of new host computers equals a number of host computers in the first cluster section of the multi-AZ cluster of host computers located in the first AZ when each of the first and cluster sections of the multi-AZ cluster of host computers includes two or three host computers.

5. The computer-implemented method of claim 1 , wherein scaling out the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes adding a number of new host computers to the second cluster section of the multi-AZ cluster of host computers, wherein the number of new host computers equals half of a number of host computers in the first cluster section of the multi-AZ cluster of host computers located in the first AZ when each of the first and cluster sections of the multi-AZ cluster of host computers includes more than three host computers.

6. The computer-implemented method of claim 1 , wherein automatically detecting the recovered state of the first AZ includes selecting a random multi-AZ cluster located partially in the first AZ to execute a recovery workflow on the random multi-AZ cluster, wherein the recovery workflow includes checking health states of all host computers in a section of the random multi-AZ cluster that is located in the first AZ to determine whether the first AZ has recovered.

7. The computer-implemented method of claim 6 , wherein the recovery workflow further includes determining that the first AZ has recovered when the health states of all the host computers in the section of the random multi-AZ cluster that is located in the first AZ are healthy.

8. The computer-implemented method of claim 7 , wherein scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ is part of scaling in all multi-AZ clusters of host computers located partially in the first AZ based the successful scale-in operation of the another multi-AZ cluster located partially in the first AZ.

9. The computer-implemented method of claim 1 , wherein scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes removing a number of host computers in the second cluster section of the multi-AZ cluster of host computers, wherein the number of host computers removed is equal to a number of host computers that were added to the second cluster section of the multi-AZ cluster of host computers located in the second AZ during the scaling out.

10. A non-transitory computer-readable storage medium containing program instructions for auto or managing multi-availability zone (AZ) clusters of host computers in a cloud computing environment, wherein execution of the program instructions by one or more processors causes the one or more processors to perform steps comprising:

automatically detecting a degraded state of a first AZ in the cloud computing environment based on host failure events for host computers in a first cluster section of a multi-AZ cluster of host computers located in the first AZ;

scaling out a second cluster section of the multi-AZ cluster of host computers located in a second AZ in response to the detecting of the degraded state of the first AZ;

automatically detecting a recovered state of the first AZ based a successful scale-in operation of another multi-AZ cluster located partially in the first AZ; and

scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ in response to the detecting of the recovered state of the first AZ.

11. The non-transitory computer-readable storage medium of claim 10 , wherein automatically detecting the degraded state of the first AZ includes receiving auto-remediation failure notifications of replacement host computers being unable to be provisioned in the first cluster section of the multi-AZ cluster of host computers located in the first AZ, wherein the auto-remediation failure notifications indicate that the first AZ is potentially a degraded AZ.

12. The non-transitory computer-readable storage medium of claim 11 , wherein automatically detecting the degraded state of the first AZ includes determining that the first AZ is a degraded AZ when a virtualization manager for the multi-AZ cluster of host computers has lost connection to all the hosts in the first cluster section of the multi-AZ cluster of host computers located in the first AZ.

13. The non-transitory computer-readable storage medium of claim 10 , wherein scaling out the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes adding a number of new host computers to the second cluster section of the multi-AZ cluster of host computers, wherein the number of new host computers equals a number of host computers in the first cluster section of the multi-AZ cluster of host computers located in the first AZ when each of the first and cluster sections of the multi-AZ cluster of host computers includes two or three host computers.

14. The non-transitory computer-readable storage medium of claim 10 , wherein scaling out the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes adding a number of new host computers to the second cluster section of the multi-AZ cluster of host computers, wherein the number of new host computers equals half of a number of host computers in the first cluster section of the multi-AZ cluster of host computers located in the first AZ when each of the first and cluster sections of the multi-AZ cluster of host computers includes more than three host computers.

15. The non-transitory computer-readable storage medium of claim 10 , wherein automatically detecting the recovered state of the first AZ includes selecting a random multi-AZ cluster located partially in the first AZ to execute a recovery workflow on the random multi-AZ cluster, wherein the recovery workflow includes checking health states of all host computers in a section of the random multi-AZ cluster that is located in the first AZ to determine whether the first AZ has recovered.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the recovery workflow further includes determining that the first AZ has recovered when the health states of all the host computers in the section of the random multi-AZ cluster that is located in the first AZ are healthy.

17. The non-transitory computer-readable storage medium of claim 16 , wherein scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ is part of scaling in all multi-AZ clusters of host computers located partially in the first AZ based the successful scale-in operation of the another multi-AZ cluster located partially in the first AZ.

18. The non-transitory computer-readable storage medium of claim 10 , wherein scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes removing a number of host computers in the second cluster section of the multi-AZ cluster of host computers, wherein the number of host computers removed is equal to a number of host computers that were added to the second cluster section of the multi-AZ cluster of host computers located in the second AZ during the scaling out.

19. A system comprising:

memory; and

at least one processor configured to:

automatically detect a degraded state of a first availability zone (AZ) in a cloud computing environment based on host failure events for host computers in a first cluster section of a multi-AZ cluster of host computers located in the first AZ;

scale out a second cluster section of the multi-AZ cluster of host computers located in a second AZ in response to the detecting of the degraded state of the first AZ;

automatically detect a recovered state of the first AZ based a successful scale-in operation of another multi-AZ cluster located partially in the first AZ; and

scale in the second cluster section of the multi-AZ cluster of host computers located in the second AZ in response to the detecting of the recovered state of the first AZ.

20. The system of claim 19 , wherein the at least one processor is configured to receive auto-remediation failure notifications of replacement host computers being unable to be provisioned in the first cluster section of the multi-AZ cluster of host computers located in the first AZ, wherein the auto-remediation failure notifications indicate that the first AZ is potentially a degraded AZ.

Assignments (2)
CHANGE OF NAME Recorded Apr 15, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 067102/0395 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2021
From: PARMAR, PIYUSH; SAXENA, PAWAN; TARASUK-LEVIN, GABRIEL; SHAH, DHAVAL; MARGI, UMESHA
To: VMWARE, INC.
Reel/Frame 058367/0303 →
Priority Claims (1)
IN 202141044762 · Oct 1, 2021 · national
Continuity (1)
Related Publication 20230107518A1 · Apr 6, 2023