IP Library Granted Patent US 12,585,521
Granted Patent B2
US 12,585,521 · App. 18/380,576 · Granted Mar 24, 2026

Degraded availability zone remediation for multi-availability zone clusters of host computers

Inventors: Piyush Parmar (Bangalore, IN); Pawan Saxena (Palo Alto, CA); Gabriel Tarasuk-Levin (Sunnyvale, CA); Dhaval Shah (Bangalore, IN); Umesha Margi (Bangalore, IN)
Assignee: VMware, Inc.
G06F11/0772G06F11/2041G06F11/3006G06F11/3072
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,521
App. No.
18/380,576
Granted
Mar 24, 2026
Kind
B2
Abstract

System and computer-implemented method for managing multi-availability zone (AZ) clusters of host computers in a cloud computing environment automatically detects a degraded state of a first AZ in the cloud computing environment based on host failure events for host computers in a first cluster section of a multi-AZ cluster of host computers located in the first AZ and a recovered state of the first AZ based a successful scale-in operation of another multi-AZ cluster located partially in the first AZ. In response to the detection of the degraded state of the first AZ, a second cluster section of the multi-AZ cluster of host computers located in a second AZ is scaled out. In response to the detection of the recovered state of the first AZ, the second cluster section of the multi-AZ cluster of host computers located in the second AZ is scaled in.

Claims (34)

1 . A computer-implemented method for managing multi-availability zone (AZ) clusters of host computers in a cloud computing environment, the method comprising:

automatically detecting a degraded state of a first AZ in the cloud computing environment having a first cluster section of a multi-AZ cluster of host computers located in the first AZ;

scaling out a second cluster section of the multi-AZ cluster of host computers located in a second AZ in response to the detecting of the degraded state of the first AZ;

automatically detecting a recovered state of the first AZ, wherein automatically detecting the recovered state of the first AZ includes selecting a random multi-AZ cluster located partially in the first AZ to execute a recovery workflow on the random multi-AZ cluster, wherein the recovery workflow includes checking health states of all host computers in a section of the random multi-AZ cluster that is located in the first AZ to determine whether the first AZ has recovered; and

scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ in response to the detecting of the recovered state of the first AZ.

2 . The computer-implemented method of claim 1 , wherein automatically detecting the degraded state of the first AZ includes receiving an auto-remediation failure notification of a replacement host computer being unable to be provisioned in a first cluster section of the multi-AZ cluster of host computers located in the first AZ, wherein the auto-remediation failure notification indicates that the first AZ is potentially a degraded AZ.

3 . The computer-implemented method of claim 2 , wherein automatically detecting the degraded state of the first AZ includes determining that the first AZ is a degraded AZ when a virtualization manager for the multi-AZ cluster of host computers has lost connection to all the hosts in the first cluster section of the multi-AZ cluster of host computers located in the first AZ.

4 . The computer-implemented method of claim 1 , wherein scaling out the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes adding a number of new host computers to the second cluster section of the multi-AZ cluster of host computers, wherein the number of new host computers equals a number of host computers in the first cluster section of the multi-AZ cluster of host computers located in the first AZ when each of the first and cluster sections of the multi-AZ cluster of host computers includes two or three host computers.

5 . The computer-implemented method of claim 1 , wherein scaling out the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes adding a number of new host computers to the second cluster section of the multi-AZ cluster of host computers, wherein the number of new host computers equals half of a number of host computers in the first cluster section of the multi-AZ cluster of host computers located in the first AZ when each of the first and cluster sections of the multi-AZ cluster of host computers includes more than three host computers.

6 . The computer-implemented method of claim 1 wherein the recovery workflow further includes determining that the first AZ has recovered when the health states of all the host computers in the section of the random multi-AZ cluster that is located in the first AZ are healthy.

7 . The computer-implemented method of claim 6 , wherein scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ is part of scaling in all multi-AZ clusters of host computers located partially in the first AZ based the successful scale-in operation of the another multi-AZ cluster located partially in the first AZ.

8 . The computer-implemented method of claim 1 , wherein scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes removing a number of host computers in the second cluster section of the multi-AZ cluster of host computers, wherein the number of host computers removed is equal to a number of host computers that were added to the second cluster section of the multi-AZ cluster of host computers located in the second AZ during the scaling out.

9 . A non-transitory computer-readable storage medium containing program instructions for auto or managing multi-availability zone (AZ) clusters of host computers in a cloud computing environment, wherein execution of the program instructions by one or more processors causes the one or more processors to perform steps comprising:

automatically detecting a degraded state of a first AZ in the cloud computing environment having a first cluster section of a multi-AZ cluster of host computers located in the first AZ, wherein automatically detecting the recovered state of the first AZ includes selecting a random multi-AZ cluster located partially in the first AZ to execute a recovery workflow on the random multi-AZ cluster, wherein the recovery workflow includes checking health states of all host computers in a section of the random multi-AZ cluster that is located in the first AZ to determine whether the first AZ has recovered;

scaling out a second cluster section of the multi-AZ cluster of host computers located in a second AZ in response to the detecting of the degraded state of the first AZ;

automatically detecting a recovered state of the first AZ; and

scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ in response to the detecting of the recovered state of the first AZ.

10 . The non-transitory computer-readable storage medium of claim 9 , wherein automatically detecting the degraded state of the first AZ includes receiving an auto-remediation failure notification of a replacement host computer being unable to be provisioned in a first cluster section of the multi-AZ cluster of host computers located in the first AZ, wherein the auto-remediation failure notification indicates that the first AZ is potentially a degraded AZ.

11 . The non-transitory computer-readable storage medium of claim 10 , wherein automatically detecting the degraded state of the first AZ includes determining that the first AZ is a degraded AZ when a virtualization manager for the multi-AZ cluster of host computers has lost connection to all the hosts in the first cluster section of the multi-AZ cluster of host computers located in the first AZ.

12 . The non-transitory computer-readable storage medium of claim 9 , wherein scaling out the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes adding a number of new host computers to the second cluster section of the multi-AZ cluster of host computers, wherein the number of new host computers equals a number of host computers in the first cluster section of the multi-AZ cluster of host computers located in the first AZ when each of the first and cluster sections of the multi-AZ cluster of host computers includes two or three host computers.

13 . The non-transitory computer-readable storage medium of claim 9 , wherein scaling out the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes adding a number of new host computers to the second cluster section of the multi-AZ cluster of host computers, wherein the number of new host computers equals half of a number of host computers in the first cluster section of the multi-AZ cluster of host computers located in the first AZ when each of the first and cluster sections of the multi-AZ cluster of host computers includes more than three host computers.

14 . The non-transitory computer-readable storage medium of claim 9 , wherein the recovery workflow further includes determining that the first AZ has recovered when the health states of all the host computers in the section of the random multi-AZ cluster that is located in the first AZ are healthy.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ is part of scaling in all multi-AZ clusters of host computers located partially in the first AZ based the successful scale-in operation of the another multi-AZ cluster located partially in the first AZ.

16 . The non-transitory computer-readable storage medium of claim 9 , wherein scaling in the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes removing a number of host computers in the second cluster section of the multi-AZ cluster of host computers, wherein the number of host computers removed is equal to a number of host computers that were added to the second cluster section of the multi-AZ cluster of host computers located in the second AZ during the scaling out.

17 . A system comprising:

memory; and

at least one processor configured to:

automatically detect a degraded state of a first availability zone (AZ) in a cloud computing environment having a first cluster section of a multi-AZ cluster of host computers located in the first AZ;

scale out a second cluster section of the multi-AZ cluster of host computers located in a second AZ in response to the detecting of the degraded state of the first AZ;

automatically detect a recovered state of the first AZ, wherein automatically detecting the recovered state of the first AZ includes selecting a random multi-AZ cluster located partially in the first AZ to execute a recovery workflow on the random multi-AZ cluster, wherein the recovery workflow includes checking health states of all host computers in a section of the random multi-AZ cluster that is located in the first AZ to determine whether the first AZ has recovered; and

scale in the second cluster section of the multi-AZ cluster of host computers located in the second AZ in response to the detecting of the recovered state of the first AZ.

18 . The system of claim 17 , wherein automatically detecting the degraded state of the first AZ includes receiving an auto-remediation failure notification of a replacement host computer being unable to be provisioned in a first cluster section of the multi-AZ cluster of host computers located in the first AZ, wherein the auto-remediation failure notification indicates that the first AZ is potentially a degraded AZ.

19 . The system of claim 17 , wherein automatically detecting the degraded state of the first AZ includes determining that the first AZ is a degraded AZ when a virtualization manager for the multi-AZ cluster of host computers has lost connection to all the hosts in the first cluster section of the multi-AZ cluster of host computers located in the first AZ.

20 . The system of claim 17 , wherein scaling out the second cluster section of the multi-AZ cluster of host computers located in the second AZ includes adding a number of new host computers to the second cluster section of the multi-AZ cluster of host computers, wherein the number of new host computers equals a number of host computers in the first cluster section of the multi-AZ cluster of host computers located in the first AZ when each of the first and cluster sections of the multi-AZ cluster of host computers includes two or three host computers.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2024
From: PARMAR, PIYUSH; SAXENA, PAWAN; TARASUK-LEVIN, GABRIEL; SHAH, DHAVAL; MARGI, UMESHA
To: VMWARE, INC.
Reel/Frame 067351/0545 →
CHANGE OF NAME Recorded May 8, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 067355/0001 →
Priority Claims (1)
IN 202141044762 · Oct 1, 2021 · national
Continuity (2)
Continuation 17548635 · Dec 13, 2021
Related Publication 20240036961A1 · Feb 1, 2024
References Cited (7)
US 8010829B1 · Chatterjee · 2011 [cited by examiner]
US 20140047264A1 · Wang · 2014 [cited by examiner]
US 20160366220A1 · Gottlieb · 2016 [cited by examiner]
US 20170286518A1 · Horowitz · 2017 [cited by examiner]
US 20170316078A1 · Funke · 2017 [cited by examiner]
US 20190342149A1 · Guo · 2019 [cited by examiner]
US 20200236159A1 · Shang · 2020 [cited by examiner]