IP Library Granted Patent US 11,824,929
Granted Patent B2
US 11,824,929 · App. 17/872,784 · Granted Nov 21, 2023

Using maintenance mode to upgrade a distributed system

Inventors: Alkesh Shah (Palo Alto, CA); Ramses V. Morales (Palo Alto, CA); Leonid Livshin (Boston, MA); Austin Kramer (Palo Alto, CA); Nitin Nagaraja (Palo Alto, CA); Brian Masao Oki (Palo Alto, CA); Sunil Vajir (Palo Alto, CA)
Assignee: VMware, Inc.
H04L67/1095H04L41/5025H04L67/148H04L67/561
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,824,929
App. No.
17/872,784
Granted
Nov 21, 2023
Kind
B2
Abstract

The present disclosure relates to using maintenance mode to upgrade a distributed system. One method includes determining that a first host of a cluster of a software-defined datacenter (SDDC) is to be upgraded as a part of a rolling upgrade of the hosts of the cluster, wherein the first host is executing a process instance of a cluster store, demoting the process instance to a proxy, creating a replica of the process instance using a different proxy on a second host of the cluster, instructing the first host to enter a maintenance mode, upgrading the first host, and instructing the first host to leave the maintenance mode.

Claims (52)

1. A method, comprising:

determining that a first host of a cluster of a datacenter is to be upgraded as a part of a rolling upgrade of the hosts of the cluster, wherein the first host is executing a process instance of a cluster store;

demoting the process instance to a proxy;

creating a replica of the process instance using a different proxy on a second host of the cluster; and

upgrading the first host while the first host is in a maintenance mode.

2. The method of claim 1 , wherein the method includes creating another replica of the process instance using the proxy on the first host subsequent to instructing the first host to leave the maintenance mode responsive to a determination that a third host is to be upgraded as part of the rolling upgrade.

3. The method of claim 1 , wherein the method includes creating the replica of the process instance using the different proxy on the second host responsive to a determination that the second host is not executing another process instance of the cluster store.

4. The method of claim 1 , wherein the method includes:

determining that the second host is to be upgraded as a part of the rolling upgrade subsequent to instructing the first host to enter the maintenance mode, wherein the second host is executing another process instance of the cluster store;

demoting the other process instance to another proxy;

creating another replica of the process instance using the proxy on the host;

instructing the second host to enter the maintenance mode;

upgrading the second host; and

instructing the second host to leave the maintenance mode.

5. The method of claim 1 , wherein the method includes tracking the different proxy to determine whether the different proxy is capable of being used to create the replica.

6. The method of claim 1 , wherein the method includes maintaining a fault tolerance threshold associated with the cluster.

7. The method of claim 1 , wherein the first host entering the maintenance mode includes stopping or migrating all processes running on the first host.

8. The method of claim 1 , wherein the maintenance mode is a partial maintenance mode, and wherein entering the partial maintenance mode includes stopping or migrating a particular subset of processes running on the first host.

9. The method of claim 8 , wherein the method includes determining that the process instance is not included in the particular subset of processes running on the first host that is to be stopped or migrated according to the partial maintenance mode.

10. The method of claim 9 , wherein the method includes determining that the process instance has a dependency on a process included in the particular subset of processes running on the first host that is to be stopped or migrated according to the partial maintenance mode.

11. The method of claim 10 , wherein the method includes:

demoting the process instance to the proxy;

creating the replica of the process instance using the different proxy on the second host of the cluster;

instructing the first host to enter the partial maintenance mode;

upgrading the first host; and

instructing the first host to leave the partial maintenance mode.

12. A non-transitory machine-readable medium having instructions stored thereon which, when executed by a processor, cause the processor to:

determine that a first host of a cluster of a datacenter is to be upgraded as a part of a rolling upgrade of the hosts of the cluster, wherein the first host is executing a process instance of a cluster store;

demote the process instance to a proxy;

responsive to a determination that a different proxy on a second host of the cluster is capable of being used to create a replica of the process instance:

create the replica of the process instance using the different proxy on the second host;

upgrade the first host while the first host in in a maintenance mode; and

upgrade the first host while the first host is in the maintenance mode responsive to a determination that there is not a proxy in the cluster that is capable of being used to create the replica of the process instance.

13. The medium of claim 12 , including instructions to decrement a fault tolerance threshold associated with the cluster responsive to the determination that there is not the proxy in the cluster that is capable of being used to create the replica of the process instance.

14. The medium of claim 12 , including instructions to, responsive to the determination that there is not the proxy in the cluster that is capable of being used to create the replica of the process instance:

instruct the first host to leave the maintenance mode and instantiate another replica.

15. The medium of claim 14 , including instructions to increment the fault tolerance threshold associated with the cluster responsive to a determination that the other replica is being executed.

16. The medium of claim 12 , wherein the determination that there is not a proxy in the cluster that is capable of being used to create the replica of the process instance includes a determination of a lack of proxy on the second host.

17. The medium of claim 12 , wherein the determination that there is not a proxy in the cluster that is capable of being used to create the replica of the process instance includes a determination that the cluster includes only one host.

18. The medium of claim 12 , including instructions to instruct a plurality of hosts to enter the maintenance mode simultaneously.

19. A system, comprising:

an upgrade engine configured to determine that a first host of a cluster of a datacenter is to be upgraded as a part of a rolling upgrade of the hosts of the cluster, wherein the first host is executing a process instance of a cluster store;

a demotion engine configured to demote the process instance to a proxy;

a proxy available engine configured to, responsive to a determination that a different proxy on a second host of the cluster is capable of being used to create a replica of the process instance:

create the replica of the process instance using the different proxy on the second host; and

upgrade the first host while the first host is in a maintenance mode;

a proxy unavailable engine configured to, responsive to a determination that there is not a proxy in the cluster that is capable of being used to create the replica of the process instance:

decrement a fault tolerance threshold associated with the cluster;

upgrade the first host while the first host is in the maintenance mode;

instruct the first host to leave the maintenance mode and instantiate another replica; and

increment the fault tolerance threshold associated with the cluster responsive to a determination that the other replica is being executed.

20. The system of claim 19 , wherein the second host is not in maintenance mode when the replica of the process instance is created using the different proxy on the second host.

Assignments (2)
CHANGE OF NAME Recorded Feb 27, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 066692/0103 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2022
From: SHAH, ALKESH; MORALES, RAMSES V.; LIVSHIN, LEONID; KRAMER, AUSTIN; NAGARAJA, NITIN; OKI, BRIAN MASAO; VAJIR, SUNIL
To: VMWARE, INC.
Reel/Frame 060609/0384 →
Continuity (2)
Continuation 17384295 · Jul 23, 2021
Related Publication 20230023625A1 · Jan 26, 2023