IP Library Granted Patent US 11,841,759
Granted Patent B2
US 11,841,759 · App. 17/657,836 · Granted Dec 12, 2023

Fault tolerance handling for services

Inventors: Santhosh Sreenivasaiah (Woodinville, WA); Mansi Shah (San Jose, CA)
Assignee: VMware, Inc.
G06F11/0754G06F11/004G06F11/0709G06F2201/81
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,841,759
App. No.
17/657,836
Granted
Dec 12, 2023
Kind
B2
Abstract

The disclosure provides an approach for fault tolerance handling. Embodiments include determining, by a management component, that a host stores data relating to a service. Embodiments include receiving, by the management component, fault tolerance information from the service, the fault tolerance information comprising first information about host failures tolerated by the service and second information about existing host failures related to the service. Embodiments include determining, by the management component, based on the fault tolerance information from the service, whether the service will tolerate the host becoming unavailable. Embodiments include performing, by the management component, one or more actions based on the determining of whether the service will tolerate the host becoming unavailable.

Claims (50)

1. A method of fault tolerance handling, comprising:

determining, by a management component, that a host stores data relating to a service;

receiving, by the management component, fault tolerance information from the service, the fault tolerance information comprising:

first information about host failures tolerated by the service; and

second information about existing host failures related to the service;

determining, by the management component, based on the fault tolerance information from the service, whether the service will tolerate the host becoming unavailable;

performing, by the management component, one or more actions based on the determining of whether the service will tolerate the host becoming unavailable, wherein the one or more actions comprise placing the host in a maintenance mode;

determining, by the management component, that a spare host is available; and

re-creating, by the management component, the data relating to the service from the host on the spare host.

2. The method of claim 1 , wherein receiving, by the management component, the fault tolerance information from the service comprises receiving the fault tolerance information from a data storage entity, wherein the fault tolerance information was published to the data storage entity by the service.

3. The method of claim 1 , further comprising determining, by the management component, that the fault tolerance information is not stale based on metadata associated with the fault tolerance information.

4. The method of claim 1 , further comprising determining, by the management component, based on the fault tolerance information from the service, whether to place the host in the maintenance mode based on the determining of whether the service will tolerate the host becoming unavailable.

5. The method of claim 1 , wherein the one or more actions comprise delaying placing the host in the maintenance mode.

6. The method of claim 5 , wherein the one or more actions further comprise:

receiving, by the management component, updated fault tolerance information from the service; and

determining, by the management component, based on the updated fault tolerance information, to place the host in the maintenance mode.

7. The method of claim 1 , further comprising:

determining, by the management component, that the host stores additional data relating to an additional service; and

receiving, by the management component, additional fault tolerance information from the additional service, wherein the performing, by the management component, the one or more actions is further based on the additional fault tolerance information.

8. The method of claim 7 , further comprising determining, by the management component, based on the fault tolerance information and the additional fault tolerance information, that the service will tolerate the host becoming unavailable and that the additional service will not tolerate the host becoming unavailable, wherein the performing, by the management component, the one or more actions comprises delaying placing the host in the maintenance mode.

9. A system for processing application programming interface (API) requests, the system comprising:

at least one memory; and

at least one processor coupled to the at least one memory, the at least one processor and the at least one memory configured to:

determine, by a management component, that a host stores data relating to a service;

receive, by the management component, fault tolerance information from the service, the fault tolerance information comprising:

first information about host failures tolerated by the service; and

second information about existing host failures related to the service;

determine, by the management component, based on the fault tolerance information from the service, whether the service will tolerate the host becoming unavailable;

perform, by the management component, one or more actions based on the determining of whether the service will tolerate the host becoming unavailable, wherein the one or more actions comprise placing the host in a maintenance mode;

determine, by the management component, that a spare host is available; and

re-create, by the management component, the data relating to the service from the host on the spare host.

10. The system of claim 9 , wherein receiving, by the management component, the fault tolerance information from the service comprises receiving the fault tolerance information from a data storage entity, wherein the fault tolerance information was published to the data storage entity by the service.

11. The system of claim 9 , wherein the at least one processor and the at least one memory are further configured to determine, by the management component, that the fault tolerance information is not stale based on metadata associated with the fault tolerance information.

12. The system of claim 9 , wherein the least one processor and the at least one memory are further configured to determine, by the management component, based on the fault tolerance information from the service, whether to place the host in the maintenance mode based on the determining of whether the service will tolerate the host becoming unavailable.

13. The system of claim 9 , wherein the one or more actions comprise delaying placing the host in the maintenance mode.

14. The system of claim 13 , wherein the one or more actions comprise:

receiving, by the management component, updated fault tolerance information from the service; and

determining, by the management component, based on the updated fault tolerance information, to place the host in the maintenance mode.

15. The system of claim 9 , wherein the at least one processor and the at least one memory are further configured to:

determine, by the management component, that the host stores additional data relating to an additional service; and

receive, by the management component, additional fault tolerance information from the additional service, wherein the performing, by the management component, the one or more actions is further based on the additional fault tolerance information.

16. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:

determine, by a management component, that a host stores data relating to a service;

receive, by the management component, fault tolerance information from the service, the fault tolerance information comprising:

first information about host failures tolerated by the service; and

second information about existing host failures related to the service;

determine, by the management component, based on the fault tolerance information from the service, whether the service will tolerate the host becoming unavailable;

perform, by the management component, one or more actions based on the determining of whether the service will tolerate the host becoming unavailable, wherein the one or more actions comprise placing the host in a maintenance mode;

determine, by the management component, that a spare host is available; and

re-create, by the management component, the data relating to the service from the host on the spare host.

Assignments (2)
CHANGE OF NAME Recorded Apr 15, 2024
From: VMWARE, INC.
To: VMWARE LLC
Reel/Frame 067102/0395 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2022
From: SREENIVASAIAH, SANTHOSH; SHAH, MANSI
To: VMWARE, INC.
Reel/Frame 059490/0342 →
Continuity (1)
Related Publication 20230315554A1 · Oct 5, 2023