IP Library Granted Patent US 12664042
Granted Patent B2
US 12664042 · App. 18/893,290 · Granted Jun 23, 2026

Analysis and remediation of service availability based on failures to tolerate

Inventors: Steven Soumpholphakdy (Austin, TX); Michael L. Burriss (Raleigh, NC)
Assignee: DELL PRODUCTS L.P.
G06F11/0793G06F11/0709
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664042
App. No.
18/893,290
Granted
Jun 23, 2026
Kind
B2
Abstract

A method facilitating analysis and remediation of service availability based on failures to tolerate includes, in response to detecting a change to an operational state of a computing cluster, comparing a number of first instances of a computing service running at respective first locations of the computing cluster to a threshold number of instances, the threshold number of instances being defined based on a number of failures, associated with the computing service, that is able to be tolerated; and in response to determining that the number of first instances is equal to the threshold number of instances, initializing at least one second location of the computing cluster for performance of the computing service, the at least one second location being distinct from the respective first locations; and instantiating at least one second instance of the computing service at the at least one second location of the computing cluster.

Claims (41)

1 . A system, comprising:

at least one processor; and

at least one memory that stores executable instructions that, when executed by the at least one processor, facilitate performance of operations, the operations comprising:

in response to detecting a change to an operational state of a computing cluster, comparing a number of first instances of a computing service running at respective first locations of the computing cluster to a threshold number of instances of the computing service, the threshold number of instances being associated with maintaining continued functionality of the computing service on the computing cluster; and

in response to determining, as a result of the comparing, that the number of first instances of the computing service is no greater than the threshold number of instances:

initializing at least one second location of the computing cluster for performance of the computing service, the at least one second location being distinct from each of the respective first locations, the initializing comprising converting a cluster node corresponding to the at least one second location from a storage node that provides storage via a clustered file system associated with the computing cluster to a diskless node that is not part of the clustered file system; and

instantiating at least one second instance of the computing service at the at least one second location of the computing cluster.

2 . The system of claim 1 , wherein the respective first locations of the computing cluster comprise respective first nodes of the computing cluster, and wherein the cluster node corresponding to the at least one second location is a second node of the computing cluster, the second node being distinct from the respective first nodes.

3 . The system of claim 2 , wherein the computing cluster is associated with a cloud computing system, and wherein the initializing of the at least one second location comprises initializing the second node as a new diskless virtual machine within the cloud computing system.

4 . The system of claim 1 , wherein the respective first locations of the computing cluster comprise respective first nodes of the computing cluster, wherein the initializing of the at least one second location of the computing cluster comprises selecting, as the at least one second location, a second node of the computing cluster that is determined to be capable of running the computing service based on a configuration of the second node, and wherein the second node is not any of the respective first nodes.

5 . The system of claim 4 , wherein the operations further comprise:

obtaining, prior to the detecting of the change to the operational state of the computing cluster, service registration information associated with the computing service, the service registration information comprising a group of node capabilities associated with running the computing service, wherein the selecting of the second node comprises selecting the second node based on a result of comparing the configuration of the second node to the group of node capabilities.

6 . The system of claim 1 , wherein the computing cluster is associated with a cloud computing system, wherein the initializing of the at least one second location comprises instantiating a new cloud native container within the cloud computing system, and wherein the instantiating of the at least one second instance of the computing service comprises instantiating the at least one second instance of the computing service within the new cloud native container.

7 . The system of claim 1 , wherein the change to the operational state of the computing cluster is a first change, and wherein the operations further comprise:

in response to detecting a second change to the operational state of the computing cluster that is after the first change, repeating the comparing of the number of first instances of the computing service running at the respective first locations of the computing cluster to the threshold number of instances of the computing service; and

in response to determining, as a result of the repeating of the comparing, that the number of first instances of the computing service is greater than the threshold number of instances, de-instantiating the at least one second instance of the computing service at the at least one second location of the computing cluster.

8 . The system of claim 7 , wherein the respective first locations of the computing cluster comprise respective first nodes of the computing cluster, wherein the at least one second location of the computing cluster comprises a second node of the computing cluster, and wherein the operations further comprise:

in further response to determining, as a result of the repeating of the comparing, that the number of first instances of the computing service is greater than the threshold number of instances, removing the second node from the computing cluster.

9 . The system of claim 1 , wherein the operations further comprise:

receiving registration information from the computing service, the registration information comprising data indicative of the threshold number of instances.

10 . The system of claim 1 , wherein the threshold number of instances is a first threshold number of instances, and wherein the operations further comprise:

in response to determining, as a result of the comparing, that the number of first instances of the computing service is no greater than a second threshold number of instances that is greater than the first threshold number of instances, generating an event notification, the event notification comprising an indication of the number of first instances of the computing service.

11 . The system of claim 1 , wherein the initializing of the second location further comprises rebooting the cluster node in response to the converting of the cluster node, and wherein the instantiating of the at least one second instance of the computing service is in response to the rebooting being determined to have completed successfully.

12 . A method, comprising:

in response to detecting a change to an operational state of a computing cluster, comparing, by a system comprising at least one processor, a number of first instances of a computing service running at respective first locations of the computing cluster to a threshold number of instances of the computing service, wherein the threshold number of instances is defined based on a number of failures, associated with the computing service, that is able to be tolerated; and

in response to determining that the number of first instances of the computing service is equal to the threshold number of instances:

initializing, by the system, at least one second location of the computing cluster for performance of the computing service, the at least one second location being distinct from each of the respective first locations, the initializing comprising converting a cluster node corresponding to the at least one second location from a storage node that provides storage via a clustered file system associated with the computing cluster to a diskless node that is not part of the clustered file system; and

instantiating, by the system, at least one second instance of the computing service at the at least one second location of the computing cluster.

13 . The method of claim 12 , wherein the respective first locations of the computing cluster comprise respective first nodes of the computing cluster, and wherein the cluster node corresponding to the at least one second location is a second node of the computing cluster.

14 . The method of claim 13 , wherein the computing cluster is associated with a cloud computing system, and wherein the initializing of the at least one second location comprises initializing the second node as a new diskless virtual machine within the cloud computing system.

15 . The method of claim 12 , wherein the respective first locations of the computing cluster comprise respective first nodes of the computing cluster, wherein the initializing of the at least one second location of the computing cluster comprises selecting, as the at least one second location, a second node of the computing cluster that is determined to be capable of running the computing service based on a configuration of the second node, and wherein the second node is not any of the respective first nodes.

16 . A non-transitory machine-readable medium comprising computer executable instructions that, when executed by at least one processor, facilitate performance of operations, the operations comprising:

comparing, in response to detecting a change to an operational state of a computing cluster, a number of first instances of a computing service running at respective first locations of the computing cluster to a threshold number of instances of the computing service, the threshold number of instances being associated with a failures-to-tolerate threshold of the computing service; and

in response to determining that the number of first instances of the computing service is less than or equal to the threshold number of instances:

initializing at least one second location of the computing cluster for performance of the computing service, the at least one second location being different from any of the respective first locations, the initializing comprising converting a cluster node corresponding to the at least one second location from a storage node that provides storage via a clustered file system associated with the computing cluster to a diskless node that is not part of the clustered file system; and

starting at least one second instance of the computing service at the at least one second location of the computing cluster.

17 . The non-transitory machine-readable medium of claim 16 , wherein the respective first locations of the computing cluster comprise respective first nodes of the computing cluster, and wherein the cluster node corresponding to the at least one second location of the computing cluster is a second node of the computing cluster.

18 . The non-transitory machine-readable medium of claim 17 , wherein the operations further comprise:

adding the second node to the computing cluster in response to the converting of the second node.

19 . The non-transitory machine-readable medium of claim 17 , wherein the computing cluster is associated with a cloud computing system, and wherein the initializing of the at least one second location comprises initializing the second node as a new service container within the cloud computing system.

20 . The non-transitory machine-readable medium of claim 16 , wherein the respective first locations of the computing cluster comprise respective first nodes of the computing cluster, wherein the initializing of the at least one second location of the computing cluster comprises selecting, as the at least one second location, a second node of the computing cluster that is determined to be capable of running the computing service based on a configuration of the second node, and wherein the second node is not any of the respective first nodes.