Machine learning assisted remediation of networked computing failure patterns
Disclosed are techniques for automatically determining whether a new disruption of service alert corresponds to a pattern of failures and automatically applying remedies based on the determined pattern. Datasets of historical disruption of service alerts on networked computing clusters are used to train a machine learning algorithm to identify patterns between alerts. When a new disruption of service alert is received, historical disruption of service alerts for the originating networked computing cluster are also received and provided as input to the machine learning model. The machine learning model then automatically determines whether the new alert fits a pattern with the historical alerts from the same cluster, and when a fit is found, remedial actions are sourced from the alerts that fit the pattern to be applied automatically to the originating networked computing cluster.
1 . A computer implemented method (CIM) comprising:
receiving a set of historical disruption of service alerts and their corresponding solutions;
training a machine learning model for determining patterns for disruption of service alerts and their corresponding solutions;
receiving (i) a new disruption of service alert for a first networked computing cluster, (ii) a first set of context information of the new disruption of service, (iii) a corresponding set of historical disruption of service events for the first networked computing cluster, and (iv) a second set of context information corresponding to each of the set of historical disruption of service events, wherein the context information includes pending instructions;
determining whether the new disruption of service alert corresponds to a pattern of disruption of service events in the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;
determining a set of automated remedial steps to remedy the new disruption of service alert based, at least in part, on the machine learning model; and
automatically executing the set of automated remedial steps on the first networked computing cluster, wherein the set of automated remedial steps includes restarting one or more computers associated with the new disruption of service alert.
2 . The CIM of claim 1 , wherein networked computing clusters are cloud computing clusters.
3 . The CIM of claim 1 , further comprising:
responsive to determining that the new disruption of service alert corresponds to a pattern of disruption of service events, outputting an alert message to a computer device, wherein the message includes information indicative of the pattern; and
responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, outputting a message to the computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user.
4 . The CIM of claim 3 , wherein automatically executing the set of automated remedial steps on the first networked computing cluster is responsive to receiving user input corresponding to a selection of the one or more subsets in the message, with the automatically executed set of automated remedial steps corresponding to the selected subset of steps selected by the user.
5 . The CIM of claim 1 , further comprising:
responsive to determining that the new disruption of service alert does not correspond to a pattern of disruption of service events, determining a set of similar disruption of service events from the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;
wherein determining the set of automated remedial steps to remedy the new disruption of service alert includes determining one or more subsets of steps to remedy the new disruption of service alert based, at least in part, on the set of similar disruption of service events.
6 . The CIM of claim 5 , further comprising:
responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, communicating a message to a computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user, including information indicative of which similar disruption of service events correspond to the subsets of steps; and
receiving user input corresponding to a selection of at least one subset of steps for automatic execution on the first networked computing cluster;
wherein automatically executing the set of automated remedial steps on the first networked computing cluster corresponds to automatically executing the selected at least one subset of steps on the first networked computing cluster.
7 . A computer program product (CPP) comprising:
one or more computer-readable storage media; and
program instructions stored on the one or more computer-readable storage media to perform operations comprising:
receiving a set of historical disruption of service alerts and their corresponding solutions;
training a machine learning model for determining patterns for disruption of service alerts and their corresponding solutions;
receiving (i) a new disruption of service alert for a first networked computing cluster, (ii) a first set of context information of the new disruption of service, (iii) a corresponding set of historical disruption of service events for the first networked computing cluster, and (iv) a second set of context information corresponding to each of the set of historical disruption of service events, wherein the context information includes pending instructions;
determining whether the new disruption of service alert corresponds to a pattern of disruption of service events in the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;
determining a set of automated remedial steps to remedy the new disruption of service alert based, at least in part, on the machine learning model; and
automatically executing the set of automated remedial steps on the first networked computing cluster, wherein the set of automated remedial steps includes restarting one or more computers associated with the new disruption of service alert.
8 . The CPP of claim 7 , wherein networked computing clusters are cloud computing clusters.
9 . The CPP of claim 7 , wherein the operations further comprise:
responsive to determining that the new disruption of service alert corresponds to a pattern of disruption of service events, outputting an alert message to a computer device, wherein the message includes information indicative of the pattern; and
responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, outputting a message to the computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user.
10 . The CPP of claim 9 , wherein automatically executing the set of automated remedial steps on the first networked computing cluster is responsive to receiving user input corresponding to a selection of the one or more subsets in the message, with the automatically executed set of automated remedial steps corresponding to the selected subset of steps selected by the user.
11 . The CPP of claim 7 , wherein the operations further comprise:
responsive to determining that the new disruption of service alert does not correspond to a pattern of disruption of service events, determining a set of similar disruption of service events from the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;
wherein determining the set of automated remedial steps to remedy the new disruption of service alert includes determining one or more subsets of steps to remedy the new disruption of service alert based, at least in part, on the set of similar disruption of service events.
12 . The CPP of claim 11 , wherein the operations further comprise:
responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, communicating a message to a computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user, including information indicative of which similar disruption of service events correspond to the subsets of steps; and
receiving user input corresponding to a selection of at least one subset of steps for automatic execution on the first networked computing cluster;
wherein automatically executing the set of automated remedial steps on the first networked computing cluster corresponds to automatically executing the selected at least one subset of steps on the first networked computing cluster.
13 . A computer system (CS) comprising:
a processor set;
one or more computer-readable storage media; and
program instructions stored on the one or more computer-readable storage media to cause the processor set to perform operations comprising:
receiving a set of historical disruption of service alerts and their corresponding solutions;
training a machine learning model for determining patterns for disruption of service alerts and their corresponding solutions;
receiving (i) a new disruption of service alert for a first networked computing cluster, (ii) a first set of context information of the new disruption of service, (iii) a corresponding set of historical disruption of service events for the first networked computing cluster, and (iv) a second set of context information corresponding to each of the set of historical disruption of service events, wherein the context information includes pending instructions;
determining whether the new disruption of service alert corresponds to a pattern of disruption of service events in the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;
determining a set of automated remedial steps to remedy the new disruption of service alert based, at least in part, on the machine learning model; and
automatically executing the set of automated remedial steps on the first networked computing cluster, wherein the set of automated remedial steps includes restarting one or more computers associated with the new disruption of service alert.
14 . The CS of claim 13 , wherein networked computing clusters are cloud computing clusters.
15 . The CS of claim 13 , wherein the operations further comprise:
responsive to determining that the new disruption of service alert corresponds to a pattern of disruption of service events, outputting an alert message to a computer device, wherein the message includes information indicative of the pattern; and
responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, outputting a message to the computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user.
16 . The CS of claim 15 , wherein automatically executing the set of automated remedial steps on the first networked computing cluster is responsive to receiving user input corresponding to a selection of the one or more subsets in the message, with the automatically executed set of automated remedial steps corresponding to the selected subset of steps selected by the user.
17 . The CS of claim 13 , wherein the operations further comprise:
responsive to determining that the new disruption of service alert does not correspond to a pattern of disruption of service events, determining a set of similar disruption of service events from the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;
wherein determining the set of automated remedial steps to remedy the new disruption of service alert includes determining one or more subsets of steps to remedy the new disruption of service alert based, at least in part, on the set of similar disruption of service events.
18 . The CS of claim 17 , wherein the operations further comprise:
responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, communicating a message to a computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user, including information indicative of which similar disruption of service events correspond to the subsets of steps; and
receiving user input corresponding to a selection of at least one subset of steps for automatic execution on the first networked computing cluster;
wherein automatically executing the set of automated remedial steps on the first networked computing cluster corresponds to automatically executing the selected at least one subset of steps on the first networked computing cluster.
19 . The CIM of claim 1 , wherein the set of automated remedial steps further comprise clearing a memory cache of non-responsive nodes.