IP Library › Granted Patent US 12,737,249
Granted Patent B2
US 12,737,249 · App. 17/648,576 · Granted Sep 15, 2026

Machine learning assisted remediation of networked computing failure patterns

Inventors: Shyamala Gowri (Bangalore, IN); Deepashree Gandhi (Bengaluru, IN); Mrudula Madiraju (Bangalore Urban, IN); Rakhi S Arora (Bangalore, IN); Jaya H Gabhane (Bangalore, IN)
Assignee: International Business Machines Corporation
G06F11/0793G06F11/0709G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,249
App. No.
17/648,576
Granted
Sep 15, 2026
Kind
B2
Abstract

Disclosed are techniques for automatically determining whether a new disruption of service alert corresponds to a pattern of failures and automatically applying remedies based on the determined pattern. Datasets of historical disruption of service alerts on networked computing clusters are used to train a machine learning algorithm to identify patterns between alerts. When a new disruption of service alert is received, historical disruption of service alerts for the originating networked computing cluster are also received and provided as input to the machine learning model. The machine learning model then automatically determines whether the new alert fits a pattern with the historical alerts from the same cluster, and when a fit is found, remedial actions are sourced from the alerts that fit the pattern to be applied automatically to the originating networked computing cluster.

Claims (63)

1 . A computer implemented method (CIM) comprising:

receiving a set of historical disruption of service alerts and their corresponding solutions;

training a machine learning model for determining patterns for disruption of service alerts and their corresponding solutions;

receiving (i) a new disruption of service alert for a first networked computing cluster, (ii) a first set of context information of the new disruption of service, (iii) a corresponding set of historical disruption of service events for the first networked computing cluster, and (iv) a second set of context information corresponding to each of the set of historical disruption of service events, wherein the context information includes pending instructions;

determining whether the new disruption of service alert corresponds to a pattern of disruption of service events in the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;

determining a set of automated remedial steps to remedy the new disruption of service alert based, at least in part, on the machine learning model; and

automatically executing the set of automated remedial steps on the first networked computing cluster, wherein the set of automated remedial steps includes restarting one or more computers associated with the new disruption of service alert.

2 . The CIM of claim 1 , wherein networked computing clusters are cloud computing clusters.

3 . The CIM of claim 1 , further comprising:

responsive to determining that the new disruption of service alert corresponds to a pattern of disruption of service events, outputting an alert message to a computer device, wherein the message includes information indicative of the pattern; and

responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, outputting a message to the computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user.

4 . The CIM of claim 3 , wherein automatically executing the set of automated remedial steps on the first networked computing cluster is responsive to receiving user input corresponding to a selection of the one or more subsets in the message, with the automatically executed set of automated remedial steps corresponding to the selected subset of steps selected by the user.

5 . The CIM of claim 1 , further comprising:

responsive to determining that the new disruption of service alert does not correspond to a pattern of disruption of service events, determining a set of similar disruption of service events from the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;

wherein determining the set of automated remedial steps to remedy the new disruption of service alert includes determining one or more subsets of steps to remedy the new disruption of service alert based, at least in part, on the set of similar disruption of service events.

6 . The CIM of claim 5 , further comprising:

responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, communicating a message to a computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user, including information indicative of which similar disruption of service events correspond to the subsets of steps; and

receiving user input corresponding to a selection of at least one subset of steps for automatic execution on the first networked computing cluster;

wherein automatically executing the set of automated remedial steps on the first networked computing cluster corresponds to automatically executing the selected at least one subset of steps on the first networked computing cluster.

7 . A computer program product (CPP) comprising:

one or more computer-readable storage media; and

program instructions stored on the one or more computer-readable storage media to perform operations comprising:

receiving a set of historical disruption of service alerts and their corresponding solutions;

training a machine learning model for determining patterns for disruption of service alerts and their corresponding solutions;

receiving (i) a new disruption of service alert for a first networked computing cluster, (ii) a first set of context information of the new disruption of service, (iii) a corresponding set of historical disruption of service events for the first networked computing cluster, and (iv) a second set of context information corresponding to each of the set of historical disruption of service events, wherein the context information includes pending instructions;

determining whether the new disruption of service alert corresponds to a pattern of disruption of service events in the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;

determining a set of automated remedial steps to remedy the new disruption of service alert based, at least in part, on the machine learning model; and

automatically executing the set of automated remedial steps on the first networked computing cluster, wherein the set of automated remedial steps includes restarting one or more computers associated with the new disruption of service alert.

8 . The CPP of claim 7 , wherein networked computing clusters are cloud computing clusters.

9 . The CPP of claim 7 , wherein the operations further comprise:

responsive to determining that the new disruption of service alert corresponds to a pattern of disruption of service events, outputting an alert message to a computer device, wherein the message includes information indicative of the pattern; and

responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, outputting a message to the computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user.

10 . The CPP of claim 9 , wherein automatically executing the set of automated remedial steps on the first networked computing cluster is responsive to receiving user input corresponding to a selection of the one or more subsets in the message, with the automatically executed set of automated remedial steps corresponding to the selected subset of steps selected by the user.

11 . The CPP of claim 7 , wherein the operations further comprise:

responsive to determining that the new disruption of service alert does not correspond to a pattern of disruption of service events, determining a set of similar disruption of service events from the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;

wherein determining the set of automated remedial steps to remedy the new disruption of service alert includes determining one or more subsets of steps to remedy the new disruption of service alert based, at least in part, on the set of similar disruption of service events.

12 . The CPP of claim 11 , wherein the operations further comprise:

responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, communicating a message to a computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user, including information indicative of which similar disruption of service events correspond to the subsets of steps; and

receiving user input corresponding to a selection of at least one subset of steps for automatic execution on the first networked computing cluster;

wherein automatically executing the set of automated remedial steps on the first networked computing cluster corresponds to automatically executing the selected at least one subset of steps on the first networked computing cluster.

13 . A computer system (CS) comprising:

a processor set;

one or more computer-readable storage media; and

program instructions stored on the one or more computer-readable storage media to cause the processor set to perform operations comprising:

receiving a set of historical disruption of service alerts and their corresponding solutions;

training a machine learning model for determining patterns for disruption of service alerts and their corresponding solutions;

receiving (i) a new disruption of service alert for a first networked computing cluster, (ii) a first set of context information of the new disruption of service, (iii) a corresponding set of historical disruption of service events for the first networked computing cluster, and (iv) a second set of context information corresponding to each of the set of historical disruption of service events, wherein the context information includes pending instructions;

determining whether the new disruption of service alert corresponds to a pattern of disruption of service events in the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;

determining a set of automated remedial steps to remedy the new disruption of service alert based, at least in part, on the machine learning model; and

automatically executing the set of automated remedial steps on the first networked computing cluster, wherein the set of automated remedial steps includes restarting one or more computers associated with the new disruption of service alert.

14 . The CS of claim 13 , wherein networked computing clusters are cloud computing clusters.

15 . The CS of claim 13 , wherein the operations further comprise:

responsive to determining that the new disruption of service alert corresponds to a pattern of disruption of service events, outputting an alert message to a computer device, wherein the message includes information indicative of the pattern; and

responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, outputting a message to the computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user.

16 . The CS of claim 15 , wherein automatically executing the set of automated remedial steps on the first networked computing cluster is responsive to receiving user input corresponding to a selection of the one or more subsets in the message, with the automatically executed set of automated remedial steps corresponding to the selected subset of steps selected by the user.

17 . The CS of claim 13 , wherein the operations further comprise:

responsive to determining that the new disruption of service alert does not correspond to a pattern of disruption of service events, determining a set of similar disruption of service events from the corresponding set of historical disruption of service events for the first networked computing cluster based, at least in part, on the machine learning model;

wherein determining the set of automated remedial steps to remedy the new disruption of service alert includes determining one or more subsets of steps to remedy the new disruption of service alert based, at least in part, on the set of similar disruption of service events.

18 . The CS of claim 17 , wherein the operations further comprise:

responsive to determining the set of automated remedial steps to remedy the new disruption of service alert, communicating a message to a computer device, with the message including the set of automated remedial steps as one or more subsets of steps for selection by a user, including information indicative of which similar disruption of service events correspond to the subsets of steps; and

receiving user input corresponding to a selection of at least one subset of steps for automatic execution on the first networked computing cluster;

wherein automatically executing the set of automated remedial steps on the first networked computing cluster corresponds to automatically executing the selected at least one subset of steps on the first networked computing cluster.

19 . The CIM of claim 1 , wherein the set of automated remedial steps further comprise clearing a memory cache of non-responsive nodes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2022
From: GOWRI, SHYAMALA; GANDHI, DEEPASHREE; MADIRAJU, MRUDULA; ARORA, RAKHI S; GABHANE, JAYA H
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058721/0425 →
Continuity (1)
Related Publication 20230236923A1 · Jul 27, 2023
References Cited (19)
US 7721152B1 · Joshi · 2010 [cited by applicant]
US 10303538B2 · Thomas · 2019 [cited by applicant]
US 11188411B1 · Shen · 2021 [cited by applicant]
US 20140310564A1 · Mallige · 2014 [cited by applicant]
US 20150081885A1 · Thomas · 2015 [cited by applicant]
US 20180173584A1 · Eckstein · 2018 [cited by applicant]
US 20190068622A1 · Lin · 2019 [cited by applicant]
US 20190155678A1 · Hsiong · 2019 [cited by examiner]
US 20190163594A1 · Hayden · 2019 [cited by examiner]
US 20200117531A1 · Sudharsana · 2020 [cited by applicant]
US 20210081842A1 · Polleri · 2021 [cited by examiner]
CN 103366312A · 2013 [cited by applicant]
CN 104125286A · 2014 [cited by applicant]
WO 2023138594A1 · 2023 [cited by applicant]
International Search Report and Written Opinion, International Application No. PCT/CN2023/072764, International Filing Date Jan. 18, 2023, 8 pages. [cited by applicant]
“Customize Business Outcomes with ZIF”, Dec. 2, 2020, Downloaded from the Internet on Apr. 14, 2021, 13 pgs., <https://zif.ai/customize-business-outcomes-with-zif/>. [cited by applicant]
“Moogsoft Observability for DevOps and SREs”, Solution Brief Moogsoft, 2020, 10 pgs., <https://www.moogsoft.com/wp-content/uploads/2020/10/Moogosft-Solution-Brief-100120.pdf>. [cited by applicant]
Chen, et al., “Failure Analysis of Jobs in Compute Clouds: A Google Cluster Case Study”, 2014 IEEE 25th International Symposium on Software Reliability Engineering, 2014, 11 pgs., doi: 10.1109/ISSRE.2014.34. [cited by applicant]
Disclosed Anonymously, “Feedback-Based Automatic Anomaly Detection for Cloud Platforms”, An IP.com Prior Art Database Technical Disclosure, IP.com No. IPCOM000260742D, Dec. 18, 2019, 4 pgs. [cited by applicant]