Health evaluation and auto remediation based on off-cluster logset
Various systems and methods are presented herein regarding identifying an operational issue are a data server, automatically identifying/implementing an action to fix the operational issue. The data server can be co-located with a collection of data servers in a server cluster. The action can be configured to be specifically implemented at the data server without affecting an operational status of the other data servers in the collection of data servers. The action can be a server reboot/reset instruction, terminate operation of an application, and suchlike. The operational issue can be compared with a prior operational issue having an associated action, wherein the associated action can be utilized as the action to fix the operational issue at the data server. Over time, respective actions implemented at the one or more data servers in the server cluster can be compiled from which a software service pack can be subsequently compiled and distributed.
1 . A system, comprising:
at least one processor, and
a memory coupled to the at least one processor and having instructions stored thereon, wherein, in response to the at least one processor, the instructions facilitate performance of operations, comprising:
receiving a first notification of a current operational issue, wherein the current operational issue is occurring at a data server, wherein the data server is remotely located from the system;
identifying a prior operational issue having at least one feature comparable to the current operational issue according to a defined similarity criterion;
identifying a first action associated with the prior operational issue;
instructing the data server to implement the first action to address the current operational issue;
receiving a second notification from the data server;
in response to determining that the second notification indicates that the first action first action did not fix the current operational issue, identifying a second action associated with the prior operational issue; and
instructing the data server to implement the second action to address the current operational issue.
2 . The system of claim 1 , wherein the data server is included in a collection of servers located in a server cluster.
3 . The system of claim 2 , wherein the first notification comprises an identifier configured to identify at least one of the data server, at least one component included in the data server, an application hosted by the data server, or a location of the server cluster.
4 . The system of claim 1 , wherein the first action comprises at least one of rebooting the data server, power cycling the data server, terminating operation of the data server, terminating operation of an application hosted by the data server, adjusting a system configuration pertaining to the data server, adjusting a configuration of an application implemented on the data server, throttle operation of an application implemented on the data server, adjust an operational threshold of an application hosted on the data server, or adjust an operational threshold of a component pertinent to operation of the data server.
5 . The system of claim 4 , wherein the data server is a first data server included in a collection of servers, wherein the collection of servers comprises an nth data server, while implementing operation of the action on the first data server, a current operational status of the nth server remains unchanged.
6 . The system of claim 1 , wherein the first action implemented at the data server is an edited action, the operations further comprise:
informing a customer support system of the first action; and
in response to the informing, receiving an edit to the first action via information received from the customer support system, to generate the edited action.
7 . The system of claim 1 , wherein the operations further comprise:
receiving a third notification from the data server, wherein the third notification indicates the second action has been implemented at the data server; and
further monitoring operation of the data server for a subsequent operational issue.
8 . The system of claim 1 , wherein the first action and second action are included in a collection of actions associated with the prior operational issue, and wherein second action is determined to have a lower probability of fixing the current operational issue than the first action.
9 . A computer-implemented method, comprising:
receiving, by a device comprising at least one processor, a first notification identifying a current operational issue identified at a data server, wherein the data server is remotely located from the device;
identifying, by the device, a prior operational issue having at least one feature comparable to the current operational issue according to a defined similarity criterion;
identifying, by the device, a first action associated with the prior operational issue;
instructing, by the device, the data server to implement the first action to address the current operational issue;
receiving, by the device, a second notification, wherein the second notification indicates the first action did not fix the current operational issue;
identifying, by the device, a second action associated with the prior operational issue; and
instructing, by the device, the data server to implement the second action to address the current operational issue.
10 . The computer-implemented method of claim 9 , further comprising:
parsing, by the device, a logset reporting operation of the data server; and
identifying, by the device, the current operational issue in the logset.
11 . The computer-implemented method of claim 10 , wherein the logset is generated in accordance with a defined schedule.
12 . The computer-implemented method of claim 9 , wherein the data server is included in a collection of data servers located in a same server cluster.
13 . The computer-implemented method of claim 12 , wherein the data server is a first data server, wherein the collection of servers further comprises a second data server, and wherein the first action is configured for implementation at the first data server, while operation of the second server remains unchanged as a function of the first action being implemented on the first data server.
14 . The computer-implemented method of claim 13 , wherein the first action comprises at least one of rebooting the first data server, power cycling the first data server, terminating operation of the first data server, terminating operation of an application hosted by the first data server, adjusting a system configuration pertaining to the first data server, adjusting a configuration of an application implemented on the first data server, throttle operation of an application implemented on the first data server, adjust an operational threshold of an application hosted on the first data server, or adjust an operational threshold of a component pertinent to operation of the first data server.
15 . The computer-implemented method of claim 9 , wherein the first action and second action are included in a collection of actions associated with the prior operational issue, and wherein second action is determined to have a lower probability of fixing the current operational issue than the first action.
16 . The computer-implemented method of claim 9 , further comprising:
receiving, by the device, a third notification from the data server, wherein the third notification indicates the second action has been implemented at the data server; and
further monitoring, by the device, operation of the data server for a subsequent operational issue.
17 . A computer program product stored on a non-transitory computer-readable medium and comprising machine-executable instructions, wherein, in response to being executed, the machine-executable instructions cause computing equipment to perform operations, comprising:
receiving a first notification of a current operational issue occurring at a data server, wherein the data server is remotely located from the computing equipment;
identifying a prior operational issue having at least one feature comparable to the current operational issue according to a defined similarity criterion;
identifying a first action associated with the prior operational issue;
instructing the data server to implement the first action;
receiving a second notification indicating the first action did not fix the current operational issue;
identifying a second action associated with the prior operational issue; and
instructing the data server to implement the second action to address the current operational issue.
18 . The computer program product according to claim 17 , wherein the first action comprises at least one of rebooting the data server, power cycling the data server, terminating operation of the data server, terminating operation of an application hosted by the data server, adjusting a system configuration pertaining to the data server, adjusting a configuration of an application implemented on the data server, throttle operation of an application implemented on the data server, adjust an operational threshold of an application hosted on the data server, or adjust an operational threshold of a component pertinent to operation of the data server.
19 . The computer program product according to claim 17 , wherein the current operational issue is a first current operational issue, wherein the data server is a first data server, and wherein the operations further comprise:
receiving a second current operational issue, wherein the second current operational issue is received from a second data server;
determining the second current operational issue is comparable to the first current operational issue according to the defined similarity criterion; and
instructing the first action be implemented on the second data server to address the second current operational issue.
20 . The computer program product according to claim 17 , wherein the operations further comprising:
receiving a third notification from the data server, wherein the third notification indicates the second action has been implemented at the data server; and
further monitoring operation of the data server for a subsequent operational issue.