Systems and methods for performing a root cause analysis utilizing parameters including device context features to troubleshoot enterprise information technology problems
Examples of troubleshooting enterprise information technology problems are described. In some examples, a root cause analysis service configures parameters for a root cause analysis. The parameters include device context features of client devices, which are specified to be analyzed. The root cause analysis service identifies a target issue affecting the client devices and performs the root cause analysis according to the parameters to generate root cause candidate patterns to provide through an interface.
1 . A method, comprising:
receiving, by a root cause analysis service from an analytics platform, timestamped device data corresponding to a plurality of client devices monitored by the analytics platform, the timestamped device data comprising a plurality of device context features of the client devices;
receiving, by the root cause analysis service, a plurality of parameters for a root cause analysis entered via a management console of the analytics platform, wherein the plurality of parameters include a subset of the plurality of device context features to consider in the root cause analysis and a minimum lift threshold value to be used in the root cause analysis;
identifying, by the root cause analysis service, a target issue affecting at least a subset of the client devices;
performing, by the root cause analysis service, the root cause analysis according to the plurality of parameters to generate a plurality of root cause candidate patterns, a respective root cause candidate pattern comprising at least a subset of the device context features, wherein performing the root cause analysis by the root cause analysis service comprises
creating a target set of samples and a normal set of samples from the timestamped device data,
computing a lift value for each of the plurality of root cause candidate patterns based on support of the root cause candidate pattern in the target set and support of the root cause candidate pattern in the normal set, and
filtering the plurality of root cause candidate patterns based on comparing the computed lift value of a respective root cause candidate pattern to the minimum lift threshold value specified via the management console of the analytics platform to produce a subset of the plurality of root cause candidate patterns; and
providing, by at least one of the root cause analysis service and the analytics platform, access to the subset of the plurality of root cause candidate patterns through an application programming interface (API) that is invoked by one or more network services, client devices or management services to cause at least one of the root cause analysis service and the analytics platform to transmit the subset of the plurality of root cause candidate patterns over a network connection,
wherein the management console of the analytics platform generates an estimated runtime indicator for the root cause analysis, wherein modifying the minimum lift threshold value causes an update to the estimated runtime indicator by recomputing an estimated runtime of the root cause analysis based on the modified minimum lift threshold value.
2 . The method according to claim 1 , wherein the device context features correspond to at least one of hardware datapoints and software datapoints for the client devices.
3 . The method according to claim 1 , wherein the target issue comprises a problem or anomaly corresponding to the device context features comprising at least one of: device crashes, device boot time, device shut down time, application hangs corresponding to temporary non-responsiveness, application crashes, increase/decrease of application usage, failed logins, failed application launch, and application launch time.
4 . The method according to claim 1 , wherein the target issue is user-defined through a user interface element generated by the management console of the analytics platform.
5 . The method according to claim 1 , wherein the target issue is programmatically identified based at least in part on an anomaly in the timestamped device data.
6 . The method according to claim 1 , further comprising: providing the subset of the plurality of root cause candidate patterns via the management console of the analytics platform.
7 . A non-transitory computer-readable medium embodying instructions executable by at least one computing device, the instructions, when executed, causing the at least one computing device to at least:
receive, by a root cause analysis service from an analytics platform, timestamped device data corresponding to a plurality of client devices monitored by the analytics platform, the timestamped device data comprising a plurality of device context features of the client devices;
receive, by the root cause analysis service, a plurality of parameters for a root cause analysis entered via a management console of the analytics platform, wherein the plurality of parameters include a subset of the plurality of device context features to consider in the root cause analysis and a minimum lift threshold value to be used in the root cause analysis;
identify, by the root cause analysis service, a target issue affecting at least a subset of the client devices;
perform, by the root cause analysis service, the root cause analysis according to the plurality of parameters to generate a plurality of root cause candidate patterns, a respective root cause candidate pattern comprising at least a subset of the device context features, wherein performing the root cause analysis by the root cause analysis service comprises
creating a target set of samples and a normal set of samples from the timestamped device data,
computing a lift value for each of the plurality of root cause candidate patterns based on support of the root cause candidate pattern in the target set and support of the root cause candidate pattern in the normal set, and
filtering the plurality of root cause candidate patterns based on comparing the computed lift value of a respective root cause candidate pattern to the minimum lift threshold value specified via the management console of the analytics platform to produce a subset of the plurality of root cause candidate patterns; and
provide, by at least one of the root cause analysis service and the analytics platform, access to the subset of the plurality of root cause candidate patterns through an application programming interface (API) that is invoked by one or more network services, client devices or management services to cause at least one of the root cause analysis service and the analytics platform to transmit the subset of the plurality of root cause candidate patterns over a network connection,
wherein the management console of the analytics platform generates an estimated runtime indicator for the root cause analysis, wherein modifying the minimum lift threshold value causes an update to the estimated runtime indicator by recomputing an estimated runtime of the root cause analysis based on the modified minimum lift threshold value.
8 . The non-transitory computer-readable medium according to claim 7 , wherein the device context features correspond to at least one of hardware datapoints and software datapoints for the client devices.
9 . The non-transitory computer-readable medium according to claim 7 , wherein the target issue comprises a problem or anomaly corresponding to the device context features comprising at least one of: device crashes, device boot time, device shut down time, application hangs corresponding to temporary non-responsiveness, application crashes, increase/decrease of application usage, failed logins, failed application launch, and application launch time.
10 . The non-transitory computer-readable medium according to claim 7 , wherein the target issue is user-defined through a user interface element generated by the management console of the analytics platform.
11 . The non-transitory computer-readable medium according to claim 7 , wherein the target issue is programmatically identified based at least in part on an anomaly in the timestamped device data.
12 . The non-transitory computer-readable medium according to claim 7 , further comprising instructions that, when executed, cause the at least one computing device to at least: provide the subset of the plurality of root cause candidate patterns via the management console of the analytics platform.
13 . A system, comprising:
at least one computing device; and
instructions accessible by the at least one computing device, wherein the instructions, when executed, cause the at least one computing device to at least:
receive, by a root cause analysis service from an analytics platform, timestamped device data corresponding to a plurality of client devices monitored by the analytics platform, the timestamped device data comprising a plurality of device context features of the client devices;
receive, by the root cause analysis service, a plurality of parameters for a root cause analysis entered via a management console of the analytics platform, wherein the plurality of parameters include a subset of the plurality of device context features to consider in the root cause analysis and a minimum lift threshold value to be used in the root cause analysis;
identify, by the root cause analysis service, a target issue affecting at least a subset of the client devices;
perform, by the root cause analysis service, the root cause analysis according to the plurality of parameters to generate a plurality of root cause candidate patterns, a respective root cause candidate pattern comprising at least a subset of the device context features, wherein performing the root cause analysis by the root cause analysis service comprises
creating a target set of samples and a normal set of samples from the timestamped device data,
computing a lift value for each of the plurality of root cause candidate patterns based on support of the root cause candidate pattern in the target set and support of the root cause candidate pattern in the normal set, and
filtering the plurality of root cause candidate patterns based on comparing the computed lift value of a respective root cause candidate pattern to the minimum lift threshold value specified via the management console of the analytics platform to produce a subset of the plurality of root cause candidate patterns; and
provide, by at least one of the root cause analysis service and the analytics platform, access to the subset of the plurality of root cause candidate patterns through an application programming interface (API) that is invoked by one or more network services, client devices or management services to cause at least one of the root cause analysis service and the analytics platform to transmit the subset of the plurality of root cause candidate patterns over a network connection,
wherein the management console of the analytics platform generates an estimated runtime indicator for the root cause analysis, wherein modifying the minimum lift threshold value causes an update to the estimated runtime indicator by recomputing an estimated runtime of the root cause analysis based on the modified minimum lift threshold value.
14 . The system of claim 13 , wherein the device context features correspond to at least one of hardware datapoints and software datapoints for the client devices.
15 . The system of claim 13 , wherein the target issue comprises a problem or anomaly corresponding to the device context features comprising at least one of: device crashes, device boot time, device shut down time, application hangs corresponding to temporary non-responsiveness, application crashes, increase/decrease of application usage, failed logins, failed application launch, and application launch time.
16 . The system of claim 13 , wherein the target issue is user-defined through a user interface element generated by the management console of the analytics platform.
17 . The system of claim 13 , wherein the target issue is programmatically identified based at least in part on an anomaly in the timestamped device data.
18 . The system of claim 13 , further comprising instructions which, when executed, cause the at least one computing device to at least: provide the subset of the plurality of root cause candidate patterns via the management console of the analytics platform.