Detecting a quality-related faulty component and predicting uncorrectable errors incurred by a component using machine learning
Aspects of the disclosure identify a pattern of features associated with a computing component indicating an increased probability that the component incurs an uncorrectable error. Historical component data, component features, server data, and server services are utilized to identify the patterns that correlate to an increased probability that the component incurs an uncorrectable error. For example, this data is used as input into a machine learning platform. Proactive and/or mitigating actions that reduce the probability of an uncorrectable error or its negative effects are presented and/or implemented to minimize or eliminate disruption in cloud computing services.
1 . A system comprising:
a processor;
a historical database comprising historical error information for a plurality of components; and
a computer storage medium comprising computer-executable instructions that, upon execution by the processor, cause the processor to perform the following operations:
identifying, from the historical error information, that a first component of the plurality of components has incurred an uncorrectable error;
extracting a first feature associated with the first component, the first feature associated with the first component comprising:
a feature of the first component,
a feature of a compute node upon which the first component is executed, or
a feature of a service provided by the compute node;
comparing the first feature associated with the first component with a second feature associated with a second component of the plurality of components;
based on the comparing, identifying a third component of the plurality of components comprising a pattern of features that increases a probability of the third component to incur the uncorrectable error;
identifying a feature of the third component from the pattern of features that, responsive to being changed, decreases the probability of the third component to incur the uncorrectable error; and
based on identifying the feature of the third component that responsive to being changed decreases the probability of the third component to incur the uncorrectable error, implementing a mitigating action for the identified feature of the third component comprising migrating a virtual machine from a compute node implementing the third component to another compute node.
2 . The system of claim 1 , wherein the computer storage medium comprises further executable instructions that, on being executed by the processor, further cause the processor to perform the following operations:
filtering out a fourth component in the plurality of components that has incurred a number of errors greater than a threshold number of errors during a period of time.
3 . The system of claim 1 , wherein the computer storage medium comprises further executable instructions that, on being executed by the processor, further cause the processor to perform the following operations:
filtering out a fourth component in the plurality of components that has incurred an error having a severity level greater than a threshold severity level.
4 . The system of claim 1 , wherein the third component is a Dual In-Line Memory Module (DIMM) or a Dynamic Random Access Memory (DRAM).
5 . The system of claim 1 , further comprising a pattern recognition platform coupled to the historical database, the pattern recognition platform generating a pattern failure predictor that provides the computer-executable instructions, the pattern failure predictor trained in part using machine learning.
6 . The system of claim 1 , wherein identifying the feature from the pattern of features that, responsive to being changed, decreases the probability of the third component to incur the uncorrectable error further comprises using a causal inference.
7 . The system of claim 1 , wherein the computer storage medium comprises further executable instructions that, upon being executed by the processor, further cause the processor to perform the following operations:
extracting the first feature associated with the first component further comprising extracting signal data from a telemetry log; and
detecting, based at least on the extracted signal data and a quality of the first component, that the third component has failed or is likely to fail.
8 . A computerized method comprising:
identifying, from historical error information, that a first component of a plurality of components has incurred an uncorrectable error;
extracting a first feature associated with the first component, the first feature associated with the first component comprising:
a feature of the first component,
a feature of a compute node upon which the first component is executed, or
a feature of a service provided by the compute node;
comparing the first feature associated with the first component with a second feature associated with a second component of the plurality of components;
based on the comparing, identifying a third component of the plurality of components comprising a pattern of features that increases a probability of the third component to incur the uncorrectable error;
identifying a feature of the third component from the pattern of features that, responsive to being changed, decreases the probability of the third component to incur the uncorrectable error; and
based on identifying the feature of the third component that responsive to being changed decreases the probability of the third component to incur the uncorrectable error, implementing a mitigating action for the identified feature of the third component comprising migrating a virtual machine from a compute node implementing the third component to another compute node.
9 . The computerized method of claim 8 , further comprising filtering out a fourth component in the plurality of components that has incurred a number of errors greater than a threshold number of errors during a period of time.
10 . The computerized method of claim 8 , further comprising filtering out a fourth component in the plurality of components that has incurred an error having a severity level greater than a threshold severity level.
11 . The computerized method of claim 8 , wherein the third component is a Dual In-Line Memory Module (DIMM) or a Dynamic Random Access Memory (DRAM).
12 . The computerized method of claim 8 , wherein the uncorrectable error is an anomaly-type error.
13 . The computerized method of claim 8 , wherein a machine learning platform identifies the third component comprising the pattern of features that increase the probability of the third component to incur the uncorrectable error prior to an actual failure of the third component.
14 . The computerized method of claim 8 , wherein identifying the feature from the pattern of features that, responsive to being changed, decreases the probability of the third component to incur the uncorrectable error further comprises using a causal inference.
15 . A computer storage medium storing computer-executable instructions that, upon execution by a processor, cause the processor to perform the following:
identifying, from historical error information, that a first component of a plurality of components has incurred an uncorrectable error:
extracting a first feature associated with the first component, the first feature associated with the first component comprising:
a feature of the first component,
a feature of a compute node upon which the first component is executed, or
a feature of a service provided by the compute node;
comparing the first feature associated with the first component with a second feature associated with a second component of the plurality of components;
based on the comparing, identifying a third component of the plurality of components comprising a pattern of features that increases a probability of the third component to incur the uncorrectable error;
identifying a feature of the third component from the pattern of features that, responsive to being changed, decreases the probability of the third component to incur the uncorrectable error; and
based on identifying the feature of the third component that responsive to being changed decreases the probability of the third component to incur the uncorrectable error, implementing a mitigating action for the identified feature of the third component comprising migrating a virtual machine from a compute node implementing the third component to another compute node.
16 . The computer storage medium of claim 15 , wherein the computer storage medium comprises further executable instructions that, upon execution by the processor, further cause the processor to perform the following operations:
filtering out a fourth component in the plurality of components that has incurred a number of errors greater than a threshold number of errors during a period of time.
17 . The computer storage medium of claim 15 , wherein the computer storage medium comprises further executable instructions that, upon execution by the processor, further cause the processor to perform the following operations:
filtering out a fourth component in the plurality of components that has incurred an error having a severity level greater than a threshold severity level.
18 . The computer storage medium of claim 15 , wherein identifying the feature from the pattern of features that, responsive to being changed, decreases the probability of the third component to incur the uncorrectable error further comprises using a causal inference.
19 . The computer storage medium of claim 15 , wherein the computer storage medium comprises further executable instructions that, upon execution by the processor, further cause the processor to perform the following operations:
using counter-factual analysis to identify the feature of the third component from the pattern of features that, if changed, decreases the probability of the third component to incur an uncorrectable error.
20 . The computer storage medium of claim 15 , wherein the uncorrectable error is an anomaly-type error.