Content-aware anomaly detection and diagnosis
Methods and systems for detecting a system fault include determining a network of broken correlations for a current timestamp, relative to a predicted set of correlations, based on a current set of sensor data. The network of broken correlations for the current timestamp is compared to networks of broken correlations for previous timestamps to determine a fault propagation pattern. It is determined whether a fault has occurred based on the fault propagation pattern. A system management action is performed if a fault has occurred.
1. A method for detecting a system fault, comprising:
determining a network of broken correlations for a current timestamp, relative to a predicted set of correlations, based on a current set of sensor data;
comparing the network of broken correlations for the current timestamp to networks of broken correlations for previous timestamps by determining a precision curve and a recall curve to determine a fault propagation pattern;
determining that a fault has occurred based on the fault propagation pattern using a processor; and
automatically shutting down one or more systems, by the processor, responsive to the determination that a fault has occurred to prevent the fault from propagating;
wherein comparing the network of broken correlations for the current timestamp to networks of broken correlations for previous timestamps further comprises determining a time range between the current timestamp and a first timestamp at which a value of the precision curve drops below a threshold.
2. The method of claim 1 , wherein comparing the network of broken correlations for the current timestamp to networks of broken correlations for previous timestamps further comprises determining a behavior of the recall curve within the time range.
3. The method of claim 2 , wherein determining whether a fault has occurred comprises determining that a fault has occurred if the recall curve increases monotonically in the time range.
4. A system for detecting a fault comprising:
a computing device including a processor and a memory operatively coupled to the processor, said memory having stored thereon computer executable instructions that when executed by the processor cause the system to execute:
an invariant graph module configured to determine a network of broken correlations for a current timestamp, relative to a predicted set of correlations, based on a current set of sensor data;
an invariant comparison module configured to compare the network of broken correlations for the current timestamp to networks of broken correlations for previous timestamps by determining a precision curve and a recall curve to determine a fault propagation pattern and to further determine that a fault has occurred based on the fault propagation pattern; and
a fault management module configured to automatically shut down one or more systems, responsive to the determination that a fault has occurred to prevent the fault from propagating;
wherein the invariant comparison module is further configured to determine a time range between the current timestamp and a first timestamp at which a value of the precision curve drops below a threshold.
5. The system of claim 4 , wherein the invariant comparison module is further configured to determine a behavior of the recall curve within the time range.
6. The system of claim 5 , wherein the invariant comparison module is further configured to determine that a fault has occurred if the recall curve increases monotonically in the time range.