Automated real-time detection, prediction and prevention of rare failures in industrial system with unlabeled sensor data
Example implementations described herein are directed to management of a system comprising a plurality of apparatuses providing unlabeled sensor data, which can involve executing feature extraction on the unlabeled sensor data to generate a plurality of features; executing failure detection by processing the plurality of features with a failure detection model to generate failure detection labels, the failure detection model generated from a machine learning framework that applies supervised machine learning on unsupervised machine learning models generated from unsupervised machine learning; and providing extracted features and the failure detection label to a failure prediction model to generate failure prediction and a sequence of features.
1 . A method for a system comprising a plurality of industrial apparatuses providing unlabeled sensor data, the method comprising:
executing feature extraction on the unlabeled sensor data to generate a plurality of features;
executing failure detection by processing the plurality of features with a failure detection model to generate failure detection labels, the failure detection model generated from a machine learning framework that applies supervised machine learning on unsupervised machine learning models generated from unsupervised machine learning;
providing extracted features and the failure detection label to a failure prediction model, the failure prediction model being a Long Short-Term Memory (LSTM) sequence prediction model that receives the extracted features and the failure detection labels and outputs both a predicted failure score indicating likelihood of failure and a predicted sequence of features having a same format as the extracted features from the feature extraction;
identifying a root cause of ensemble failures and generating automated remediation recommendations to address the ensemble failures;
generating alerts from the ensemble failures;
executing an alert suppression process with cost-sensitive optimization technique to suppress ones of the alerts based on urgency level;
providing remaining ones of the alerts to one or more operators of the plurality of systems; and
executing processes to control one or more of the plurality of systems based on the remediation recommendations,
wherein the executing processes to control comprises at least one of shutting down, rebooting, or triggering warning indicators on the industrial apparatuses based on the predicted failure and the remediation recommendations.
2 . The method of claim 1 , wherein the machine learning framework generates the failure detection model from applying the supervised machine learning on the unsupervised machine learning models generated from the unsupervised machine learning by:
executing the unsupervised machine learning to generate the unsupervised machine learning models based on the features;
executing supervised machine learning on results from each of the unsupervised machine learning models to generate supervised ensembled machine learning models, each of the supervised ensemble machine learning models corresponding to each of the unsupervised machine learning models; and
selecting ones of the unsupervised machine learning models as the failure detection model based on an evaluation of the results of the unsupervised machine models against predictions generated by the supervised ensemble machine learning models.
3 . The method of claim 1 , wherein the executing feature extraction comprises applying an Exponential Moving Average (EMA) to smooth time series data from the unlabeled sensor data, wherein the EMA places greater weight on most recent data points with weight reducing in exponential order to prior time points.
4 . The method of claim 1 , wherein the executing feature extraction comprises applying a differencing technique to the unlabeled sensor data to generate stationary time series signals, the differencing technique comprising calculating at least one of a first order derivation representing change of sensor values or a second order derivation representing change in the change of sensor values.
5 . The method of claim 1 , wherein the failure detection model comprises an Isolation Forest model configured to perform anomaly detection on the plurality of features, wherein an output of the anomaly detection comprises anomaly scores for instances in the features, the anomaly scores being in a range of [ 0 , 1 ] with higher scores indicating higher likelihood of anomaly.
6 . The method of claim 1 , wherein the LSTM sequence prediction model receives encoded features from an LSTM, the LSTM comprising an encoder component and a decoder component configured to remove redundant information from time series data while retaining signals for failure prediction.
7 . The method of claim 1 , further comprising extracting features from a feature window and failures from a failure window, wherein a lead time window separates an end of the feature window from a start of the failure window, the lead time window providing response time for an operator to respond to predicted failures.
8 . The method of claim 1 , wherein the ensemble failures are generated by ensembling predicted failures from the LSTM sequence prediction model and detected failures from applying the failure detection model to the predicted sequence of features, the ensembling comprising at least one of calculating an average value, a weighted average, a maximum value, or a minimum value of the predicted failures and the detected failures.
9 . The method of claim 1 , wherein the alert suppression process comprises controlling alert generation based on a threshold T for predicted failure score, a count N of predicted failures, and a time period E, wherein a first alert is generated after N predicted failures appear within time period E and the predicted failure score exceeds threshold T.
10 . The method of claim 1 , wherein the identifying the root cause comprises applying a supervised decision tree model to outputs of the failure detection model, wherein the decision tree is followed from a tree root to a leaf to obtain a list of conditions associated with features that lead to the predicted failure.
11 . The method of claim 1 , wherein the cost-sensitive optimization technique comprises minimizing a cost function based on false positive cost associated with predicting a failure when no actual failure exists and false negative cost associated with predicting no failure when an actual failure exists, wherein the false negative cost is larger than the false positive cost.
12 . The method of claim 1 , further comprising maintaining an alert queue storing alerts, wherein alerts with same asset and failure mode are aggregated into an alert group, and the alert groups are ordered by urgency level in descending order, the urgency level determined based on at least one of importance of asset, aggregated failure scores, failure mode, or remediation time and cost.
13 . A method for a system comprising a plurality of apparatuses providing unlabeled sensor data, the method comprising:
executing feature extraction on the unlabeled sensor data to generate a plurality of features;
executing failure detection by processing the plurality of features with a failure detection model to generate failure detection labels, the failure detection model generated from a machine learning framework that applies supervised machine learning on unsupervised machine learning models generated from unsupervised machine learning;
providing extracted features and the failure detection label to a failure prediction model to generate failure prediction and a sequence of features;
identifying a root cause of ensemble failures and generating automated remediation recommendations to address the ensemble failures;
generating alerts from the ensemble failures;
executing an alert suppression process with cost-sensitive optimization technique to suppress ones of the alerts based on urgency level;
providing remaining ones of the alerts to one or more operators of the plurality of systems;
executing processes to control one or more of the plurality of systems based on the remediation recommendations; and
generating the failure prediction model, the generating the failure prediction model comprising:
extracting features from an optimized feature window from historical sensor data;
determining an optimized failure window and a lead time window based on failures from the historical sensor data;
encoding the features with Long Short-Term Memory (LSTM);
training a LSTM sequence prediction model configured to learn patterns in feature sequences from the feature window to derive failure in the failure window;
providing die LSTM sequence prediction model as the failure prediction model; and
ensembling failures from detected failures from the failure detection model and predicted failures from the failure prediction model; wherein the failure prediction is ensemble failures from detected failures and predicted failures.
14 . A method for a system comprising a plurality of industrial apparatuses providing unlabeled data, the method comprising:
executing feature extraction on the unlabeled data to generate a plurality of features;
executing a machine learning framework that transforms unsupervised learning tasks into supervised learning tasks through applying supervised machine learning on unsupervised machine learning models generated from unsupervised machine learning, the executing the machine learning framework comprising:
executing the unsupervised machine learning to generate the unsupervised machine learning models based on the features;
executing supervised machine learning on results from each of the unsupervised machine learning models to generate supervised ensembled machine learning models, each of the supervised ensemble machine learning models corresponding to each of the unsupervised machine learning models;
selecting ones of the unsupervised machine learning models based on an evaluation of the results of the unsupervised machine learning models against predictions by the supervised ensemble machine learning models;
selecting features based on the evaluation results of the unsupervised learning models; and
converting the selected ones of unsupervised learning models to supervised learning models for facilitating explainable artificial intelligence (AI);
identifying a root cause of ensemble failures and generating automated remediation recommendations to address the ensemble failures;
generating alerts from the ensemble failures;
executing an alert suppression process with cost-sensitive optimization technique to suppress ones of the alerts based on urgency level;
providing remaining ones of the alerts to one or more operators of the plurality of systems; and
executing processes to control one or more of the plurality of systems based on the remediation recommendations, wherein the executing processes to control comprises at least one of shutting down, rebooting, or triggering warning indicators on the industrial apparatuses based on the predicted failure and the remediation recommendations.
15 . A non-transitory computer readable medium, storing instructions for management of a system comprising a plurality of industrial apparatuses providing unlabeled sensor data, the instructions comprising:
executing feature extraction on the unlabeled sensor data to generate a plurality of features;
executing failure detection by processing the plurality of features with a failure detection model to generate failure detection labels, the failure detection model generated from a machine learning framework that applies supervised machine learning on unsupervised machine learning models generated from unsupervised machine learning;
providing extracted features and the failure detection label to a failure prediction model, the failure prediction model being a Long Short-Term Memory (LSTM) sequence prediction model that receives the extracted features and the failure detection labels and outputs both a predicted failure score indicating likelihood of failure and a predicted sequence of features having a same format as the extracted features from the feature extraction;
identifying a root cause of ensemble failures and generating automated remediation recommendations to address the ensemble failures;
generating alerts from the ensemble failures;
executing an alert suppression process with cost-sensitive optimization technique to suppress ones of the alerts based on urgency level;
providing remaining ones of the alerts to one or more operators of the plurality of systems; and
executing processes to control one or more of the plurality of systems based on the remediation recommendations,
wherein the executing processes to control comprises at least one of shutting down, rebooting, or triggering warning indicators on the industrial apparatuses based on the predicted failure and the remediation recommendations.
16 . The non-transitory computer readable medium of claim 15 , wherein the machine learning framework generates the failure detection model from applying the supervised machine learning on the unsupervised machine learning models generated from the unsupervised machine learning by:
executing the unsupervised machine learning to generate the unsupervised machine learning models based on the features;
executing supervised machine learning on results from each of the unsupervised machine learning models to generate supervised ensembled machine learning models, each of the supervised ensemble machine learning models corresponding to each of the unsupervised machine learning models; and
selecting ones of the unsupervised machine learning models as the failure detection model based on an evaluation of the results of the unsupervised machine learning models against predictions generated by the supervised ensemble machine learning models.
17 . The non-transitory computer readable medium of claim 15 , wherein generating the LSTM sequence prediction model comprises:
extracting features from an optimized feature window from historical sensor data;
determining an optimized failure window and a lead time window based on failures from the historical sensor data;
encoding the features with Long Short-Term Memory (LSTM);
training the LSTM sequence prediction model configured to learn patterns in feature sequences from the feature window to derive failure in the failure window; and
providing the LSTM sequence prediction model as the failure prediction model; and
ensembling failures from detected failures from the failure detection model and predicted failures from the failure prediction model; wherein the failure prediction is ensemble failures from detected failures and predicted failures.