IP Library › Granted Patent US 10,338,982
Granted Patent B2
US 10,338,982 · App. 15/397,627 · Granted Jul 2, 2019

Hybrid and hierarchical outlier detection system and method for large scale data protection

Inventors: Mu Qiao (Belmont, CA); Ramani R. Routray (San Jose, CA); Quan Zhang (Detroit, MI)
Assignee: International Business Machines Corporation
G06F11/0751G06F11/0727G06F11/1448G06F11/1458G06F17/10G06F17/18G06K9/0053G06K9/0055G06K9/6247G06K9/6284G06K9/6292G06N3/00G06N7/00G06F2212/1032
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,338,982
App. No.
15/397,627
Granted
Jul 2, 2019
Kind
B2
Abstract

One embodiment provides a method comprising receiving metadata comprising univariate time series data for each variable of a multivariate time series. The method comprises, for each variable of the multivariate time series, applying a hybrid and hierarchical model selection process to select an anomaly detection model suitable for the variable based on corresponding univariate time series data for the variable and covariations and interactions between the variable and at least one other variable of the multivariate time series, and detecting an anomaly on the variable utilizing the anomaly detection model selected for the variable. Based on each anomaly detection model selected for each variable of the multivariate time series, the method further comprises performing ensemble learning to determine whether the multivariate time series is anomalous at a particular time point.

Claims (78)

1. A method comprising:

receiving metadata associated with one or more data backup jobs performed on one of more storage devices, wherein the metadata comprises univariate time series data for each variable of a multivariate time series, and the multivariate time series comprises different variables that exhibit different characteristics over time; and

decreasing likelihood of a failure in data protection involving the one or more data backup jobs by:

for each variable of the multivariate time series:

selecting, from different anomaly detection models with different performance costs, an anomaly detection model suitable for the variable based on one or more characteristics exhibited by corresponding univariate time series data for the variable and covariations and interactions between the variable and at least one other variable of the multivariate time series; and

detecting an anomaly on the variable utilizing the anomaly detection model selected for the variable; and

based on each anomaly detection model selected for each variable of the multivariate time series,

determining whether the multivariate time series is anomalous at a particular time point, and generating data indicative of whether the multivariate time series is anomalous at the particular time point.

2. The method of claim 1 , further comprising:

for each variable of the multivariate time series:

determining whether corresponding univariate time series data for the variable exhibits seasonality;

in response to determining the corresponding univariate time series data for the variable exhibits seasonality, grouping the variable into a first variable set; and

in response to determining the corresponding univariate time series data for the variable does not exhibit seasonality, grouping the variable into a second variable set.

3. The method of claim 2 , further comprising:

for the first variable set:

applying dynamic factor analysis to identify one or more common trends among the first variable set; and

for each variable of the first variable set, detecting an anomaly on the variable using seasonal-trend decomposition and the common trends.

4. The method of claim 2 , further comprising:

for the second variable set:

applying vector autoregression (VAR) modeling to all the variables;

for each variable of the second variable set, determining a goodness-of-fit value for the variable based on the VAR modeling; and

determining whether VAR fits the multivariate time series based on an average of each goodness-of-fit value determined for each variable of the second variable set.

5. The method of claim 4 , further comprising:

in response to determining VAR fits the multivariate time series, detecting an anomaly on each variable of the second variable set using VAR.

6. The method of claim 4 , further comprising:

in response to determining VAR does not fit the multivariate time series, detecting an anomaly on variables of the second variable set using local outlier factor.

7. A system comprising:

at least one processor; and

a non-transitory processor-readable memory device storing instructions that when executed by the at least one processor causes the at least one processor to perform operations including:

receiving metadata associated with one or more data backup jobs performed on one of more storage devices, wherein the metadata comprises univariate time series data for each variable of a multivariate time series, and the multivariate time series comprises different variables that exhibit different characteristics over time; and

decreasing likelihood of a failure in data protection involving the one or more data backup jobs by:

for each variable of the multivariate time series:

selecting, from different anomaly detection models with different performance costs, an anomaly detection model suitable for the variable based on one or more characteristics exhibited by corresponding univariate time series data for the variable and covariations and interactions between the variable and at least one other variable of the multivariate time series; and

detecting an anomaly on the variable utilizing the anomaly detection model selected for the variable; and

based on each anomaly detection model selected for each variable of the multivariate time series, determining whether the multivariate time series is anomalous at a particular time point, and generating data indicative of whether the multivariate time series is anomalous at the particular time point.

8. The system of claim 7 , the operations further comprising:

for each variable of the multivariate time series:

determining whether corresponding univariate time series data for the variable exhibits seasonality;

in response to determining the corresponding univariate time series data for the variable exhibits seasonality, grouping the variable into a first variable set; and

in response to determining the corresponding univariate time series data for the variable does not exhibit seasonality, grouping the variable into a second variable set.

9. The system of claim 8 , the operations further comprising:

for the first variable set:

applying dynamic factor analysis to identify one or more common trends among the first variable set; and

for each variable of the first variable set, detecting an anomaly on the variable using seasonal-trend decomposition and the common trends.

10. The system of claim 8 , the operations further comprising:

for the second variable set:

applying vector autoregression (VAR) modeling to all the variables;

for each variable of the second variable set, determining a goodness-of-fit value for the variable based on the VAR modeling; and

determining whether VAR fits the multivariate time series based on an average of each goodness-of-fit value determined for each variable of the second variable set.

11. The system of claim 10 , the operations further comprising:

in response to determining VAR fits the multivariate time series, detecting an anomaly on each variable of the second variable set using VAR.

12. The system of claim 10 , the operations further comprising:

in response to determining VAR does not fit the multivariate time series, detecting an anomaly on variables of the second variable set using local outlier factor.

13. A computer program product comprising a computer-readable hardware storage medium having program code embodied therewith, the program code being executable by a computer to implement a method comprising:

receiving metadata associated with one or more data backup jobs performed on one of more storage devices, wherein the metadata comprises univariate time series data for each variable of a multivariate time series, and the multivariate time series comprises different variables that exhibit different characteristics over time; and

decreasing likelihood of a failure in data protection involving the one or more data backup jobs by:

for each variable of the multivariate time series:

selecting, from different anomaly detection models with different performance costs, an anomaly detection model suitable for the variable based on one or more characteristics exhibited by corresponding univariate time series data for the variable and covariations and interactions between the variable and at least one other variable of the multivariate time series; and

detecting an anomaly on the variable utilizing the anomaly detection model selected for the variable; and

based on each anomaly detection model selected for each variable of the multivariate time series, determining whether the multivariate time series is anomalous at a particular time point, and generating data indicative of whether the multivariate time series is anomalous at the particular time point.

14. The computer program product of claim 13 , the operations further comprising:

for each variable of the multivariate time series:

determining whether corresponding univariate time series data for the variable exhibits seasonality;

in response to determining the corresponding univariate time series data for the variable exhibits seasonality, grouping the variable into a first variable set; and

in response to determining the corresponding univariate time series data for the variable does not exhibit seasonality, grouping the variable into a second variable set.

15. The computer program product of claim 14 , the operations further comprising:

for the first variable set:

applying dynamic factor analysis to identify one or more common trends among the first variable set; and

for each variable of the first variable set, detecting an anomaly on the variable using seasonal-trend decomposition and the common trends.

16. The computer program product of claim 14 , the operations further comprising:

for the second variable set:

applying vector autoregression (VAR) modeling to all the variables;

for each variable of the second variable set, determining a goodness-of-fit value for the variable based on the VAR modeling; and

determining whether VAR fits the multivariate time series based on an average of each goodness-of-fit value determined for each variable of the second variable set.

17. The computer program product of claim 16 , the operations further comprising:

in response to determining VAR fits the multivariate time series, detecting an anomaly on each variable of the second variable set using VAR.

18. The computer program product of claim 16 , the operations further comprising:

in response to determining VAR does not fit the multivariate time series, detecting an anomaly on variables of the second variable set using local outlier factor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2017
From: QIAO, MU; ROUTRAY, RAMANI R.; ZHANG, QUAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 040832/0988 →
Continuity (1)
Related Publication 20180189128A1 · Jul 5, 2018