ANOMALY DETECTION BY CORRELATED METRICS
A system is configured to detect a small, but meaningful, anomaly within one or more metrics associated with a platform being monitored. The system displays visuals of the metrics so that a user monitoring the platform can effectively notice a problem associated with the anomaly and take appropriate action to remediate the problem. Moreover, the system uses an ensemble of machine learning algorithms, with a multi-agent voting system, to detect the anomaly. Therefore, via the display of the visuals and the implementation of the machine learning algorithms, the techniques described herein provide an improved way of representing a large number of metrics (e.g., hundreds, thousands, etc.) being monitored for a platform. Moreover, the techniques are configured to expose actionable and useful information associated with the platform in a manner that can be effectively interpreted by a user.
1 . A method comprising:
determining that a first metric is correlated to a second metric;
generating, by one or more devices, a prediction model for the first metric that is correlated to the second metric;
obtaining errors of the prediction model;
determining an upper bound and a lower bound on the errors of the prediction model;
using the prediction model to predict a data value for the second metric from an actual data value for the first metric;
comparing an actual data value for the second metric to the predicted data value for the second metric to determine a difference;
determining that the difference is outside either the upper bound or the lower bound resulting in a voting agent signaling an anomaly associated with a voted metric; and
displaying the anomaly associated with the voted metric.
2 . The method of claim 1 , further comprising:
determining a total number of agents that share an attribute with the voted metric;
determining a number of voting agents from the total number of agents;
generating a percentage for the attribute based on the number of voting agents and the total number of agents;
determining that the percentage is greater than or equal to a threshold percentage; and
determining that a problem associated with the anomaly is localized to the attribute based on the percentage being greater than or equal to the threshold percentage.
3 . The method of claim 2 , wherein the attribute comprises one of a specific location, a type of device, or a type of payment method.
4 . The method of claim 2 , wherein the attribute is related to a list of items being sold by a user of an electronic commerce site.
5 . The method of claim 1 , further comprising:
determining a total number of agents that share an attribute with the voted metric;
determining a number of voting agents from the total number of agents;
generating a percentage for the attribute based on the number of voting agents and the total number of agents;
determining that the percentage is less than a threshold percentage; and
determining that a problem associated with the anomaly is not localized to the attribute based on the percentage being less than the threshold percentage.
6 . The method of claim 1 , wherein:
a first Quantile-Loss Gradient Boosted Tree (QLGBT) error thresholding model is used to determine the upper bound on the errors of the prediction model; and
a second QLGBT error thresholding model is used to determine the lower bound on the errors of the prediction model.
7 . The method of claim 6 , further comprising categorizing one of the first QLGBT error thresholding model or the second QLGBT error thresholding model as the voting agent which signals the anomaly associated with the voted metric.
8 . The method of claim 1 , further comprising evaluating a plurality of metrics to determine that the first metric is correlated to the second metric.
9 . A system comprising:
one or more processing units; and
computer-readable storage media storing instructions that, when executed by the one or more processing units, cause the system to perform operations comprising:
determining that a first metric is correlated to a second metric;
generating a prediction model for the first metric that is correlated to the second metric;
obtaining errors of the prediction model;
determining an upper bound and a lower bound on the errors of the prediction model;
using the prediction model to predict a data value for the second metric from an actual data value for the first metric;
comparing an actual data value for the second metric to the predicted data value for the second metric to determine a difference;
determining that the difference is outside either the upper bound or the lower bound resulting in a voting agent signaling an anomaly associated with a voted metric; and
displaying the anomaly associated with the voted metric.
10 . The system of claim 9 , wherein the operations further comprise:
determining a total number of agents that share an attribute with the voted metric;
determining a number of voting agents from the total number of agents;
generating a percentage for the attribute based on the number of voting agents and the total number of agents;
determining that the percentage is greater than or equal to a threshold percentage; and
determining that a problem associated with the anomaly is localized to the attribute based on the percentage being greater than or equal to the threshold percentage.
11 . The system of claim 10 , wherein the attribute comprises one of a specific location, a type of device, or a type of payment method.
12 . The system of claim 10 , wherein the attribute is related to a list of items being sold by a user of an electronic commerce site.
13 . The system of claim 9 , wherein the operations further comprise:
determining a total number of agents that share an attribute with the voted metric;
determining a number of voting agents from the total number of agents;
generating a percentage for the attribute based on the number of voting agents and the total number of agents;
determining that the percentage is less than a threshold percentage; and
determining that a problem associated with the anomaly is not localized to the attribute based on the percentage being less than the threshold percentage.
14 . The system of claim 9 , wherein:
a first Quantile-Loss Gradient Boosted Tree (QLGBT) error thresholding model is used to determine the upper bound on the errors of the prediction model; and
a second QLGBT error thresholding model is used to determine the lower bound on the errors of the prediction model.
15 . The system of claim 14 , wherein the operations further comprise categorizing one of the first QLGBT error thresholding model or the second QLGBT error thresholding model as the voting agent which signals the anomaly associated with the voted metric.
16 . The system of claim 9 , further comprising evaluating a plurality of metrics to determine that the first metric is correlated to the second metric.
17 . Computer-readable storage media comprising instructions that, when executed by one or more processing units, cause a system to perform operations comprising:
determining that a first metric is correlated to a second metric;
generating a prediction model for the first metric that is correlated to the second metric;
obtaining errors of the prediction model;
determining an upper bound and a lower bound on the errors of the prediction model;
using the prediction model to predict a data value for the second metric from an actual data value for the first metric;
comparing an actual data value for the second metric to the predicted data value for the second metric to determine a difference;
determining that the difference is outside either the upper bound or the lower bound resulting in a voting agent signaling an anomaly associated with a voted metric; and
displaying the anomaly associated with the voted metric.
18 . The computer-readable storage media of claim 17 , wherein the operations further comprise:
determining a total number of agents that share an attribute with the voted metric;
determining a number of voting agents from the total number of agents;
generating a percentage for the attribute based on the number of voting agents and the total number of agents;
determining that the percentage is greater than or equal to a threshold percentage; and
determining that a problem associated with the anomaly is localized to the attribute based on the percentage being greater than or equal to the threshold percentage.
19 . The computer-readable storage media of claim 18 , wherein the attribute comprises one of a specific location, a type of device, or a type of payment method.
20 . The computer-readable storage media of claim 17 , wherein:
a first Quantile-Loss Gradient Boosted Tree (QLGBT) error thresholding model is used to determine the upper bound on the errors of the prediction model;
a second QLGBT error thresholding model is used to determine the lower bound on the errors of the prediction model; and
the operations further comprise categorizing one of the first QLGBT error thresholding model or the second QLGBT error thresholding model as the voting agent which signals the anomaly associated with the voted metric.