Anomaly score normalisation based on extreme value theory
The present invention relates to threshold estimation and calibration for anomaly detection. Herein are machine learning (ML) and extreme value theory (EVT) techniques for normalizing and thresholding anomaly scores without presuming a values distribution. In an embodiment, a computer receives many unnormalized anomaly scores and, according to peak over threshold (POT), selects a highest subset of the unnormalized anomaly scores that exceed a tail threshold. Based on the highest subset of the unnormalized anomaly scores, parameters of a probability density function are trained according to EVT. After training and in a production environment, a normalized anomaly score is generated based on an unnormalized anomaly score and the trained parameters of the probability density function. Anomaly detection compares the normalized anomaly score to an optimized anomaly threshold.
1 . A method comprising:
receiving a plurality of unnormalized anomaly scores;
performing for each predefined tail threshold in a plurality of predefined tail thresholds:
a) selecting a highest subset of the plurality of unnormalized anomaly scores that exceed the predefined tail threshold;
b) training, based on the highest subset of the plurality of unnormalized anomaly scores, parameters of a probability density function; and
c) measuring a fitness score of the parameters of the probability density function for the highest subset of the plurality of unnormalized anomaly scores;
generating an unnormalized anomaly score based on a feature vector that contains i) at least a portion of a database statement and ii) at least one selected from the group consisting of: an identifier of a database session, a network address of a database client, and an identifier of an operating system (OS) of a database client;
generating, by the probability density function for the predefined tail threshold with a highest fitness score, a normalized anomaly score from the unnormalized anomaly score; and
detecting, based on the normalized anomaly score, that the database statement is anomalous;
wherein the method is performed by one or more computers.
2 . The method of claim 1 wherein:
the method further comprises configuring, based on the parameters of the probability density function, a cumulative density function;
said generating the normalized anomaly score based on the parameters of the probability density function comprises applying the cumulative density function to the unnormalized anomaly score.
3 . The method of claim 1 wherein said measuring the fitness score of the parameters of the probability density function comprises applying at least one selected from the group consisting of: a Kolmogorov-Smirnov test, an Anderson-Darling test, and a quantile-quantile (QQ) plot.
4 . The method of claim 1 wherein the plurality of predefined tail thresholds consists of at least two selected from the group consisting of 0.9, 0.99, 0.999, and 0.9999.
5 . The method of claim 1 wherein each predefined tail threshold of the plurality of predefined tail thresholds has a distinct numeric precision.
6 . The method of claim 1 further comprising:
unsupervised training an anomaly scoring model without an anomaly threshold;
said generating the plurality of unnormalized anomaly scores is based on the anomaly scoring model.
7 . The method of claim 1 wherein said training the parameters of the probability density function comprises maximum likelihood estimating.
8 . The method of claim 1 wherein the probability density function is a generalized Pareto distribution.
9 . A method comprising:
receiving a plurality of unnormalized anomaly scores;
selecting a highest subset of the plurality of unnormalized anomaly scores that exceed a tail threshold;
training, based on the highest subset of the plurality of unnormalized anomaly scores, parameters of a probability density function;
first detecting, after said training, whether that a first unnormalized anomaly score for a first database statement does not exceed the tail threshold;
second detecting, in response to said first detecting, that the first database statement is not anomalous;
third detecting, after said training, that a second unnormalized anomaly score for a second database statement exceeds the tail threshold;
generating, in response to said third detecting, a normalized anomaly score based on: the second unnormalized anomaly score and the parameters of the probability density function; and
fourth detecting that the second database statement is anomalous, including detecting that the normalized anomaly score exceeds the tail threshold;
wherein the method is performed by one or more computers.
10 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:
receiving a plurality of unnormalized anomaly scores;
performing for each predefined tail threshold in a plurality of predefined tail thresholds:
a) selecting a highest subset of the plurality of unnormalized anomaly scores that exceed the predefined tail threshold;
b) training, based on the highest subset of the plurality of unnormalized anomaly scores, parameters of a probability density function; and
c) measuring a fitness score of the parameters of the probability density function for the highest subset of the plurality of unnormalized anomaly scores;
generating an unnormalized anomaly score based on a feature vector that contains i) at least a portion of a database statement and ii) at least one selected from the group consisting of: an identifier of a database session, a network address of a database client, and an identifier of an operating system (OS) of a database client;
generating, by the probability density function for the predefined tail threshold with a highest fitness score, a normalized anomaly score from the unnormalized anomaly score; and
detecting, based on the normalized anomaly score, that the database statement is anomalous.
11 . The one or more non-transitory computer-readable media of claim 10 wherein:
the instructions further cause configuring, based on the parameters of the probability density function, a cumulative density function;
said generating the normalized anomaly score based on the parameters of the probability density function comprises applying the cumulative density function to the unnormalized anomaly score.
12 . The one or more non-transitory computer-readable media of claim 10 wherein each predefined tail threshold of the plurality of predefined tail thresholds has a distinct numeric precision.
13 . The one or more non-transitory computer-readable media of claim 10 wherein the instructions further cause:
unsupervised training an anomaly scoring model without an anomaly threshold;
said generating the plurality of unnormalized anomaly scores is based on the anomaly scoring model.
14 . The one or more non-transitory computer-readable media of claim 10 wherein said training the parameters of the probability density function comprises maximum likelihood estimating.
15 . The one or more non-transitory computer-readable media of claim 10 wherein the probability density function is a generalized Pareto distribution.
16 . The one or more non-transitory computer-readable media of claim 10 wherein:
the instructions further cause detecting whether the unnormalized anomaly score exceeds the tail threshold;
said generating the normalized anomaly score occurs only when the unnormalized anomaly score exceeds the tail threshold.
17 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:
receiving a plurality of unnormalized anomaly scores;
selecting a highest subset of the plurality of unnormalized anomaly scores that exceed a tail threshold;
training, based on the highest subset of the plurality of unnormalized anomaly scores, parameters of a probability density function;
first detecting, after said training, whether that a first unnormalized anomaly score for a first database statement does not exceed the tail threshold;
second detecting, in response to said first detecting, that the first database statement is not anomalous;
third detecting, after said training, that a second unnormalized anomaly score for a second database statement exceeds the tail threshold;
generating, in response to said third detecting, a normalized anomaly score based on: the second unnormalized anomaly score and the parameters of the probability density function; and
fourth detecting that the second database statement is anomalous, including detecting that the normalized anomaly score exceeds the tail threshold.