IP Library › Granted Patent US 12,614,096
Granted Patent B2
US 12,614,096 · App. 17/745,103 · Granted Apr 28, 2026

Anomaly score normalisation based on extreme value theory

Inventors: Marija Nikolic (Zurich, CH); Matteo Casserini (Zurich, CH); Arno Schneuwly (Effretikon, CH); Nikola Milojkovic (Dietikon, CH); Milos Vasic (Zurich, CH); Renata Khasanova (Zurich, CH); Felix Schmidt (Baden-Dattwil, CH)
Assignee: Oracle International Corporation
G06N7/01G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,096
App. No.
17/745,103
Granted
Apr 28, 2026
Kind
B2
Abstract

The present invention relates to threshold estimation and calibration for anomaly detection. Herein are machine learning (ML) and extreme value theory (EVT) techniques for normalizing and thresholding anomaly scores without presuming a values distribution. In an embodiment, a computer receives many unnormalized anomaly scores and, according to peak over threshold (POT), selects a highest subset of the unnormalized anomaly scores that exceed a tail threshold. Based on the highest subset of the unnormalized anomaly scores, parameters of a probability density function are trained according to EVT. After training and in a production environment, a normalized anomaly score is generated based on an unnormalized anomaly score and the trained parameters of the probability density function. Anomaly detection compares the normalized anomaly score to an optimized anomaly threshold.

Claims (61)

1 . A method comprising:

receiving a plurality of unnormalized anomaly scores;

performing for each predefined tail threshold in a plurality of predefined tail thresholds:

a) selecting a highest subset of the plurality of unnormalized anomaly scores that exceed the predefined tail threshold;

b) training, based on the highest subset of the plurality of unnormalized anomaly scores, parameters of a probability density function; and

c) measuring a fitness score of the parameters of the probability density function for the highest subset of the plurality of unnormalized anomaly scores;

generating an unnormalized anomaly score based on a feature vector that contains i) at least a portion of a database statement and ii) at least one selected from the group consisting of: an identifier of a database session, a network address of a database client, and an identifier of an operating system (OS) of a database client;

generating, by the probability density function for the predefined tail threshold with a highest fitness score, a normalized anomaly score from the unnormalized anomaly score; and

detecting, based on the normalized anomaly score, that the database statement is anomalous;

wherein the method is performed by one or more computers.

2 . The method of claim 1 wherein:

the method further comprises configuring, based on the parameters of the probability density function, a cumulative density function;

said generating the normalized anomaly score based on the parameters of the probability density function comprises applying the cumulative density function to the unnormalized anomaly score.

3 . The method of claim 1 wherein said measuring the fitness score of the parameters of the probability density function comprises applying at least one selected from the group consisting of: a Kolmogorov-Smirnov test, an Anderson-Darling test, and a quantile-quantile (QQ) plot.

4 . The method of claim 1 wherein the plurality of predefined tail thresholds consists of at least two selected from the group consisting of 0.9, 0.99, 0.999, and 0.9999.

5 . The method of claim 1 wherein each predefined tail threshold of the plurality of predefined tail thresholds has a distinct numeric precision.

6 . The method of claim 1 further comprising:

unsupervised training an anomaly scoring model without an anomaly threshold;

said generating the plurality of unnormalized anomaly scores is based on the anomaly scoring model.

7 . The method of claim 1 wherein said training the parameters of the probability density function comprises maximum likelihood estimating.

8 . The method of claim 1 wherein the probability density function is a generalized Pareto distribution.

9 . A method comprising:

receiving a plurality of unnormalized anomaly scores;

selecting a highest subset of the plurality of unnormalized anomaly scores that exceed a tail threshold;

training, based on the highest subset of the plurality of unnormalized anomaly scores, parameters of a probability density function;

first detecting, after said training, whether that a first unnormalized anomaly score for a first database statement does not exceed the tail threshold;

second detecting, in response to said first detecting, that the first database statement is not anomalous;

third detecting, after said training, that a second unnormalized anomaly score for a second database statement exceeds the tail threshold;

generating, in response to said third detecting, a normalized anomaly score based on: the second unnormalized anomaly score and the parameters of the probability density function; and

fourth detecting that the second database statement is anomalous, including detecting that the normalized anomaly score exceeds the tail threshold;

wherein the method is performed by one or more computers.

10 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

receiving a plurality of unnormalized anomaly scores;

performing for each predefined tail threshold in a plurality of predefined tail thresholds:

a) selecting a highest subset of the plurality of unnormalized anomaly scores that exceed the predefined tail threshold;

b) training, based on the highest subset of the plurality of unnormalized anomaly scores, parameters of a probability density function; and

c) measuring a fitness score of the parameters of the probability density function for the highest subset of the plurality of unnormalized anomaly scores;

generating an unnormalized anomaly score based on a feature vector that contains i) at least a portion of a database statement and ii) at least one selected from the group consisting of: an identifier of a database session, a network address of a database client, and an identifier of an operating system (OS) of a database client;

generating, by the probability density function for the predefined tail threshold with a highest fitness score, a normalized anomaly score from the unnormalized anomaly score; and

detecting, based on the normalized anomaly score, that the database statement is anomalous.

11 . The one or more non-transitory computer-readable media of claim 10 wherein:

the instructions further cause configuring, based on the parameters of the probability density function, a cumulative density function;

said generating the normalized anomaly score based on the parameters of the probability density function comprises applying the cumulative density function to the unnormalized anomaly score.

12 . The one or more non-transitory computer-readable media of claim 10 wherein each predefined tail threshold of the plurality of predefined tail thresholds has a distinct numeric precision.

13 . The one or more non-transitory computer-readable media of claim 10 wherein the instructions further cause:

unsupervised training an anomaly scoring model without an anomaly threshold;

said generating the plurality of unnormalized anomaly scores is based on the anomaly scoring model.

14 . The one or more non-transitory computer-readable media of claim 10 wherein said training the parameters of the probability density function comprises maximum likelihood estimating.

15 . The one or more non-transitory computer-readable media of claim 10 wherein the probability density function is a generalized Pareto distribution.

16 . The one or more non-transitory computer-readable media of claim 10 wherein:

the instructions further cause detecting whether the unnormalized anomaly score exceeds the tail threshold;

said generating the normalized anomaly score occurs only when the unnormalized anomaly score exceeds the tail threshold.

17 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

receiving a plurality of unnormalized anomaly scores;

selecting a highest subset of the plurality of unnormalized anomaly scores that exceed a tail threshold;

training, based on the highest subset of the plurality of unnormalized anomaly scores, parameters of a probability density function;

first detecting, after said training, whether that a first unnormalized anomaly score for a first database statement does not exceed the tail threshold;

second detecting, in response to said first detecting, that the first database statement is not anomalous;

third detecting, after said training, that a second unnormalized anomaly score for a second database statement exceeds the tail threshold;

generating, in response to said third detecting, a normalized anomaly score based on: the second unnormalized anomaly score and the parameters of the probability density function; and

fourth detecting that the second database statement is anomalous, including detecting that the normalized anomaly score exceeds the tail threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2022
From: NIKOLIC, MARIJA; CASSERINI, MATTEO; SCHNEUWLY, ARNO; MILOJKOVIC, NIKOLA; VASIC, MILOS; KHASANOVA, RENATA; SCHMIDT, FELIX
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 059918/0967 →
Continuity (1)
Related Publication 20230368054A1 · Nov 16, 2023
References Cited (42)
US 10069900B2 · Poola et al. · 2018 [cited by applicant]
US 10270788B2 · Faigon · 2019 [cited by examiner]
US 20140108314A1 · Chen · 2014 [cited by applicant]
US 20180095004A1 · Ide · 2018 [cited by examiner]
US 20200045064A1 · Bindal · 2020 [cited by examiner]
US 20200210393A1 · Beaver · 2020 [cited by examiner]
US 20210328969A1 · Gaddam · 2021 [cited by examiner]
US 20220188694A1 · Suzani et al. · 2022 [cited by applicant]
Gao et al. (“Converting Output Scores from Outlier Detection Algorithms into Probability Estimates,” Sixth International Conference on Data Mining (ICDM'06), Hong Kong, China, 2006, pp. 212-221) (Year: 2006). [cited by examiner]
Bourguignon et al. (“The Kumaraswamy Pareto distribution”, arXiv, 2012) (Year: 2012). [cited by examiner]
Siffer et al. (“Anomaly Detection in Streams with Extreme Value Theory”, KDD 2017 Research Paper) (Year: 2017). [cited by examiner]
Zhao et al. (“Anomaly Detection with Score functions based on Nearest Neighbor Graphs”, 2009) (Year: 2009). [cited by examiner]
Qian et al. (“A Rank-SVM Approach to Anomaly Detection”, A New One-Class Svm for Anomaly Detection, 2014) (Year: 2014). [cited by examiner]
Yao et al., “Rethinking Class-Prior Estimation for Positive-Unlabeled Learning”, in International Conference on Learning Representations, dated Sep. 28, 2021, 12 pages. [cited by applicant]
Perini et al., “Class Prior Estimation in Active Positive and Unlabeled Learning”, in Proceedings of the 29th IJCAI and the 17th PRICAI, dated Jul. 2020, 7 pages. [cited by applicant]
Kriegel at al., “Interpreting and Unifying Outlier Scores”, in Proceedings of the 2011 SIAM International Conference on Data Mining, dated 2011, 12 pages. [cited by applicant]
Christoffel et al., “Class-prior Estimation for Learning from Positive and Unlabeled Data”, Asian Conference on Machine Learning, PMLR, vol. 45, dated Feb. 2016, 16 pages. [cited by applicant]
Eskin, Eleazar, “Anomaly Detection Over Noisy Data Using Learned Probability Distributions”, https://academiccommons.columbia.edu/doi/10.7916/D8C53SKF, dated 2000, 8 pages. [cited by applicant]
An et al., “Variational Autoencoder Based Anomaly Detection Using Reconstruction Probability”, SNU Data Mining Center, Special Lecture on IE 2.1 (2015), http://dm.snu.ac.kr/static/docs/TR/SNUDM-TR-2015-03.pdf, dated Dec… [cited by applicant]
Angiulli et al., “Distance-Based Detection and Prediction of Outliers”, IEEE Transactions on Knowledge and Data Engineering, vol. 18, No. 2, dated Feb. 2006, 16 pages. [cited by applicant]
Batt et al., “Extreme Events in Lake Ecosystem Time Series”, Limnology and Oceanography Letters, vol. 2, No. 3, doi: 10.1002/lol2.10037, dated Feb. 2017, 7 pages. [cited by applicant]
Breunig et al., “LOF: Identifying Density-Based Local Outliers”, Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data, https://dl.acm.org/doi/pdf/10.1145/342009.335388, dated 2000, 12 pages. [cited by applicant]
Chakraborty et al., “Early detection of faults in HVAC systems using an XGBoost model with a dynamic threshold”, Energy and Buildings, dated 2019, pp. 326-344. [cited by applicant]
Chalapathy et al., “Deep Learning for Anomaly Detection: A Survey”, https://arxiv.org/pdf/1901.03407.pdf%20http://arxiv.org/abs/1901.03407.pdf, dated Jan. 24, 2019, 50 pages. [cited by applicant]
Chandola et al., “Anomaly Detection: A Survey”, ACM Computing Surveys (CSUR) vol. 41, No. 3, https://conservancy.umn.edu/bitstream/handle/11299/215731/07-017.pdf?sequence=1, dated Aug. 15, 2007, 74 pages. [cited by applicant]
Chang et al., “A Dynamic Threshold Decision System for Stock Trading Signal Detection”, Applied Soft Computing 11, dated 2011, 13 pages. [cited by applicant]
Davis et al., “LSTM-Based Anomaly Detection: Detection Rules from Extreme Value Theory”, EPIA Conference on Artificial Intelligenc https://arxiv.org/pdf/1909.06041.pdf, dated Sep. 13, 2019, 12 pages. [cited by applicant]
Agarwal, Deepak, “An Empirical Bayes Approach to Detect Anomalies in Dynamic Multidimensional Arrays”, Fifth IEEE International Conference on Data Mining (ICDM'05), dated 2005, 9 pages. [cited by applicant]
Dykes, Sandra, “Poster: An Extreme Value Theory Approach to Anomaly Detection (EVT-AD)”, https://www.ieee-security.org/TC/SP2012/posters/An%20Extreme%20Value%20Theory%20Approach.pdf, dated 2012, 2 pages. [cited by applicant]
Tippett et al., “More Tornadoes in the Most Extreme U.S. Tornado Outbreaks”, https://www.science.org/doi/epdf/10.1126/science.aah7393, dated Oct. 2016, 6 pages. [cited by applicant]
Haan et al., “Extreme Value Theory: An Introduction”, vol. 21, New York: Springer, DOI:10.1007/0-387-34471-3, dated Jan. 2006, 15 pages. [cited by applicant]
Jain et al., “Score normalization in multimodal biometric systems”, Pattern Recognition 38, dated Jan. 2005, 16 pages. [cited by applicant]
Kratz et al., “The QQ—Estimator and Heavy Tails”, Stochastic Models, vol. 12, No. 4, https://ecommons.cornell.edu/bitstream/handle/1813/9004/TR001122.pdf, dated Jan. 1995, 25 pages. [cited by applicant]
Le Cam, Lucien, “Maximum Likelihood: An Introduction”, International Statistical Review/Revue Internationale de Statistique, vol. 58, No. 2, http://www.jstor.org/stable/1403464, dated Aug. 1990, 20 pages. [cited by applicant]
Massey Jr, Frank, “The Kolmogorov-Smirnov Test for Goodness of Fit”, Journal of the American Statistical Association vol. 46, No. 253, https://www.jstor.org/stable/2280095, dated Mar. 1951, 12 pages. [cited by applicant]
Rocco, Marco, “Extreme Value Theory for Finance: A Survey”, Journal of Economic Surveys, vol. 28, No. 1, http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.650.8776&rep=rep1&type=pdf, dated Jul. 2011, 74 pages. [cited by applicant]
Rudd et al., “The Extreme Value Machine”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, No. 3, https://ieeexplore.ieee.org/ielaam/34/8281092/7932895-aam.pdf, dated Mar. 2017, 8 pages. [cited by applicant]
Siffer et al., “Anomaly Detection in Streams with Extreme Value Theory”, Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, https://hal.archives-ouvertes.fr/hal-01640325,… [cited by applicant]
Singh et al., “Quantitative Evaluation of Normalization Techniques of Matching Scores in Multimodal Biometric Systems”, Springer-Verlag Berlin Heidelberg, dated 2007, 10 pages. [cited by applicant]
Dong et al., “Quality-Based Dynamic Threshold for Iris Matching”, IEEE, dated 2009, 4 pages. [cited by applicant]
Neuberg, Richard, et al., “Detective Relative Anomaly”, 2015 18th Intl Conf, on Mach Learning and Data Mining in Pattern Recognition, Lecture Notes in Computer Science, LNAI 10358. Springer, doi.org/10.1007/978-3-319-62… [cited by applicant]
Gao, Jing, et al., “Converting Output Scores from Outlier Detection Algorithms into Probability Estimates”, 6th Intl Conf on Data Mining (ICDM'06), pp. 212-221, doi: 10.1109/ICDM.2006.43, Dec. 18, 2006, 10pgs. [cited by applicant]