IP Library › Granted Patent US 12,614,121
Granted Patent B2
US 12,614,121 · App. 18/075,784 · Granted Apr 28, 2026

Learning hyper-parameter scaling models for unsupervised anomaly detection

Inventors: Fatjon Zogaj (Zurich, CH); Yasha Pushak (Vancouver, CA); Hesam Fathi Moghadam (Sunnyvale, CA); Sungpack Hong (Palo Alto, CA); Hassan Chafi (San Mateo, CA)
Assignee: Oracle International Corporation
G06N20/20G06F16/2365G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,121
App. No.
18/075,784
Granted
Apr 28, 2026
Kind
B2
Abstract

A computer sorts empirical validation scores of validated training scenarios of an anomaly detector. Each training scenario has a dataset to train an instance of the anomaly detector that is configured with values for hyperparameters. Each dataset has values for metafeatures. For each predefined ranking percentage, a subset of best training scenarios is selected that consists of the ranking percentage of validated training scenarios having the highest empirical validation scores. Linear optimizers train to infer a value for a hyperparameter. Into many distinct unvalidated training scenarios, a scenario is generated that has metafeatures values and hyperparameters values that contains the value inferred for that hyperparameter by a linear optimizer. For each unvalidated training scenario, a validation score is inferred. A best linear optimizer is selected having a highest combined inferred validation score. For a new dataset, the best linear optimizer infers a value of that hyperparameter.

Claims (95)

1 . A method comprising:

sorting respective empirical validation scores of a plurality of validated training scenarios of an anomaly detector, wherein:

each validated training scenario has a respective training dataset to train a respective instance of the anomaly detector that is respectively configured with a respective plurality of values for a plurality of hyperparameters of the anomaly detector, and

said respective training dataset has a respective set of values for a set of metafeatures;

for each ranking percentage of a plurality of predefined distinct ranking percentages:

a) selecting a subset of best training scenarios that consists of the ranking percentage of the plurality of validated training scenarios of the anomaly detector having the highest empirical validation scores;

b) training a respective linear optimizer that infers a respective value for a particular hyperparameter of the plurality of hyperparameters of the anomaly detector, wherein the training is based on the sets of values of the set of metafeatures of the training datasets of the subset of best training scenarios and the empirical validation scores of the subset of best training scenarios,

c) for each set of metafeatures values in a plurality of unvalidated sets of values of the set of metafeatures, wherein the set of metafeatures values does not correspond to an actual dataset:

i) inferring, by the linear optimizer, a particular value for the particular hyperparameter of the plurality of hyperparameters of the anomaly detector;

ii) generating, into a plurality of distinct unvalidated training scenarios, a distinct unvalidated training scenario of the anomaly detector that has: the set of metafeatures values and a plurality of values for the plurality of hyperparameters of the anomaly detector that contains the particular value for the particular hyperparameter;

inferring, for each unvalidated training scenario of the plurality of distinct unvalidated training scenarios of the anomaly detector, a respective inferred validation score;

selecting a best linear optimizer of a ranking percentage of the plurality of predefined distinct ranking percentages having a highest combined inferred validation score based on said inferred validation scores for the plurality of distinct unvalidated training scenarios of the anomaly detector;

inferring, by the best linear optimizer of the ranking percentage of the plurality of predefined distinct ranking percentages, for a new dataset, an inferred value of the particular hyperparameter of the plurality of hyperparameters of the anomaly detector; and

detecting, by a new instance of the anomaly detector that is based on the inferred value of the particular hyperparameter of the plurality of hyperparameters of the anomaly detector, an anomaly;

wherein the method is performed by one or more computers.

2 . The method of claim 1 further comprising training a metamodel to infer a validation score of a training scenario of the anomaly detector, wherein the metamodel performs said inferring the validation scores, and the training the metamodel is based on:

the empirical validation scores of the plurality of validated training scenarios of the anomaly detector,

the sets of values of the set of metafeatures of the training datasets of the plurality of validated training scenarios, and

the pluralities of values of the plurality of hyperparameters of the anomaly detector.

3 . The method of claim 2 wherein one selected from the group consisting of:

said training the metamodel occurs after said training the linear optimizers, and

said training the metamodel and said training the linear optimizers concurrently occur.

4 . The method of claim 2 further comprising:

inferring, by a second optimizer for a second particular hyperparameter of the plurality of hyperparameters of the anomaly detector, a particular value for the second particular hyperparameter;

generating, into a second plurality of distinct unvalidated training scenarios, a distinct unvalidated training scenario of the anomaly detector that has a plurality of values for the plurality of hyperparameters of the anomaly detector that contains the particular value for the second particular hyperparameter;

inferring, by the metamodel, a respective inferred validation score for each unvalidated training scenario of the second plurality of distinct unvalidated training scenarios of the anomaly detector.

5 . The method of claim 4 wherein said second optimizer is one selected from the group consisting of a linear regression and a logistic regression.

6 . The method of claim 2 wherein:

the plurality of validated training scenarios of the anomaly detector contains:

a first training scenario that has a first training dataset that consists of representations of instances of a first kind of object, and

a second training scenario that has a second training dataset that consists of representations of instances of a second kind of object;

the first kind of object and the second kind of object are distinct kinds of objects selected from the group consisting of: a network packet, a communicated message, a log entry, a logic statement, a semi-structured document, a photograph, a database record, a logical graph, and a parse tree.

7 . The method of claim 2 wherein:

said anomaly detector is a first anomaly detector;

said plurality of hyperparameters of the first anomaly detector is disjoint from a plurality of hyperparameters of a second anomaly detector;

the method further comprises:

inferring, by a linear optimizer for a particular hyperparameter of the plurality of hyperparameters of the second anomaly detector, a particular value for the particular hyperparameter of the plurality of hyperparameters of the second anomaly detector;

inferring, by a second metamodel, a respective inferred validation score for an unvalidated training scenario of the second anomaly detector that has the particular value for the particular hyperparameter of the plurality of hyperparameters of the second anomaly detector.

8 . The method of claim 7 further comprising combining, in an ensemble, an instance of the first anomaly detector and an instance of the second anomaly detector.

9 . The method of claim 2 wherein the metamodel does not comprise an artificial neural network.

10 . The method of claim 1 wherein at least one selected from the group consisting of:

said set of metafeatures consists only of a neutral metafeature,

said set of metafeatures contains a dataset cardinality metafeature,

said best linear optimizer linearly scales the particular hyperparameter based on only one metafeature, and

said best linear optimizer monotonically scales the particular hyperparameter.

11 . The method of claim 10 further comprising:

before said inferring the inferred value of the particular hyperparameter of the plurality of hyperparameters of the anomaly detector, tuning the particular hyperparameter based on a subset of the new dataset to provide an initial value of the particular hyperparameter;

training, based on the entire new dataset, an instance of the anomaly detector that is based on the inferred value of the particular hyperparameter.

12 . The method of claim 1 wherein the training datasets of the plurality of validated training scenarios of the anomaly detector contain at least one selected from the group consisting of:

two training datasets that have distinct respective dimensionalities, and

two training datasets that have respective distinct pluralities of metafeatures that include said set of metafeatures.

13 . The method of claim 1 wherein:

a step is based on an unlabeled training dataset;

the step is at least one selected from the group consisting of:

said inferring the inferred value of the particular hyperparameter of the plurality of hyperparameters of the anomaly detector, and

said detecting the anomaly by said new instance of the anomaly detector that is based on the inferred value of the particular hyperparameter of the plurality of hyperparameters of the anomaly detector.

14 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

sorting respective empirical validation scores of a plurality of validated training scenarios of an anomaly detector, wherein:

each validated training scenario has a respective training dataset to train a respective instance of the anomaly detector that is respectively configured with a respective plurality of values for a plurality of hyperparameters of the anomaly detector, and

said respective training dataset has a respective set of values for a set of metafeatures;

for each ranking percentage of a plurality of predefined distinct ranking percentages:

a) selecting a subset of best training scenarios that consists of the ranking percentage of the plurality of validated training scenarios of the anomaly detector having the highest empirical validation scores;

b) training a respective linear optimizer that infers a respective value for a particular hyperparameter of the plurality of hyperparameters of the anomaly detector, wherein the training is based on the sets of values of the set of metafeatures of the training datasets of the subset of best training scenarios and the empirical validation scores of the subset of best training scenarios,

c) for each set of metafeatures values in a plurality of unvalidated sets of values of the set of metafeatures, wherein the set of metafeatures values does not correspond to an actual dataset:

i) inferring, by the linear optimizer, a particular value for the particular hyperparameter of the plurality of hyperparameters of the anomaly detector;

ii) generating, into a plurality of distinct unvalidated training scenarios, a distinct unvalidated training scenario of the anomaly detector that has: the set of metafeatures values and a plurality of values for the plurality of hyperparameters of the anomaly detector that contains the particular value for the particular hyperparameter;

inferring, for each unvalidated training scenario of the plurality of distinct unvalidated training scenarios of the anomaly detector, a respective inferred validation score;

selecting a best linear optimizer of a ranking percentage of the plurality of predefined distinct ranking percentages having a highest combined inferred validation score based on said inferred validation scores for the plurality of distinct unvalidated training scenarios of the anomaly detector;

inferring, by the best linear optimizer of the ranking percentage of the plurality of predefined distinct ranking percentages, for a new dataset, an inferred value of the particular hyperparameter of the plurality of hyperparameters of the anomaly detector; and

detecting, by a new instance of the anomaly detector that is based on the inferred value of the particular hyperparameter of the plurality of hyperparameters of the anomaly detector, an anomaly.

15 . The one or more non-transitory computer-readable media of claim 14 wherein the instructions further cause training a metamodel to infer a validation score of a training scenario of the anomaly detector, wherein the metamodel performs said inferring the validation scores, and the training the metamodel is based on:

the empirical validation scores of the plurality of validated training scenarios of the anomaly detector,

the sets of values of the set of metafeatures of the training datasets of the plurality of validated training scenarios, and

the pluralities of values of the plurality of hyperparameters of the anomaly detector.

16 . The one or more non-transitory computer-readable media of claim 15 wherein one selected from the group consisting of:

said training the metamodel occurs after said training the linear optimizers, and

said training the metamodel and said training the linear optimizers concurrently occur.

17 . The one or more non-transitory computer-readable media of claim 15 wherein the instructions further cause:

inferring, by a second optimizer for a second particular hyperparameter of the plurality of hyperparameters of the anomaly detector, a particular value for the second particular hyperparameter;

generating, into a second plurality of distinct unvalidated training scenarios, a distinct unvalidated training scenario of the anomaly detector that has a plurality of values for the plurality of hyperparameters of the anomaly detector that contains the particular value for the second particular hyperparameter;

inferring, by the metamodel, a respective inferred validation score for each unvalidated training scenario of the second plurality of distinct unvalidated training scenarios of the anomaly detector.

18 . The one or more non-transitory computer-readable media of claim 15 wherein:

the plurality of validated training scenarios of the anomaly detector contains:

a first training scenario that has a first training dataset that consists of representations of instances of a first kind of object, and

a second training scenario that has a second training dataset that consists of representations of instances of a second kind of object;

the first kind of object and the second kind of object are distinct kinds of objects selected from the group consisting of: a network packet, a communicated message, a log entry, a logic statement, a semi-structured document, a photograph, a database record, a logical graph, and a parse tree.

19 . The one or more non-transitory computer-readable media of claim 15 wherein:

said anomaly detector is a first anomaly detector;

said plurality of hyperparameters of the first anomaly detector is disjoint from a plurality of hyperparameters of a second anomaly detector;

the instructions further cause:

inferring, by a linear optimizer for a particular hyperparameter of the plurality of hyperparameters of the second anomaly detector, a particular value for the particular hyperparameter of the plurality of hyperparameters of the second anomaly detector;

inferring, by a second metamodel, a respective inferred validation score for an unvalidated training scenario of the second anomaly detector that has the particular value for the particular hyperparameter of the plurality of hyperparameters of the second anomaly detector.

20 . The one or more non-transitory computer-readable media of claim 14 wherein the training datasets of the plurality of validated training scenarios of the anomaly detector contain at least one selected from the group consisting of:

two training datasets that have distinct respective dimensionalities, and

two training datasets that have respective distinct pluralities of metafeatures that include said set of metafeatures.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2022
From: ZOGAJ, FATJON; PUSHAK, YASHA; FATHI MOGHADAM, HESAM; HONG, SUNGPACK; CHAFI, HASSAN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 061995/0416 →
Continuity (2)
Provisional Application 63408602 · Sep 21, 2022
Related Publication 20240095604A1 · Mar 21, 2024
References Cited (52)
US 11392859B2 · Basu et al. · 2022 [cited by applicant]
US 11868854B2 · Moharrer et al. · 2024 [cited by applicant]
US 20140188768A1 · Bonissone et al. · 2014 [cited by applicant]
US 20180060330A1 · Clinton et al. · 2018 [cited by applicant]
US 20180336453A1 · Merity et al. · 2018 [cited by applicant]
US 20190095756A1 · Agrawal · 2019 [cited by applicant]
US 20190121919A1 · Rafaila · 2019 [cited by applicant]
US 20190244139A1 · Varadarajan · 2019 [cited by examiner]
US 20200038378A1 · Crew · 2020 [cited by examiner]
US 20200242400A1 · Perkins · 2020 [cited by applicant]
US 20200334569A1 · Moghadam · 2020 [cited by applicant]
EP 3101599A2 · 2016 [cited by applicant]
WO WO2008133509A1 · 2008 [cited by applicant]
Soenen et al., “The effect of hyperparameter tuning on the comparative evaluation of unsupervised anomaly detection methods”, KDD'21, 2021 (Year: 2021). [cited by examiner]
Purohit et al., “Deep autoencoding GMM-based unsupervised anomaly detection in acoustic signals and its hyper-parameter optimization”, Detection and Classification of Acoustic Scenes and Events 2020, Nov. 2-3, 2020, Tok… [cited by examiner]
Xu et al., “Automatic hyperparameter tuning method for local outlier factor, with applications to anomaly detection”, 2019 IEEE International conference on big data, 2019 (Year: 2019). [cited by examiner]
Malkomes, Gustavo, et al.,“Bayesian optimization for automated model selection”, NIPS 2016, (https://proceedings.neurips.cc/paper/2016/file/3bbfdde8842a5c44a0323518eec97cbe-Paper.pdf) Dec. 2016, 9 pages. [cited by applicant]
Le Clei et al.,“N-1 Experts: Unsupervised Anomaly Detection Model Selection”, in First Conference on Automated Machine Learning (Late-Breaking Workshop), dated May 2022, 14 pages. [cited by applicant]
Bergstra et al., “Hyperopt: A Python Library for Model Selection and Hyperparameter Optimization”, Computational Science & Discovery, vol. 8, No. 1, DOI: 10.1088/1749-4699/8/1/014008, dated Jul. 2015, 24 pages. [cited by applicant]
Butakov, Nikita, “How to Build Robust Anomaly Detectors with Machine Learning”, Ericsson, https://www.ericsson.com/en/blog/2020/4/anomaly-detection-with-machine-learning, dated Apr. 2020, 11 pages. [cited by applicant]
Caruana et al., “Ensemble Selection From Libraries of Models”, in Proceedings of the 21st International Conference on Machine Learning, dated Jul. 2004, 8 pages. [cited by applicant]
Christoffel et al., “Class-prior Estimation for Learning from Positive and Unlabeled Data”, Asian Conference on Machine Learning, PMLR, vol. 45, dated Feb. 2016, 16 pages. [cited by applicant]
De Souto et al., “Ranking and Selecting Clustering Algorithms using a Meta-Learning Approach”, IEEE International Joint Conference on Neural Networks, dated Jun. 2008, 7 pages. [cited by applicant]
De Souza et al., “Improved Regression Models for Algorithm Configuration”, in Proceedings of the Genetic and Evolutionary Computation Conference, DOI: 10.1145/3512290.3528750, dated Jul. 2022, 10 pages. [cited by applicant]
Doan et al., “Algorithm Selection using Performance and Run Time Behavior”, International Conference on Artificial Intelligence: Methodology, Systems, and Applications, Lecture Notes in Computer Science, vol. 9883, date… [cited by applicant]
Doan et al., “Selecting Machine Learning Algorithms using Regression Models”, IEEE International Conference on Data Mining Workshop, https://www.researchgate.net/publication/304298580, dated Nov. 2015, 8 pages. [cited by applicant]
Feurer et al., “Hyperparameter Optimization”, in Automated Machine Learning, The Springer Series on Challenges in Machine Learning, DOI: 10.1007/978-3-030-05318-5_1, dated May 2019, 31 pages. [cited by applicant]
Guo et al., “A New Approach Towards the Combined Algorithm Selection and Hyper-parameter Optimization Problem”, IEEE Symposium Series on Computational Intelligence, dated Dec. 2019, 57 pages. [cited by applicant]
Aldave et al., “Systematic Ensemble Learning for Regression”, Machine Learning, DOI: 10.48550/arXiv.1403.7267 dated Mar. 2014, 38 pages. [cited by applicant]
Kuck et al., “Meta-Learning with Neural Networks and Landmarking for Forecasting Model Selection an Empirical Evaluation of Different Feature Sets Applied to Industry Data”, International Joint Conference on Neural Netw… [cited by applicant]
Zhao et al., “Automatic Unsupervised Outlier Model Selection”, in Advances in Neural Information Processing Systems, 34, dated 2021, 14 pages. [cited by applicant]
Molina et al., “Meta-Learning Approach for Automatic Parameter Tuning: A Case Study with Educational Datasets”, Proceedings of the 5th International Conference on Educational Data Mining, dated Jun. 2012, 4 pages. [cited by applicant]
Perini et al., “Transferring the Contamination Factor between Anomaly Detection Domains by Shape Similarity”, in Proceedings of the Thirty-Sixth AAAI Conference on Artificial Intelligence, vol. 36, No. 4, dated Jun. 202… [cited by applicant]
Simpson et al, “Automatic Algorithm Selection in Computational Software using Machine Learning”, 15th IEEE International Conference on Machine Learning and Applications, dated Dec. 2016, 10 pages. [cited by applicant]
Swierczewski et al., “Use the Built-in Amazon SageMaker Random Cut Forest Algorithm for Anomaly Detection”, https://aws.amazon.com/blogs/machine-learning/use-the-built-in-amazon-sagemaker-random-cut-forest-algorithm- fo… [cited by applicant]
Wichard, J.D., “Model Selection in an Ensemble Framework”, The 2013 International Joint Conference on Neural Networks, dated Jan. 1, 2006, pp. 2187-2192. [cited by applicant]
Wistuba et al., “Learning Hyperparameter Optimization Initializations”, IEEE, 2015, 10 pages. [cited by applicant]
Wistuba et al., “Scalable Gaussian process-based transfer surrogates for hyperparameter optimization”, Machine Learning, 107(1), 2017, 36 pages. [cited by applicant]
Xing, Tony, “Introducing Azure Anomaly Detector API”, Microsoft, https://techcommunity.microsoft.com/t5/ai-customer-engineering-team/introducing-azure-anomaly-detector-api/ba-p/490162, dated Apr. 2019, 22 pages. [cited by applicant]
Xing, Tony, “Introducing Multivariate Anomaly Detection”, Microsoft, https://techcommunity.microsoft.com/t5/ai-cognitive-services-blog/introducing-multivariate-anomaly-detection/ba-p/2260679, dated Apr. 2021, 9 pages. [cited by applicant]
Yakovlev et al., “Oracle AutoML: A Fast and Predictive AutoML Pipeline”, in Proceedings of the VLDB Endowment, vol. 13, No. 12, DOI: https://doi.org/10.14778/3415478.3415542, dated Aug. 2020, 15 pages. [cited by applicant]
Kriegel at al., “Interpreting and Unifying Outlier Scores”, in Proceedings of the 2011 SIAM International Conference on Data Mining, dated 2011, 12 pages. [cited by applicant]
Zhao, Yue, et al., “Towards Unsupervised HPO for Outlier Detection”, https://arxiv.org/pdf/2208.11727.pdf, Aug. 24, 2022, 15pgs. [cited by applicant]
Yogatama et al., “Efficient transfer learning method for automatic hyperparameter tuning.”, Artificial intelligence and statistics. PMLR, 2014, 9 pages. [cited by applicant]
Wistuba et al., “Learning hyperparameter optimization initializations.” 2015 IEEE international conference on data science and advanced analytics (DSAA). IEEE, 2015. [cited by applicant]
Thornton et al., “Auto-WEKA: Combined selection and hyperparameter optimization of classification algorithms.”, Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining. 2013. [cited by applicant]
Maher et al., “Smartml: A meta learning-based framework for automated selection and hyperparameter tuning for machine learning algorithms.” EDBT: 22nd International conference on extending database technology. 2019. [cited by applicant]
Jomaa et al., “Dataset2vec: Learning dataset meta-features.” Data Mining and Knowledge Discovery 35.3 (2021): 964-985. [cited by applicant]
Hutter et al., “Sequential model-based optimization for general algorithm configuration.” International conference on learning and intelligent optimization. Berlin, Heidelberg: Springer Berlin Heidelberg, 2011. [cited by applicant]
Kim et al., “Learning to transfer initializations for bayesian hyperparameter optimization.”, arXiv preprint arXiv:1710.06219, 2017. [cited by applicant]
Feurer et al., ; “Initializing bayesian hyperparameter optimization via meta-learning”; Proceedings of the AAAI conference on artificial intelligence. vol. 29. No. 1. 2015, 8 pages. [cited by applicant]
Bardenet et al., “Collaborative hyperparameter tuning.” International conference on machine learning. PMLR, 2013, 9 pages. [cited by applicant]