IP Library Granted Patent US 12,468,780
Granted Patent B2
US 12,468,780 · App. 17/380,731 · Granted Nov 11, 2025

Balancing feature distributions using an importance factor

Inventors: Matteo Casserini (Zurich, CH); Saeid Allahdadian (Vancouver, CA); Felix Schmidt (Baden-Daettwil, CH); Andrew Brownsword (Vancouver, CA)
Assignee: Oracle International Corporation
G06F18/217G06F18/2113G06F40/279G06N3/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,780
App. No.
17/380,731
Granted
Nov 11, 2025
Kind
B2
Abstract

Herein are machine learning techniques that adjust reconstruction loss of a reconstructive model such as an autoencoder based on importances of values of features. In an embodiment and before, during, or after training, the reconstructive model that more or less accurately reconstructs its input, a computer measures, for each distinct value of each feature, a respective importance that is not based on the reconstructive model. For example, importance may be based solely on a training corpus. For each feature during or after training, a respective original loss from the reconstructive model measures a difference between a value of the feature in an input and a reconstructed value of the feature generated by the reconstructive model. For each feature, the respective importance of the input value of the feature is applied to the respective original loss to generate a respective weighted loss. The weighted losses of the features of the input are collectively detected as anomalous or non-anomalous.

Claims (57)

1 . A method comprising:

measuring without using a reconstructive model:

a) for a distinct first categorical value of a particular feature of a plurality of features, a distinct first importance and

b) for a distinct second categorical value of the particular feature of the plurality of features, a distinct second importance; and

performing by a network switch:

i) measuring, for a network packet, a loss from the reconstructive model for the particular feature of the plurality of features;

ii) increasing the loss of the particular feature based on the second importance for the second categorical value of the particular feature;

iii) detecting, based on said increasing, that the network packet is anomalous; and

iv) rejecting, in response to said detecting, the network packet.

2 . The method of claim 1 wherein said measuring the first importance is based on an inverse document frequency that is based solely on categorical values of the particular feature.

3 . The method of claim 2 wherein:

the particular feature is one-hot encoded;

the inverse document frequency comprises a ratio having a denominator that is based solely on a count of the first categorical value of the particular feature.

4 . The method of claim 3 wherein each of the denominator and said count of the first categorical value of the particular feature is a count of the first categorical value of the particular feature in a training corpus.

5 . The method of claim 3 wherein:

a training corpus comprises a plurality of tuples;

each tuple in the plurality of tuples contains the plurality of features;

the ratio has a numerator that is not based on tuples in the plurality of tuples that lack a categorical value of the particular feature.

6 . The method of claim 1 further comprising manually assigning:

a first predefined multiplier to the first categorical value of the particular feature and

a second predefined multiplier to the second categorical value of the particular feature.

7 . The method of claim 1 further comprising at least one selected from a group consisting of:

training the reconstructive model based on the first importance,

training the reconstructive model before said measuring the first importance, and

training the reconstructive model after said measuring the first importance.

8 . The method of claim 1 wherein one is selected from a group consisting of:

said increasing occurs while training the reconstructive model and after training the reconstructive model,

said increasing occurs while training the reconstructive model but not after training the reconstructive model, and

said increasing occurs after training the reconstructive model but not while training the reconstructive model.

9 . The method of claim 1 wherein said increasing comprises a loss function accepting the first importance without the second importance.

10 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

measuring without using a reconstructive model:

a) for a distinct first categorical value of a particular feature of a plurality of features, a distinct first importance and

b) for a distinct second categorical value of the particular feature of the plurality of features, a distinct second importance;

performing by a network switch:

i) measuring, for a network packet, a loss from the reconstructive model for the particular feature of the plurality of features;

ii) increasing the loss of the particular feature based on the second importance for the second categorical value of the particular feature;

iii) detecting, based on said increasing, that the network packet is anomalous; and

iv) rejecting, in response to said detecting, the network packet.

11 . The one or more non-transitory computer-readable media of claim 10 wherein said measuring the first importance is based on an inverse document frequency that is based solely on categorical values of the particular feature.

12 . The one or more non-transitory computer-readable media of claim 11 wherein:

the particular feature is one-hot encoded;

the inverse document frequency comprises a ratio having a denominator that is based solely on a count of the first categorical value of the particular feature.

13 . The one or more non-transitory computer-readable media of claim 12 wherein each of the denominator and said count of the first categorical value of the particular feature is a count of the first categorical value of the particular feature in a training corpus.

14 . The one or more non-transitory computer-readable media of claim 12 wherein:

a training corpus comprises a plurality of tuples;

each tuple in the plurality of tuples contains the plurality of features;

the ratio has a numerator that is not based on tuples in the plurality of tuples that lack a categorical value of the particular feature.

15 . The one or more non-transitory computer-readable media of claim 10 wherein the instructions further cause at least one selected from a group consisting of:

training the reconstructive model based on the first importance,

training the reconstructive model before said measuring the first importance, and

training the reconstructive model after said measuring the first importance.

16 . The one or more non-transitory computer-readable media of claim 10 wherein one is selected from a group consisting of:

said increasing occurs while training the reconstructive model and after training the reconstructive model,

said increasing occurs while training the reconstructive model but not after training the reconstructive model, and

said increasing occurs after training the reconstructive model but not while training the reconstructive model.

17 . The one or more non-transitory computer-readable media of claim 10 wherein said increasing comprises a loss function accepting the first importance without the second importance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2021
From: CASSERINI, MATTEO; ALLAHDADIAN, SAEID; SCHMIDT, FELIX; BROWNSWORD, ANDREW
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 056936/0260 →
Continuity (1)
Related Publication 20230024884A1 · Jan 26, 2023
References Cited (66)
US 6963831B1 · Epstein · 2005 [cited by examiner]
US 11537902B1 · Aydore · 2022 [cited by examiner]
US 20110185233A1 · Belluomini · 2011 [cited by examiner]
US 20170169360A1 · Veeramachaneni et al. · 2017 [cited by applicant]
US 20180109589A1 · Ozaki et al. · 2018 [cited by applicant]
US 20180158078A1 · Hsieh · 2018 [cited by examiner]
US 20190370695A1 · Chandwani et al. · 2019 [cited by applicant]
US 20200004616A1 · Natsumeda · 2020 [cited by examiner]
US 20200076841A1 · Hajimirsadeghi · 2020 [cited by applicant]
US 20200175610A1 · Pikle · 2020 [cited by examiner]
US 20200278408A1 · Sung · 2020 [cited by examiner]
US 20200364585A1 · Chandrashekar · 2020 [cited by applicant]
US 20200410289A1 · Arunmozhi · 2020 [cited by examiner]
US 20210011832A1 · Togawa · 2021 [cited by applicant]
US 20220138504A1 · Moghadam et al. · 2022 [cited by applicant]
US 20220188410A1 · Allahdadian et al. · 2022 [cited by applicant]
US 20220303288A1 · Wang · 2022 [cited by examiner]
Zhou, Junlin, et al., “Unsupervised Learning Based Distributed Detection of Global Anomalies”, 2010, International Journal of Information Technology and Decision Making, vol. 2010, Nov. 2010, pp. 1-11. [cited by applicant]
Vikram, Adiya, et al., “Anomaly detection in Network Traffic Using Unsupervised Machine learning Approach”, 2020 5th Intl Conf on Communication and Electronics Systems (ICCES), vol. 5 (2020), pp. 476-479, Jul. 10, 2020,… [cited by applicant]
Angiulli, Fabrizio, “Concentration Free Outlier Detection”, 2017, Machine Learning and Knowledge Discovery in Databases, vol. 2017, pp. 3-19. [cited by applicant]
Du et al. DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning, CCS'17, Oct. 30-Nov. 3, 2017, 14 pages. [cited by applicant]
Altmann et al., “Permutation Importance: A corrected Feature Importance Measure”, Department of Computational Biology and Applied Algorithmics, Max Planck Institute for Informatics, vol. 00, No. 00 2009, Year 2009, 8 pa… [cited by applicant]
Kim et al., “Behavior-based anomaly detection on big data”, Edith Cowan University, Research Online, dated 2015, 9 pages. [cited by applicant]
Kamei et al., “The Effects of Over and Under Sampling on Fault-Prone Module Detection”, Nara Institute of Science and Technology Academic Repository, ESEM 2007, dated Sep. 2007, 10 pages. [cited by applicant]
Hu et al., “Anomalous User Activity Detection in Enterprise Multi-Source Logs”, dated Nov. 2017, 8 pages. [cited by applicant]
Hasan et al., “Attack and Anomaly Detection in IoT Sensors in IoT Sites Using Machine Learning Approaches”, Department of Computer Science and Engineering, Khulna University of Engineering & Technology, published by Els… [cited by applicant]
HaddadPajouh et al, “A Two-layer Dimension Reduction and Two-tier Classification Model for Anomaly-Based Intrusion Detection in IoT Backbone Networks”, dated 2016, 12 pages. [cited by applicant]
Golan et al., “Deep Anomaly Detection Using Geometric Transformations”, 32nd Conference on Neural Information Processing Systems (NeurIPS dated 2018), Montréal, Canada, 12 pages. [cited by applicant]
Goix, Nicolas, “How to Evaluate the Quality of Unsupervised Anomaly Detection Algorithms?”, Presented at ICML2016 Anomaly Detection Workshop, New York, NY, USA, 2016, 13 pages. [cited by applicant]
Kolter et al., “Dynamic Weighted Majority: An Ensemble Method for Drifting Concepts”, Journal of Machine Learning Research 8, dated 2007, 36 pages. [cited by applicant]
Ghosh et al., “Detecting Anomalous and Unknown Intrusions Against Programs”, dated 1998, 9 pages. [cited by applicant]
Larasati et al., “Improve the Accuracy of Support Vector Machine Using Chi Square Statistic and Term Frequency Inverse Document Frequency on Movie Review Sentiment Analysis”, Scientific Journal of Informatics, vol. 6, N… [cited by applicant]
Clemencon et al., “Scoring Anomalies: A M-estimation Formulation”, Proceedings of the 16th International Conference on Artifical Intelligence and Statistics (AISTATS) 2013, vol. 31 of JMLR, 9 pgs. [cited by applicant]
Chawla, “Data Mining for Imbalanced Datasets: An Overview”, Department of Computer Science and Engineering University of Notre Dame, https://www.researchgate.net/ publication/226755026, dated Jan. 2005, 16 pages. [cited by applicant]
Chandola et al., “Anomaly Detection : A Survey”, ACM Computing Surveys, dated Sep. 2009, 75 pages. [cited by applicant]
Carta et al., “A Local Feature Engineering Strategy to Improve Network Anomaly Detection”, Department of Mathematics and Computer Science, University of Cagliari, dated Oct. 21, 2020, 30 pages. [cited by applicant]
Buczak et al., “A Survey of Data Mining and Machine Learning Methods for Cyber Security Intrusion Detection”, IEEE Communications Surveys & Tutorials, vol. 18, No. 2, Second Quarter 2016, 24 pages. [cited by applicant]
Brownlee_Jason, “A Gentle Introduction to Cross-Entropy for Machine Learning”, dated Oct. 21, 2019, https://machinelearningmastery.com/cross-entropy-for-machine-learning/, 34 pages. [cited by applicant]
Bornelöv et al., “Selection of Significant Features Using Monte Carlo Feature Selection”, Springer International Publishing, Year 2016, 14 pages. [cited by applicant]
Bontemps et al., “Collective Anomaly Detection Based on Long Short Term Memory Recurrent Neural Network”, dated 2016, 12 pages. [cited by applicant]
Anonymous authors, “Versatile Outlier Detection With Outlier Preserving Distribution Mapping Autoencoders”, conference paper at ICLR 2020, dated 2019, 13 pages. [cited by applicant]
Godoy, Daniel, “Understanding Binary cross-Entropy / Log Loss: a Visual explanation”, dated Nov. 21, 2018, 13 pages. [cited by applicant]
Panday et al., “Feature Weighting as a Tool for Unsupervised Feature Selection”, Information Processing Letters, https://doi.org/10.1016/j.ipl.2017.09.005, Year 2017, 12 pages. [cited by applicant]
Yousefi-Azar et al., “Autoencoder-based Feature Learning for Cyber Security Applications”, dated 2017, 8 pages. [cited by applicant]
Wang et al., “Towards a Hierarchical Bayesian Model of Multi-View Anomaly Detection”, Twenty-Ninth International Joint Conference on Artificial Intelligence, dated Jul. 11, 2020, 7 pages. [cited by applicant]
Tuor et al., “Deep Learning for Unsupervised Insider Threat Detection in Structured Cybersecurity Data Streams”, dated Dec. 15, 2017, 9 pages. [cited by applicant]
Strobl et al., “Bias in Random Forest Variable Importance Measures: Illustrations,Sources and a Solution”, Research Report Series, Report 40, http://epub.wu.ac.at/1274/, Sep. 2006, 22 pages. [cited by applicant]
Shipmon et al., “Time Series Anomaly Detection”, Detection of Anomalous Drops with Limited Features and Sparse Examples in Noisy Highly Periodic Data, dated 2017, 9 pages. [cited by applicant]
Seleznyov et al., “Anomaly Intrusion Detection Systems: Handling Temporal Relations between Events”, dated 1999, 12 pages. [cited by applicant]
Schubert et al., “On Evaluation of Outlier Rankings and Outlier Scores”, dated 2012, 12 pages. [cited by applicant]
Sakurada et al., “Anomaly Detection Using Autoencoders with Nonlinear Dimensionality Reduction”, MLSDA '14, Dec. 2, 2014, Gold Coast, QLD, Australia Copyright 2014 ACM, 8 pages. [cited by applicant]
Kolosnjaji et al., “Deep Learning for Classification of Malware System Call Sequences”, dated 2016, 12 pages. [cited by applicant]
Rayana et al., “Sequential Ensemble Learning for Outlier Detection: A Bias-Variance Perspective”, dated Sep. 18, 2016, 11 pages. [cited by applicant]
YuanZhong, Zhu, “Intrusion Detection Method based on Improved BP Neural Network Research”, International Journal of Security and Its Applications vol. 10, No. 5 (2016) pp. 193-202. [cited by applicant]
Nguyen et al., “GEE: A Gradient-Based Explainable Variational Autoencoder for Network Anomaly Detection”, https://ieeexplore.ieee.org/abstract/document/8802833, Mar. 2019, 10 pages. [cited by applicant]
Nguyen et al., “An Evaluation Method for Unsupervised Anomaly Detection Algorithms”, Journal of Computer Science and Cybernetics, V.32, N.3, dated 2016, 14 pages. [cited by applicant]
Naseer et al., “Enhanced Network Anomaly Detection Based on Deep Neural Networks”, Journal of Latex Class Files, vol. 14, No. 8, Aug. 2015, 16 pages. [cited by applicant]
Moustafa et al., “A Holistic Review of Network Anomaly Detection Systems: A comprehensive Survey”, Journal of Network and Computer Applications, vol. 128, Feb. 15, 2019, pp. 33-55. [cited by applicant]
Mirza et al., “Computer Network Intrusion Detection Using Sequential LSTM Neural Networks Autoencoders”, IEEE, dated 2018, 4 pages. [cited by applicant]
Malhotra et al., “LSTM-based Encoder-Decoder for Multi-sensor Anomaly Detection”, Presented at ICML 2016 Anomaly Detection Workshop, New York, NY, USA, 2016. Copyright 2016—5 pages. [cited by applicant]
Malhotra et al., “Long Short Term Memory Networks for Anomaly Detection in Time Series”, ESANN dated Apr. 2015 proceedings, 6 pages. [cited by applicant]
Luo et al., “A Revisit of Sparse Coding Based Anomaly Detection in Stacked RNN Framework”, dated Oct. 2017, 9 pages. [cited by applicant]
Loganathan Gobinath et al., “Sequence to Sequence Pattern Learning Algorithm for Real-Time Anomaly Detection in Network Traffic”, dated 2018 IEEE, dated May 13, 2018, pp. 1-4. [cited by applicant]
Sabokrou et al., “Real-Time Anomaly Detection and Localization in Crowded Scenes”, dated 2015, 7 pages. [cited by applicant]
Hajimirsadeghi, U.S. Appl. No. 16/122,505, filed Sep. 5, 2018, Office Action, May 13, 2021. [cited by applicant]
Hajimirsadeghi, U.S. Appl. No. 16/122,505, filed Sep. 5, 2018, Notice of Allowance and Fees Due , Sep. 23, 2021. [cited by applicant]