IP Library › Granted Patent US 12,688,424
Granted Patent B2
US 12,688,424 · App. 17/394,193 · Granted Jul 21, 2026

Anomaly detection performance enhancement using gradient-based feature importance

Inventors: Saeid Allahdadian (Vancouver, CA); Yuting Sun (Bellevue, WA); Navaneeth Jamadagni (South San Francisco, CA); Felix Schmidt (Niederweningen, CH); Maria Vlachopoulou (Bellevue, WA)
Assignee: Oracle International Corporation
G06N3/084G06F17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,424
App. No.
17/394,193
Filed
Aug 4, 2021
Granted
Jul 21, 2026
Kind
B2
Art Unit
2128
USPC
706/20
Abstract

Herein are machine learning techniques that adjust reconstruction loss of a reconstructive model, such as a principal component analysis (PCA), based on importances of features. In an embodiment having a reconstructive model that more or less accurately reconstructs its input, a computer measures, for each feature, a respective importance that is based on the reconstructive model. For example, importance may be based on grading samples that the reconstructive model correctly or incorrectly inferenced. For each feature during production inferencing, a respective original loss from the reconstructive model measures a difference between a value of the feature in an input and a reconstructed value of the feature generated by the reconstructive model. For each feature, the respective importance of the feature is applied to the respective original loss to generate a respective weighted loss, which compensates for concept drift. The weighted losses of the features of the input are collectively detected as anomalous or non-anomalous.

Claims (51)

1 . A method comprising:

measuring, for each feature of a plurality of features, a respective importance that is based on a) a reconstructive model, b) a gradient of the feature, and c) at least one selected from the group consisting of: a count of true positives, a count of false positives, an exponential term, and a z-score;

wherein the gradients of the plurality of features are based on a first plurality of log entries of a console log or a server log;

measuring, for each feature of the plurality of features in a second plurality of log entries of the console log or the server log, a respective original loss from the reconstructive model;

generating, for each feature of the plurality of features, a respective weighted loss of a plurality of weighted losses, wherein the weighted loss is based on: the original loss of the feature and the importance of the feature;

detecting, based on the plurality of weighted losses, that an anomalous subset from the second plurality of log entries of the console log or the server log are anomalous;

recording, for each log entry from the anomalous subset, a respective label that indicates a true positive or a false positive;

replacing, in the first plurality of log entries of the console log or the server log, a subset of the first plurality of log entries of the console log or the server log with the anomalous subset;

measuring, based on said replacing, recalculated importances of the plurality of features;

generating, based on the recalculated importances of the plurality of features and a plurality of original losses for a new log entry, a new plurality of weighted losses; and

detecting, based on the new plurality of weighted losses, that the new log entry is anomalous;

wherein the method is performed by one or more computers after the reconstructive model was trained.

2 . The method of claim 1 wherein said measuring the importance of the feature comprises measuring a plurality of gradients for the feature.

3 . The method of claim 2 wherein the plurality of gradients for the feature is based on the reconstructive model.

4 . The method of claim 2 wherein:

the method further comprises the reconstructive model generating a respective inference for each tuple of a plurality of tuples;

each gradient of the plurality of gradients for the feature is based on a respective tuple of the plurality of tuples.

5 . The method of claim 1 wherein the respective importance is further based on an eigenvector.

6 . The method of claim 5 wherein the respective importance is further based on at least one selected from the group consisting of:

a loss gradient,

a gradient for the feature of the importance,

a backpropogation, and

a sum or average of original losses.

7 . The method of claim 1 further comprising measuring, for each feature of a plurality of features, a respective second importance that is based on said labels of the log entries from the anomalous subset that indicate a true positive or a false positive.

8 . The method of claim 1 further comprising measuring, for each feature of a plurality of features, a respective second importance that is based on said plurality of weighted losses.

9 . The method of claim 1 wherein said reconstructive model is one selected from the group consisting of a principal component analysis (PCA), a recurrent neural network (RNN), a variational autoencoder (VAE), and an incremental PCA (IPCA).

10 . One or more computer-readable non-transitory media storing instructions that, when executed by one or more processors, cause after a reconstructive model was trained:

measuring, for each feature of a plurality of features, a respective importance that is based on a) the reconstructive model, b) a gradient of the feature, and c) at least one selected from the group consisting of: a count of true positives, a count of false positives, an exponential term, and a z-score;

wherein the gradients of the plurality of features are based on a first plurality of log entries of a console log or a server log;

measuring, for each feature of the plurality of features in a second plurality of log entries of the console log or the server log, a respective original loss from the reconstructive model;

generating, for each feature of the plurality of features, a respective weighted loss of a plurality of weighted losses, wherein the weighted loss is based on: the original loss of the feature and the importance of the feature;

detecting, based on the plurality of weighted losses, that an anomalous subset from the second plurality of log entries of the console log or the server log are anomalous;

recording, for each log entry from the anomalous subset, a respective label that indicates a true positive or a false positive;

replacing, in the first plurality of log entries of the console log or the server log, a subset of the first plurality of log entries of the console log or the server log with the anomalous subset;

measuring, based on said replacing, recalculated importances of the plurality of features;

generating, based on the recalculated importances of the plurality of features and a plurality of original losses for a new log entry, a new plurality of weighted losses; and

detecting, based on the new plurality of weighted losses, that the new log entry is anomalous.

11 . The one or more computer-readable non-transitory media of claim 10 wherein said measuring the importance of the feature comprises measuring a plurality of gradients for the feature.

12 . The one or more computer-readable non-transitory media of claim 11 wherein the plurality of gradients for the feature is based on the reconstructive model.

13 . The one or more computer-readable non-transitory media of claim 11 wherein:

the instructions further cause the reconstructive model generating a respective inference for each tuple of a plurality of tuples;

each gradient of the plurality of gradients for the feature is based on a respective tuple of the plurality of tuples.

14 . The one or more computer-readable non-transitory media of claim 10 wherein the respective importance is further based on an eigenvector.

15 . The one or more computer-readable non-transitory media of claim 14 wherein the respective importance is further based on at least one selected from the group consisting of:

a loss gradient,

a gradient for the feature of the importance,

a backpropogation, and

a sum or average of original losses.

16 . The one or more computer-readable non-transitory media of claim 10 wherein the instructions further cause measuring, for each feature of a plurality of features, a respective second importance that is based on said labels of the log entries from the anomalous subset that indicate a true positive or a false positive.

17 . The one or more computer-readable non-transitory media of claim 10 wherein the instructions further cause measuring, for each feature of a plurality of features, a respective second importance that is based on said plurality of weighted losses.

18 . The one or more computer-readable non-transitory media of claim 10 wherein said reconstructive model is one selected from the group consisting of a principal component analysis (PCA), a recurrent neural network (RNN), a variational autoencoder (VAE), and an incremental PCA (IPCA).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2021
From: ALLAHDADIAN, SAEID; SUN, YUTING; JAMADAGNI, NAVANEETH; SCHMIDT, FELIX; VLACHOPOULOU, MARIA
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 057097/0826 →
Continuity (1)
Related Publication 20230043993A1 · Feb 9, 2023
References Cited (43)
US 11537902B1 · Aydore · 2022 [cited by examiner]
US 11798090B1 · Nazir · 2023 [cited by examiner]
US 20170169360A1 · Veeramachaneni et al. · 2017 [cited by applicant]
US 20180109589A1 · Ozaki et al. · 2018 [cited by applicant]
US 20180158078A1 · Hsieh · 2018 [cited by applicant]
US 20200004616A1 · Natsumeda · 2020 [cited by applicant]
US 20200381084A1 · Kawas · 2020 [cited by applicant]
US 20210124981A1 · Kim · 2021 [cited by applicant]
US 20220188410A1 · Allahdadian et al. · 2022 [cited by applicant]
US 20220303288A1 · Wang · 2022 [cited by examiner]
US 20230024884A1 · Casserini · 2023 [cited by examiner]
CA 3037326 · 2020 [cited by applicant]
CA 3037326A1 · 2020 [cited by applicant]
Shilin He, Pinjia He, Zhuangbin Chen, Tianyi Yang, Yuxin Su, and Michael R. Lyu. 2021. A Survey on Automated Log Analysis for Reliability Engineering. ACM Comput. Surv. 54, 6, Article 130 (Jul. 2022), 37 pages. https://… [cited by examiner]
Chen, Zhuangbin & Liu, Jinyang & Gu, Wenwei & Su, Yuxin & Lyu, Michael. (2021). Experience Report: Deep Learning-based System Log Analysis for Anomaly Detection. 10.48550/arXiv.2107.05908. (Year: 2021). [cited by examiner]
Min Du, Feifei Li, Guineng Zheng, and Vivek Srikumar. 2017. DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communication… [cited by examiner]
Xu Zhang, Yong Xu, Qingwei Lin, Bo Qiao, Hongyu Zhang, Yingnong Dang, Chunyu Xie, Xinsheng Yang, Qian Cheng, Ze Li, Junjie Chen, Xiaoting He, Randolph Yao, Jian-Guang Lou, Murali Chintalapati, Furao Shen, and Dongmei Zh… [cited by examiner]
Ross, Gordon J., et al. “Exponentially weighted moving average charts for detecting concept drift”, Pattern Recognition Letters 33.2 (Year: 2012). [cited by applicant]
Jaworski et al., “Concept drift detection using autoencoders in data streams processing”, In International Conference on Artificial Intelligence and Soft Computing (Year: 2020). [cited by applicant]
Kamei et al., “The Effects of Over and Under Sampling on Fault-Prone Module Detection”, Nara Institute of Science and Technology Academic Repository, ESEM 2007, dated Sep. 2007, 10 pages. [cited by applicant]
Altmann et al., “Permutation Importance: A corrected Feature Importance Measure”, Department of Computational Biology and Applied Algorithmics, Max Planck Institute for Informatics, vol. 00, No. 00 2009, Year 2009, 8 pa… [cited by applicant]
Bach et al., A Bayesian Approach to Concept Drift, dated 2010, 9 pages. [cited by applicant]
Bontemps et al., “Collective Anomaly Detection Based on Long Short Term Memory Recurrent Neural Network”, dated 2016, 12 pages. [cited by applicant]
Bornelöv et al., “Selection of Significant Features Using Monte Carlo Feature Selection”, Springer International Publishing, Year 2016, 14 pages. [cited by applicant]
Carta et al., “A Local Feature Engineering Strategy to Improve Network Anomaly Detection”, Department of Mathematics and Computer Science, University of Cagliari, dated Oct. 21, 2020, 30 pages. [cited by applicant]
Chawla, “Data Mining for Imbalanced Datasets: An Overview”, Department of Computer Science and Engineering University of Notre Dame, https://www.researchgate.net/ publication/226755026, dated Jan. 2005, 16 pages. [cited by applicant]
Du et al. DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning, CCS'17, Oct. 30-Nov. 3, 2017, 14 pages. [cited by applicant]
Gama et al., “A Survey on Concept Drift Adaptation”, ACM Computing Surveys, vol. 1, No. 1, Article 1, Publication date: Jan. 2013, 44 pages. [cited by applicant]
Golan et al., “Deep Anomaly Detection Using Geometric Transformations”, 32nd Conference on Neural Information Processing Systems (NeurIPS dated 2018), Montréal, Canada, 12 pages. [cited by applicant]
Alexey Tsymbal, “The Problem of Concept Drift: Definitions and Related Work”, dated Apr. 29, 2004, 7 pages. [cited by applicant]
Hasan et al., “Attack and Anomaly Detection in IoT Sensors in IoT Sites Using Machine Learning Approaches”, Department of Computer Science and Engineering, Khulna University of Engineering & Technology, published by Els… [cited by applicant]
YuanZhong, Zhu, “Intrusion Detection Method based on Improved BP Neural Network Research”, International Journal of Security and Its Applications vol. 10, No. 5 (2016) pp. 193-202. [cited by applicant]
Larasati et al., “Improve the Accuracy of Support Vector Machine Using Chi Square Statistic and Term Frequency Inverse Document Frequency on Movie Review Sentiment Analysis”, Scientific Journal of Informatics, vol. 6, N… [cited by applicant]
Liu et al., “Imbalanced Text Classification: A Term Weighting Approach”, Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University, dated 2007, 12 pages. [cited by applicant]
Luo et al., “A Revisit of Sparse Coding Based Anomaly Detection in Stacked RNN Framework”, dated Oct. 2017, 9 pages. [cited by applicant]
Naseer et al., “Enhanced Network Anomaly Detection Based on Deep Neural Networks”, Journal of Latex Class Files, vol. 14, No. 8, Aug. 2015, 16 pages. [cited by applicant]
Nguyen et al., “GEE: A Gradient-Based Explainable Variational Autoencoder for Network Anomaly Detection”, https://ieeexplore.ieee.org/abstract/document/8802833, Mar. 2019, 10 pages. [cited by applicant]
Panday et al., “Feature Weighting as a Tool for Unsupervised Feature Selection”, Information Processing Letters, https://doi.org/10.1016/j.ipl.2017.09.005, Year 2017, 12 pages. [cited by applicant]
Tuor et al., “Deep Learning for Unsupervised Insider Threat Detection in Structured Cybersecurity Data Streams”, dated Dec. 15, 2017, 9 pages. [cited by applicant]
Webb et al., Characterizing Concept Drift, Data Mining and Knowledge Discovery, dated Jul. 2016, 30 pages. [cited by applicant]
Yousefi-Azar et al., “Autoencoder-based Feature Learning for Cyber Security Applications”, dated 2017, 8 pages. [cited by applicant]
HaddadPajouh et al, “A Two-layer Dimension Reduction and Two-tier Classification Model for Anomaly-Based Intrusion Detection in IoT Backbone Networks”, dated 2016, 12 pages. [cited by applicant]
Ross et al., “Exponentially weighted moving average charts for detecting concept drift.” Pattern Recognition Letters 33.2 (Year: 2012). [cited by applicant]