IP Library Granted Patent US 12,647,435
Granted Patent B2
US 12,647,435 · App. 17/660,164 · Granted Jun 2, 2026

Detection of user anomalies for software as a service application traffic with high and low variance feature modeling

Inventor: Muhammad Aurangzeb Akhtar (Santa Clara, CA)
Assignee: Palo Alto Networks, Inc.
H04L63/1425H04L63/1433
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,647,435
App. No.
17/660,164
Granted
Jun 2, 2026
Kind
B2
Abstract

Low variance clustering models and high variance clustering models comprising low and high variance features of user Software as a Service application traffic detect anomalous user behavior and, when risk thresholds are exceeded, trigger behavioral alerts. The low and high variance clustering models are trained with feature vectors that are dimension reduced using principal component analysis and clusters therein are classified as normal, benign, or malicious. Models are trained repeatedly in a sliding time window of training data to detect recent and potentially malicious user behavior. Behavioral alerts are triggered according to criterion specific to each of the low and high variance clustering models that account for increased risk associated with anomalous changes in low variance features.

Claims (37)

1 . A method comprising:

at least one of identifying a first feature as low variance that was previously identified as high variance and identifying a second feature as high variance that was previously identified as low variance based, at least in part, on statistics for feature values generated from network traffic data in a sliding window of time intervals for features identified as high variance features and features identified as low variance features, wherein the sliding window of time intervals precedes a current time interval;

determining, from network traffic data of a first user and of the current time interval, feature values corresponding to Anything-as-a-Service usage behavior of the first user;

generating a first feature vector with a first subset of the feature values corresponding to the features identified as high variance features and a second feature vector with a second subset of the feature values corresponding to the features identified as low variance features;

inputting the first feature vector into a low variance clustering model for a first behavioral scope associated with the first user and the second feature vector into a high variance clustering model for the first behavioral scope, wherein the low variance clustering model and the high variance clustering model comprise clusters that were generated with one or more clustering algorithms, wherein the low variance clustering model and the high variance clustering model output classifications of user behavior based on proximity of feature vectors to clusters of low and high variance feature vectors, respectively, labeled as normal or abnormal, and wherein the high and low variance clustering models are repeatedly trained with the network traffic data in the sliding window of time intervals preceding the current time interval; and

indicating detection of anomalous behavior for the first user based, at least in part, on output of the high and low variance clustering models.

2 . The method of claim 1 , further comprising incrementing a risk score for the first user based on at least one of the high and low variance clustering models generating outputs comprising indications of anomalous behavior for the first user.

3 . The method of claim 2 , wherein the risk score is incremented by an amount based, at least in part, on risk associated with features corresponding to inputs to the high and low variance clustering models for outputs comprising the indications of anomalous behavior for the first user.

4 . The method of claim 2 , further comprising generating an alert indicating anomalous behavior for the first user based, at least in part, on the risk score for the first user exceeding a threshold risk score.

5 . The method of claim 1 , wherein the first behavioral scope comprises a scope of network traffic data for at least the first user used to train the high and low variance clustering models.

6 . The method of claim 1 , wherein the one or more clustering algorithms comprise at least one of k-means clustering, density-based spatial clustering of applications with noise, and hierarchical clustering.

7 . The method of claim 1 , further comprising adjusting risk score weights for the features identified as high variance and the features identified as low variance.

8 . A non-transitory machine-readable medium having program code stored thereon, the program code comprising instructions to:

at least one of identify a first feature as low variance that was previously identified as high variance and identify a second feature as high variance that was previously identified as low variance based, at least in part, on statistics for feature values generated from network traffic data in a sliding window of time intervals for features identified as high variance features and features identified as low variance features, wherein the sliding window of time intervals precedes a current time interval;

determine, from network traffic data of a first user and of the current time interval, feature values corresponding to Anything-as-a-Service usage behavior of the first user;

generate a first feature vector with a first subset of the feature values corresponding to the features identified as high variance features and a second feature vector with a second subset of the feature values corresponding to the features identified as low variance features;

input the first feature vector into a low variance clustering model for a first behavioral scope associated with the first user and the second feature vector into a high variance clustering model for the first behavioral scope, wherein the low variance clustering model and the high variance clustering model comprise clusters that were generated with one or more clustering algorithms, wherein the low variance clustering model and the high variance clustering model output classifications of user behavior based on proximity of feature vectors to clusters of low and high variance feature vectors, respectively, labeled as normal or abnormal, and wherein the high and low variance clustering models are repeatedly trained with the network traffic data of the sliding window of time intervals preceding the current time interval; and

indicate detection of anomalous behavior for the first user based, at least in part, on output of the high and low variance clustering models.

9 . The non-transitory machine-readable medium of claim 8 , wherein the program code further comprises instructions to increment a risk score for the first user based on at least one of the high and low variance clustering models generating outputs comprising indications of anomalous behavior for the first user.

10 . The non-transitory machine-readable medium of claim 9 , wherein the risk score is incremented by an amount based, at least in part, on risk associated with features corresponding to inputs to the high and low variance clustering models for outputs comprising the indications of anomalous behavior for the first user.

11 . The non-transitory machine-readable medium of claim 9 , wherein the program code further comprises instructions to generate an alert indicating anomalous behavior for the first user based, at least in part, on the risk score for the first user exceeding a threshold risk score.

12 . The non-transitory machine-readable medium of claim 8 , wherein the first behavioral scope comprises a scope of network traffic data for at least the first user used to train the high and low variance clustering models.

13 . The non-transitory machine-readable medium of claim 8 , wherein the one or more clustering algorithms comprise at least one of k-means clustering, density-based spatial clustering of applications with noise, and hierarchical clustering.

14 . The non-transitory machine-readable medium of claim 8 , wherein the program code further comprises instructions to adjust risk score weights for the features identified as high variance and the features identified as low variance.

15 . An apparatus comprising:

a processor; and

a machine-readable medium having instructions stored thereon that are executable by the processor to cause the apparatus to,

at least one of identify a first feature as low variance that was previously identified as high variance and identify a second feature as high variance that was previously identified as low variance based, at least in part, on statistics for feature values generated from network traffic data in a sliding window of time intervals for features identified as high variance features and features identified as low variance features, wherein the sliding window of time intervals precedes a current time interval;

determine, from network traffic data of a first user and of the current time interval, feature values corresponding to Anything-as-a-Service usage behavior of the first user;

generate a first feature vector with a first subset of the feature values corresponding to the features identified as high variance features and a second feature vector with a second subset of the feature values corresponding to the features identified as low variance features;

input the first feature vector into a low variance clustering model for a first behavioral scope associated with the first user and the second feature vector into a high variance clustering model for the first behavioral scope, wherein the low variance clustering model and the high variance clustering model comprise clusters that were generated with one or more clustering algorithms, wherein the low variance clustering model and the high variance clustering model output classifications of user behavior based on proximity of feature vectors to clusters of low and high variance feature vectors, respectively, labeled as normal or abnormal, and wherein the high and low variance clustering models are repeatedly trained with the network traffic data of the sliding window of time intervals preceding the current time interval; and

indicate detection of anomalous behavior for the first user based, at least in part, on output of the high and low variance clustering models.

16 . The apparatus of claim 15 , wherein the machine-readable medium further has stored thereon instructions executable by the processor to cause the apparatus to increment a risk score for the first user based on at least one of the high and low variance clustering models generating outputs comprising indications of anomalous behavior for the first user.

17 . The apparatus of claim 16 , wherein the risk score is incremented by an amount based, at least in part, on risk associated with features corresponding to inputs to the high and low variance clustering models for outputs comprising the indications of anomalous behavior for the first user.

18 . The apparatus of claim 16 , wherein the machine-readable medium further has stored thereon instructions executable by the processor to cause the apparatus to generate an alert indicating anomalous behavior for the first user based, at least in part, on the risk score for the first user exceeding a threshold risk score.

19 . The apparatus of claim 15 , wherein the first behavioral scope comprises a scope of network traffic data for at least the first user used to train the high and low variance clustering models.

20 . The apparatus of claim 15 , wherein the one or more clustering algorithms comprise at least one of k-means clustering, density-based spatial clustering of applications with noise, and hierarchical clustering.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2022
From: AKHTAR, MUHAMMAD AURANGZEB
To: PALO ALTO NETWORKS, INC.
Reel/Frame 059670/0746 →
Continuity (1)
Related Publication 20230344842A1 · Oct 26, 2023
References Cited (11)
US 11706234B2 · Saunders · 2023 [cited by examiner]
US 20160050224A1 · Ricafort · 2016 [cited by examiner]
US 20170185758A1 · Oliker · 2017 [cited by examiner]
US 20170345268A1 · Cho · 2017 [cited by examiner]
US 20180322004A1 · Jain · 2018 [cited by examiner]
US 20190286242A1 · Ionescu · 2019 [cited by examiner]
US 20210049413A1 · Ma · 2021 [cited by examiner]
US 20210119941A1 · Hoole · 2021 [cited by examiner]
Moon et. al. “An ensemble approach to anomaly detection using high- and low-variance principal components”, Computers and Electrical Engineering 99 (2022) 107773, Available online Feb. 15, 2022. (Year: 2022). [cited by examiner]
Kind, et al., “Histogram-Based Traffic Anomaly Detection”, IEEE Transactions on Network and Service Management, vol. 6, No. 2, pp. 110-121, Jun. 2009. [cited by applicant]
Ringberg, et al., “Sensitivity of PCA for Traffic Anomaly Detection”, ACM Sigmetrics Performance Evaluation Review, vol. 35, Issue 1, pp. 109-120, Jun. 2007. [cited by applicant]