IP Library › Granted Patent US 11,531,676
Granted Patent B2
US 11,531,676 · App. 16/942,520 · Granted Dec 20, 2022

Method and system for anomaly detection based on statistical closed-form isolation forest analysis

Inventors: Yair Horesh (Hod Hasharon, IL); Nir Keret (Even-Yehuda, IL); Yehezkel Shraga Resheff (Hod Hasharon, IL)
Assignee: INTUIT INC.
G06F16/2462G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,531,676
App. No.
16/942,520
Filed
Jul 29, 2020
Granted
Dec 20, 2022
Kind
B2
Art Unit
2167
USPC
707/755
Abstract

Certain embodiments of the present disclosure provide techniques for detecting anomalous activity in a computing system. The method generally includes receiving a request to perform an action in a computing system. The request is added to a historical time-series data set. A portion of the historical time-series data set is selected for use in determining whether the received request is an anomalous request, and a set of previously identified outliers are removed from the selected portion of the historical time-series data set. An anomaly score is calculated based on a statistical analysis of the received request and the selected portion of the historical time-series data set, wherein the anomaly score comprises a predicted number of operations executed to isolate the received request from the selected portion of the historical time-series data set. One or more actions are taken to process the received request based on the calculated anomaly score.

Claims (74)

1. A method for detecting anomalous activity in a computing system, comprising:

receiving a request to perform an action in a computing system;

adding the request to a historical time-series data set;

selecting a portion of the historical time-series data set to be used to determine whether the received request is an anomalous request;

removing a set of previously identified outliers from the selected portion of the historical time-series data set;

calculating an anomaly score for the received request based on a statistical analysis of the received request and the selected portion of the historical time-series data set, wherein the anomaly score comprises a predicted number of operations executed to isolate the received request from the selected portion of the historical time-series data set; and

taking one or more actions to process the received request based on the calculated anomaly score.

2. The method of claim 1 , wherein calculating the anomaly score for the received request comprises:

determining that a matching data point exists in the selected portion of the historical time-series data set for the received request; and

based on determining that the matching data point exists, outputting one of a reserved value or a highest supported value for a data type associated with the anomaly score as the calculated anomaly score.

3. The method of claim 1 , wherein calculating the anomaly score for the received request comprises:

determining that the received request corresponds to a largest value in the selected portion of the historical time-series data set; and

based on determining that the received request corresponds to a largest value in the selected portion of the historical time-series data set, calculating the anomaly score as a geometric distribution based on a value of the received request and a difference between the value of the received request and a next largest value in the selected portion of the historical time-series data set.

4. The method of claim 1 , wherein calculating the anomaly score for the received request comprises:

determining that the received request corresponds to a smallest value in the selected portion of the historical time-series data set; and

based on determining that the received request corresponds to a smallest value in the selected portion of the historical time-series data set, calculating the anomaly score as a geometric distribution based on a value of the received request and a difference between the value of the received request and a next smallest value in the selected portion of the historical time-series data set.

5. The method of claim 1 , wherein the calculated anomaly score represents a number of lines expected to be drawn by an isolation forest algorithm to isolate a value of the received request within the selected portion of the historical time-series data set.

6. The method of claim 5 , wherein calculating the anomaly score comprises:

calculating a first value corresponding to an expected number of lines to be drawn between the value of the received request and a first data point in the selected portion of the historical time-series data set corresponding to a closest data point in the historical time-series data set having a value greater than the value of the received request; and

calculating a second value corresponding to an expected number of lines to be drawn between the value of the received request and a second data point in the selected portion of the historical time-series data set corresponding to a closest data point in the historical time-series data set having a value less than the value of the received request, wherein the predicted number of operations executed to isolate the received request from the selected portion of the historical time-series data set comprises a value calculated using a geometric distribution based on the first value and the second value.

7. The method of claim 6 , wherein calculating the anomaly score comprises calculating a score based on a largest value in the selected portion of the historical time-series data set, the first value, and the second value.

8. The method of claim 7 , wherein the anomaly score is calculated as a sum of:

an expected number of lines drawn in the selected portion of the data set before a line is drawn between the first data point and the second data point;

a product of the second value and a probability that a line has already been drawn above a data point representing the received request and the first data point; and

a product of the first value and a probability that a line has already been drawn below a data point representing the received request and the second data point.

9. The method of claim 1 , wherein the taking one or more actions to process the received request comprises:

determining, based on a comparison of the anomaly score and a threshold score, that the received request corresponds to an anomalous request; and

adding the received request to the set of previously identified outliers.

10. The method of claim 1 , wherein:

the received request comprises a request to access a service in a computing environment, and

the method further comprises: featurizing one or more parameters in the received request to access the service corresponding to information about a point of origin of the received request into one or more numerical values.

11. A system, comprising:

a processor; and

a memory having instructions stored thereon which, when executed by the processor, performs an operation for detecting anomalous activity in a computing system, the operation comprising:

receiving a request to perform an action in a computing system;

adding the request to a historical time-series data set;

selecting a portion of the historical time-series data set to be used to determine whether the received request is an anomalous request;

removing a set of previously identified outliers from the selected portion of the historical time-series data set;

calculating an anomaly score for the received request based on a statistical analysis of the received request and the selected portion of the historical time-series data set, wherein the anomaly score comprises a predicted number of operations executed to isolate the received request from the selected portion of the historical time-series data set; and

taking one or more actions to process the received request based on the calculated anomaly score.

12. The system of claim 11 , wherein calculating the anomaly score for the received request comprises:

determining that a matching data point exists in the selected portion of the historical time-series data set for the received request; and

based on determining that the matching data point exists, outputting one of a reserved value or a highest supported value for a data type associated with the anomaly score as the calculated anomaly score.

13. The system of claim 11 , wherein calculating the anomaly score for the received request comprises:

determining that the received request corresponds to a largest value in the selected portion of the historical time-series data set; and

based on determining that the received request corresponds to a largest value in the selected portion of the historical time-series data set, calculating the anomaly score as a geometric distribution based on a value of the received request and a difference between the value of the received request and a next largest value in the selected portion of the historical time-series data set.

14. The system of claim 11 , wherein calculating the anomaly score for the received request comprises:

determining that the received request corresponds to a smallest value in the selected portion of the historical time-series data set; and

based on determining that the received request corresponds to a smallest value in the selected portion of the historical time-series data set, calculating the anomaly score as a geometric distribution based on a value of the received request and a difference between the value of the received request and a next smallest value in the selected portion of the historical time-series data set.

15. The system of claim 11 , wherein:

the calculated anomaly score represents a number of lines expected to be drawn by an isolation forest algorithm to isolate a value of the received request within the selected portion of the historical time-series data set; and

calculating the anomaly score comprises:

calculating a first value corresponding to an expected number of lines to be drawn between the value of the received request and a first data point in the selected portion of the historical time-series data set corresponding to a closest data point in the historical time-series data set having a value greater than the value of the received request; and

calculating a second value corresponding to an expected number of lines to be drawn between the value of the received request and a second data point in the selected portion of the historical time-series data set corresponding to a closest data point in the historical time-series data set having a value less than the value of the received request, wherein the predicted number of operations executed to isolate the received request from the selected portion of the historical time-series data set comprises a value calculated using a geometric distribution based on the first value and the second value.

16. The system of claim 15 , wherein calculating the anomaly score comprises calculating a score based on a largest value in the selected portion of the historical time-series data set, the first value, and the second value.

17. The system of claim 16 , wherein the anomaly score is calculated as a sum of:

an expected number of lines drawn in the selected portion of the data set before a line is drawn between the first data point and the second data point;

a product of the second value and a probability that a line has already been drawn above a data point representing the received request and the first data point; and

a product of the first value and a probability that a line has already been drawn below a data point representing the received request and the second data point.

18. The system of claim 11 , wherein the taking one or more actions to process the received request comprises:

determining, based on a comparison of the anomaly score and a threshold score, that the received request corresponds to an anomalous request; and

adding the received request to the set of previously identified outliers.

19. The system of claim 11 , wherein:

the received request comprises a request to access a service in a computing environment, and

the operation further comprises featurizing one or more parameters in the received request to access the service corresponding to information about a point of origin of the received request into one or more numerical values.

20. A system, comprising:

a plurality of computing resources; and

a request gateway configured to:

receive a request to perform an action on one or more of the plurality of computing resources;

add the request to a historical time-series data set;

select a portion of the historical time-series data set to be used to determine whether the received request is an anomalous request;

remove a set of previously identified outliers from the selected portion of the historical time-series data set;

calculate an anomaly score for the received request based on a statistical analysis of the received request and the selected portion of the historical time-series data set, wherein the anomaly score comprises a predicted number of operations executed to isolate the received request from the selected portion of the historical time-series data set; and

take one or more actions to process the received request based on the calculated anomaly score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2020
From: HORESH, YAIR; KERET, NIR; RESHEFF, YEHEZKEL SHRAGA
To: INTUIT INC.
Reel/Frame 053346/0647 →
Continuity (1)
Related Publication 20220035806A1 · Feb 3, 2022