Data science anomaly detection
A system, method, and a computer program product for determining anomalies in a dataset are provided. The dataset, which stores data structures of parameters, is received at an anomaly detection framework. The anomaly detection framework selects multiple models, including time series models, historical models, artificial intelligence models and isolation forest models to analyze the parameters in the dataset and determine anomalies in the data structures.
1 . A system, comprising:
a memory configured to store a plurality of anomaly detection models; and
a processor coupled to the memory and configured to perform operations, the operations comprising:
receiving data structures in a dataset transmitted by a computing network from a plurality of data sources coupled to the computer network, wherein the data structures include parameters storing transaction data generated by the plurality of data sources;
determining a first anomaly detection output for a data structure in the data structures using a first anomaly detection model in the plurality of anomaly detection models by:
comparing first parameters associated with the data structure to other first parameters associated with other data structures;
generating a parameter score for each first parameter in the first parameters based on the comparison;
aggregating parameter scores for the first parameters into an aggregated parameter score; and
generating the first anomaly detection output based on the aggregated parameter score;
determining a second anomaly detection output for the data structure using a second anomaly detection model in the plurality of anomaly detection models by:
providing a random forest of trees generated using minimum and maximum values of second parameters in the dataset;
isolating second parameters in the data structure using the random forest of trees by passing the second parameters in the data structure through trees in the random forest of trees;
determining path lengths of the second parameters of the data structure through the trees; and
generating the second anomaly detection output based on the path lengths;
determining a third anomaly detection output for the data structure using a third anomaly detection model in the plurality of anomaly detection models by:
generating third anomaly detection outputs over a historical time period; and
determining the third anomaly detection output based on a number of times second anomaly detection outputs indicated an anomaly in the data structure over the historical time period; and
determining a fourth anomaly detection output for the data structure in the dataset using a fourth anomaly detection model in the plurality of anomaly detection models by:
identifying fourth parameters from parameters associated with the dataset;
generating distributions of the fourth parameters over a second historical time period;
determining expected range of values in the distributions using the fourth parameters; and
generating the fourth anomaly detection output based on fourth parameters in the data structure and the expected range of values;
determining that the data structure includes the anomaly based on the first anomaly detection output, the second anomaly detection output, the third anomaly detection output, and the fourth anomaly detection output; and
generating an alert for the data structure based on the anomaly; and
transmitting the alert to a computing device, wherein the alert, upon receipt at the computing device, activates an icon on a display screen of the computing device indicating that the alert with the anomaly has been received at the computing device.
2 . The system of claim 1 , wherein the generating first anomaly detection output is based on the aggregated parameter score being above a parameter score threshold.
3 . The system of claim 1 , wherein determining the second anomaly detection output further comprises:
determining the minimum and maximum values for the second parameters in the dataset; and
generating the trees in the random forest of trees by randomly splitting the minimum and maximum values to generate nodes in the trees.
4 . The system of claim 1 , wherein determining the second anomaly detection output further comprises:
determining path lengths of the second parameters in the dataset;
averaging the path lengths of the second parameters in the dataset; and
further generating the second anomaly detection output based on the path lengths of the data structure being less than the average path lengths.
5 . The system of claim 1 , wherein generating the fourth anomaly detection output further comprises:
determining that a value of at least one fourth parameter in the data structure is outside of the expected range of values.
6 . The system of claim 1 , further comprising:
generating cohorts from the dataset using fifth parameters in data structures of the dataset; and
processing anomalies in data structures of each cohort.
7 . The system of claim 1 , wherein the dataset includes computer network data.
8 . The system of claim 1 , wherein the dataset includes funds available for lending.
9 . The system of claim 1 , wherein the dataset includes transactions and anomalies include fraudulent transactions in the dataset.
10 . A method, comprising:
receiving a dataset including data structures storing parameters over a computer network, wherein a ground truth or expected parameter values for the parameters are unknown;
selecting, based on an available memory of a computing device, a first anomaly detection model, a second anomaly detection model, a third anomaly detection model, and a fourth anomaly detection model from a plurality of anomaly detection models;
determining, using at least one processor and the available memory, a first anomaly detection output for a data structure in the dataset using the first anomaly detection model in the plurality of anomaly detection models and first parameters in the data structure;
determining, using the at least one processor and the available memory, a second anomaly detection output for the data structure using the second anomaly detection model in the plurality of anomaly detection models and second parameters in the data structure;
determining, using the at least one processor and the available memory, a third anomaly detection output using the third anomaly detection model in the plurality of anomaly detection models and third parameters in the data structure;
determining, using the at least one processor and the available memory, a fourth anomaly detection output using the fourth anomaly detection model in the plurality of anomaly detection models and fourth parameters in the data structure; and
determining that the data structure is an anomalous data structure based on the first anomaly detection output, the second anomaly detection output, the third anomaly detection output, and the fourth anomaly detection output,
wherein determining using the first anomaly detection model, the second anomaly detection model, the third anomaly detection model, and the fourth anomaly detection model is performed in parallel.
11 . The method of claim 10 , further comprising:
selecting the first anomaly detection model, the second anomaly detection model, the third anomaly detection model, and the fourth anomaly detection model from a pool of the plurality of anomaly detection models.
12 . The method of claim 10 , wherein the first anomaly detection output for the data structure in the dataset is determined by:
comparing the first parameters associated with the data structure to other first parameters associated with other data structures in the dataset;
generating a parameter score for each first parameter in the first parameters in the data structure based on the comparison;
aggregating parameter scores into an aggregated parameter score; and
generating the first anomaly detection output based on the aggregated parameter score.
13 . The method of claim 10 , wherein the second anomaly detection output for the data structure in the data structures is determined by:
providing a random forest of trees generated using minimum and maximum values of a training dataset;
isolating the second parameters in the data structure using the random forest of trees by passing the second parameters in the data structure through trees in the random forest of trees;
determining path lengths of the second parameters of the data structure through the trees; and
generating the second anomaly detection output based on the path lengths.
14 . The method of claim 10 , wherein the third anomaly detection output for the data structure in the data structures is determined by:
accessing second anomaly detection outputs stored over a historical time interval; and
determining the third anomaly detection output based on a number of times the second anomaly detection outputs indicated an anomaly in the data structure.
15 . The method of claim 10 , wherein the fourth anomaly detection output for the data structure in the dataset is determined by:
generating distributions for fourth parameters in datasets over a historical time period;
determining expected range of values in the distributions using the fourth parameters in the datasets;
comparing the fourth parameters in the data structure to the expected range of values; and
generating the fourth anomaly detection output based on the comparison.
16 . The method of claim 10 , wherein the dataset is a fund dataset received over a network from multiple data sources.
17 . The method of claim 10 , further comprising:
dividing the dataset into cohorts according to fifth parameters.
18 . The method of claim 10 , further comprising:
generating an alert that includes information associated with the anomalous data structure, wherein the alert includes one or more of the first anomaly detection output, the second anomaly detection output, the third anomaly detection output, and the fourth anomaly detection output.
19 . The method of claim 18 , further comprising:
transmitting the alert in an email that activates an icon on a computing display as an indication of the anomalous data structure.
20 . A non-transitory computer readable medium storing instruction thereon, that when executed by a processor cause the processor to perform operations for detecting an anomaly in at least one data structure of a computer network, the operations comprising:
receiving a dataset including data structures over the computer network, the data structures storing parameters wherein a ground truth or expected parameter values for the parameters are unknown;
dividing the data structures in the dataset into a plurality of cohorts using a subset of the parameters in the data structures, wherein at least one parameter in the subset of parameters is an Internet Protocol address;
selecting a plurality of anomaly detection models;
processing, using the plurality of anomaly detection models executing in parallel, data structures in each cohort in the plurality of cohorts, wherein each anomaly detection model generates anomaly detection scores for the data structures of each cohort in the plurality of cohorts;
aggregating an anomaly detection score in the anomaly detection scores for each data structure;
generating an alert for the at least one data structure based on the corresponding aggregated anomaly detection score; and
transmitting the alert, wherein the alert, upon receipt at a computing device, activates an icon on a display screen of the computing device indicating that the alert with the anomaly detection score has been received at the computing device.