Computer-based systems configured to select a monitored data segmentation and methods of use thereof
In some embodiments, the present disclosure provides an exemplary system and method that may include steps of identifying a device capable of processing a data stream; calculating a plurality of hash keys for a plurality of monitored segmentations associated with the device capable of the data stream; generating an increment data counter that corresponds to each hash key in a plurality of counting structures; calculating an anomaly score associated for the plurality of monitored segmentations; selecting a monitored segmentation based on the anomaly score; determining that a selected monitored segmentation meets a predetermined threshold associated with the anomaly score; and automatically marking the device capable of the data stream with a pre-generated label.
1 . A computer-implemented method comprising:
identifying, by at least one processor, at least one device capable of processing a multi-dimensional data stream;
dynamically calculating, by the at least one processor, a plurality of hash keys respectively for a plurality of monitored segmentations associated with the at least one device capable of processing the multi-dimensional data stream, the plurality of monitored segmentations generated based on at least one feature;
generating, by the at least one processor, at least one increment data counter that corresponds to each hash value of the plurality of hash keys in a plurality of counting structures,
wherein each of the plurality of counting structures comprises:
a current counting structure, counting the at least one incremental data counter during a current time period, and
a total counting structure, counting the at least one incremental data counter during a total time period;
dynamically calculating, by the at least one processor, an anomaly score for each of the plurality of monitored segmentations based on at least one value of corresponding one of the plurality of counting structures to generate a plurality of anomaly scores for the plurality of monitored segmentations;
automatically marking, by the at least one processor, the at least one device capable of processing the multi-dimensional data stream with a pre-generated label based on a determination that an anomaly score, from the plurality of anomaly scores and corresponding to a selected monitored segmentation of the plurality of monitored segmentations, meets or exceeds a predetermined threshold.
2 . The method of claim 1 , further comprising simultaneously detecting a plurality of anomalies associated with multiple segmentations of the plurality of monitored segmentations.
3 . The method of claim 2 , wherein the multi-dimensional data stream comprises a plurality of feature combinations associated with the plurality of monitored segmentations.
4 . The method of claim 1 , wherein the plurality of hash keys comprise a plurality of feature values associated with a count estimate for each feature of each hash key.
5 . The method of claim 4 , wherein the plurality of hash keys comprise a plurality of feature combinations.
6 . The method of claim 5 , wherein the plurality of feature combinations comprises a plurality of individual features and combinations of the individual features.
7 . The method of claim 1 , wherein the plurality of counting structures comprise a current counting structure for a current time period and a total counting structure for all time periods since a commencement of monitoring.
8 . The method of claim 1 , wherein the anomaly score comprises a result of a chi-squared goodness of fit statistics calculated for a current time period and any past time periods.
9 . The method of claim 1 , wherein the anomaly score comprises a mean anomaly score determined as a sum of anomaly scores for each monitored segmentation divided by a total number of anomalies within each monitored segmentation.
10 . The method of claim 1 , wherein the selected monitored segmentation is selected based on a comparison of representative anomaly scores for each monitored segmentation.
11 . The method of claim 1 , wherein the predetermined threshold comprises a predetermined threshold of risk associated with an occurrence of an anomaly within the monitored segmentation.
12 . The method of claim 1 , further comprising utilizing a machine learning module to predict a modification to a calculated anomaly score associated with the selected monitored segmentation based on receiving additional information.
13 . The method of claim 1 , further comprising utilizing a graphical user interface to display the pre-generated label associated with the marking of the device capable of the multi-dimensional data stream.
14 . The method of claim 1 , further comprising enabling real-time processing of anomaly detection for high-dimensionality data streams by employing parallel processing of the calculation of anomaly scores for each segmentation of the plurality of monitored segmentations.
15 . The method of claim 1 , further comprising dynamically calculating the anomaly score and automatically marking the device with the pre-generated labels associated with the calculated anomaly score based on a batch model for when a plurality of anomalous events are detected over a fixed period of time.
16 . The method of claim 1 , further comprising:
constructing a first segmentation based on each individual feature and a representative anomaly score;
selecting a particular feature for subsequent segmentation construction based on the representative anomaly score;
constructing a second segmentation based on a pairwise feature combination and a different anomaly score associated with the pairwise feature combination; and
performing an exhaustive combinatorial search for a hierarchical approach based on selected features of each segmentation.
17 . A computer-implemented method comprising:
identifying, by at least one processor, at least one device capable of processing a multi-dimensional data stream;
dynamically calculating, by the at least one processor, a plurality of hash keys respectively for a plurality of monitored segmentations associated with the at least one device capable of processing the multi-dimensional data stream, the plurality of monitored segmentations generated based on at least one feature;
generating, by the at least one processor, at least one increment data counter that corresponds to each hash key of the plurality of hash keys in a plurality of counting structures,
wherein each of the plurality of counting structures comprises:
a current counting structure, counting the at least one incremental data counter during a current time period, and
a total counting structure, counting the at least one incremental data counter during a total time period;
dynamically calculating, by the at least one processor, an anomaly score for each of the plurality of monitored segmentations based on at least one value of corresponding one of the plurality of counting structures to generate a plurality of anomaly scores for the plurality of monitored segmentations;
utilizing, by the at least one processor, a machine learning module to predict a modification to a calculated anomaly score associated with a selected monitored segmentation of the plurality of monitored segmentations based on receiving additional information; and
automatically marking, by the at least one processor, the at least one device capable of processing the multi-dimensional data stream with a pre-generated label based on a determination that the modified anomaly score, from the plurality of anomaly scores and corresponding to a selected monitored segmentation of the plurality of monitored segmentations, meets or exceeds a predetermined threshold.
18 . The method of claim 17 , further comprising:
simultaneously analyzing the plurality of monitored segmentations;
selecting a monitored segmentation based on a comparison of representative anomaly scores for each monitored segmentation of the plurality of monitored segmentations; and
dynamically determining an identity of an anomaly based on a presence of the identified anomaly within the plurality of monitored segmentations.
19 . The method of claim 17 , further comprising:
constructing a first segmentation based on each individual feature and a representative anomaly score;
selecting a particular feature for subsequent segmentation construction based on the representative anomaly score;
constructing a second segmentation based on a pairwise feature combination and a different anomaly score associated with the pairwise feature combination; and
performing an exhaustive combinatorial search for a hierarchical approach based on selected features of each segmentation.
20 . A system comprises:
a non-transient computer memory, storing software instructions;
at least one processor of a first computing device associated with a user;
wherein, when the processor executes the software instructions, the first computing device is programmed to:
identify at least one device capable of processing a multi-dimensional data stream;
dynamically calculate a plurality of hash keys respectively for a plurality of monitored segmentations associated with the at least one device capable of processing the multi-dimensional data stream, the plurality of monitored segmentations generated based on at least one feature;
generate at least one increment data counter that corresponds to each hash key of the plurality of hash keys in a plurality of counting structures,
wherein each of the plurality of counting structures comprises,
a current counting structure, counting the at least one incremental data counter during a current time period, and
a total counting structure, counting the at least one incremental data counter during a total time period;
dynamically calculate an anomaly score for each of the plurality of monitored segmentations based on at least one value of corresponding one of the plurality of counting structures to generate a plurality of anomaly scores for the plurality of monitored segmentations; and
automatically mark the at least one device capable of processing the multi-dimensional data stream with a pre-generated label based on a determination that an anomaly score, from the plurality of anomaly scores and corresponding to a selected monitored segmentation of the plurality of monitored segmentations, meets or exceeds a predetermined threshold.