IP Library › Granted Patent US 12,579,128
Granted Patent B1
US 12,579,128 · App. 18/422,670 · Granted Mar 17, 2026

High-speed anomaly detection for large datasets

Inventor: Gleb Esman (San Francisco, CA)
Assignee: Cisco Technology, Inc.
G06F16/2365G06F16/22G06F16/288
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,128
App. No.
18/422,670
Filed
Jan 25, 2024
Granted
Mar 17, 2026
Kind
B1
Art Unit
2164
USPC
707/690
Abstract

The various implementations provide for techniques to ingest raw data into a set of events and sample a subset of the events for initial processing via high-speed anomaly detection. The high-speed anomaly detection techniques use a first set of sensitivities that enable fast detection of potential anomalies within the subset of events. The subset of events is also included in a full set of events that are stored in an index and processed via full-scale anomaly detection. The full-scale anomaly detection uses a second set of sensitivities that enable the system to detect anomalies from the full set of events accurately. Upon detecting an anomaly using the high-speed or full-scale detection techniques, the system performs actions to notify entities or automatically mitigate the effects of the anomaly such as mitigation actions to quell potentially fraudulent activity.

Claims (70)

1 . A computer-implemented method, comprising:

generating, by an edge processor, a set of events based on data received from a data source;

sampling, from the set of events, a subset of events, wherein the subset of events is selected by the edge processor based on a sample flag value associated with each event of the set of events;

identifying, via a set of high speed anomaly detection operations processing the subset of events using a first set of sensitivity thresholds, a first set of anomalies associated with the subset of events wherein the first set of sensitivity thresholds corresponds to a first quantity of events that match predefined search criteria;

identifying, via a set of full scale anomaly detection operations processing the set of events using a second set of sensitivity thresholds that correspond to thresholds of a greater quantity than the first set of sensitivity thresholds, a second set of anomalies associated with the set of events, wherein the first set of anomalies comprises a subset of the second set of anomalies, and wherein the first set of anomalies are identified prior to the second set of anomalies, wherein at least one of the first set of anomalies or the second set of anomalies includes identifying an event or series of events that deviate from a standard, user accounts with anomalous metadata, anomalous combinations of transactions, or a high rate of errors from a particular internet protocol (IP) address;

forwarding, by the edge processor, a reduced dataset relative to the data received from the data source to a data intake and query system, wherein the reduced dataset includes the first set of anomalies and the second set of anomalies;

generating a user interface that includes an indication of a likelihood of a presence of either the first set of anomalies or the second set of anomalies, wherein the user interface is configured for display on a computing device, and

blocking an identified internet protocol (IP) address associated with either the first set of anomalies or the second set of anomalies.

2 . The computer-implemented method of claim 1 , further comprising: for each event in the set of events, adding a sampling field value,

wherein sampling the subset of events comprises identifying each event in the set of events having a sampling field value that meets sampling criteria.

3 . The computer-implemented method of claim 2 , wherein:

sampling the subset of event comprises executing an object query specifying the sampling criteria, and

the object query is associated with a data model representing a view of the set of events.

4 . The computer-implemented method of claim 1 , wherein: an edge processor receives the data from the data source, the edge processor executes at least a portion of the first set of anomaly detection operations, and the edge processor executes at least a portion of the second set of anomaly detections operations.

5 . The computer-implemented method of claim 1 , further comprising:

indexing the subset of events into a first index of a data store, wherein the first set of anomaly detection operations is performed on the subset of events in the first index; and

indexing the set of events in a second index of the data store, wherein the second set of anomaly detection operations is performed on the subset of events in the second index.

6 . The computer-implemented method of claim 1 , further comprising:

upon identifying the first set of anomalies or the second set of anomalies, generating a notification report.

7 . The computer-implemented method of claim 6 , wherein the notification report includes a label for each of the first set of anomalies.

8 . The computer-implemented method of claim 1 , wherein:

the first set of anomaly detection operations processing the subset of events comprises a trained machine learning (ML) model receiving the subset of events as inputs,

the trained ML model generates a sensitivity threshold from a training set of data, and the trained ML model analyzes the subset of events for an anomaly compared to the sensitivity threshold.

9 . The computer-implemented method of claim 1 , wherein identifying the first set of anomalies comprises:

detecting a number of events in the subset of events that meets a first search criterion; and identifying a first anomaly included in the first set of anomalies upon determining that the number of events exceeds a sensitivity threshold associated with the first anomaly.

10 . The computer-implemented method of claim 1 , wherein:

each event in the set of events includes a portion of unstructured raw machine data reflecting activity in an information technology environment, and

the first set of anomalies or the second set of anomalies indicate a likelihood of fraudulent activity associated with the information technology environment.

11 . The computer-implemented method of claim 1 , wherein:

the first set of anomaly detection operations uses a first set of sensitivity thresholds on the subset of events to identify the first set of anomalies;

the second set of anomaly detection operations uses a second set of sensitivity thresholds on the set of events to identify the second set of anomalies, and

the second set of sensitivity thresholds is greater than the first set of sensitivity thresholds.

12 . The computer-implemented method of claim 1 , wherein the subset of events comprises 10% or fewer of the set of events.

13 . A computing device, comprising: a processor; and

a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations including:

generating, by an edge processor, a set of events based on data received from a data source;

sampling, from the set of events, a subset of events, wherein the subset of events is selected by the edge processor based on a sample flag value associated with each event of the set of events;

identifying, via a set of high speed anomaly detection operations processing the subset of events using a first set of sensitivity thresholds, a first set of anomalies associated with the subset of events wherein the first set of sensitivity thresholds corresponds to a first quantity of events that match predefined search criteria;

identifying, via a set of full scale anomaly detection operations processing the set of events using a second set of sensitivity thresholds that correspond to thresholds of a greater quantity than the first set of sensitivity thresholds, a second set of anomalies associated with the set of events, wherein the first set of anomalies comprises a subset of the second set of anomalies, and wherein the first set of anomalies are identified prior to the second set of anomalies, wherein at least one of the first set of anomalies or the second set of anomalies includes identifying an event or series of events that deviate from a standard, user accounts with anomalous metadata, anomalous combinations of transactions, or a high rate of errors from a particular internet protocol (IP) address;

forwarding, by the edge processor, a reduced dataset relative to the data received from the data source to a data intake and query system, wherein the reduced dataset includes the first set of anomalies and the second set of anomalies;

generating a user interface that includes an indication of a likelihood of a presence of either the first set of anomalies or the second set of anomalies, wherein the user interface is configured for display on a computing device, and

blocking an identified internet protocol (IP) address associated with either the first set of anomalies or the second set of anomalies.

14 . The computer device of claim 13 , wherein:

the operations further include, for each event in the set of events, adding a sampling field value,

sampling the subset of events comprises executing an object query specifying sampling criteria, where each event in the set of events has a sampling field value that meets the sampling criteria, and

the object query is associated with a data model representing a view of the set of events.

15 . The computer device of claim 13 , wherein:

an edge processor receives the data from the data source,

the edge processor executes at least a portion of the first set of anomaly detection operations, and

the edge processor executes at least a portion of the second set of anomaly detections operations.

16 . The computer device of claim 13 , wherein further the operations further include:

indexing the subset of events into a first index of a data store, wherein the first set of anomaly detection operations is performed on the subset of events in the first index; and

indexing the set of events in a second index of the data store, wherein the second set of anomaly detection operations is performed on the subset of events in the second index.

17 . The computer device of claim 13 , wherein the first set of anomalies and the second set of anomalies each includes one or more anomalies.

18 . One or more non-transitory computer-readable media having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform operations including:

generating, by an edge processor, a set of events based on data received from a data source;

sampling, from the set of events, a subset of events, wherein the subset of events is selected by the edge processor based on a sample flag value associated with each event of the set of events;

identifying, via a set of high speed anomaly detection operations processing the subset of events using a first set of sensitivity thresholds, a first set of anomalies associated with the subset of events wherein the first set of sensitivity thresholds corresponds to a first quantity of events that match predefined search criteria;

identifying, via a set of full scale anomaly detection operations processing the set of events using a second set of sensitivity thresholds that correspond to thresholds of a greater quantity than the first set of sensitivity thresholds, a second set of anomalies associated with the set of events, wherein the first set of anomalies comprises a subset of the second set of anomalies, and wherein the first set of anomalies are identified prior to the second set of anomalies, wherein at least one of the first set of anomalies or the second set of anomalies includes identifying an event or series of events that deviate from a standard, user accounts with anomalous metadata, anomalous combinations of transactions, or a high rate of errors from a particular internet protocol (IP) address;

forwarding, by the edge processor, a reduced dataset relative to the data received from the data source to a data intake and query system, wherein the reduced dataset includes the first set of anomalies and the second set of anomalies;

generating a user interface that includes an indication of a likelihood of a presence of either the first set of anomalies or the second set of anomalies, wherein the user interface is configured for display on a computing device, and

blocking an identified internet protocol (IP) address associated with either the first set of anomalies or the second set of anomalies.

19 . The one or more non-transitory computer-readable media of claim 18 , further comprising instructions that, when executed by the one or more processors, cause the one or more processors to further perform the operations including, for each event in the set of events, adding a sampling field value, wherein:

sampling the subset of events comprises executing an object query specifying sampling criteria, where each event in the set of events has a sampling field value that meets the sampling criteria, and

the object query is associated with a data model representing a view of the set of events.

20 . The one or more non-transitory computer-readable media of claim 18 , wherein: an edge processor receives the data from the data source,

the edge processor executes at least a portion of the first set of anomaly detection operations, and

the edge processor executes at least a portion of the second set of anomaly detections operations.

21 . The computer-implemented method of claim 1 , wherein the first set of anomalies and the second set of anomalies each includes one or more anomalies.

22 . The one or more non-transitory computer-readable media of claim 18 , wherein the first set of anomalies and the second set of anomalies each includes one or more anomalies.

Assignments (3)
CHANGE OF NAME Recorded Jul 22, 2025
From: SPLUNK INC.
To: SPLUNK LLC
Reel/Frame 072170/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2025
From: SPLUNK LLC
To: CISCO TECHNOLOGY, INC.
Reel/Frame 072173/0058 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2024
From: ESMAN, GLEB
To: SPLUNK INC.
Reel/Frame 066259/0343 →
References Cited (22)
US 7937344B2 · Baum et al. · 2011 [cited by applicant]
US 8112425B2 · Baum et al. · 2012 [cited by applicant]
US 8751529B2 · Zhang et al. · 2014 [cited by applicant]
US 8788525B2 · Neels et al. · 2014 [cited by applicant]
US 8874526B2 · Hsieh · 2014 [cited by examiner]
US 9215240B2 · Merza et al. · 2015 [cited by applicant]
US 9286413B1 · Coates et al. · 2016 [cited by applicant]
US 9516053B1 · Muddu · 2016 [cited by examiner]
US 10127258B2 · Lamas et al. · 2018 [cited by applicant]
US 20170220632A1 · Miller · 2017 [cited by examiner]
US 20180219888A1 · Apostolopoulos · 2018 [cited by examiner]
US 20190098106A1 · Mungel et al. · 2019 [cited by applicant]
US 20190235941A1 · Bath · 2019 [cited by examiner]
US 20200177611A1 · Bharrat · 2020 [cited by examiner]
US 20210286874A1 · Jin · 2021 [cited by examiner]
US 20240022583A1 · Miserendino · 2024 [cited by examiner]
US 20250004868A1 · Altman · 2025 [cited by examiner]
Splunk Enterprise 8.0.0 Overview, available online, retrieved May 20, 2020 from docs.splunk.com., 17 pages. [cited by applicant]
Splunk Cloud 8.0.2004 User Manual, available online, retrieved May 20, 2020 from docs.splunk.com., 66 pages. [cited by applicant]
Splunk Quick Reference Guide, updated 2019, available online at https://www.splunk.com/pdfs/solution-guides/splunk-quick-reference-guide.pdf. retrieved May 20, 2020, 6 pages. [cited by applicant]
Carraso, David, “Exploring Splunk,” published by CITO Research, New York, NY., Apr. 2012, 156 pages. [cited by applicant]
Bitincka, Ledion et al., “Optimizing Data Analysis with a Semi-structured Time Series Database”. self-published, first presented at “Workshop on Managing Systems via Log Analysis and Machine Learning Techniques (SLAML)”… [cited by applicant]