Methods and systems for identifying anomalous computer events to detect security incidents
A method includes receiving, from a plurality of sources, data associated with a plurality of events at the plurality of sources, standardizing the data based on a set of predefined standardization rules to define standardized data, and defining a vector representation for each event from the plurality of events based on the standardized data. The method includes assigning each event from the plurality of events to at least one cohort from a plurality of cohorts based on a similarity associated with the vector representation for that event and each cohort from the plurality of cohorts, generating, using at least one machine learning model, an ontology based on a set of cohorts from the plurality of cohorts and associated with the plurality of events, and storing the plurality of events as associated with the ontology such that the plurality of events can be filtered based on the ontology.
1 . A non-transitory processor-readable medium storing code representing instructions to be executed by one or more processors, the instructions comprising code to cause the one or more processors to:
receive, from a plurality of sources, data associated with a plurality of events at the plurality of sources;
standardize the data based on a set of predefined standardization rules to define standardized data;
define a vector representation for each event from the plurality of events based on the standardized data;
calculate a similarity associated with the vector representation for each event from the plurality of events and a plurality of cohorts using a coarse sorting process and a fine sorting process;
assign each event from the plurality of events to at least one cohort from the plurality of cohorts based on the similarity associated with the vector representation for that event and each cohort from the plurality of cohorts;
assign each cohort from the plurality of cohorts to an ontology;
generate a confidence score associated with each event from the plurality of events based on the ontology, the confidence score associated with assigning the plurality of cohorts to the ontology; and
identify an anomalous event from the plurality of events based on the confidence score associated with that event not meeting a criterion.
2 . The non-transitory processor-readable medium of claim 1 , wherein the coarse sorting process includes using a locality sensitive hashing (LSH) function.
3 . The non-transitory processor-readable medium of claim 1 , wherein the code to cause the one or more processors to calculate includes code to cause the one or more processors to calculate the similarity based on at least one of a cosine similarity, a hamming distance, a nearest neighbor search, a dot product similarity, or a Euclidean distance.
4 . The non-transitory processor-readable medium of claim 1 , wherein the ontology includes a plurality of categories.
5 . The non-transitory processor-readable medium of claim 1 , wherein the code to cause the one or more processors to define the vector representation includes code to cause the one or more processors to define the vector representation for each event from the plurality of events using a hybrid vector space based on both a dense vector search and a sparse vector search.
6 . The non-transitory processor-readable medium of claim 1 , wherein the ontology is a first ontology and the confidence score is a first score, the instructions further comprising code to cause the one or more processors to:
generate a second score for an event based on the first ontology;
based on the second score for the event being below a threshold for the first ontology, generate, using a machine learning model and based on the event, a second ontology; and
assign the event to the second ontology.
7 . A method, comprising:
receiving, from a plurality of sources, data associated with a plurality of events at the plurality of sources;
standardizing the data based on a set of predefined standardization rules to define standardized data;
defining a vector representation for each event from the plurality of events based on the standardized data;
calculating a similarity associated with the vector representation for each event from the plurality of events and a plurality of cohorts using a coarse sorting process and a fine sorting process;
assigning each event from the plurality of events to at least one cohort from the plurality of cohorts based on the similarity associated with the vector representation for that event and each cohort from the plurality of cohorts;
assigning each cohort from the plurality of cohorts to an ontology;
generating a confidence score associated with each event from the plurality of events based on the ontology, the confidence score associated with assigning the plurality of cohorts to the ontology; and
identifying an anomalous event from the plurality of events based on the confidence score associated with that event not meeting a criterion.
8 . The method of claim 7 , wherein the coarse sorting process includes using a locality sensitive hashing (LSH) function.
9 . The method of claim 7 , wherein the calculating includes calculating the similarity based on at least one of a cosine similarity, a hamming distance, a nearest neighbor search, a dot product similarity, or a Euclidean distance.
10 . The method of claim 7 , wherein the ontology includes a plurality of categories.
11 . The method of claim 7 , wherein the defining includes defining the vector representation for each event from the plurality of events using a hybrid vector space based on both a dense vector search and a sparse vector search.
12 . The method of claim 7 , wherein the ontology is a first ontology, the method further comprising:
generating a second score for an event based on the first ontology;
based on the second score for the event being below a threshold for the first ontology, generating, using a machine learning model and based on the event, a second ontology; and
assigning the event to the second ontology.
13 . An apparatus, comprising:
a processor, and
a non-transitory, processor-readable medium storing instructions that, when executed by the processor, cause the processor to:
receive, from a plurality of sources, data associated with a plurality of events at the plurality of sources,
standardize the data based on a set of predefined standardization rules to define standardized data,
define a vector representation for each event from the plurality of events based on the standardized data,
calculate a similarity associated with the vector representation for each event from the plurality of events and a plurality of cohorts using a coarse sorting process and a fine sorting process,
assign each event from the plurality of events to at least one cohort from the plurality of cohorts based on the similarity associated with the vector representation for that event and each cohort from the plurality of cohorts,
assign each cohort from the plurality of cohorts to an ontology,
generate a confidence score associated with each event from the plurality of events based on the ontology, the confidence score associated with assigning the plurality of cohorts to the ontology, and
identify an anomalous event from the plurality of events based on the confidence score associated with that event not meeting a criterion.
14 . The apparatus of claim 13 , wherein the coarse sorting process includes using a locality sensitive hashing (LSH) function.
15 . The apparatus of claim 13 , wherein the instructions to cause the processor to calculate include instructions to cause the processor to calculate the similarity based on at least one of a cosine similarity, a hamming distance, a nearest neighbor search, a dot product similarity, or a Euclidean distance.
16 . The apparatus of claim 13 , wherein the ontology includes a plurality of categories.
17 . The apparatus of claim 13 , wherein the instructions to cause the processor to define the vector representation include instructions to cause the processor to define the vector representation for each event from the plurality of events using a hybrid vector space based on both a dense vector search and a sparse vector search.