System and method for automatically associating cybersecurity intelligence to cyberthreat actors
A computerized method for associating cyberthreat actor groups responsible for different cyberthreats is described. The method involves generating a similarity matrix based on content from received clusters of cybersecurity information. Each received cluster of cybersecurity information is assumed to be associated with a cyberthreat. The similarity matrix is composed via an optimized equation combining separate similarity metrics, where each similarity metric of the plurality of similarity metrics represents a level of correlation between at least two clusters of cybersecurity information, with respect to a particular aspect of operations described in the clusters. The method further involves that, in response to queries directed to the similarity matrix, generating a listing of a subset of the clusters of cybersecurity information having a greater likelihood of being associated with cyberthreats caused by the same cyberthreat actor group.
1 . A computerized method for associating cyberthreat actor groups responsible for different cyberthreats, comprising:
receiving, by a computing system comprising one or more computing devices, a plurality of clusters of cybersecurity information respectively associated with a plurality of cyberthreats, wherein the cybersecurity information comprises a profile associated with at least a first cluster of the plurality of clusters;
generating, by the computing system, a plurality of feature vectors respectively for the plurality of clusters of cybersecurity information;
processing, by the computing system, the plurality of feature vectors with a machine-learned model to generate one or more similarity metrics, wherein each of the one or more similarity metrics describes a similarity between the first cluster and at least one additional cluster of the plurality of clusters of cybersecurity information;
associating, by the computing system, at least the first cluster of the clusters of cybersecurity information with a particular cyberthreat actor group based at least in part on the one or more similarity metrics; and
updating, by the computing system, based on associating the first cluster with the particular cyberthreat actor group, the profile associated with the first cluster to include one or more characteristics associated with the first cluster.
2 . The method of claim 1 , wherein generating the respective feature vector for each of the clusters of cybersecurity information comprises:
extracting a respective set of forensically-related indicia from each of the clusters of cybersecurity information; and
generating the respective feature vector for each cluster of cybersecurity information from the set of forensically-related indicia associated with the cluster of cybersecurity information.
3 . The method of claim 2 , wherein at least one indicium in at least one of the sets of forensically-related indicia comprises a frequency of occurrence of a type of content within the cluster of cybersecurity information.
4 . The method of claim 3 , wherein the type of content comprises one or more of the following: (i) known aliases, (ii) malware names, (iii) methods of installation and/or operation for the malware, (iv) targeted industries, (v) targeted countries, or (vi) infrastructure.
5 . The method of claim 1 , wherein processing, by the computing system, the plurality of feature vectors with a machine-learned model to generate one or more similarity metrics comprises determining, by the computing system, a cosine similarity between two of the plurality of feature vectors.
6 . The method of claim 1 , wherein the machine-learned model applies a plurality of learned weight values to generate the one or more similarity metrics.
7 . The method of claim 1 , wherein associating, by the computing system, at least the first cluster of the clusters of cybersecurity information with the particular cyberthreat actor group based at least in part on the one or more similarity metrics comprises merging, by the computing system, the first cluster with a second cluster of the clusters of cybersecurity information based on the one or more similarity metrics between the first cluster and the second cluster, and wherein the second cluster has previously been associated with the particular cyberthreat actor group.
8 . The method of claim 1 , wherein the machine-learned model has been trained to generate similarity vectors that correlate cybersecurity data to a pre-categorized profile.
9 . A computing system for associating cyberthreat actor groups responsible for different cyberthreats, the computing system comprising:
one or more processors; and
one or more non-transitory computer-readable media that store:
a machine-learned model; and
computer-executable instructions for performing operations, the operations comprising:
receiving, by the computing system, a plurality of clusters of cybersecurity information respectively associated with a plurality of cyberthreats, wherein the cybersecurity information comprises a profile associated with at least a first cluster of the plurality of clusters;
generating, by the computing system, a plurality of feature vectors respectively for the plurality of clusters of cybersecurity information;
processing, by the computing system, the plurality of feature vectors with the machine-learned model to generate one or more similarity metrics, wherein each of the one or more similarity metrics describes a similarity between the first cluster and at least one additional cluster of the plurality of clusters of cybersecurity information;
associating, by the computing system, at least the first cluster of the clusters of cybersecurity information with a particular cyberthreat actor group based at least in part on the one or more similarity metrics; and
updating, by the computing system, based on associating the first cluster with the particular cyberthreat actor group, the profile associated with the first cluster to include one or more characteristics associated with the first cluster.
10 . The computing system of claim 9 , wherein generating the respective feature vector for each of the clusters of cybersecurity information comprises:
extracting a respective set of forensically-related indicia from each of the clusters of cybersecurity information; and
generating the respective feature vector for each cluster of cybersecurity information from the set of forensically-related indicia associated with the cluster of cybersecurity information.
11 . The computing system of claim 10 , wherein at least one indicium in at least one of the sets of forensically-related indicia comprises a frequency of occurrence of a type of content within the cluster of cybersecurity information.
12 . The computing system of claim 11 , wherein the type of content comprises one or more of the following: (i) known aliases, (ii) malware names, (iii) methods of installation and/or operation for the malware, (iv) targeted industries, (v) targeted countries, or (vi) infrastructure.
13 . The computing system of claim 9 , wherein processing, by the computing system, the plurality of feature vectors with a machine-learned model to generate one or more similarity metrics comprises determining, by the computing system, a cosine similarity between two of the plurality of feature vectors.
14 . The computing system of claim 9 , wherein the machine-learned model applies a plurality of learned weight values to generate the one or more similarity metrics.
15 . The computing system of claim 9 , wherein associating, by the computing system, at least the first cluster of the clusters of cybersecurity information with the particular cyberthreat actor group based at least in part on the one or more similarity metrics comprises merging, by the computing system, the first cluster with a second cluster of the clusters of cybersecurity information based on the one or more similarity metrics between the first cluster and the second cluster, and wherein the second cluster has previously been associated with the particular cyberthreat actor group.
16 . The computing system of claim 9 , wherein the machine-learned model has been trained to generate similarity vectors that correlate cybersecurity data to a pre-categorized profile.
17 . One or more non-transitory computer-readable media that store:
a machine-learned model; and
computer-executable instructions for performing operations, the operations comprising:
receiving, by a computing system, a plurality of clusters of cybersecurity information respectively associated with a plurality of cyberthreats, wherein the cybersecurity information comprises a profile associated with at least a first cluster of the plurality of clusters;
generating, by the computing system, a plurality of feature vectors respectively for the plurality of clusters of cybersecurity information;
processing, by the computing system, the plurality of feature vectors with the machine-learned model to generate one or more similarity metrics, wherein each of the one or more similarity metrics describes a similarity between the first cluster and at least one additional cluster of the plurality of clusters of cybersecurity information;
associating, by the computing system, at least the first cluster of the clusters of cybersecurity information with a particular cyberthreat actor group based at least in part on the one or more similarity metrics; and
updating, by the computing system, based on associating the first cluster with the particular cyberthreat actor group, the profile associated with the first cluster to include one or more characteristics associated with the first cluster.
18 . The one or more non-transitory computer-readable media of claim 17 , wherein generating the respective feature vector for each of the clusters of cybersecurity information comprises:
extracting a respective set of forensically-related indicia from each of the clusters of cybersecurity information; and
generating the respective feature vector for each cluster of cybersecurity information from the set of forensically-related indicia associated with the cluster of cybersecurity information.
19 . The one or more non-transitory computer-readable media of claim 18 , wherein at least one indicium in at least one of the sets of forensically-related indicia comprises a frequency of occurrence of a type of content within the cluster of cybersecurity information.
20 . The one or more non-transitory computer-readable media of claim 17 , wherein processing, by the computing system, the plurality of feature vectors with a machine-learned model to generate one or more similarity metrics comprises determining, by the computing system, a cosine similarity between two of the plurality of feature vectors.