System for extracting, classifying, and enriching cyber criminal communication data
An apparatus, including systems and methods, for classifying, mapping, and predicting cybercriminal activity is disclosed herein. For example, in some embodiments, an apparatus is configured to: receive cybercriminal communication (CCC) data of postings from a source forum; identify, classify, and rank a threat topic for each posting; identify a first subset of postings that includes postings assigned the threat topic classification with the greatest threat topic rank; for each posting of the first subset of postings: identify and rank the threat actor; identify a second subset of postings that includes postings associated with the threat actor assigned the greatest threat actor rank; and send, to a cybersecurity data exchange module, the CCC data of the second subset of postings and associated enriched data including the source forum, the threat topic classifications, the threat actor, the threat actor rank, or the other threat actors that mentioned the threat actor.
1. At least one non-transitory computer-readable medium comprising one or more instructions that when executed by a processor, cause the processor to:
receive CCC data of a source forum, wherein the CCC data includes a plurality of postings made on the source forum;
store the CCC data;
extract artifacts from the CCC data, wherein the extracted artifacts indicate the source forum, a threat topic, or a threat actor of a posting;
store the extracted artifacts according to a pre-defined taxonomy;
for each posting of the plurality of postings:
identify the threat topic;
assign a threat topic classification; and
assign a threat topic rank based on the threat topic classification;
identify a first subset of postings, wherein the first subset of postings includes postings assigned the threat topic classification with the greatest threat topic rank;
for each posting of the first subset of postings:
identify the threat actor; and
assign a threat actor rank based at least in part on an official source forum rank (OFR), a posting activity score, or a number of times the threat actor is mentioned by other threat actors;
identify a second subset of postings from the first subset of postings, wherein the second subset of postings includes postings made by and associated with the threat actor assigned the greatest threat actor rank; and
send, to a cybersecurity data exchange module, the CCC data of the second subset of postings and associated enriched data, wherein the associated enriched data includes one or more of the source forum, the threat topic classification, the threat topic rank, the threat actor, the threat actor rank, and the other threat actors that mentioned the threat actor.
2. The at least one non-transitory computer-readable medium of claim 1 , wherein assign the threat topic classification includes analyzing each posting using a keyword topic list.
3. The at least one non-transitory computer-readable medium of claim 2 , further comprising one or more instructions that when executed by the processor, cause the processor to:
calculate a threat topic score based on the keyword topic list analysis; and
assign the threat topic classification based on the threat topic score.
4. The at least one non-transitory computer-readable medium of claim 1 , wherein the source forum is one of a plurality of source forums and the CCC data is received from the plurality of source forums, and further comprising one or more instructions that when executed by the processor, cause the processor to:
for each posting of the plurality of postings from each source forum:
identify the source forum;
for each source forum of the plurality of source forums:
determine a source forum rank for the source forum based on the threat topic rank assigned to each posting of the plurality of postings from the source forum;
identify the source forum having the greatest source forum rank; and
wherein the first subset of postings is identified from the source forum having the greatest forum rank.
5. The at least one non-transitory computer-readable medium of claim 1 , wherein assign the threat topic classification includes analyzing each posting using a Natural Language Processing (NLP) algorithm.
6. The at least one non-transitory computer-readable medium of claim 1 , wherein the posting activity score is determined using a weighted average based on a number of postings and a date of postings.
7. The at least one non-transitory computer-readable medium of claim 1 , wherein the threat actor rank is equal to the OFR multiplied by the posting activity score added to the number of times the threat actor is mentioned by other threat actors.
8. The at least one non-transitory computer-readable medium of claim 1 , further comprising one or more instructions that when executed by the processor, cause the processor to:
receive a request from a requestor to query the CCC data;
query the CCC data for the request; and
send the query results to the requestor, wherein the query results include a portion of the CCC data and the associated enriched data, wherein the associated enriched data includes one or more of the source forum, the threat topic classification, the threat topic rank, the threat actor, the threat actor rank, and the other threat actors that mentioned the threat actor.
9. An apparatus, comprising:
memory operable to store instructions; and
one or more processors operable to execute the instructions, such that the apparatus is configured to:
receive CCC data of a plurality of source forums, wherein the CCC data of each source forum of the plurality of source forums includes a plurality of postings made on the respective source forum;
store the CCC data;
extract artifacts from the CCC data, wherein the extracted artifacts indicate the source forum, a threat topic, or a threat actor of a posting;
store the extracted artifacts according to a pre-defined taxonomy;
for each posting of the plurality of postings of each source forum:
identify the source forum;
identify the threat topic;
assign a threat topic classification; and
assign a threat topic rank based on the threat topic classification;
for each source forum of the plurality of source forums:
determine a source forum rank for the source forum based on the threat topic rank assigned to each posting of the plurality of postings from the source forum;
identify the source forum having the greatest source forum rank;
identify a first subset of postings from the source forum having the greatest forum rank, wherein the first subset of postings includes postings assigned the threat topic classification with the greatest threat topic rank;
for each posting of the first subset of postings:
identify the threat actor; and
assign a threat actor rank based at least in part on an official source forum rank (OFR), a posting activity score, or a number of times the threat actor is mentioned by other threat actors;
identify a second subset of postings from the first subset of postings, wherein the second subset of postings includes postings associated with the threat actor assigned the greatest threat actor rank; and
send, to a cybersecurity data exchange module, the CCC data of the second subset of postings and associated enriched data, wherein the associated enriched data includes one or more of the source forum, the threat topic classification, the threat topic rank, the threat actor, the threat actor rank, and the other threat actors that mentioned the threat actor.
10. The apparatus of claim 9 , wherein assign the threat topic classification includes analyzing each posting using a keyword topic list.
11. The apparatus of claim 10 , further configured to:
calculate a threat topic score based on the keyword topic list analysis; and
assign the threat topic classification based on the threat topic score.
12. The apparatus of claim 9 , wherein assign the threat topic classification includes analyzing each posting using a Natural Language Processing (NLP) algorithm.
13. The apparatus of claim 9 , wherein the posting activity score is determined using a weighted average based on a number of postings and a date of postings made by the threat actor on the source forum.
14. The apparatus of claim 9 , wherein the threat actor rank is equal to the OFR multiplied by the posting activity score added to the number of times the threat actor is mentioned by other threat actors on the source forum.
15. The apparatus of claim 9 , further configured to:
receive a request from a requestor to query the CCC data;
query the CCC data for the request; and
send the query results to the requestor, wherein the query results include a portion of the CCC data and the associated enriched data, wherein the associated enriched data includes one or more of the source forum, the threat topic classification, the threat topic rank, the threat actor, the threat actor rank, and the other threat actors that mentioned the threat actor.
16. A method, comprising:
receiving CCC data of a plurality of source forums, wherein the CCC data of each source forum of the plurality of source forums includes a plurality of postings made on the respective source forum;
storing the CCC data;
extracting artifacts from the CCC data, wherein the extracted artifacts indicate the source forum, a threat topic, or a threat actor of a posting;
storing the extracted artifacts according to a pre-defined taxonomy;
for each posting of the plurality of postings:
identifying the source forum;
assigning a first threat topic classification based on a first analysis; and
assigning a first threat topic rank based on the first threat topic classification;
for each source forum of the plurality of source forums:
determining a source forum rank for the source forum based on the first threat topic rank assigned to each posting of the plurality of postings from the source forum;
identifying the source forum having the greatest source forum rank;
for each posting of the plurality of postings of the source forum having the greatest source forum rank:
assigning a second threat topic classification based on a second analysis; and
assigning a second threat topic rank based on the second threat topic classification;
identifying a first subset of postings from the source forum having the greatest forum rank, wherein the first subset of postings includes postings assigned the second threat topic classification with the greatest second threat topic rank;
for each posting of the first subset of postings:
identifying the threat actor; and
assigning a threat actor rank based at least in part on an official source forum rank (OFR), a posting activity score, or a number of times the threat actor is mentioned by other threat actors on the source forum;
identifying a second subset of postings from the first subset of postings, wherein the second subset of postings includes postings made by the threat actor assigned the greatest threat actor rank; and
sending, to a cybersecurity data exchange module, the CCC data of the second subset of postings and associated enriched data, wherein the associated enriched data includes one or more of the source forum, the threat topic classification, the threat topic rank, the threat actor, the threat actor rank, and the other threat actors that mentioned the threat actor.
17. The method of claim 16 , wherein the first analysis includes analyzing each posting using a Natural Language Processing (NLP) algorithm and the second analysis includes analyzing each posting using a keyword topic list.
18. The method of claim 17 , further comprising:
calculating a threat topic score based on the keyword topic list analysis; and
assigning the threat topic classification based on the threat topic score.
19. The method of claim 16 , wherein the posting activity score is determined using a weighted average based on a number of postings and a date of postings made by the threat actor on the source forum.
20. The method of claim 16 , wherein the threat actor rank is equal to the threat actor's OFR multiplied by the threat actor's posting activity score added to the number of times the threat actor is mentioned by other threat actors on the source forum.