System and method for electronic message data analysis and neutralization via machine learning
Systems, computer program products, and methods are described herein for electronic message data analysis and neutralization via machine learning. The present disclosure includes receiving message data at a gateway server, analyzing, using a coarse message data filter, the message data, receiving, upon a condition where the message data does not meet the at least one coarse message data filtering condition, the message data in a categorization engine, scoring, using the categorization engine, the message data based on contents of the message data, categorizing the message data in a category of malicious or non-malicious, applying metadata to the message data based on the category, transmitting the message data to an endpoint device having an email exchange application with a malicious email filtering plug-in, and displaying the message data in a predetermined folder of the email exchange application based on the category.
1 . A system for electronic message data analysis and neutralization via machine learning, the system comprising:
a processing device; and
a non-transitory storage device containing instructions, when executed by the processing device, the instructions cause the processing device to perform the steps of:
receive message data at a gateway server;
analyze, using a coarse message data filter, the message data, and remove the message data upon a condition where the message data meets at least one coarse message data filter filtering condition;
receive, upon a condition where the message data does not meet the at least one coarse message data filtering condition, the message data in a categorization engine, the categorization engine comprising a machine learning model;
score, using the categorization engine, the message data based on contents of the message data;
categorize, using the machine learning model, the message data in a category of malicious or non-malicious, wherein the machine learning model is configured to further categorize the message data categorized as malicious as being in a first category or a second category,
wherein upon a first condition where the message data is in the first category, the categorization engine removes one or more hyperlinks in the message data,
wherein upon a second condition where the message data is in the second category, the categorization engine alters one or more hyperlinks in the message data such that a hyperlink of the one or more hyperlinks, upon interaction therewith, presents a pop-up message comprising an indicator as to a reasoning that the message data was categorized as malicious, wherein the reasoning is an output from the machine learning model;
apply metadata, at the gateway server, to the message data based on the category;
transmit, upon receiving a transmission request, the message data to an endpoint device comprising an email exchange application comprising a malicious email filtering plug-in; and
display the message data in a predetermined folder of the email exchange application based on the category.
2 . The system of claim 1 , wherein the instructions further cause the processing device to perform the steps of:
receive, at the malicious email filtering plug-in, a signal comprising a misclassification signal;
append the message data with the misclassification signal; and
ingest, upon receiving the misclassification signal, the message data at the machine learning model as new training data.
3 . The system of claim 1 , wherein further categorizing the message data categorized as malicious as being in the first category or the second category comprises comparing a similarity distance to a predetermined threshold for the first category and the second category.
4 . The system of claim 1 , wherein, upon clicking of the altered hyperlink, the instructions further cause the processing device to perform the steps of:
display a warning message via the malicious email filtering plug-in.
5 . A computer program product for electronic message data analysis and neutralization via machine learning, the computer program product comprising a non-transitory computer-readable medium comprising code causing an apparatus to:
receive message data at a gateway server;
analyze, using a coarse message data filter, the message data, and remove the message data upon a condition where the message data meets at least one coarse message data filter filtering condition;
receive, upon a condition where the message data does not meet the at least one coarse message data filtering condition, the message data in a categorization engine, the categorization engine comprising a machine learning model;
score, using the categorization engine, the message data based on contents of the message data;
categorize, using the machine learning model, the message data in a category of malicious or non-malicious, wherein the machine learning model is configured to further categorize the message data categorized as malicious as being in a first category or a second category,
wherein upon a first condition where the message data is in the first category, the categorization engine removes one or more hyperlinks in the message data,
wherein upon a second condition where the message data is in the second category, the categorization engine alters one or more hyperlinks in the message data such that a hyperlink of the one or more hyperlinks, upon interaction therewith, presents a pop-up message comprising an indicator as to a reasoning that the message data was categorized as malicious, wherein the reasoning is an output from the machine learning model;
apply metadata, at the gateway server, to the message data based on the category;
transmit, upon receiving a transmission request, the message data to an endpoint device comprising an email exchange application comprising a malicious email filtering plug-in; and
display the message data in a predetermined folder of the email exchange application based on the category.
6 . The computer program product of claim 5 , wherein the code further causes the apparatus to:
receive, at the malicious email filtering plug-in, a signal comprising a misclassification signal;
append the message data with the misclassification signal; and
ingest, upon receiving the misclassification signal, the message data at the machine learning model as new training data.
7 . The computer program product of claim 5 , wherein further categorizing the message data categorized as malicious as being in the first category or the second category comprises comparing a similarity distance to a predetermined threshold for the first category and the second category.
8 . The computer program product of claim 5 , wherein, upon clicking of the altered hyperlink, the code further causes the apparatus to:
display a warning message via the malicious email filtering plug-in.
9 . A method for electronic message data analysis and neutralization via machine learning, the method comprising:
receiving message data at a gateway server;
analyzing, using a coarse message data filter, the message data, and remove the message data upon a condition where the message data meets at least one coarse message data filter filtering condition;
receiving, upon a condition where the message data does not meet the at least one coarse message data filtering condition, the message data in a categorization engine, the categorization engine comprising a machine learning model;
scoring, using the categorization engine, the message data based on contents of the message data;
categorizing, using the machine learning model, the message data in a category of malicious or non-malicious, wherein the machine learning model is configured to further categorize the message data categorized as malicious as being in a first category or a second category,
wherein upon a first condition where the message data is in the first category, the categorization engine removes one or more hyperlinks in the message data,
wherein upon a second condition where the message data is in the second category, the categorization engine alters one or more hyperlinks in the message data such that a hyperlink of the one or more hyperlinks, upon interaction therewith, presents a pop-up message comprising an indicator as to a reasoning that the message data was categorized as malicious, wherein the reasoning is an output from the machine learning model;
applying metadata, at the gateway server, to the message data based on the category;
transmitting, upon receiving a transmission request, the message data to an endpoint device comprising an email exchange application comprising a malicious email filtering plug-in; and
displaying the message data in a predetermined folder of the email exchange application based on the category.
10 . The method of claim 9 , the method further comprising:
receiving, at the malicious email filtering plug-in, a signal comprising a misclassification signal;
appending the message data with the misclassification signal; and
ingesting, upon receiving the misclassification signal, the message data at the machine learning model as new training data.
11 . The method of claim 9 , wherein further categorizing the message data categorized as malicious as being in the first category or the second category comprises comparing a similarity distance to a predetermined threshold for the first category and the second category.
12 . The method of claim 9 , wherein, upon clicking of the altered hyperlink, the method further comprises:
displaying a warning message via the malicious email filtering plug-in.