Automatically rating the product's security during software development
According to an aspect, a method is provided that includes: receiving a first report from at least a first vulnerability evaluation tool; pre-processing the first report by at least tokenizing the first report and generating a first vector for a first text portion of the first report; providing, to a machine learning model, the first vector as an input; classifying, by the machine learning model, the first vector based on a plurality of vulnerability vectors generated from a database of vulnerability policies required for an evaluation of the application; and outputting, by the machine learning model, a first indication of a first match between the first vector and a first vulnerability vector of the plurality of vulnerability vectors, the first indication representing a presence in the application of a first vulnerability mapped to the first vulnerability vector of the plurality of vulnerability vectors generated from the database of vulnerability policies.
1. A system, comprising:
at least one data processor; and
at least one memory storing instructions which, when executed by the at least one data processor, result in operations comprising:
receiving a first report from at least a first vulnerability evaluation tool, the first report including text indicating at least one vulnerability of an application being evaluated;
pre-processing the first report by at least tokenizing the first report and generating a first vector for a first text portion of the first report;
providing, to a machine learning model, the first vector as an input;
classifying, by the machine learning model, the first vector based on a plurality of vulnerability vectors generated from a database of vulnerability policies required for an evaluation of the application;
outputting, by the machine learning model, a first indication of a first match between the first vector and a first vulnerability vector of the plurality of vulnerability vectors, the first indication representing a presence in the application of a first vulnerability mapped to the first vulnerability vector of the plurality of vulnerability vectors generated from the database of vulnerability policies; and
generating, for the application, a vulnerability score based on a quantity of indications classified by the machine learning model, wherein the vulnerability score is determined by reducing a pre-determined score by the quantity of the indications including the first indication and one or more other indications.
2. The system of claim 1 further comprising:
receiving a second report from at least a second vulnerability evaluation tool, the second report including text indicating at least a second vulnerability of the application being evaluated;
pre-processing the second report by at least tokenizing the second report and generating a second vector for a second text portion of the second report;
providing, to the machine learning model, the second vector as the input,
wherein the classifying, by the machine learning model, further comprises classifying the first vector and the second vector based on the plurality of vulnerability vectors generated from the database of vulnerability policies required for the evaluation of the application, and
wherein the outputting, by the machine learning model, further comprises outputting a second indication of a second match between the second vector and a second vulnerability vector of the plurality of vulnerability vectors, the second indication representing a presence in the application of the second vulnerability mapped to the second vulnerability vector of the plurality of vulnerability vectors generated from the database of vulnerability policies.
3. The system of claim 2 , wherein the first vulnerability evaluation tool is a cloud-based service, and the second vulnerability evaluation tool is on premise with the machine learning model.
4. The system of claim 1 further comprising:
generating a user interface including the vulnerability score to enable display.
5. The system of claim 1 , wherein the first vulnerability comprises an SQL injection vulnerability.
6. The system of claim 2 , wherein the second vulnerability comprises a no cross site scripting vulnerability or a no remote code injection vulnerability.
7. The system of claim 1 , wherein the first report includes the text indicating the at least one vulnerability of the application being evaluated, a version of the application being evaluated, a location where a portion of code having the first vulnerability was detected in the application, and a criticality indication of the first vulnerability.
8. The system of claim 1 , wherein the machine learning model comprises a neural network.
9. The system of claim 1 , wherein the classifying includes comparing the first vector to the plurality of vulnerability vectors, wherein the first vector matches the first vulnerability vector within a similarity threshold.
10. A method, comprising:
receiving a first report from at least a first vulnerability evaluation tool, the first report including text indicating at least one vulnerability of an application being evaluated;
pre-processing the first report by at least tokenizing the first report and generating a first vector for a first text portion of the first report;
providing, to a machine learning model, the first vector as an input;
classifying, by the machine learning model, the first vector based on a plurality of vulnerability vectors generated from a database of vulnerability policies required for an evaluation of the application;
outputting, by the machine learning model, a first indication of a first match between the first vector and a first vulnerability vector of the plurality of vulnerability vectors, the first indication representing a presence in the application of a first vulnerability mapped to the first vulnerability vector of the plurality of vulnerability vectors generated from the database of vulnerability policies;
generating, for the application, a vulnerability score based on a quantity of indications classified by the machine learning model, wherein the vulnerability score is determined by reducing a pre-determined score by the quantity of the indications including the first indication and one or more other indications.
11. The method of claim 10 further comprising:
receiving a second report from at least a second vulnerability evaluation tool, the second report including text indicating at least a second vulnerability of the application being evaluated;
pre-processing the second report by at least tokenizing the second report and generating a second vector for a second text portion of the second report;
providing, to the machine learning model, the second vector as the input,
wherein the classifying, by the machine learning model, further comprises classifying the first vector and the second vector based on the plurality of vulnerability vectors generated from the database of vulnerability policies required for the evaluation of the application, and
wherein the outputting, by the machine learning model, further comprises outputting a second indication of a second match between the second vector and a second vulnerability vector of the plurality of vulnerability vectors, the second indication representing a presence in the application of the second vulnerability mapped to the second vulnerability vector of the plurality of vulnerability vectors generated from the database of vulnerability policies.
12. The method of claim 11 , wherein the first vulnerability evaluation tool is a cloud-based service, and the second vulnerability evaluation tool is on premise with the machine learning model.
13. The method of claim 10 further comprising:
generating a user interface including the vulnerability score to enable display.
14. The method of claim 11 , wherein the first vulnerability comprises an SQL injection vulnerability, and wherein the second vulnerability comprises a no cross site scripting vulnerability or a no remote code injection vulnerability.
15. The method of claim 10 , wherein the first report includes the text indicating the at least one vulnerability of the application being evaluated, a version of the application being evaluated, a location where a portion of code having the first vulnerability was detected in the application, and a criticality indication of the first vulnerability.
16. A non-transitory computer readable storage medium including instructions which, when executed by at least one data processor, result in operations comprising:
receiving a first report from at least a first vulnerability evaluation tool, the first report including text indicating at least one vulnerability of an application being evaluated;
pre-processing the first report by at least tokenizing the first report and generating a first vector for a first text portion of the first report;
providing, to a machine learning model, the first vector as an input;
classifying, by the machine learning model, the first vector based on a plurality of vulnerability vectors generated from a database of vulnerability policies required for an evaluation of the application;
outputting, by the machine learning model, a first indication of a first match between the first vector and a first vulnerability vector of the plurality of vulnerability vectors, the first indication representing a presence in the application of a first vulnerability mapped to the first vulnerability vector of the plurality of vulnerability vectors generated from the database of vulnerability policies; and
generating, for the application, a vulnerability score based on a quantity of indications classified by the machine learning model, wherein the vulnerability score is determined by reducing a pre-determined score by the quantity of the indications including the first indication and one or more other indications.