Source code vulnerability detection using deep learning
Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for detecting and locating vulnerabilities in source code. The method comprises receiving one or more source code files, matching source code from the one or more source code files to one or more program slices by parsing the source code and mapping one or more portions of the source code to the one or more program slices, wherein each of the one or more program slices comprises one or more program statements associated with one or more vulnerabilities, and generating, using a predictive machine learning model, a vulnerability prediction for each of the one or more source code files, the vulnerability prediction comprising one or more locations of vulnerable code in the source code based on the matching and a vulnerability class associated with each location of vulnerable code.
1 . A computer-implemented method comprising:
receiving, by one or more processors, a source code file;
matching, by the one or more processors, source code from the source code file to a plurality of program slices by parsing the source code and mapping a portion of the source code to a program slice of the plurality of program slices, wherein the program slice comprises a program statement associated with a vulnerability;
generating, by the one or more processors and using a predictive machine learning model, a vulnerability prediction for the source code file, the vulnerability prediction comprising: (a) a location of vulnerable code in the source code based on the matching, and (b) a vulnerability class, of a plurality of vulnerability classes, associated with the vulnerable code in the source code, wherein: (i) the vulnerability prediction indicates whether the source code file increases a susceptibility to a malicious attack, (ii) the predictive machine learning model comprises a multiclass classification machine learning model and is trained based on a training dataset, and (iii) the training dataset is generated by:
(1) receiving a plurality of training source code files and the plurality of vulnerability classes associated with the plurality of training source code files,
(2) receiving a plurality of syntax features corresponding to the plurality of vulnerability classes,
(3) determining a program slicing criterion based on the plurality of syntax features,
(4) extracting a set of program slices from the plurality of training source code files based on the program slicing criterion, wherein one or more program slices of the set of program slices is extracted by performing a forward slice or a backward slice, and
(5) labeling the set of program slices with the plurality of vulnerability classes; and
initiating, by the one or more processors, one or more prediction-based actions based on the vulnerability prediction.
2 . The computer-implemented method of claim 1 , wherein:
(i) determining the program slicing criterion further comprises determining a plurality of potential vulnerability candidates by performing static analysis on a plurality of program statements associated with the plurality of training source code files and matching the plurality of program statements associated with the plurality of training source code files with the plurality of syntax features,
(ii) the program slicing criterion comprises a set of variables corresponding to a plurality of values that are required to be preserved,
(iii) the forward slice comprises a program statement affected by the program slicing criterion, and
(iv) the backward slice comprises a program statement that affects the program slicing criterion.
3 . The computer-implemented method of claim 2 , wherein the static analysis comprises generating, for a training source code file of the plurality of training source code files, at least one of: a program dependency graph, a data dependency graph, and a control dependency graph.
4 . The computer-implemented method of claim 3 , wherein the program dependency graph comprises a first set of edges representative of data dependencies between one or more program statements in the training source code file and a second set of edges representative of one or more control dependencies between the one or more program statements in the training source code file.
5 . The computer-implemented method of claim 1 , wherein the plurality of syntax features comprises application programming interface (API) or library calls, array declarations, pointer declarations, or operators in an expression.
6 . The computer-implemented method of claim 1 , wherein extracting the plurality of program slices comprises generating a source code subset, the source code subset comprising the program statement from the plurality of training source code files contributing to the vulnerability.
7 . The computer-implemented method of claim 1 , wherein the training dataset comprises the plurality of program slices assigned with labels associated with the plurality of vulnerability classes.
8 . The computer-implemented method of claim 1 , wherein the location of vulnerable code in the source code further comprises one or more of an indication of a class or a function associated with the vulnerable code or an identifier associated with a source code file comprising the source code.
9 . The computer-implemented method of claim 1 , wherein initiating the one or more prediction-based actions further comprises:
performing one or more load balancing operations to set a number of allowed computing entities used by a post-prediction system based on the vulnerability prediction.
10 . A system comprising:
one or more processors; and
at least one memory storing processor-executable instructions that, when executed by any one or more of the one or more processors, causes the one or more processors to perform operations comprising:
receive a source code file;
match source code from the source code file to a plurality of program slices by parsing the source code and mapping a portion of the source code to a program slice of the plurality of program slices, wherein the program slice comprises a program statement associated with a vulnerability;
generate, using a predictive machine learning model, a vulnerability prediction for the source code file, the vulnerability prediction comprising: (a) a location of vulnerable code in the source code based on the matching, and (b) a vulnerability class, of a plurality of vulnerability classes, associated with the vulnerable code in the source code, wherein: (i) the vulnerability prediction indicates whether the source code file increases a susceptibility to a malicious attack, (ii) the predictive machine learning model comprises a multiclass classification machine learning model and is trained based on a training dataset, and (iii) the training dataset is generated by:
(1) receiving a plurality of training source code files and the plurality of vulnerability classes associated with the plurality of training source code files,
(2) receiving a plurality of syntax features corresponding to the plurality of vulnerability classes,
(3) determining a program slicing criterion based on the plurality of syntax features,
(4) extracting a set of program slices from the plurality of training source code files based on the program slicing criterion, wherein one or more program slices of the set of program slices is extracted by performing a forward slice or a backward slice, and
(5) labeling the set of program slices with the plurality of vulnerability classes; and
initiate one or more prediction-based actions based on the vulnerability prediction.
11 . The system of claim 10 , wherein
(i) determining the program slicing criterion further comprises determining a plurality of potential vulnerability candidates by performing static analysis on a plurality of program statements associated with the plurality of training source code files and matching the plurality of program statements associated with the plurality of training source code files with the plurality of syntax features,
(ii) the program slicing criterion comprises a set of variables corresponding to a plurality of values that are required to be preserved,
(iii) the forward slice comprises a program statement affected by the program slicing criterion, and
(iv) the backward slice comprises a program statement that affects the program slicing criterion.
12 . The system of claim 11 , wherein the static analysis comprises generating, for a training source code file of the plurality of training source code files, at least one of: a program dependency graph, a data dependency graph, and a control dependency graph.
13 . The system of claim 10 , wherein the plurality of syntax features comprises application programming interface (API) or library calls, array declarations, pointer declarations, or operators in an expression.
14 . The system of claim 10 , wherein extracting the plurality of program slices comprises generating a source code subset, the source code subset comprising the program statement from the plurality of training source code files contributing to the vulnerability.
15 . The system of claim 10 , wherein the training dataset is further generated by replacing names of functions and variables in the plurality of program slices with symbolic names.
16 . One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:
receive a source code file;
match source code from the source code file to a plurality of program slices by parsing the source code and mapping a portion of the source code to a program slice of the plurality of program slices, wherein the program slice comprises a program statement associated with a vulnerability;
generate, using a predictive machine learning model, a vulnerability prediction for the source code file, the vulnerability prediction comprising: (a) a location of vulnerable code in the source code based on the matching, and (b) a vulnerability class, of a plurality of vulnerability classes, associated with the vulnerable code in the source code, wherein: (i) the vulnerability prediction indicates whether the source code file increases a susceptibility to a malicious attack, (ii) the predictive machine learning model comprises a multiclass classification machine learning model and is trained based on a training dataset, and (iii) the training dataset is generated by:
(1) receiving a plurality of training source code files and the plurality of vulnerability classes associated with the plurality of training source code files,
(2) receiving a plurality of syntax features corresponding to the plurality of vulnerability classes,
(3) determining a program slicing criterion based on the plurality of syntax features,
(4) extracting a set of program slices from the plurality of training source code files based on the program slicing criterion, wherein one or more program slices of the set of program slices is extracted by performing a forward slice or a backward slice, and
(5) labeling the set of program slices with the plurality of vulnerability classes; and
initiate one or more prediction-based actions based on the vulnerability prediction.
17 . The one or more non-transitory computer-readable storage media of claim 16 , wherein:
(i) determining the program slicing criterion further comprises determining a plurality of potential vulnerability candidates by performing static analysis on a plurality of program statements associated with the plurality of training source code files and matching the plurality of program statements associated with the plurality of training source code files with the plurality of syntax features,
(ii) the program slicing criterion comprises a set of variables corresponding to a plurality of values that are required to be preserved,
(iii) the forward slice comprises a program statement affected by the program slicing criterion, and
(iv) the backward slice comprises a program statement that affects the program slicing criterion.
18 . The one or more non-transitory computer-readable storage media of claim 17 , wherein the static analysis comprises generating, for a training source code file of the plurality of training source code files, at least one of: a program dependency graph, a data dependency graph, and a control dependency graph.
19 . The one or more non-transitory computer-readable storage media of claim 16 , wherein extracting the plurality of program slices comprises generating a source code subset, the source code subset comprising the program statement from the plurality of training source code files contributing to the vulnerability.
20 . The one or more non-transitory computer-readable storage media of claim 16 , wherein the training dataset is further generated by replacing names of functions and variables in the plurality of program slices with symbolic names.