Adversial deep neural network fuzzing
View Patent ↗A method for detecting security vulnerabilities, comprising: generating a corpus of input samples each labeled to indicate a threat level when executed by an input processing code; training a neural network (NN) using the plurality of input samples to classify inputs according to a plurality of labels of the plurality of input samples; for each input sample: iteratively altering the input sample to correspond to a process of gradient change of the NN, until the NN classifies the altered input sample to a different label than a respective label of the input sample; assigning the different label to the altered input sample; using the plurality of relabeled altered input samples to further train the NN and augment the corpus of input samples.
1 . A method for detecting security vulnerabilities, comprising:
generating a corpus of a plurality of input samples each labeled to indicate a threat level when executed by an input processing code;
training a neural network (NN) using the plurality of input samples to classify inputs according to a plurality of labels of the plurality of input samples;
augmenting the corpus of input samples by, for each input sample:
iteratively alternating between altering the input sample to correspond to a process of gradient change of the NN, and altering the input sample by a fuzzing algorithm until the NN classifies the altered input sample to a classification that is different from a respective label of the input sample before altering;
executing, after the NN classifies the altered input sample to a classification that is different from a respective label of the input sample, the altered input by the input processing code to observe an effect, wherein the input processing code is distinct from the NN; and
assigning a new label to the altered input sample based on the observed effect of executing the altered input by the input processing code; and
using the plurality of labeled altered input samples to further train the NN.
2 . The method of claim 1 , wherein each of the plurality of labels is a benign label, an error label, or a malicious label.
3 . The method of claim 2 , wherein the threat level indicated by the benign label corresponds to a benign effect produced by processing a respective input sample by the input processing code,
wherein the threat level indicated by the error label corresponds to an error effect produced by processing a respective input sample by the input processing code;
wherein the threat level indicated by the malicious label corresponds to a malicious effect produced by processing a respective input sample by the input processing code.
4 . The method of claim 1 , wherein the fuzzing algorithm is selected from the group consisting of:
generation based fuzzing of the input samples;
genetic fuzzing of the input samples; and
concolic testing of the input samples.
5 . The method of claim 4 , further comprising:
extracting a plurality of features from internal layers of the NN;
creating a dictionary based on the plurality of features; and
using the dictionary for embedding dictionary entries within the corpus of input samples.
6 . The method of claim 4 , further comprising training the NN using the plurality of input samples and respective feature extraction data obtained during processing the plurality of input samples by the fuzzing algorithm.
7 . The method of claim 1 , further comprising discarding a plurality of labeled altered input samples which do not trigger unique execution paths when processed by the input processing code.
8 . The method of claim 1 , wherein an amount of iterations applied for the gradient change of the NN on altered input samples of a plurality of altered input samples is equal to an amount of iterations needed until a respective internal layer of the NN reaches a state of maximal activation.
9 . The method of claim 1 , further comprising, during iteratively altering each input sample:
performing a plurality of intermediate alterations of the input sample; and
incorporating the plurality of intermediate alterations into the corpus of input samples.
10 . The method of claim 1 wherein an output of an internal layer of the NN is used to produce at least one of a token or a string related to security vulnerabilities of the input samples.
11 . The method of claim 1 wherein the altering by gradient change produces an alteration value to be added to the input sample.
12 . The method of claim 1 wherein the gradient change is constrained to follow paths which transform input samples classified to a given label to be classified to a respective label according to a mapping between labels.
13 . A system for detecting security vulnerabilities, comprising:
at least one processor adapted to execute a code for:
generating a corpus of a plurality of input samples each labeled to indicate a threat level when executed by an input processing code;
training a neural network (NN) using the plurality of input samples to classify inputs according to a plurality of labels of the plurality of input samples;
augmenting the corpus of input samples by, for each input sample:
iteratively alternating between altering the input sample to correspond to a process of gradient change of the NN, and altering the input sample by a fuzzing algorithm until the NN classifies the altered input sample to a classification that is different from a respective label of the input sample before altering;
executing, after the NN classifies the altered input sample to a classification that is different from a respective label of the input sample, the altered input by the input processing code to observe an effect, wherein the input processing code is distinct from the NN; and
assigning a new label to the altered input sample based on the observed effect of executing the altered input by the input processing code; and
using the plurality of labeled altered input samples to further train the NN.