IP Library › Granted Patent US 11,514,297
Granted Patent B2
US 11,514,297 · App. 16/885,177 · Granted Nov 29, 2022

Post-training detection and identification of human-imperceptible backdoor-poisoning attacks

Inventors: David Jonathan Miller (State College, PA); George Kesidis (State College, PA)
Assignee: Anomalee Inc.
G06N3/0454G06N3/0481G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,514,297
App. No.
16/885,177
Filed
May 27, 2020
Granted
Nov 29, 2022
Kind
B2
Art Unit
2633
USPC
706/20
Abstract

This patent concerns novel technology for detecting backdoors of neural network, particularly deep neural network (DNN), classifiers. The backdoors are planted by suitably poisoning the training dataset, i.e., a data-poisoning attack. Once added to input samples from a source class (or source classes), the backdoor pattern causes the decision of the neural network to change to a target class. The backdoors under consideration are small in norm so as to be imperceptible to a human, but this does not limit their location, support or manner of incorporation. There may not be components (edges, nodes) of the DNN which are dedicated to achieving the backdoor function. Moreover, the training dataset used to learn the classifier may not be available. In one embodiment of the present invention which addresses such challenges, if the classifier is poisoned then the backdoor pattern is determined through a feasible optimization process, followed by an inference process, so that both the backdoor pattern itself and the associated source class(es) and target class are determined based only on the classifier parameters and a set of clean (unpoisoned attacked) samples from the different classes (none of which may be training samples).

Claims (51)

1. A computer-implemented method for detecting backdoor poisoning of a trained classifier, comprising:

receiving the trained classifier, wherein the trained classifier maps input data samples to one of a plurality of predefined classes based on a decision rule that leverages a set of parameters that are learned from a training dataset that may be backdoor-poisoned;

receiving a set of clean (unpoisoned) data samples that includes members from each of the plurality of predefined classes;

using the trained classifier and the clean data samples, estimating for each possible source-target class pair in the plurality of predefined classes one or more potential backdoor perturbations that when incorporated into the clean data samples induce the trained classifier to misclassify the perturbed data samples from a respective source class to a respective target class;

comparing the set of potential backdoor perturbations for the possible source-target class pairs to determine a candidate backdoor perturbation based on at least one of perturbation sizes and misclassification rates; and

determining from the candidate backdoor perturbation whether the trained classifier has been backdoor-poisoned.

2. The computer-implemented method of claim 1 ,

wherein backdoor-poisoning the trained classifier comprises influencing the trained classifier so that a class decision for an input data sample changes from the input data sample's class of origin (source class) to a backdoor-attacker's target class when a backdoor-attacker's backdoor perturbation is incorporated into the input data sample; and

wherein backdoor-poisoning the training set comprises including one or more additional data samples in the training set that each include the backdoor perturbation and are mislabeled to the backdoor-attacker's target class.

3. The computer-implemented method of claim 1 , wherein determining whether the trained classifier has been backdoor-poisoned further comprises, upon determining that the size of the candidate backdoor perturbation is not smaller, by at least a pre-specified margin, than the size of the potential backdoor perturbations for any other source-target class pairs, determining that the trained classifier is not backdoor-poisoned.

4. The computer-implemented method of claim 1 , wherein determining whether the trained classifier has been backdoor-poisoned further comprises, upon determining that the size of the candidate backdoor perturbation is smaller, by at least a pre-specified margin, than the size of the potential backdoor perturbations for any other source-target class pairs, determining that the trained classifier is backdoor-poisoned.

5. The computer-implemented method of claim 4 , wherein determining that the trained classifier is backdoor-poisoned further comprises:

identifying the source class and the target class that are associated with the candidate backdoor perturbation; and

using the candidate backdoor perturbation to estimate a backdoor perturbation that was applied to data samples of the training dataset to perform a backdoor-poisoning attack on the trained classifier.

6. The computer-implemented method of claim 5 , wherein the method further comprises using the estimated backdoor perturbation associated with the backdoor-poisoning attack and its identified source and target classes to detect an unlabeled test sample that includes characteristics of the backdoor perturbation that is associated with the backdoor-poisoning attack.

7. The computer-implemented method of claim 1 ,

wherein the trained classifier is a neural network that was trained using the training dataset; and

wherein the training set is unknown and inaccessible to backdoor-poisoning detection efforts that leverage the trained classifier.

8. The computer-implemented method of claim 7 ,

wherein the neural network comprises internal neurons that are activated when the clean data samples are input to the neural network;

wherein the potential backdoor perturbations are applied to a subset of the internal neurons rather than being applied directly to the clean data samples; and

wherein applying potential backdoor perturbations to the internal neurons facilitates applying the computer-implemented method to any application domain regardless of how a backdoor-poisoning attack is incorporated by an attacker.

9. The computer-implemented method of claim 1 ,

wherein the set of clean data samples are unlabeled; and

wherein class labels are obtained for the set of clean data samples by applying the decision rule of the trained classifier upon the set of clean data samples.

10. The computer-implemented method of claim 1 , wherein estimating the set of potential backdoor perturbations further comprises ensuring that potential backdoor perturbations for each source-target class pair achieve a pre-specified minimum misclassification rate among perturbed clean samples.

11. The computer-implemented method of claim 1 ,

wherein the data samples are images; and

wherein the backdoor perturbation comprises at least one of imperceptible modification of a few pixels in the images, most of the pixels in the images, and all of the pixels in the images.

12. The computer-implemented method of claim 1 , wherein determining whether the trained classifier has been backdoor-poisoned is based on statistical significance assessment, such as p-values of null distributions based on the set of sizes of the estimated potential backdoor perturbations.

13. The computer-implemented method of claim 1 , wherein estimating a potential backdoor perturbation comprises using a gradient ascent technique to maximize a differentiable objective function, with respect to the potential backdoor perturbations, that is an approximation of the non-differentiable count of misclassified perturbed clean samples.

14. The computer-implemented method of claim 1 , wherein each potential backdoor perturbation constitutes a vector whose size is measured using a p-norm, including the Euclidean norm (2-norm).

15. The computer-implemented method of claim 1 , wherein estimating for each possible source-target class pair the one or more potential backdoor perturbations comprises determining one common perturbation for all clean source class samples.

16. The computer-implemented method of claim 1 , wherein estimating potential backdoor perturbations for each possible source-target class pair comprises perturbing only a portion of a clean source class sample.

17. A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for detecting backdoor poisoning of a trained classifier, the method comprising:

receiving the trained classifier, wherein the trained classifier maps input data samples to one of a plurality of predefined classes based on a decision rule that leverages a set of parameters that are learned from a training dataset that may be backdoor-poisoned;

receiving a set of clean (unpoisoned) data samples that includes members from each of the plurality of predefined classes;

using the trained classifier and the clean data samples, estimating for each possible source-target class pair in the plurality of predefined classes one or more potential backdoor perturbations that when incorporated into the clean data samples induce the trained classifier to misclassify the perturbed data samples from a respective source class to a respective target class;

comparing the set of potential backdoor perturbations for the possible source-target class pairs to determine a candidate backdoor perturbation based on at least one of perturbation sizes and misclassification rates; and

determining from the candidate backdoor perturbation whether the trained classifier has been backdoor-poisoned).

18. A backdoor-detection system that performs backdoor-detection on a trained classifier, comprising:

a processor;

a memory; and

a backdoor-detection mechanism;

wherein at least one of the processor and the backdoor-detection mechanism are configured to receive the trained classifier and store parameters for the trained classifier and program instructions that operate upon the trained classifier in the memory;

wherein the trained classifier maps input data samples to one of a plurality of predefined classes based on a decision rule that leverages a set of parameters that are learned from a training dataset that may be backdoor-poisoned;

wherein the backdoor-detection system is configured to:

load from the memory a set of clean (unpoisoned) data samples that includes members from each of the plurality of predefined classes;

execute instructions that, using the trained classifier and the clean data samples, estimate for each possible source-target class pair in the plurality of predefined classes one or more potential backdoor perturbations that when incorporated into the clean data samples induce the trained classifier to misclassify the perturbed data samples from a respective source class to a respective target class;

compare the set of potential backdoor perturbations for the possible source-target class pairs to determine a candidate backdoor perturbation based on at least one of perturbation sizes and misclassification rates; and

determine from the candidate backdoor perturbation whether the trained classifier has been backdoor-poisoned.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2020
From: MILLER, DAVID JONATHAN; KESIDIS, GEORGE
To: ANOMALEE INC.
Reel/Frame 052931/0779 →
Continuity (2)
Provisional Application 62854078 · May 29, 2019
Related Publication 20200380118A1 · Dec 3, 2020