IP Library › Granted Patent US 11,455,515
Granted Patent B2
US 11,455,515 · App. 16/580,650 · Granted Sep 27, 2022

Efficient black box adversarial attacks exploiting input data structure

Inventors: Jeremy Zieg Kolter (Pittsburgh, PA); Anit Kumar Sahu (Pittsburgh, PA)
Assignee: Robert Bosch GmbH
G06N3/0454G06T7/0002G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,455,515
App. No.
16/580,650
Granted
Sep 27, 2022
Kind
B2
Abstract

Markov random field parameters are identified to use for covariance modeling of correlation between gradient terms of a loss function of the classifier. A subset of images are sampled, from a dataset of images, according to a normal distribution to estimate the gradient terms. Black-box gradient estimation is used to infer values of the parameters of the Markov random field according to the sampling. Fourier basis vectors are generated from the inferred values. An original image is perturbed using the Fourier basis vectors to obtain loss function values. An estimate of a gradient is obtained from the loss function values. An image perturbation is created using the estimated gradient. The image perturbation is added to an original input to generate a candidate adversarial input that maximizes loss in identifying the image by the classifier. The neural network classifier is queried to determine a classifier prediction for the candidate adversarial input.

Claims (57)

1. A method for performing a single-step adversarial attack on a neural network classifier, comprising:

identifying Markov random field parameters to use for covariance modeling of correlation between gradient terms of a loss function of the classifier;

sampling a subset of images, from a dataset of images, according to a normal distribution to estimate the gradient terms;

using black-box gradient estimation to infer values of the parameters of the Markov random field according to the sampling;

generating Fourier basis vectors from the inferred values;

perturbing an original image using the Fourier basis vectors to obtain loss function values;

obtaining an estimate of a gradient from the loss function values;

creating an image perturbation using the estimated gradient;

adding the image perturbation to an original input to generate a candidate adversarial input that maximizes loss in identifying the image by the classifier;

querying the neural network classifier to determine a classifier prediction for the candidate adversarial input;

computing a score for the classifier prediction; and

accepting the candidate adversarial input as a successful adversarial attack responsive to the classifier prediction being incorrect.

2. The method of claim 1 , further comprising adding the image perturbation to the original input using a fast gradient sign method.

3. The method of claim 2 , wherein the fast gradient sign method utilizes a pre-specified L ∞ perturbation bound to adding the image perturbation to the original input.

4. The method of claim 1 , wherein the Markov random field parameters include a first parameter that governs diagonal terms, and a second parameter that governs adjacent pixels corresponding to indices that are neighboring in an image.

5. The method of claim 1 , wherein the values of the parameters are inferred using Newton's method.

6. The method of claim 1 , wherein a predefined maximum number of queries of the classifier are utilized to estimate the gradients.

7. The method of claim 1 , wherein the images include video image data.

8. The method of claim 1 , wherein the images include still image data.

9. A computational system for performing a single-step adversarial attack on a neural network classifier, the system comprising:

a memory storing instructions of black box adversarial attack algorithms of a software program; and

a processor programmed to execute the instructions to perform operations including to

identify Markov random field parameters to use for covariance modeling of correlation between gradient terms of a loss function of the classifier;

sample a subset of images, from a dataset of images, according to a normal distribution to estimate the gradient terms;

use black-box gradient estimation to infer values of the parameters of the Markov random field according to the sampling;

generate Fourier basis vectors from the inferred values;

perturb an original image using the Fourier basis vectors to obtain loss function values;

obtain an estimate of a gradient from the loss function values;

create an image perturbation using the estimated gradient;

add the image perturbation to an original input to generate a candidate adversarial input that maximizes loss in identifying the image by the classifier;

query the neural network classifier to determine a classifier prediction for the candidate adversarial input;

compute a score for the classifier prediction;

accept the candidate adversarial input as a successful adversarial attack responsive to the classifier prediction being incorrect; and

reject the candidate adversarial input responsive to the classifier prediction being correct.

10. The computational system of claim 9 , wherein the processor is further programmed to add the image perturbation to the original input using a fast gradient sign method.

11. The computational system of claim 10 , wherein the fast gradient sign method utilizes a pre-specified L ∞ perturbation bound to adding the image perturbation to the original input.

12. The computational system of claim 9 , wherein the Markov random field parameters include a first parameter that governs diagonal terms, and a second parameter that governs adjacent pixels corresponding to indices that are neighboring in an image.

13. The computational system of claim 9 , wherein the values of the parameters are inferred using Newton's method.

14. The computational system of claim 9 , wherein a predefined maximum number of queries of the classifier are utilized to estimate the gradients.

15. A non-transitory computer-readable medium comprising instructions for performing a single-step adversarial attack on a neural network classifier that, when executed by a processor, cause the processor to:

identify Markov random field parameters to use for covariance modeling of correlation between gradient terms of a loss function of the classifier;

sample a subset of images, from a dataset of images, according to a normal distribution to estimate the gradient terms;

use black-box gradient estimation to infer values of the parameters of the Markov random field according to the sampling;

generate Fourier basis vectors from the inferred values;

perturb an original image using the Fourier basis vectors to obtain loss function values;

obtain an estimate of a gradient from the loss function values;

create an image perturbation using the estimated gradient;

add the image perturbation to an original input to generate a candidate adversarial input that maximizes loss in identifying the image by the classifier;

query the neural network classifier to determine a classifier prediction for the candidate adversarial input;

compute a score for the classifier prediction;

accept the candidate adversarial input as a successful adversarial attack responsive to the classifier prediction being incorrect; and

reject the candidate adversarial input responsive to the classifier prediction being correct.

16. The medium of claim 15 , further comprising instructions that, when executed by the processor, cause the processor to add the image perturbation to the original input using a fast gradient sign method.

17. The medium of claim 16 , wherein the fast gradient sign method utilizes a pre-specified L ∞ perturbation bound to adding the image perturbation to the original input.

18. The medium of claim 15 , wherein the Markov random field parameters include a first parameter that governs diagonal terms, and a second parameter that governs adjacent pixels corresponding to indices that are neighboring in an image.

19. The medium of claim 15 , wherein the values of the parameters are inferred using Newton's method.

20. The medium of claim 15 , wherein a predefined maximum number of queries of the classifier are utilized to estimate the gradients.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2019
From: KOLTER, JEREMY ZIEG; SAHU, ANIT KUMAR
To: ROBERT BOSCH GMBH
Reel/Frame 050476/0627 →
Continuity (1)
Related Publication 20210089866A1 · Mar 25, 2021
Cited By (1)
US 12,632,541