IP Library › Granted Patent US 11,948,054
Granted Patent B2
US 11,948,054 · App. 17/083,928 · Granted Apr 2, 2024

Masked projected gradient transfer attacks

Inventors: Luke Edward Richards (Ellicott City, MD); Andre Tai Nguyen (Columbia, MD); Ryan Joseph Capps (Atlanta, GA); Edward Simon Paster Raff (Jamesville, NY)
Assignee: BOOZ ALLEN HAMILTON INC.
G06N20/00H04L63/1416H04L63/1466
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,948,054
App. No.
17/083,928
Granted
Apr 2, 2024
Kind
B2
Abstract

A system and method for transferring an adversarial attack involving generating a surrogate model having an architecture and a dataset that mirrors at least one aspect of a target model of a target module, wherein the surrogate model includes a plurality of classes. The method involves generating a masked version of the surrogate model having fewer classes than the surrogate model by randomly selecting at least one class of the plurality of classes for removal. The method involves attacking the masked surrogate model to create a perturbed sample. The method involves generalizing the perturbed sample for use with the target module. The method involves transferring the perturbed sample to the target module to alter an operating parameter of the target model.

Claims (109)

1. A system for generating a transfer adversarial attack, the system comprising:

an attack module configured to generate an adversarial attack on a target module, wherein the attack module is configured to:

generate a surrogate model having an architecture and a dataset that mirrors at least one aspect of a target model of the target module, the surrogate model including a plurality of classes;

generate a masked version of the surrogate model having fewer classes than the surrogate model by randomly selecting at least one class of the plurality of classes for removal;

attack the masked surrogate model to create a perturbed sample;

generalize the perturbed sample for use with the target module; and

transfer the perturbed sample to the target module to alter an operating parameter of the target model.

2. The system of claim 1 , wherein:

the attack module includes program instructions stored within a memory having an algorithm as follows:

for n = 1...S do

  for i = 1...N do

   δ avg = 0

    for j = 1...T do

     // Compute a random mask for logits of f θ

     M = random mask

     // Apply mask, M to model output

     x p = f θ (x) · M

     δ = δ + α · sign(∇ δ    (x p + δ), y)) δ avg = δ avg + max(min(δ,

     ε), −ε)

    end

  end

 // Average the deltas found from random masks

 δ avg = δ avg /T

end

wherein:

S=total number of data points to attack;

N=number of times to perturb each data point;

δ avg =one of N perturbations to a data point;

T=number of mask iterations;

x p =the result or output of the surrogate model;

f θ =a machine learning model; θ=parameters of the machine learning model;

δ=a current perturbation;

α=a constant;

ε=a constant;

∇ δ =the computed gradient with respect to δ perturbation;

l=a loss function; and

Y=a class of data.

3. The system of claim 2 , wherein:

N>=1,000; and

T>=100.

4. The system of claim 2 , wherein:

α=represents a learning rate.

5. The system of claim 2 , wherein:

ε is a value that bounds the algorithm to limit perturbation of data for the perturbed sample.

6. The system of claim 5 , wherein:

ε is set to limit perturbation of data so that the perturbed sample is undetectable but causes the altered the operating parameter in the target model.

7. The system of claim 1 , wherein:

the attack module is configured to generate the surrogate model having a characteristic of the surrogate model architecture and dataset that matches a characteristic of the target model architecture and dataset; and

the degree of matching between the architecture of the surrogate model and the architecture of the target model is greater than the degree of matching between the dataset of the surrogate model and the dataset of the target model.

8. The system of claim 1 , wherein:

the attack on the masked surrogate model is a projected gradient descent attack.

9. The system of claim 1 , wherein:

the target model has an architecture and a dataset, the dataset including a training dataset based on machine learning or deep learning.

10. A method for transferring an adversarial attack, the method comprising:

generating a surrogate model having an architecture and a dataset that mirrors at least one aspect of a target model of a target module, the surrogate model including a plurality of classes;

generating a masked version of the surrogate model having fewer classes than the surrogate model by randomly selecting at least one class of the plurality of classes for removal;

attacking the masked surrogate model to create a perturbed sample;

generalizing the perturbed sample for use with the target module; and

transferring the perturbed sample to the target module to alter an operating parameter of the target model.

11. The method of claim 10 , wherein:

generating the masked version of the surrogate model involves implementation of the following algorithm:

for n = 1...S do

  for i = 1...N do

   δ avg = 0

    for j = 1...T do

     // Compute a random mask for logits of f θ

     M = random mask

     // Apply mask, M to model output

     x p = f θ (x) · M

     δ = δ + α · sign(∇ δ    (x p + δ), y)) δ avg = δ avg + max(min(δ,

     ε), −ε)

    end

  end

 // Average the deltas found from random masks

 δ avg = δ avg /T

end

wherein:

S=total number of data points to attack;

N=number of times to perturb each data point;

δ avg =one of N perturbations to a data point;

T=number of mask iterations;

x p =the result or output of the surrogate model;

f θ =a machine learning model; θ=parameters of the machine learning model;

δ=a current perturbation;

α=a constant;

ε=a constant;

∇ δ =the computed gradient with respect to δ perturbation;

l=a loss function; and

Y=a class of data.

12. The method of claim 11 , wherein:

N>=1,000; and

T>=100.

13. The method of claim 11 , wherein:

α=represents a learning rate.

14. The method of claim 11 , wherein:

ε is a value that bounds the algorithm to limit perturbation of data for the perturbed sample.

15. The method of claim 11 , wherein:

ε is set to limit perturbation of data so that the perturbed sample is undetectable but causes the altered the operating parameter in the target model.

16. The method of claim 10 , wherein:

the surrogate model has a characteristic of the surrogate model architecture and dataset that matches a characteristic of the target model architecture and dataset; and

the degree of matching between the architecture of the surrogate model and the architecture of the target model is greater than the degree of matching between the dataset of the surrogate model and the dataset of the target model.

17. The method of claim 10 , wherein:

attacking the masked surrogate model is via a projected gradient descent attack.

18. The method of claim 17 , comprising:

developing a defense model based on the attack on the target module.

19. The method of claim 10 , wherein:

the target model has an architecture and a dataset, the dataset including a training dataset based on machine learning or deep learning.

20. The method of claim 10 , comprising:

estimating a likelihood of success for the attack on the target module.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2020
From: RICHARDS, LUKE EDWARD; NGUYEN, ANDRE TAI; CAPPS, RYAN JOSEPH; RAFF, EDWARD SIMON PASTER
To: BOOZ ALLEN HAMILTON INC.
Reel/Frame 054213/0557 →
Continuity (1)
Related Publication 20220141251A1 · May 5, 2022