IP Library Granted Patent US 11,568,049
Granted Patent B2
US 11,568,049 · App. 16/586,476 · Granted Jan 31, 2023

Methods and apparatus to defend against adversarial machine learning

Inventors: Sherin M. Mathews (Santa Clara, CA); Celeste R. Fralick (Plano, TX)
Assignee: McAfee, LLC
G06F21/56G06F16/9027G06K9/6215G06K9/6267G06N3/0481
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,568,049
App. No.
16/586,476
Granted
Jan 31, 2023
Kind
B2
Abstract

Methods, apparatus, systems and articles of manufacture to defend against adversarial machine learning are disclosed. An example apparatus includes a model trainer to train a classification model based on files with expected classifications; and a model modifier to select a convolution layer of the trained classification model based on an analysis of the convolution layers of the trained classification model; and replace the convolution layer with a tree-based structure to generate a modified classification model.

Claims (39)

1. An apparatus including:

a model trainer to train a classification model based on files with expected classifications; and

a model modifier to:

select a convolution layer of the trained classification model based on an analysis of the convolution layers of the trained classification model; and

replace the selected convolution layer with a tree-based structure to generate a modified classification model.

2. The apparatus of claim 1 , wherein the model trainer is to train the trained classification model to output a score for an input, the score corresponding to whether the input is non-malware or a malware.

3. The apparatus of claim 1 , wherein the model modifier is to select the convolution layer by:

providing an input into the trained classification model;

providing an adversarial input into the trained classification model; and

when the difference between (A) a first distance between the input and the adversarial input at a first layer of the trained classification model and (B) a second distance between the input and the adversarial input at a second layer preceding the first layer satisfies a threshold value, select the first layer to be the convolution layer.

4. The apparatus of claim 3 , wherein the model modifier is to generate the adversarial input by adding a perturbation to the input.

5. The apparatus of claim 3 , wherein the first and second distances are Euclidean distances.

6. The apparatus of claim 1 , wherein the model modifier is to replace convolution layers subsequent the selected convolution layer with tree-based structures.

7. The apparatus of claim 1 , further including an interface to deploy the modified classification model.

8. The apparatus of claim 1 , wherein the modified classification model protects against gradient-based adversarial attacks.

9. A non-transitory computer readable storage medium comprising instructions which, when executed, cause one or more processors to at least:

train a classification model based on files with expected classifications;

select a convolution layer of the trained classification model based on an analysis of the convolution layers of the trained classification model; and

replace the selected convolution layer with a tree-based structure to generate a modified classification model.

10. The computer readable storage medium of claim 9 , wherein the instructions cause the one or more processors to train the trained classification model to output a score for an input, the score corresponding to whether the input is non-malware or a malware.

11. The computer readable storage medium of claim 9 , wherein the instructions cause the one or more processors to select the convolution layer by:

providing an input into the trained classification model;

providing an adversarial input into the trained classification model; and

when the difference between (A) a first distance between the input and the adversarial input at a first layer of the trained classification model and (B) a second distance between the input and the adversarial input at a second layer preceding the first layer satisfies a threshold value, select the first layer to be the convolution layer.

12. The computer readable storage medium of claim 11 , wherein the instructions cause the one or more processors to generate the adversarial input by adding a perturbation to the input.

13. The computer readable storage medium of claim 11 , wherein the first and second distances are Euclidean distances.

14. The computer readable storage medium of claim 9 , wherein the instructions cause the one or more processors to replace convolution layers subsequent the selected convolution layer with tree-based structures.

15. The computer readable storage medium of claim 9 , wherein the instructions cause the one or more processors to deploy the modified classification model.

16. The computer readable storage medium of claim 9 , wherein the modified classification model protects against gradient-based adversarial attacks.

17. A method comprising:

training, by executing an instruction with a processor, a classification model based on files with expected classifications;

selecting, by executing an instruction with the processor, a convolution layer of the trained classification model based on an analysis of the convolution layers of the trained classification model; and

replacing, by executing an instruction with the processor, the selected convolution layer with a tree-based structure to generate a modified classification model.

18. The method of claim 17 , wherein the training of the trained classification model includes outputting a score for an input, the score corresponding to whether the input is non-malware or a malware.

19. The method of claim 17 , wherein the selecting of the convolution layer includes:

providing an input into the trained classification model;

providing an adversarial input into the trained classification model; and

when the difference between (A) a first distance between the input and the adversarial input at a first layer of the trained classification model and (B) a second distance between the input and the adversarial input at a second layer preceding the first layer satisfies a threshold value, select the first layer to be the convolution layer.

20. The method of claim 19 , wherein the generating of the adversarial input includes adding a perturbation to the input.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE PATENT TITLES AND REMOVE DUPLICATES IN THE SCHEDULE PREVIOUSLY RECORDED AT REEL: 059354 FRAME: 0335. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 23, 2022
From: MCAFEE, LLC
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 060792/0307 →
SECURITY INTEREST Recorded Mar 3, 2022
From: MCAFEE, LLC
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
Reel/Frame 059354/0335 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2019
From: MATHEWS, SHERIN M.; FRALICK, CELESTE R.
To: MCAFEE, LLC
Reel/Frame 050867/0471 →
Continuity (1)
Related Publication 20210097176A1 · Apr 1, 2021