IP Library Granted Patent US 12,045,713
Granted Patent B2
US 12,045,713 · App. 16/950,684 · Granted Jul 23, 2024

Detecting adversary attacks on a deep neural network (DNN)

Inventors: Jialong Zhang (White Plains, NY); Zhongshu Gu (Ridgewood, NJ); Jiyong Jang (Chappaqua, NY); Marc Philippe Stoecklin (Zurich, CH); Ian Michael Molloy (Chappaqua, NY)
Assignee: International Business Machines Corporation
G06N3/063G06F18/214G06F18/2433G06N3/082G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,045,713
App. No.
16/950,684
Granted
Jul 23, 2024
Kind
B2
Abstract

A method, apparatus and computer program product to protect a deep neural network (DNN) having a plurality of layers including one or more intermediate layers. In this approach, a training data set is received. During training of the DNN using the received training data set, a representation of activations associated with an intermediate layer is recorded. For at least one or more of the representations, a separate classifier (model) is trained. The classifiers, collectively, are used to train an outlier detection model. Following training, the outliner detection model is used to detect an adversarial input on the deep neural network. The outlier detection model generates a prediction, and an indicator whether a given input is the adversarial input. According to a further aspect, an action is taken to protect a deployed system associated with the DNN in response to detection of the adversary input.

Claims (38)

1. A method to protect a deep neural network (DNN) having a plurality of layers including one or more intermediate layers, comprising:

recording one or more representations of activations, wherein each recorded representation of activations is associated with a respective intermediate layer of the one or more intermediate layers;

for each of the one or more representations of activations, training a separate classifier;

following training of separate classifiers for the one or more representations of activations, using the separate classifiers trained from the one or more representations of activations to detect an adversarial input on the deep neural network;

de-activating one or more neurons at a last DNN layer;

using a set of values from one or more remaining neurons in the last DNN layer to generate an additional classifier; and

using the additional classifier to confirm detection of the adversarial input.

2. The method as described in claim 1 , wherein training a classifier, of the separate classifiers, generates a set of label arrays, a label array being a set of labels for a representation of activations associated with an intermediate layer.

3. The method as described in claim 2 , wherein using the separate classifiers further includes aggregating respective sets of label arrays into an outlier detection model.

4. The method as described in claim 3 , wherein the outlier detection model generates a prediction, together with an indicator, that indicates whether a given input is the adversarial input.

5. The method as described in claim 4 , further including taking an action in response to detection of the adversarial input.

6. The method as described in claim 5 , wherein the action is one of: issuing a notification, preventing an adversary from providing one or more additional inputs that are determined to be adversarial inputs, taking an action to protect a deployed system associated with the DNN, taking an action to retrain or harden the DNN.

7. An apparatus, comprising:

a processor;

computer memory holding computer program instructions executed by the processor to protect a deep neural network (DNN) having a plurality of layers including one or more intermediate layers, the computer program instructions configured to:

record one or more representations of activations, wherein each recorded representation of activations is associated with a respective intermediate layer of the one or more intermediate layers;

for each of the one or more representations of activations, train a separate classifier;

following training of separate classifiers for the one or more representations of activations, use the separate classifiers trained from the one or more representations of activations to detect an adversarial input on the deep neural network;

de-activate one or more neurons at a last DNN layer;

use a set of values from one or more remaining neurons in the last DNN layer to generate an additional classifier; and

use the additional classifier to confirm detection of the adversarial input.

8. The apparatus as described in claim 7 , wherein training a classifier, of the separate classifiers, generates a set of label arrays, a label array being a set of labels for a representation of activations associated with an intermediate layer.

9. The apparatus as described in claim 8 , wherein the computer program instructions configured to use the separate classifiers further includes computer program instruction configured to aggregate respective sets of label arrays into an outlier detection model.

10. The apparatus as described in claim 9 , wherein the computer program instructions further include computer program instructions configured using the outlier detection model to generate a prediction, together with an indicator, that indicates whether a given input is the adversarial input.

11. The apparatus as described in claim 10 , wherein the computer program instructions include computer program instructions further configured to take an action in response to detection of the adversarial input.

12. The apparatus as described in claim 11 , wherein the action is one of: issuing a notification, preventing an adversary from providing one or more additional inputs that are determined to be adversarial inputs, taking an action to protect a deployed system associated with the DNN, taking an action to retrain or harden the DNN.

13. A computer program product in a non-transitory computer readable medium for use in a data processing system to protect a deep neural network (DNN) having a plurality of layers including one or more intermediate layers, the computer program product holding computer program instructions that, when executed by the data processing system, are configured to:

record one or more representations of activations, wherein each recorded representation of activations is associated with a respective intermediate layer of the one or more intermediate layers;

for each of the one or more representations of activations, train a separate classifier;

following training of separate classifiers for the one or more representations of activations, use the separate classifiers trained from the one or more representations of activations to detect an adversarial input on the deep neural network;

de-activate one or more neurons at a last DNN layer;

use a set of values from one or more remaining neurons in the last DNN layer to generate an additional classifier; and

use the additional classifier to confirm detection of the adversarial input.

14. The computer program product as described in claim 13 , wherein training a classifier, of the separate classifiers, generates a set of label arrays, a label array being a set of labels for a representation of activations associated with an intermediate layer.

15. The computer program product as described in claim 14 , wherein the computer program instructions configured to use the separate classifiers further includes computer program instruction configured to aggregate respective sets of label arrays into an outlier detection model.

16. The computer program product as described in claim 15 , wherein the computer program instructions further include computer program instructions configured using the outlier detection model to generate a prediction, together with an indicator, that indicates whether a given input is the adversarial input.

17. The computer program product as described in claim 16 , wherein the computer program instructions include computer program instructions further configured to take an action in response to detection of the adversarial input.

18. The computer program instructions as described in claim 17 , wherein the action is one of: issuing a notification, preventing an adversary from providing one or more additional inputs that are determined to be adversarial inputs, taking an action to protect a deployed system associated with the DNN, taking an action to retrain or harden the DNN.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE RECEIVING PARTY DATA STREET ADDRESS PREVIOUSLY RECORDED ON REEL 54395 FRAME 399. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 6, 2024
From: ZHANG, JIALONG; GU, ZHONGSHU; JANG, JIYONG; STOECKLIN, MARC PHILIPPE; MOLLOY, IAN MICHAEL
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 069314/0568 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2020
From: ZHANG, JIALONG; GU, ZHONGSHU; JANG, JIYONG; STOECKLIN, MARC PHILIPPE; MOLLOY, IAN MICHAEL
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054395/0399 →
Continuity (1)
Related Publication 20220156563A1 · May 19, 2022
Cited By (2)
US 12,530,881 US 12,597,234