IP Library Granted Patent US 12,050,671
Granted Patent B2
US 12,050,671 · App. 17/858,775 · Granted Jul 30, 2024

Methods and systems for watermarking neural networks

Inventors: Nandish Chattopadhyay (Singapore, SG); Anupam Chattopadhyay (Singapore, SG)
Assignee: NANYANG TECHNOLOGICAL UNIVERSITY
G06F21/16G06F18/2431G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,050,671
App. No.
17/858,775
Granted
Jul 30, 2024
Kind
B2
Abstract

Disclosed herein is a system for watermarking a neural network, comprising memory; and at least one processor in communication with the memory; wherein the memory stores instructions for causing the at least one processor to carry out a method comprising: generating a trigger set by obtaining examples from a training set by random sampling from the training set, respective examples being associated with respective true classes of a plurality of classes; generating a set of adversarial examples by structured perturbation of the examples; generating, for each adversarial example, one or more adversarial class labels by passing the adversarial example to the neural network; and applying one or more trigger labels to each said adversarial example, wherein the one or more trigger labels are selected randomly from the plurality of classes, and wherein each trigger label is not a said true class label for the corresponding example or a said adversarial class label for the corresponding adversarial example; and storing the adversarial examples and corresponding trigger labels as the trigger set; and performing a tuning process to adjust parameters at each layer of the neural network using the trigger set, to thereby generate a watermarked neural network.

Claims (42)

1. A system for watermarking a neural network, comprising:

memory; and

at least one processor in communication with the memory;

wherein the memory stores instructions for causing the at least one processor to carry out a method comprising:

generating a trigger set by:

obtaining examples from a training set by random sampling from the training set, respective examples being associated with respective true classes of a plurality of classes;

generating a set of adversarial examples by structured perturbation of the examples, by the fast gradient sign method;

generating, for each adversarial example, one or more adversarial class labels by passing the adversarial example to the neural network; and

applying one or more trigger labels to each said adversarial example, wherein the one or more trigger labels are selected randomly from the plurality of classes, and wherein each trigger label is not a said true class label for the corresponding example or a said adversarial class label for the corresponding adversarial example; and

storing the adversarial examples and corresponding trigger labels as the trigger set; and

performing a tuning process to adjust parameters at each layer of the neural network using the trigger set, to thereby generate a watermarked neural network.

2. A system according to claim 1 , wherein the random sampling is performed evenly across the plurality of classes.

3. A system according to claim 1 wherein said tuning process comprises a sequence of epochs, and wherein each epoch comprises generating an updated neural network by (i) updating parameters for each layer separately using the trigger set while keeping parameters for all other layers fixed.

4. A system according to claim 3 , wherein each epoch comprises (ii) determining a classification accuracy S acc of the updated neural network for a test set and a classification accuracy T acc of the updated neural network for the trigger set.

5. A system according to claim 4 , wherein (i) and (ii) are performed for each epoch in sequence until T acc starts to saturate and/or S acc begins to decrease.

6. A system according claim 1 , wherein the watermarked neural network is configured to be verified based on the adversarial examples and the trigger labels.

7. A system according to claim 1 , wherein the instructions cause the at least one processor to:

obtain the examples from the training set by obtaining out-of-domain examples; and

generate the adversarial examples by structured perturbation of the examples and of the out-of-domain examples.

8. A system according to claim 1 , wherein performing the tuning process to adjust the parameters at each layer of the neural network uses both the trigger set and one or more clean samples.

9. A method of watermarking a neural network that is trained using a training set of samples, the neural network being configured to classify samples into one of a plurality of classes, the method comprising:

generating a trigger set by:

obtaining examples from the training set by random sampling from the training set, respective examples being associated with respective true classes of a plurality of classes;

generating a set of adversarial examples by structured perturbation of the examples, by the fast gradient sign method;

generating, for each adversarial example, one or more adversarial class labels by passing the adversarial example to the neural network; and

applying one or more trigger labels to each said adversarial example, wherein the one or more trigger labels are selected randomly from the plurality of classes, and wherein each trigger label is not a said true class label for the corresponding example or a said adversarial class label for the corresponding adversarial example; and

storing the adversarial examples and corresponding trigger labels as the trigger set; and

performing a tuning process to adjust parameters at each layer of the neural network using the trigger set, to thereby generate a watermarked neural network.

10. A method according to claim 9 , wherein the random sampling is performed evenly across the plurality of classes.

11. A method according to claim 9 , wherein said tuning process comprises a sequence of epochs, and wherein each epoch comprises generating an updated neural network by (i) updating parameters for each layer separately using the trigger set while keeping parameters for all other layers fixed.

12. A method according to claim 11 , wherein each epoch comprises (ii) determining a classification accuracy S acc of the updated neural network for a test set and a classification accuracy T acc of the updated neural network for the trigger set.

13. A method according to claim 12 , wherein (i) and (ii) are performed for each epoch in sequence until T acc starts to saturate and/or S acc begins to decrease.

14. A method according to claim 9 , wherein the watermarked neural network is configured to be verified based on the adversarial examples and the trigger labels.

15. A method according to claim 10 , wherein obtaining the examples from the training set comprises obtaining out-of-domain examples.

16. A method according to claim 10 , wherein performing the tuning process to adjust the parameters at each layer of the neural network to thereby generate a watermarked neural network is by using the trigger set and clean samples.

17. Non-transitory computer-readable storage having machine-readable instructions stored thereon for watermarking a neural network, causing at least one processor to: generating a trigger set by:

obtaining examples from the training set by random sampling from the training set, respective examples being associated with respective true classes of a plurality of classes;

generating a set of adversarial examples by structured perturbation of the examples, by the fast gradient sign method;

generating, for each adversarial example, one or more adversarial class labels by passing the adversarial example to the neural network; and

applying one or more trigger labels to each said adversarial example, wherein the one or more trigger labels are selected randomly from the plurality of classes, and wherein each trigger label is not a said true class label for the corresponding example or a said adversarial class label for the corresponding adversarial example; and

storing the adversarial examples and corresponding trigger labels as the trigger set; and

performing a tuning process to adjust parameters at each layer of the neural network using the trigger set, to thereby generate a watermarked neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2022
From: CHATTOPADHYAY, NANDISH; CHATTOPADHYAY, ANUPAM
To: NANYANG TECHNOLOGICAL UNIVERSITY
Reel/Frame 060812/0135 →
Priority Claims (1)
SG 10202107447S · Jul 7, 2021 · national
Continuity (1)
Related Publication 20230012871A1 · Jan 19, 2023
Cited By (3)
US 12,346,417 US 12,738,281 US 12,738,282