IP Library › Granted Patent US 12,210,619
Granted Patent B2
US 12,210,619 · App. 17/035,173 · Granted Jan 28, 2025

Method and system for breaking backdoored classifiers through adversarial examples

Inventors: Mingjie Sun (Pittsburgh, PA); Jeremy Kolter (Pittsburgh, PA); Filipe J. Cabrita Condessa (Pittsburgh, PA)
Assignee: Robert Bosch GmbH
G06F21/554G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,619
App. No.
17/035,173
Granted
Jan 28, 2025
Kind
B2
Abstract

A computer-implemented method for training a machine learning network includes receiving an input data from one or more sensors, selecting one or more batch samples from the input data, wherein the batch samples include one or more perturbed samples from a source class configured to be misclassified into a target class, identifying the one or more perturbed samples from the one or more batch samples, determining a trigger event in response to identification of a trigger pattern of the one or more batch samples, wherein the trigger pattern induces a pre-determined response on a classifier, outputting a classification in response to identification of the trigger pattern via the classifier, and outputting a set of trigger patterns extracted from the machine-learning network.

Claims (44)

1. A computer-implemented method for identifying backdoor triggers from a machine learning network, comprising:

receiving an input data from one or more sensors, wherein the input data includes information indicative of image information or sound information;

selecting one or more batch samples from the input data, wherein the batch samples include one or more perturbed samples from a source class configured to be misclassified into a target class;

identifying the one or more perturbed samples from the one or more batch samples;

determining a trigger event in response to identification of a trigger pattern of the one or more batch samples, wherein the trigger pattern induces a pre-determined response on a classifier;

determining that the classifier requires robustification, wherein the classifier is trained with robustification data augmentation to robustify the classifier in response to the determination indicating the classifier requires robustification;

outputting a classification in response to identification of the trigger pattern via the classifier; and

outputting a set of trigger patterns extracted from the machine learning network.

2. The computer-implemented method of claim 1 , wherein the method includes training the classifier via operating Gaussian data augmentation of the classifier.

3. The computer-implemented method of claim 1 , wherein the method includes prepending a denoiser to the classifier.

4. The computer-implemented method of claim 3 , wherein the denoiser is a custom-trained denoiser.

5. The computer-implemented method of claim 1 , wherein a perturbation associated with the input data and machine learning network facilitates the identification of the trigger event.

6. The computer-implemented method of claim 1 , wherein the classification associated with input data facilitates in identification of the trigger event.

7. The computer-implemented method of claim 1 , wherein the method includes cropping colors and patterns associated with the input data to initiate the trigger event.

8. The computer-implemented method of claim 1 , wherein the identifying is in response to a project gradient descent attack process.

9. A system including a machine learning network, comprising:

an input interface configured to receive input data, wherein the input interface is connected to one or more sensors, wherein the one or more sensors includes a video, radar, LiDAR, sound, sonar, ultrasonic, motion, or thermal imaging sensor; and

a processor, in communication with the input interface, wherein the processor is programmed to:

receive an input data from one or more sensors, wherein the input data includes information indicative of image information or sound information;

select one or more batch samples from the input data, wherein the batch samples include one or more perturbed samples from a source class configured to be misclassified into a target class;

identify the one or more perturbed samples from the one or more batch samples;

determine a trigger event in response to identification of a trigger pattern of the one or more batch samples, wherein the trigger pattern induces a pre-determined response on a classifier and the trigger pattern occupy a portion of the image; and

output a classification in response to identification of the trigger pattern via the classifier;

determining if the classifier requires robustification, wherein the classifier is trained with data augmentation in response to the determination indicating the classifier requires robustification;

outputting a set of trigger patterns extracted from the machine learning network; and

operating a physical system based on output data including the set of trigger patterns, wherein the physical system is a computer-controlled machine, a robot, a vehicle, a domestic appliance, a power tool, a manufacturing machine, a personal assistant or an access control system.

10. The system of claim 9 , wherein the processor is programmed to identify in response to a project gradient descent attack process.

11. The system of claim 9 , wherein the processor is further programmed to operate a Gaussian data augmentation of the classifier.

12. The system of claim 9 , wherein processor is further programmed to remove Gauissian noise associated with the input data via a denoiser applied to the classifier.

13. The system of claim 9 , wherein a perturbation associated with the input data and machine learning network facilities in identification of the trigger event.

14. The system of claim 9 , wherein the classification associated with input data facilitates in identification of the trigger event.

15. A computer-program product storing instructions on non-transitory memory, the instructions which, when executed by a computer, cause the computer to:

receive an input data from one or more sensors, wherein the input data includes information indicative of image information or sound information;

select one or more batch samples from the input data, wherein the batch samples include one or more perturbed samples from a source class configured to be misclassified into a target class of a machine-learning network;

identify the one or more perturbed samples from the one or more batch samples;

determine a trigger event in response to identification of a trigger pattern of the one or more batch samples, wherein the trigger pattern induces a pre-determined response on a classifier;

determine if the classifier requires robustification, wherein the classifier is trained with data augmentation to robustify the classifier in response to the determination indicating the classifier requires robustification;

output a classification in response to identification of the trigger pattern via the classifier; and

output a set of trigger patterns extracted from the machine-learning network.

16. The computer-program product of claim 15 , wherein a perturbation associated with the input data facilities in identification of the trigger event.

17. The computer-program product of claim 15 , wherein the trigger pattern is a cropped trigger pattern.

18. The computer-program product of claim 15 , wherein the trigger pattern is a synthetic trigger pattern.

19. The computer-program product of claim 15 , wherein the instructions that cause the computer to determine the trigger event is responsive to a Tikhov regularization.

20. The computer-program product of claim 15 , wherein the trigger pattern is associated with an adversarial attack.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 10, 2022
From: SUN, MINGJIE; KOLTER, JEREMY; CABRITA CONDESSA, FILIPE J.
To: ROBERT BOSCH GMBH
Reel/Frame 060157/0468 →
Continuity (1)
Related Publication 20220100850A1 · Mar 31, 2022
References Cited (18)
US 10469508B2 · Muddu · 2019 [cited by examiner]
US 10569199B2 · Lin · 2020 [cited by examiner]
US 10706113B2 · Lundin · 2020 [cited by examiner]
US 11209813B2 · Cella · 2021 [cited by examiner]
US 20140254920A1 · Xu · 2014 [cited by examiner]
US 20190130110A1 · Lee · 2019 [cited by examiner]
US 20190156198A1 · Mars · 2019 [cited by examiner]
US 20190318099A1 · Carvalho · 2019 [cited by examiner]
US 20200265271A1 · Zhang · 2020 [cited by examiner]
US 20210019399A1 · Miller · 2021 [cited by examiner]
US 20210157912A1 · Kruthiveti Subrahmanyeswara Sai · 2021 [cited by examiner]
JP 2016114992A · 2014 [cited by examiner]
Title: Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks Author(s): Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao Date: 2019 Publisher: IEE… [cited by examiner]
Title: Detecting Backdoor Attacks on Deep Neural Networks byActivation Clustering Author(s): Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Benjamin Edwards, Taesung Lee, and Bipla Srivastava Date: 2018 … [cited by examiner]
Cohen et al., “Certified Adversarial Robustness via Randomized Smoothing”, Proceedings of the 36th International Conference on Machine Learning, Long Beach, California, PMLR 97, Jun. 2019, 36 pages. [cited by applicant]
Salman et al., “Provably Robust Deep Learning via Adversarially Trained Smoothed Classifiers”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada, arXiv:1906.04584v5, Jan. 2020, 3… [cited by applicant]
Salman et al., “Denoised Smoothing: A Provable Defense for Pretrained Classifiers”, arXiv:2003.01908v2, Sep. 2020, 29 pages. [cited by applicant]
Chen et al., “Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning”, arXiv:1712.05526v1, Dec. 2017, 18 pages. [cited by applicant]