IP Library › Granted Patent US 12,736,933
Granted Patent B2
US 12,736,933 · App. 18/331,044 · Granted Sep 15, 2026

Device and method for determining adversarial perturbations of a machine learning system

Inventors: Nicole Ying Finnie (Renningen, DE); Jan Hendrik Metzen (Boeblingen, DE); Robin Hutmacher (Renningen, DE)
Assignee: ROBERT BOSCH GMBH
G05B13/045G05B13/0265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,736,933
App. No.
18/331,044
Granted
Sep 15, 2026
Kind
B2
Abstract

A computer-implemented method for determining an adversarial perturbation for input signals, especially sensor signals or features of sensor signals, of a machine learning system. A best perturbation is determined iteratively, wherein the best perturbation is provided as adversarial perturbation after a predefined amount of iterations, wherein at least one iteration includes: sampling a perturbation; applying the sampled perturbation to an input signal thereby determining a potential adversarial example; determining an output signal from the machine learning system for the potential adversarial example, determining a loss value characterizing a deviation of the output signal to a desired output signal, wherein the desired output signal corresponds to the input signal, if the loss value is larger than a previous loss value setting the best perturbation to the sampled perturbation.

Claims (46)

1 . A computer-implemented method for determining an adversarial perturbation for input signals of a machine learning system, the method comprising the following steps:

iteratively determining a best perturbation, wherein the best perturbation is provided as adversarial perturbation after a predefined amount of iterations, wherein at least one iteration includes the following steps:

sampling a perturbation;

applying the sampled perturbation to an input signal to determine a potential adversarial example;

determining an output signal from the machine learning system for the potential adversarial example;

determining a loss value characterizing a deviation of the output signal to a desired output signal, wherein the desired output signal corresponds to the input signal;

based on the loss value being larger than a previous loss value, setting the best perturbation to the sampled perturbation;

wherein in each iteration, elements of the sampled perturbation are set to zero, wherein a number of elements set to zero is proportional to how many iterations have passed.

2 . The method according to claim 1 , wherein the input signals are sensor signals or features of sensor signals.

3 . The method according to claim 1 , wherein at least one element of the input signal characterizes an integer and the sampled perturbation includes a corresponding element characterizing an integer.

4 . The method according to claim 1 , wherein the adversarial perturbation is sampled by sampling a random perturbation for each input signal of a dataset and combining the sampled random perturbations.

5 . The method according to claim 1 , wherein the output signal characterizes a classification and/or regression result and/or a density value and/or a probability value, based on the input signal.

6 . The method according to claim 1 , further comprising:

applying the adversarial perturbation to a training input signal to determine an adversarial example for training the machine learning system.

7 . A method for training a machine learning system, the method comprising the following steps:

training the machine learning system including:

determining for a training input signal of the machine learning system an adversarial perturbation by:

iteratively determining a best perturbation, wherein the best perturbation is provided as adversarial perturbation after a predefined amount of iterations, wherein at least one iteration includes the following steps:

sampling a perturbation,

applying the sampled perturbation to an input signal to determine a potential adversarial example,

determining an output signal from the machine learning system for the potential adversarial example,

determining a loss value characterizing a deviation of the output signal to a desired output signal, wherein the desired output signal corresponds to the input signal,

based on the loss value being larger than a previous loss value, setting the best perturbation to the sampled perturbation,

wherein in each iteration, elements of the sampled perturbation are set to zero, wherein a number of elements set to zero is proportional to how many iterations have passed;

applying the adversarial perturbation to the training input signal to determining an adversarial example and training the machine learning system to predict a desired output signal corresponding to the training input signal for the adversarial example.

8 . A training system configured to train a machine learning system, the training system configured to:

train the machine learning system including:

determining for a training input signal of the machine learning system an adversarial perturbation by:

iteratively determining a best perturbation, wherein the best perturbation is provided as adversarial perturbation after a predefined amount of iterations, wherein at least one iteration includes the following steps:

sampling a perturbation,

applying the sampled perturbation to an input signal to determine a potential adversarial example,

determining an output signal from the machine learning system for the potential adversarial example,

determining a loss value characterizing a deviation of the output signal to a desired output signal, wherein the desired output signal corresponds to the input signal,

based on the loss value being larger than a previous loss value, setting the best perturbation to the sampled perturbation,

wherein in each iteration, elements of the sampled perturbation are set to zero, wherein a number of elements set to zero is proportional to how many iterations have passed;

apply the adversarial perturbation to the training input signal to determining an adversarial example and training the machine learning system to predict a desired output signal corresponding to the training input signal for the adversarial example.

9 . A non-transitory machine-readable storage medium on which is stored a computer program for determining an adversarial perturbation for input signals of a machine learning system, the computer program, when executed by a computer, causing the computer to perform the following steps comprising:

iteratively determining a best perturbation, wherein the best perturbation is provided as adversarial perturbation after a predefined amount of iterations, wherein at least one iteration includes the following steps:

sampling a perturbation;

applying the sampled perturbation to an input signal to determine a potential adversarial example;

determining an output signal from the machine learning system for the potential adversarial example;

determining a loss value characterizing a deviation of the output signal to a desired output signal, wherein the desired output signal corresponds to the input signal;

based on the loss value being larger than a previous loss value, setting the best perturbation to the sampled perturbation,

wherein in each iteration, elements of the sampled perturbation are set to zero, wherein a number of elements set to zero is proportional to how many iterations have passed.

10 . The non-transitory machine-readable storage medium according to claim 9 , wherein the computer program, when executed by a computer, further causing the computer to perform the following step comprising:

applying the adversarial perturbation to a training input signal to determine an adversarial example for training the machine learning system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: FINNIE, NICOLE YING; METZEN, JAN HENDRIK; HUTMACHER, ROBIN
To: ROBERT BOSCH GMBH
Reel/Frame 064032/0944 →
Priority Claims (1)
EP 22180551 · Jun 22, 2022 · regional
Continuity (1)
Related Publication 20230418246A1 · Dec 28, 2023
References Cited (21)
US 10521718B1 · Szegedy · 2019 [cited by examiner]
US 11893111B2 · Kruthiveti Subrahmanyeswara Sai · 2024 [cited by examiner]
US 12069047B2 · Wu · 2024 [cited by examiner]
US 12481874B2 · Liu · 2025 [cited by examiner]
US 12505513B2 · Kapoor · 2025 [cited by examiner]
US 20190130110A1 · Lee · 2019 [cited by examiner]
US 20190295721A1 · Madabhushi · 2019 [cited by examiner]
US 20190370683A1 · Metzen · 2019 [cited by examiner]
US 20200257978A1 · Roy · 2020 [cited by examiner]
US 20200285952A1 · Liu · 2020 [cited by examiner]
US 20210089957A1 · Ermans · 2021 [cited by examiner]
US 20220207304A1 · Hashimoto · 2022 [cited by examiner]
US 20240005209A1 · Schmidt · 2024 [cited by examiner]
US 20250307703A1 · Kakizaki · 2025 [cited by examiner]
CN 113033822 · 2021 [cited by examiner]
CN 113033822A · 2021 [cited by applicant]
CN 113326356A · 2021 [cited by applicant]
Seyed-Mohsen Moosavi-Dezfooli; Alhussein Fawzi; Universal Adversarial Perturbations, 2017, IEEE Conference on Computer Vision and Pattern Recognition, vol. 2. 2017, pp. 1-9, 2017. [cited by examiner]
He et al. (“Adversarial Personalized Ranking for Recommendation”, Published: Jun. 27, 2018, p. 355-364) (Year: 2018). [cited by examiner]
Ballet et al., “Imperceptible Adversarial Attacks on Tabular Data,” Neurips 2019 Workshop on Robust AI in Financial Services: Data, Fairness, Explainability, Trustworthiness and Privacy (Robust AI in FS 2019), 2019, pp.… [cited by applicant]
Brendel et al., “Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models,” Sixth International Conference on Learning Representations (ICLR 2018), 2018, pp. 1-12. <https://arxiv.or… [cited by applicant]