IP Library › Granted Patent US 12,524,574
Granted Patent B2
US 12,524,574 · App. 17/937,242 · Granted Jan 13, 2026

Defense against XAI adversarial attacks by detection of computational resource footprints

Inventors: Iam Palatnik de Sousa (Rio de Janeiro, BR); Adriana Bechara Prado (Niterói, BR)
Assignee: Dell Products L.P.
G06F21/64G06F21/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,574
App. No.
17/937,242
Granted
Jan 13, 2026
Kind
B2
Abstract

One example method includes initiating an audit of a machine learning model, providing input data to the machine learning model as part of the audit, while the audit is running, receiving information regarding operation of the machine learning model, wherein the information comprises a computational resource footprint, analyzing the computational resource footprint, and determining, based on the analyzing, that the computational resource footprint is characteristic of an adversarial attack on the machine learning model.

Claims (36)

1 . A method, comprising:

initiating an audit of a machine learning model for deployment;

providing input data to the machine learning model as part of the audit;

while the audit is running, receiving information regarding operation of the machine learning model, wherein the information comprises a computational resource footprint;

analyzing the computational resource footprint to calculate a deviation of the computational resource footprint from baseline parameters; and

in a case where the deviation is greater than known benchmarks:

asking the machine learning model for access to architectures internals; and

in a case where the access is denied, determining that the computational resource footprint is characteristic of an adversarial attack on the machine learning model and marking the machine learning model as unacceptable,

wherein the computational resource footprint indicates an extent to which the adversarial attack has used, or attempted to use, computing resources including communication bandwidth, data storage, and memory.

2 . The method as recited in claim 1 , wherein the audit comprises an explainable artificial intelligence audit.

3 . The method as recited in claim 1 , wherein the adversarial attack is associated with an adversary that controls the machine learning model.

4 . The method as recited in claim 1 , wherein the computational resource footprint comprises information indicating a number of function calls made by the adversarial attack.

5 . The method as recited in claim 1 , wherein the adversarial attack changes an output of the machine learning model from what the output would have been had the adversarial attack not occurred.

6 . The method as recited in claim 1 , wherein the adversarial attack introduces a bias into an output generated by the machine learning model.

7 . The method as recited in claim 1 , wherein the audit is performed without knowledge of a structure or operation of the machine learning model.

8 . The method as recited in claim 1 , wherein the audit comprises a gradient method explainable artificial intelligence audit.

9 . The method as recited in claim 1 , wherein the audit comprises a perturbation method explainable artificial intelligence audit.

10 . The method as recited in claim 1 , wherein the adversarial attack cannot be optimized in a single pass through the machine learning model and/or the adversarial attack relies for its effectiveness on second order gradients of an output of the machine learning model as a function of the input data.

11 . A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising:

initiating an audit of a machine learning model for deployment;

providing input data to the machine learning model as part of the audit;

while the audit is running, receiving information regarding operation of the machine learning model, wherein the information comprises a computational resource footprint;

analyzing the computational resource footprint to calculate a deviation of the computational resource footprint from baseline parameters; and

in a case where the deviation is greater than known benchmarks:

asking the machine learning model for access to architectures internals; and

in a case where the access is denied, determining that the computational resource footprint is characteristic of an adversarial attack on the machine learning model and marking the machine learning model as unacceptable,

wherein the computational resource footprint indicates an extent to which the adversarial attack has used, or attempted to use, computing resources including communication bandwidth, data storage, and memory.

12 . The non-transitory storage medium as recited in claim 11 , wherein the audit comprises an explainable artificial intelligence audit.

13 . The non-transitory storage medium as recited in claim 11 , wherein the adversarial attack is associated with an adversary that controls the machine learning model.

14 . The non-transitory storage medium as recited in claim 11 , wherein the computational resource footprint comprises information indicating a number of function calls made by the adversarial attack.

15 . The non-transitory storage medium as recited in claim 11 , wherein the adversarial attack changes an output of the machine learning model from what the output would have been had the adversarial attack not occurred.

16 . The non-transitory storage medium as recited in claim 11 , wherein the adversarial attack introduces a bias into an output generated by the machine learning model.

17 . The non-transitory storage medium as recited in claim 11 , wherein the audit is performed without knowledge of a structure or operation of the machine learning model.

18 . The non-transitory storage medium as recited in claim 11 , wherein the audit comprises a gradient method explainable artificial intelligence audit.

19 . The non-transitory storage medium as recited in claim 11 , wherein the audit comprises a perturbation method explainable artificial intelligence audit.

20 . The non-transitory storage medium as recited in claim 11 , wherein the adversarial attack cannot be optimized in a single pass through the machine learning model and/or the adversarial attack relies for its effectiveness on second order gradients of an output of the machine learning model as a function of the input data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2022
From: PALATNIK DE SOUSA, IAM; PRADO, ADRIANA BECHARA
To: DELL PRODUCTS L.P.
Reel/Frame 061276/0721 →
Continuity (1)
Related Publication 20240111902A1 · Apr 4, 2024
References Cited (13)
US 10783401B1 · Jiang · 2020 [cited by examiner]
US 10839268B1 · Ardulov · 2020 [cited by examiner]
US 20210097176A1 · Mathews · 2021 [cited by examiner]
US 20210319784A1 · Le Roux · 2021 [cited by examiner]
US 20220012591A1 · Dalli · 2022 [cited by examiner]
US 20220114399A1 · Castiglione · 2022 [cited by examiner]
US 20220156376A1 · dos Santos Silva · 2022 [cited by examiner]
Adaptive iterative attack towards explainable adversarial robustness (Year: 2020). [cited by examiner]
Adversarial attacks and defenses in XAI (Year: 2023). [cited by examiner]
Explaining Vulnerabilities to Adversarial Machine Learning through Visual Analytics (Year: 2020). [cited by examiner]
Detect Adversarial Attacks Against Deep Neural Networks With GPU Monitoring (Year: 2021). [cited by examiner]
Dombrowski, A. K., Alber, M., Anders, C., Ackermann, M., Müller, K. R., & Kessel, P. (Sep. 2019 ), “Explanations can be manipulated and geometry is to blame,” Advances in Neural Information Processing Systems, 32. [cited by applicant]
Slack, D., Hilgard, S., Jia, E., Singh, S., & Lakkaraju, H. (Feb. 2020), “Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods,” In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society… [cited by applicant]