IP Library › Granted Patent US 12,361,174
Granted Patent B2
US 12,361,174 · App. 18/459,122 · Granted Jul 15, 2025

Detecting possible attacks on artificial intelligence models using strengths of causal relationships in training data

Inventors: Ofir Ezrielev (Beer Sheva, IL); Tomer Kushnir (Omer, IL); Amihai Savir (Newton, MA)
Assignee: Dell Products L.P.
G06F21/64G06F21/55G06F21/56G06N20/20H04L63/1408H04L63/1425
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,174
App. No.
18/459,122
Granted
Jul 15, 2025
Kind
B2
Abstract

Methods and systems for managing artificial intelligence (AI) models are disclosed. To manage AI models, an instance of an AI model may not be re-trained using training data determined to be potentially poisoned. By doing so, malicious attacks intending to influence the AI model using poisoned training data may be prevented. To do so, a first level of strength of a first causal relationship present in historical training data may be compared to a second level of strength of a second causal relationship present in a candidate training data set. The first level of strength and the second level of strength may be expected to be similar within a threshold. If a difference between the first level of strength and the second level of strength is not within the threshold, the candidate training data may be treated as including poisoned training data.

Claims (52)

1. A method of managing an artificial intelligence (AI) model, performed by one or more hardware processors, the method comprising:

obtaining, by the one or more hardware processors through a communication system, a candidate training data set usable to update an instance of the AI model;

identifying a historical training data set, the historical training data set being obtained prior to the candidate training data set and the historical training data set already having been considered as trustworthy;

obtaining a quantification of a difference between levels of strengths of similar causal relationships in the candidate training data set and the historical training data set, wherein a level of strength of the levels of strengths is based on a goodness of fit of a function to a portion of the historical training data set, the function defining a causal relationship of the similar causal relationships;

making a determination regarding whether the quantification is within a threshold for the quantification;

in a first instance of the determination in which the quantification is within the threshold: obtaining a second instance of the AI model using at least the candidate training data set; and

in a second instance of the determination in which the quantification is not within the threshold: treating the candidate training data set as comprising poisoned training data.

2. The method of claim 1 , wherein obtaining the quantification comprises:

identifying a first causal relationship of the similar causal relationships in the historical training data set;

identifying a first level of strength of the first causal relationship;

identifying a second causal relationship of the similar causal relationships in the candidate training data set; and

identifying a second level of strength of the second causal relationship.

3. The method of claim 2 , wherein the first causal relationship and the second causal relationship relate same features and same labels.

4. The method of claim 3 , wherein the first causal relationship is based on a first feature present in the historical training data set and a first label present in the historical training data set.

5. The method of claim 4 , wherein the second causal relationship is based on a second feature present in the historical training data set and a second label present in the historical training data set.

6. The method of claim 5 , wherein the first feature is based on first measurements of a quantity during a first period of time and the second feature is based on second measurements of the quantity during a second period of time, the first period of time being prior to the second period of time.

7. The method of claim 6 , wherein the first label is based on third measurements of a second quantity during the first period of time and the second label is based on fourth measurements of the second quantity during the second period of time.

8. The method of claim 1 , wherein the quantification of the difference is based on the goodness of fit and a second goodness of a second fit of a second function to a second portion of the candidate training data set.

9. The method of claim 1 , wherein the threshold is based on a level of tolerance for use of poisoned training data in the AI model.

10. A non-transitory machine-readable medium having instructions stored therein, which when executed by a hardware processor, cause the hardware processor to perform operations for managing an artificial intelligence (AI) model, the operations comprising:

obtaining, through a communication system, a candidate training data set usable to update an instance of the AI model;

identifying a historical training data set, the historical training data set being obtained prior to the candidate training data set and the historical training data set already having been considered as trustworthy;

obtaining a quantification of a difference between levels of strengths of similar causal relationships in the candidate training data set and the historical training data set, wherein a level of strength of the levels of strengths is based on a goodness of fit of a function to a portion of the historical training data set, the function defining a causal relationship of the similar causal relationships;

making a determination regarding whether the quantification is within a threshold for the quantification;

in a first instance of the determination in which the quantification is within the threshold: obtaining a second instance of the AI model using at least the candidate training data set; and

in a second instance of the determination in which the quantification is not within the threshold: treating the candidate training data set as comprising poisoned training data.

11. The non-transitory machine-readable medium of claim 10 , wherein obtaining the quantification comprises:

identifying a first causal relationship of the similar causal relationships in the historical training data set;

identifying a first level of strength of the first causal relationship;

identifying a second causal relationship of the similar causal relationships in the candidate training data set; and

identifying a second level of strength of the second causal relationship.

12. The non-transitory machine-readable medium of claim 11 , wherein the first causal relationship and the second causal relationship relate same features and same labels.

13. The non-transitory machine-readable medium of claim 12 , wherein the first causal relationship is based on a first feature present in the historical training data set and a first label present in the historical training data set.

14. The non-transitory machine-readable medium of claim 13 , wherein the second causal relationship is based on a second feature present in the historical training data set and a second label present in the historical training data set.

15. The non-transitory machine-readable medium of claim 14 , wherein the first feature is based on first measurements of a quantity during a first period of time and the second feature is based on second measurements of the quantity during a second period of time, the first period of time being prior to the second period of time.

16. The non-transitory machine-readable medium of claim 15 , wherein the first label is based on third measurements of a second quantity during the first period of time and the second label is based on fourth measurements of the second quantity during the second period of time.

17. A data processing system, comprising:

a hardware processor; and

a memory coupled to the hardware processor to store instructions, which when executed by the hardware processor, cause the hardware processor to perform operations for managing an artificial intelligence (AI) model, the operations comprising:

obtaining, through a communication system, a candidate training data set usable to update an instance of the AI model;

identifying a historical training data set, the historical training data set being obtained prior to the candidate training data set and the historical training data set already having been considered as trustworthy;

obtaining a quantification of a difference between levels of strengths of similar causal relationships in the candidate training data set and the historical training data set, wherein a level of strength of the levels of strengths is based on a goodness of fit of a function to a portion of the historical training data set, the function defining a causal relationship of the similar causal relationships;

making a determination regarding whether the quantification is within a threshold for the quantification;

in a first instance of the determination in which the quantification is within the threshold: obtaining a second instance of the AI model using at least the candidate training data set; and

in a second instance of the determination in which the quantification is not within the threshold: treating the candidate training data set as comprising poisoned training data.

18. The data processing system of claim 17 , wherein obtaining the quantification comprises:

identifying a first causal relationship of the similar causal relationships in the historical training data set;

identifying a first level of strength of the first causal relationship;

identifying a second causal relationship of the similar causal relationships in the candidate training data set; and

identifying a second level of strength of the second causal relationship.

19. The data processing system of claim 18 , wherein the first causal relationship and the second causal relationship relate same features and same labels.

20. The data processing system of claim 19 , wherein the first causal relationship is based on a first feature present in the historical training data set and a first label present in the historical training data set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2023
From: EZRIELEV, OFIR; KUSHNIR, TOMER; SAVIR, AMIHAI
To: DELL PRODUCTS L.P.
Reel/Frame 064872/0215 →
Continuity (1)
Related Publication 20250077713A1 · Mar 6, 2025
References Cited (36)
US 6922816B1 · Amin et al. · 2005 [cited by applicant]
US 10936173B2 · Ubillos et al. · 2021 [cited by applicant]
US D917546S · Mayler et al. · 2021 [cited by applicant]
US 11030709B2 · McLinden et al. · 2021 [cited by applicant]
US 11449942B2 · Basu et al. · 2022 [cited by applicant]
US 11460997B2 · Homma et al. · 2022 [cited by applicant]
US 11783025B2 · Molloy · 2023 [cited by examiner]
US 20180293516A1 · Lavid Ben Lulu · 2018 [cited by examiner]
US 20210209512A1 · Gaddam · 2021 [cited by applicant]
US 20220027792A1 · Cummings · 2022 [cited by examiner]
US 20220078637A1 · Tullberg · 2022 [cited by examiner]
US 20220300857A1 · Lavid Ben Lulu · 2022 [cited by examiner]
US 20230030136A1 · Lancioni · 2023 [cited by examiner]
US 20230032822A1 · Wang · 2023 [cited by examiner]
US 20230198855A1 · Ganesan · 2023 [cited by examiner]
US 20230368074A1 · Mallya · 2023 [cited by examiner]
US 20240134972A1 · Boue · 2024 [cited by examiner]
US 20240185090A1 · Rafferty · 2024 [cited by applicant]
US 20240195826A1 · Mathews et al. · 2024 [cited by applicant]
US 20240250975A1 · Pourahmadi · 2024 [cited by examiner]
US 20240273184A1 · Yarabolu et al. · 2024 [cited by applicant]
US 20250013913A1 · Boué · 2025 [cited by examiner]
EP 4481632A1 · 2024 [cited by examiner]
KR 20250011365A · 2025 [cited by examiner]
WO 2020123985A · 2020 [cited by applicant]
WO WO2021167733A1 · 2021 [cited by examiner]
WO WO2022115623A1 · 2022 [cited by examiner]
Anastasovski, Goce, “Classification of Malicious Web Traffic” (2013). Graduate Theses, Dissertations, and Problem Reports. 153. (118 Pages). [cited by applicant]
Joshi, Naveen, “Is The Data Used for Training Your Machine Learning Model Safe?” Technology for You, Jul. 28, 2022, Web Page <https://www.technologyforyou.org/is-the-data-used-for-training-your-machine-learning-model-sa… [cited by applicant]
Di, Jimmy Z., et al. “Hidden poison: Machine unlearning enables camouflaged poisoning attacks.” NeurIPS ML Safety Workshop. 2022. (28 Pages). [cited by applicant]
Wang, Siruo, et al., “Methods for correcting inference based on outcomes predicted by machine learning.” Proceedings of the National Academy of Sciences 117.48 (2020): 30266-30275. (10 Pages). [cited by applicant]
Rauschmayr, Nathalie, et al., “Detecting and analyzing incorrect model predictions with Amazon SageMaker Model Monitor and Debugger,” Amazon Web Services, Jul. 9, 2020, Web Page <https://aws.amazon.com/blogs/machine-lea… [cited by applicant]
Jackson Higgins, Kelly, “Honeypot Stings Attackers With Counterattacks,” DarkReading, Mar. 26, 2013, Web Page <https://www.darkreading.com/vulnerabilities-threats/honeypot-stings-attackers-with-counterattacks> accessed … [cited by applicant]
Susmelj, Igor, “The Data You Don't Need: Removing Redundant Samples,” Towards Data Science, Mar. 19, 2020, Web Page <https://towardsdatascience.com/the-data-you-don-t-need-removing-redundant-samples-6bfd07c1516c> access… [cited by applicant]
Aghakhani, Hojjat, et al. “Bullseye polytope: A scalable clean-label poisoning attack with improved transferability.” 2021 IEEE European symposium on security and privacy (EuroS&P). IEEE, 2021. (20 Pages). [cited by applicant]
Zhibo Wang et al., “Threats to Training: A Survey of Poisoning Attacks and Defenses on Machine Learning Systems.” ACM Computing Surveys, vol. 55, No. 7, Article 134. Dec. 2022: pp. 1-36. (Year: 2022). [cited by applicant]