IP Library › Granted Patent US 12,481,793
Granted Patent B2
US 12,481,793 · App. 18/147,757 · Granted Nov 25, 2025

System and method for proactively identifying poisoned training data used to train artificial intelligence models

Inventors: Ofir Ezrielev (Be'er Sheva, IL); Amihai Savir (Newton, MA); Tomer Kushnir (Omer, IL)
Assignee: Dell Products L.P.
G06F21/64
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,793
App. No.
18/147,757
Granted
Nov 25, 2025
Kind
B2
Abstract

Methods and systems for identifying poisoned training data used for training artificial intelligence (AI) models are disclosed. To identify poisoned training data in a proposed training dataset, a causal model may be obtained. The causal model may include relationships relating data elements. The proposed training dataset may be identified as poisoned when data elements within the proposed training dataset do not satisfy the relationships set forth by the causal model. When the identification of poisoned training data is made, the AI model may not be updated using the proposed training dataset and the proposed training dataset may be discarded. If poisoned training data is not identified prior to training an AI model, methods and systems are disclosed for the remediation of the poisoned training dataset and subsequent tainted AI models. By doing so, the effect of poisoned training data may be prevented and/or efficiently computationally mitigated.

Claims (73)

1 . A method for identifying poisoned training data used for training an artificial intelligence (AI) model, comprising:

making an identification that a second training dataset usable to retrain a first instance of the AI model is available; and

determining, based on the identification and before using the second training dataset to retrain the first instance of the AI model, whether the second training dataset is poisoned by, at least:

identifying a causal relationship within a first training dataset, the first training dataset being used to train the first instance of the AI model, and

identifying a third variable and a fourth variable from a second plurality of variables of the second training dataset, the third variable and the fourth variable being analogous to a first variable and a second variable of the first training dataset, respectively; and

transforming data elements of the third variable based on the causal relationship to obtained transformed data elements of the third variable

comparing the transformed data elements of the third variable to data elements of the fourth variable to make a determination regarding whether the third variable and the fourth variable satisfy the causal relationship,

in a first instance of the determination where the third variable and the fourth variable satisfy the causal relationship:

treating the second training dataset as not being poisoned, and

retraining the first instance of AI model using the second training dataset to obtain a second instance of the of the AI model; and

in a second instance of the determination where the third variable and the fourth variable do not satisfy the causal relationship:

treating the second training dataset as being poisoned.

2 . The method of claim 1 , further comprising, after making the identification and before using the second training dataset to retrain the first instance of the AI model:

obtaining the first instance of the AI model; and

obtaining a causal model based on the first training dataset, the causal model comprising the causal relationship.

3 . The method of claim 1 , wherein identifying the causal relationship within the first training dataset comprises:

identifying a first variable from a first plurality of variables of the first training dataset and a second variable from the first plurality of variables of the first training dataset; and

reading the causal relationship from a causal model, the causal relationship defining a functional relationship between the first variable and the second variable.

4 . The method of claim 1 , wherein treating the second training dataset as being poisoned comprises passing on a retraining opportunity for the AI model presented by the second training dataset.

5 . The method of claim 4 , wherein passing on the retraining opportunity comprises discarding the second training dataset.

6 . The method of claim 1 , wherein the causal relationship is identified using a causal model comprising nodes and edges between the nodes, the nodes correspond to portions of the first training data set, and the edges between the nodes correspond to relationships between the portions of the first training data set corresponding to the nodes.

7 . A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for identifying poisoned training data used for training an artificial intelligence (AI) model, the operations comprising:

making an identification that a second training dataset usable to retrain a first instance of the AI model is available; and

determining, based on the identification and before using the second training dataset to retrain the first instance of the AI model, whether the second training dataset is poisoned by, at least:

identifying a causal relationship within a first training dataset, the first training dataset being used to train the first instance of the AI model, and

identifying a third variable and a fourth variable from a second plurality of variables of the second training dataset, the third variable and the fourth variable being analogous to a first variable and a second variable of the first training dataset, respectively; and

transforming data elements of the third variable based on the causal relationship to obtained transformed data elements of the third variable

comparing the transformed data elements of the third variable to data elements of the fourth variable to make a determination regarding whether the third variable and the fourth variable satisfy the causal relationship,

in a first instance of the determination where the third variable and the fourth variable satisfy the causal relationship:

treating the second training dataset as not being poisoned, and

retraining the first instance of AI model using the second training dataset to obtain a second instance of the of the AI model; and

in a second instance of the determination where the third variable and the fourth variable do not satisfy the causal relationship:

treating the second training dataset as being poisoned.

8 . The non-transitory machine-readable medium of claim 7 , the operations further comprising:

obtaining the first instance of the AI model; and

obtaining a causal model based on the first training dataset, the causal model comprising the causal relationship.

9 . The non-transitory machine-readable medium of claim 7 , wherein identifying the causal relationship within the first training dataset comprises:

identifying a first variable from a first plurality of variables of the first training dataset and a second variable from the first plurality of variables of the first training dataset; and

reading the causal relationship from a causal model, the causal relationship defining a functional relationship between the first variable and the second variable.

10 . The non-transitory machine-readable medium of claim 7 , wherein treating the second training dataset as being poisoned comprises passing on a retraining opportunity for the AI model presented by the second training dataset.

11 . The non-transitory machine-readable medium of claim 10 , wherein passing on the retraining opportunity comprises discarding the second training dataset.

12 . The non-transitory machine-readable medium of claim 7 , wherein the causal relationship is identified using a causal model comprising nodes and edges between the nodes, the nodes correspond to portions of the first training data set, and the edges between the nodes correspond to relationships between the portions of the first training data set corresponding to the nodes.

13 . The non-transitory machine-readable medium of claim 7 , wherein making the identification comprises, at least:

determining that a predetermined amount of new training data for retraining the first instance of the AI model has been collected from data sources; and

using the predetermined amount of the new training data as the second training dataset.

14 . A data processing system, comprising:

a processor; and

a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for identifying poisoned training data used for training an artificial intelligence (AI) model, the operations comprising:

making an identification that a second training dataset usable to retrain a first instance of the AI model is available, and

determining, based on the identification and before using the second training dataset to retrain the first instance of the AI model, whether the second training dataset is poisoned by, at least:

identifying a causal relationship within a first training dataset, the first training dataset being used to train the first instance of the AI model, and

identifying a third variable and a fourth variable from a second plurality of variables of the second training dataset, the third variable and the fourth variable being analogous to a first variable and a second variable of the first training dataset, respectively; and

transforming data elements of the third variable based on the causal relationship to obtained transformed data elements of the third variable

comparing the transformed data elements of the third variable to data elements of the fourth variable to make a determination regarding whether the third variable and the fourth variable satisfy the causal relationship,

in a first instance of the determination where the third variable and the fourth variable satisfy the causal relationship:

treating the second training dataset as not being poisoned, and

retraining the first instance of AI model using the second training dataset to obtain a second instance of the of the AI model; and

in a second instance of the determination where the third variable and the fourth variable satisfy the causal relationship:

treating the second training dataset as being poisoned.

15 . The data processing system of claim 14 , the operations further comprising:

obtaining the first instance of the AI model; and

obtaining a causal model based on the first training dataset, the causal model comprising the causal relationship.

16 . The data processing system of claim 14 , wherein identifying the causal relationship within the first training dataset comprises:

identifying a first variable from a first plurality of variables of the first training dataset and a second variable from the first plurality of variables of the first training dataset; and

reading the causal relationship from a causal model, the causal relationship defining a functional relationship between the first variable and the second variable.

17 . The data processing system of claim 14 , wherein treating the second training dataset as being poisoned comprises passing on a retraining opportunity for the AI model presented by the second training dataset.

18 . The data processing system of claim 17 , wherein passing on the retraining opportunity comprises discarding the second training dataset.

19 . The method of claim 1 , wherein making the identification comprises, at least:

determining that a predetermined amount of new training data for retraining the first instance of the AI model has been collected from data sources; and

using the predetermined amount of the new training data as the second training dataset.

20 . The data processing system of claim 14 , wherein making the identification comprises, at least:

determining that a predetermined amount of new training data for retraining the first instance of the AI model has been collected from data sources; and

using the predetermined amount of the new training data as the second training dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2022
From: EZRIELEV, OFIR; SAVIR, AMIHAI; KUSHNIR, TOMER
To: DELL PRODUCTS L.P.
Reel/Frame 062244/0990 →
Continuity (1)
Related Publication 20240220663A1 · Jul 4, 2024
References Cited (71)
US 10936969B2 · Patel et al. · 2021 [cited by applicant]
US 11087170B2 · Malaya · 2021 [cited by examiner]
US 11487963B2 · Angel · 2022 [cited by examiner]
US 11544501B2 · Dong · 2023 [cited by examiner]
US 11636726B2 · Purohit · 2023 [cited by applicant]
US 11645515B2 · Angel · 2023 [cited by examiner]
US 11689566B2 · Baracaldo-Angel · 2023 [cited by examiner]
US 11785024B2 · Karam · 2023 [cited by applicant]
US 11797672B1 · Beveridge · 2023 [cited by applicant]
US 11829193B2 · Shukla · 2023 [cited by examiner]
US 11847217B2 · Healy · 2023 [cited by applicant]
US 11921903B1 · Beveridge · 2024 [cited by applicant]
US 12032541B2 · Hasabnis · 2024 [cited by applicant]
US 12126640B2 · Woodworth · 2024 [cited by applicant]
US 12143405B2 · Chen Kaidi · 2024 [cited by applicant]
US 20170177860A1 · Suarez · 2017 [cited by applicant]
US 20180255023A1 · Whaley · 2018 [cited by applicant]
US 20190377873A1 · Murphy · 2019 [cited by applicant]
US 20190384790A1 · Bequet · 2019 [cited by applicant]
US 20200019821A1 · Baracaldo-Angel · 2020 [cited by applicant]
US 20200050945A1 · Chen · 2020 [cited by examiner]
US 20200057857A1 · Roytman · 2020 [cited by applicant]
US 20200082097A1 · Poliakov · 2020 [cited by examiner]
US 20200082270A1 · Gu · 2020 [cited by examiner]
US 20200134374A1 · Oros · 2020 [cited by examiner]
US 20200244674A1 · Arzani · 2020 [cited by applicant]
US 20210073685A1 · Veshchikov · 2021 [cited by applicant]
US 20210081831A1 · Angel · 2021 [cited by applicant]
US 20210209512A1 · Gaddam · 2021 [cited by examiner]
US 20210303695A1 · Grosse · 2021 [cited by applicant]
US 20210374247A1 · Sultana · 2021 [cited by applicant]
US 20210398020A1 · Ahmad et al. · 2021 [cited by applicant]
US 20220166782A1 · Zoldi · 2022 [cited by applicant]
US 20220179840A1 · Chatterjee · 2022 [cited by applicant]
US 20220368706A1 · Tang · 2022 [cited by applicant]
US 20220414492A1 · Jezewski · 2022 [cited by applicant]
US 20230004654A1 · Jurzak · 2023 [cited by applicant]
US 20230079112A1 · Cheruvu · 2023 [cited by applicant]
US 20230134218A1 · Semenov · 2023 [cited by applicant]
US 20230148116A1 · Stokes, III · 2023 [cited by applicant]
US 20230164162A1 · Lee · 2023 [cited by applicant]
US 20230222385A1 · Shimizu · 2023 [cited by applicant]
US 20230274003A1 · Liu · 2023 [cited by applicant]
US 20230274192A1 · Wang · 2023 [cited by applicant]
US 20240015019A1 · Sneider · 2024 [cited by applicant]
US 20240020580A1 · Brower · 2024 [cited by applicant]
US 20240048977A1 · Marzban · 2024 [cited by applicant]
US 20240119153A1 · Ludmir · 2024 [cited by applicant]
US 20240364534A1 · Ezrielev · 2024 [cited by applicant]
US 20250053664A1 · Cameron · 2025 [cited by applicant]
US 20250055762A1 · Walker · 2025 [cited by applicant]
WO 2020040777A1 · 2020 [cited by applicant]
WO 2021213626A1 · 2021 [cited by applicant]
WO 2022216142A1 · 2022 [cited by applicant]
WO 2023111287A1 · 2023 [cited by applicant]
Anastasovski, Goce, “Classification of Malicious Web Traffic” (2013), Graduate Theses, Dissertations, and Problem Reports 153 (118 Pages). [cited by applicant]
Joshi, Naveen, “Is The Data Used For Training Your Machine Learning Model Safe?”, Technology For You, Jul. 28, 2022, <https://www.technologyforyou.org/is-the-data-used-for-training-your-machine-learning-model-safe/> (3 … [cited by applicant]
Wang, Siruo et al., “Methods for correcting inference based on outcomes predicted by machine learning.” Proceedings of the National Academy of Sciences 117.48 (2020): 30266-30275. (10 Pages). [cited by applicant]
Rauschmayr, Nathalie et al., “Detecting and analyzing incorrect model predictions with Amazon SageMaker Model Monitor and Debugger”, AWS Machine Learning Blog, Jul. 9, 2020, <https://aws.amazon.com/blogs/machine-learnin… [cited by applicant]
Higgins, Kelly Jackson, “Honeypot Stings Attackers With Counterattacks”, Dark Reading, Mar. 26, 2013, <https://www.darkreading.com/vulnerabilities-threats/honeypot-stings-attackers-with-counterattacks> (4 Pages). [cited by applicant]
Susmelj, Igor, “The Data You Don't Need: Removing Redundant Samples”, Towards Data Science, Mar. 19, 2020, <https://towardsdatascience.com/the-data-you-don-t-need-removing-redundant-samples-6bfd07c1516c> (10 Pages). [cited by applicant]
Paduraru, Ciprian, Marius-Constantin Melemciuc, and Bogdan Ghimis. “Fuzz Testing with Dynamic Taint Analysis based Tools for Faster Code Coverage.” ICSOFT 19 (2019): 82-93. (Year: 2019). [cited by applicant]
Jiang, Bingchen, and Zhao Li. “Defending Against Backdoor Attack on Graph Nerual Network by Explainability.” arXiv preprint arXiv: 2209.02902, 10 pages, (Year: 2022). [cited by applicant]
Raghavan, Vijay, Thomas Mazzuchi, and Shahram Sarkani. “Discover Artificial Intelligence: An improved real time detection of data poisoning attacks in deep learning vision systems”, 17 pages, Discover 2022, (Year: 2022). [cited by applicant]
Albert Cheng, “The Machine Learning Minefield—How to Avoid Getting Hit by Machine Learning Poisoning”, Mar. 22, 2022, retrieved from <https://ayc-data.com/data_science/2022/03/22/data-poisoning.html> on May 1, 2025 (10 … [cited by applicant]
Zhang et al., “FL Detector: Defending Federated Learning Against Model Poisoning Attacks via Detecting Malicious Clients”, Available at https://arxiv.org/abs/2207.092009 (Year: 2022), (11 pages). [cited by applicant]
Tran et al., “Manipulating Machine Learning Poisoning Attacks and Countermeasures for Regression Learning”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montreal, Canada; 2018, pp. 1-11 (Year… [cited by applicant]
Zeng et al., “CNNComparator: Comparative Analytics of Convolutional Neural Networks”, arXiv: 1710.05285v1 [cs.LG] Oct. 15, 2017, pp. 1-5 (Year: 2017), (5 pages). [cited by applicant]
Hendrycks et al., “Natural Adversarial Examples”, arXiv:1907.07174v4[cs.LG] Mar. 4, 2021; pp. 1-16 (Year:2021), (16 pages). [cited by applicant]
Xu et al., “Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks”, In Network and Distributed Systems Security Symposium (NDSS) 2018, San Diego, Feb. 2018; arXiv:1704.01155v2 [cs.CV] Dec. 5, 2017; p… [cited by applicant]
Lao; “Reorienting Machine Learning Education Towards Tinkerers and ML-Engaged Citizens”, Doctoral Dissertation; Massachusetts Institute of Technology, Department of Electrical Engineering and Computer Science; 2020; pp.… [cited by applicant]
Cited By (1)
US 12,699,882