IP Library Granted Patent US 12,388,860
Granted Patent B2
US 12,388,860 · App. 18/174,342 · Granted Aug 12, 2025

AI-based trojans for evading machine learning detection

Inventors: Prabhat Kumar Mishra (Gainesville, FL); Zhixin Pan (Gainesville, FL)
Assignee: University of Florida Research Foundation, Incorporated
H04L63/1433H04L41/16H04L63/145
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,388,860
App. No.
18/174,342
Granted
Aug 12, 2025
Kind
B2
Abstract

Various embodiments provide a robust backdoor attack on machine learning (ML)-based detection systems that can be applied to demonstrate and identify vulnerabilities thereof. In various embodiments, an artificial intelligence (AI)-based Trojan attack is generated and implanted inside a ML model trained for classification and/or detection tasks, and the AI-based Trojan attack can be triggered by specific inputs to manipulate the expected outputs of the ML model. Analysis of the behavior of an ML model having the AI-based Trojan implanted (and/or triggered) then enables identification of vulnerabilities of the ML model and further enables the design of ML models with improved security. Various embodiments of the present disclosure provide a fast and cost-effective solution in achieving 100% attack success rate that significantly outperforms adversarial attacks on ML models, thereby improving applicability and depth in testing ML-based detection systems.

Claims (41)

1. A computer-implemented method comprising:

generating, by one or more processors and using a benign machine learning model, a first classification of an input circuit with respect to a hardware-based cyberattack;

generating, by the one or more processors and using a payload machine learning model, a perturbation value for the first classification based on a trigger condition, wherein:

(i) the payload machine learning model is trained based on a plurality of backdoor features and

(ii) the trigger condition corresponds to a backdoor feature of the plurality of backdoor features;

providing, by the one or more processors, a second classification as an output of an outsourced backdoored machine learning model, wherein the outsourced backdoored machine learning model (i) comprises the benign machine learning model inserted with the payload machine learning model and (ii) is trained to generate the second classification based on the first classification and the perturbation value; and

determining, by the one or more processors, a vulnerability associated with the outsourced backdoored machine learning model based on the second classification.

2. The computer-implemented method of claim 1 , further comprising:

performing a vulnerability analysis for the benign machine learning model based at least in part on providing an input circuit sample comprising a plurality of input backdoor features to the outsourced backdoored machine learning model.

3. The computer-implemented method of claim 2 , wherein the plurality of input backdoor features comprises a modification of at least one logic gate of an original circuit sample associated with the input circuit sample, the modification having logical equivalence with the at least one logic gate.

4. The computer-implemented method of claim 1 , wherein the payload machine learning model is trained to output zero values responsive to the plurality of backdoor features not satisfying one or more constraints.

5. The computer-implemented method of claim 1 , wherein the benign machine learning model is trained according to supervised learning to generate a classification output based at least in part on a plurality of normal features extracted from a plurality of circuit samples.

6. The computer-implemented method of claim 5 , wherein the plurality of normal features and the plurality of backdoor features are orthogonally-variant with respect to each other.

7. The computer-implemented method of claim 1 , wherein an output layer of the benign machine learning model is combined with an output layer of the payload machine learning model.

8. The computer-implemented method of claim 1 , wherein at least one of one or more nodes or one or more edges associated with the benign machine learning model is merged with at least one of one or more respective nodes or one or more respective edges associated with the payload machine learning model.

9. The computer-implemented method of claim 1 , wherein the outsourced backdoored machine learning model comprises a partially-outsourced backdoored machine learning model that is configured to be retrainable.

10. One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

generating, using a benign machine learning model, a first classification of an input circuit with respect to a hardware-based cyberattack;

generating, using a payload machine learning model, a perturbation value for the first classification based on a trigger condition, wherein:

(i) the payload machine learning model is trained based on a plurality of backdoor features and

(ii) the trigger condition corresponds to a backdoor feature of the plurality of backdoor features;

providing a second classification as an output of an outsourced backdoored machine learning model, wherein the outsourced backdoored machine learning model (i) comprises the benign machine learning model inserted with the payload machine learning model and (ii) is trained to generate the second classification based on the first classification and the perturbation value; and

determining a vulnerability associated with the outsourced backdoored machine learning model based on the second classification.

11. The one or more non-transitory computer-readable storage media of claim 10 , wherein the operations further comprise:

performing a vulnerability analysis for the benign machine learning model based at least in part on providing an input circuit sample comprising a plurality of input backdoor features to the outsourced backdoored machine learning model.

12. The one or more non-transitory computer-readable storage media of claim 11 , wherein the plurality of input backdoor features comprises a modification of at least one logic gate of an original circuit sample associated with the input circuit sample, the modification having logical equivalence with the at least one logic gate.

13. The one or more non-transitory computer-readable storage media of claim 10 , wherein the benign machine learning model is trained according to supervised learning to generate a classification output based at least in part on a plurality of normal features extracted from a plurality of circuit samples.

14. The one or more non-transitory computer-readable storage media of claim 13 , wherein the plurality of normal features and the plurality of backdoor features are orthogonally-variant with respect to each other.

15. The one or more non-transitory computer-readable storage media of claim 10 , wherein an output layer of the benign machine learning model is combined with an output layer of the payload machine learning model.

16. The one or more non-transitory computer-readable storage media of claim 10 , wherein at least one of one or more nodes or one or more edges associated with the benign machine learning model is merged with at least one of one or more respective nodes or one or more respective edges associated with the payload machine learning model.

17. The one or more non-transitory computer-readable storage media of claim 10 , wherein the outsourced backdoored machine learning model comprises a partially-outsourced backdoored machine learning model that is configured to be retrainable.

18. A system comprising one or more processors, a memory, and one or more programs stored in the memory, the one or more programs comprising instructions configured to cause the one or more processors to:

generate, using a benign machine learning model, a first classification of an input circuit with respect to a hardware-based cyberattack;

generate, using a payload machine learning model, a perturbation value for the first classification based on a trigger condition, wherein:

(i) the payload machine learning model is trained based on a plurality of backdoor features and

(ii) the trigger condition corresponds to a backdoor feature of the plurality of backdoor features;

provide a second classification as an output of an outsourced backdoored machine learning model, wherein the outsourced backdoored machine learning model (i) comprises the benign machine learning model inserted with the payload machine learning model and (ii) is trained to generate the second classification based on the first classification and the perturbation value; and

determine a vulnerability associated with the outsourced backdoored machine learning model based on the second classification.

19. The system of claim 18 , wherein the one or more programs further comprise instructions that are configured to cause the one or more processors to:

perform a vulnerability analysis for the benign machine learning model based at least in part on providing an input circuit sample comprising a plurality of input backdoor features to the outsourced backdoored machine learning model.

20. The system of claim 18 , wherein the outsourced backdoored machine learning model comprises a partially-outsourced backdoored machine learning model that is configured to be retrainable.

Assignments (2)
CONFIRMATORY LICENSE Recorded Feb 13, 2025
From: UNIVERSITY OF FLORIDA
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 070205/0146 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2023
From: MISHRA, PRABHAT KUMAR; PAN, ZHIXIN
To: UNIVERSITY OF FLORIDA RESEARCH FOUNDATION, INCORPORATED
Reel/Frame 062899/0349 →
Continuity (2)
Provisional Application 63315219 · Mar 1, 2022
Related Publication 20230421596A1 · Dec 28, 2023
References Cited (17)
US 10225277B1 · Shintre · 2019 [cited by examiner]
US 20200380123A1 · Reimann · 2020 [cited by examiner]
US 20210204011A1 · Jain · 2021 [cited by examiner]
US 20220030013A1 · Crouch · 2022 [cited by examiner]
US 20220164437A1 · Crouch · 2022 [cited by examiner]
US 20230274003A1 · Liu · 2023 [cited by examiner]
Zhixin, Pan Z et al. “Automated Test Generation For Hardware Trojan Detection Using Reinforcement Learning”, [cited by applicant]
Gu, Tianyu et al. “BadNets: Evaluating Backdooring Attacks On Deep Neural Networks”, [cited by applicant]
Liu, Kang et al. “Fine-Pruning: Defending Against Backdooring Attacks On Deep Neural Networks”, arXiv preprint arXiv: 1805.12185vl [cs.CR], May 30, 2018, (21 pages), available online: <URL: https://arxiv.org/pdf/1805.12… [cited by applicant]
Chen, Xiaoming et al. “Hardware Trojan Detection In Third-Party Digital Intellectual Property Cores By Multilevel Feature Analysis”, [cited by applicant]
Lyu, Yangdi et al. “MaxSense: Side-Channel Sensitivity Maximization For Trojan Detection Using Statistical Test Patterns,” [cited by applicant]
Wang, Bolun et al. “Neural Cleanse: Identifying and Mitigating Backdoor Attacks In Neural Networks”, In 2019 IEEE Symposium on Security and Privacy (SP), May 19, 2019, (17 pages), IEEE, available online: <URL: https://p… [cited by applicant]
Lyu, Yangdi et al. “Scalable Activation Of Rare Triggers In Hardware Trojans By Repeated Maximal Clique Sampling”, [cited by applicant]
Gao, Yansong et al. “STRIP: A Defence Against Trojan Attacks On Deep Neural Networks”, [cited by applicant]
Pan, Zhixin et al. “Test Generation Using Reinforcement Learning For Delay-Based Side-Channel Analysis”, [cited by applicant]
Hasegawa, Kento et al. “Trojan-Feature Extraction At Gage-Level Netlists and Its Application To Hardware-Trojan Detection Using Random Forest Classifier”, [cited by applicant]
Liu, Yingqi et al. “Trojaning Attack On Neural Networks”, Purdue University, [cited by applicant]
Cited By (1)
US 12,737,463