IP Library › Granted Patent US 12,524,677
Granted Patent B2
US 12,524,677 · App. 17/157,077 · Granted Jan 13, 2026

Generating unsupervised adversarial examples for machine learning

Inventors: Pin-Yu Chen (White Plains, NY); Chia-Yi Hsu (Taipei, TW); Songtao Lu (White Plains, NY); Sijia Liu (Somerville, MA); Chuang Gan (Cambridge, MA); Chia-Mu Yu (Hsinchu, TW)
Assignees: International Business Machines Corporation; NATIONAL CHUNG HSING UNIVERSITY
G06N3/088G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,677
App. No.
17/157,077
Granted
Jan 13, 2026
Kind
B2
Abstract

A trained machine learning model and a training dataset used to train the trained machine learning model can be received. Based on the training dataset, unsupervised adversarial examples can be generated. Robustness of the trained machine learning model can be determined using the generated unsupervised adversarial examples. The training dataset can be augmented with the generated unsupervised adversarial examples. The trained machine learning model can be retrained using the augmented training dataset.

Claims (38)

1 . A computer-implemented method comprising:

receiving a trained machine learning model and a training dataset used to train the trained machine learning model;

based on the training dataset and a loss function of the trained machine learning model, generating images representing adversarial examples, the adversarial examples being perturbed samples of the training dataset, the generating creating an adversarial example which is least similar to an original sample in the training dataset, and also satisfying an adversarial criterion that a loss associated with the adversarial example is less than that of the original sample in the training dataset;

determining robustness of the trained machine learning model using the generated adversarial examples;

augmenting the training dataset with the generated images representing adversarial examples;

retraining the trained machine learning model using the augmented training dataset;

performing by the retrained machine learning model, image classification; and

displaying the original sample, a generated image representing the adversarial example, and a reconstructed image reconstructed using the retrained machine learning model.

2 . The method of claim 1 , wherein the adversarial example is randomly sampled.

3 . The method of claim 1 , wherein the adversarial example is sampled using an output of a convolutional layer of the trained machine learning model.

4 . The method of claim 1 , wherein the generating of the adversarial examples includes solving a minmax algorithm which finds the adversarial example that has minimum training loss and least similarity to the original sample.

5 . The method of claim 1 , wherein the trained machine learning model includes a neural network model.

6 . The method of claim 1 , wherein the trained machine learning model includes an autoencoder.

7 . The method of claim 1 , wherein the trained machine learning model includes a representation learning model.

8 . The method of claim 1 , wherein the trained machine learning model includes a contrastive learning model.

9 . The method of claim 1 , wherein the trained machine learning model includes an unsupervised machine learning model.

10 . A computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions readable by a device to cause the device to:

receive a trained machine learning model and a training dataset used to train the trained machine learning model;

based on the training dataset and a loss function of the trained machine learning model, generate mages representing adversarial examples, the adversarial examples being perturbed samples of the training dataset, the generating creating an adversarial example which is least similar to an original sample in the training dataset, and also satisfying an adversarial criterion that a loss associated with the adversarial example is less than that of the original sample in the training dataset;

augment the training dataset with the generated images representing adversarial examples;

retrain the trained machine learning model using the augmented training dataset;

perform by the retrained machine learning model, image classification; and

display the original sample, a generated image representing the adversarial example, and a reconstructed image reconstructed using the retrained machine learning model.

11 . The computer program product of claim 10 , wherein the device is caused to determine robustness of the trained machine learning model using the generated adversarial examples.

12 . The computer program product of claim 10 , wherein the adversarial sample is randomly sampled.

13 . The computer program product of claim 10 , wherein the adversarial sample is sampled using an output of a convolutional layer of the trained machine learning model.

14 . The computer program product of claim 10 , wherein the generating adversarial examples includes solving a minmax algorithm which finds the adversarial example that has minimum training loss and least similarity to the original sample.

15 . A system comprising:

a hardware processor; and

a memory device coupled with the hardware processor;

the hardware processor configured to:

receive a trained machine learning model and a training dataset used to train the trained machine learning model;

based on the training dataset and a loss function of the trained machine learning model, generate images representing adversarial examples, the adversarial examples being perturbed samples of the training dataset, the generating creating an adversarial example which is least similar to an original sample in the training dataset, and also satisfying an adversarial criterion that a loss associated with the adversarial example is less than that of the original sample in the training dataset;

determine robustness of the trained machine learning model using the generated adversarial examples;

augment the training dataset with the generated images representing adversarial examples;

retrain the trained machine learning model using the augmented training dataset;

perform by the retrained machine learning model, image classification; and

display the original sample, a generated image representing the adversarial example, and a reconstructed image reconstructed using the retrained machine learning model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2021
From: HSU, CHIA-YI; YU, CHIA-MU
To: NATIONAL CHUNG HSING UNIVERSITY
Reel/Frame 057127/0570 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2021
From: CHEN, PIN-YU; LU, SONGTAO; LIU, SIJIA; GAN, CHUANG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055019/0724 →
Continuity (1)
Related Publication 20220253714A1 · Aug 11, 2022
References Cited (60)
US 10474929B2 · Choi et al. · 2019 [cited by applicant]
US 10643602B2 · Li et al. · 2020 [cited by applicant]
US 10698063B2 · Braun et al. · 2020 [cited by applicant]
US 10706534B2 · Middlebrooks et al. · 2020 [cited by applicant]
US 11494667B2 · Carbune · 2022 [cited by examiner]
US 20190197368A1 · Madani · 2019 [cited by examiner]
US 20190385018A1 · Ngo Dinh · 2019 [cited by examiner]
US 20200293890A1 · Baker · 2020 [cited by applicant]
US 20200320371A1 · Baker · 2020 [cited by applicant]
US 20200372301A1 · Kearney · 2020 [cited by examiner]
US 20210067549A1 · Chen · 2021 [cited by examiner]
US 20210150310A1 · Wu · 2021 [cited by examiner]
US 20220077944A1 · Zou · 2022 [cited by examiner]
US 20230260083A1 · Edlund · 2023 [cited by examiner]
EP 4149110A1 · 2023 [cited by examiner]
WO WO2020093987A1 · 2020 [cited by examiner]
Belghazi MI, Baratin A, Rajeshwar S, Ozair S, Bengio Y, Courville A, Hjelm D. Mutual information neural estimation. InInternational conference on machine learning Jul. 3, 2018 (pp. 531-540). PMLR. (Year: 2018). [cited by examiner]
Kwak H, Zhang BT. Ways of conditioning generative adversarial networks. arXiv preprint arXiv:1611.01455. Nov. 4, 2016. (Year: 2016). [cited by examiner]
Wu E, Cui H, Welsch RE. Dual autoencoders generative adversarial network for imbalanced classification problem. IEEE Access. May 13, 2020;8:91265-75. (Year: 2020). [cited by examiner]
Duan Y, Tao X, Xu M, Han C, Lu J. GAN-NL: Unsupervised representation learning for remote sensing image classification. In2018 IEEE Global Conference on Signal and Information Processing (GlobalSIP) Nov. 26, 2018 (pp. 3… [cited by examiner]
Kang, Minguk, and Jaesik Park. “Contragan: Contrastive learning for conditional image generation.” Advances in Neural Information Processing Systems 33 (2020): 21357-21369. (Year: 2020). [cited by examiner]
Dutta IK, Ghosh B, Carlson A, Totaro M, Bayoumi M. Generative adversarial networks in security: A survey. In2020 11th IEEE Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON) Oct. 28, 202… [cited by examiner]
Spampinato, C., et al., “Adversarial Framework for Unsupervised Learning of Motion Dynamics in Videos”, Computer Science, International Journal of Computer Vision, 2019, 19 pages. [cited by applicant]
Donahue, J., et al., “Large Scale Adversarial Representation Learning”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 11 pages. [cited by applicant]
Shrivastava, A., et al., “Learning from Simulated and Unsupervised Images through Adversarial Training”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21-26, 2017, pp. 2107-2116. [cited by applicant]
Anonymous, “Serving Multiple Anatomies and Protocols with an Unsupervised Gating Mechanism for Deep Learning Models”, An IP.com Prior Art Database Technical Disclosure, IPCOM000263846D, Oct. 9, 2020, 7 pages. [cited by applicant]
Anonymous, “Ranking and Automatic Selection of Machine Learning Models”, An IP.com Prior Art Database Technical Disclosure, IPCOM000252275D, Jan. 3, 2018, 34 pages. [cited by applicant]
Anonymous, Automatically Scaling Multi-Tenant Machine Learning, an IP.com Prior Art Database Technical Disclosure, IPCOM000252098D, Dec. 15, 2017, 35 pages. [cited by applicant]
Athalye, A., et al., “Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples”, Proceedings of the 35 th International Conference on Machine Learning, PMLR 80, 2018, arXiv:180… [cited by applicant]
Abid, A., et al., “Concrete Autoencoders: Differentiable Feature Selection and Reconstruction”, arXiv:1901.09346v2, Jan. 31, 2019, 16 pages. [cited by applicant]
Belghazi, M.I., et al., “Mutual Information Neural Estimation”, Proceedings of the 35 th International Conference on Machine Learning, PMLR 80, 2018, arXiv:1801.04062v4, Jun. 7, 2018, 18 pages. [cited by applicant]
Biggio, B., et al., “Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning”, arXiv:1712.03141v2, Jul. 19, 2018, 17 pages. [cited by applicant]
Binkowski, M., et al., “Demystifying MMD GANs”, Published as a conference paper at ICLR 2018, arXiv:1801.01401v5, Jan. 14, 2021, 36 pages. [cited by applicant]
Brendel, W., et al., “Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models”, Published as a conference paper at ICLR 2018, 2018, 12 pages. [cited by applicant]
Candes, E. J., et al., “An Introduction to Compressive Sampling”, IEEE Signal Processing Magazine, Mar. 2008, pp. 21-30, 25(2). [cited by applicant]
Carlini, N., et al., “Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods”, AlSec'17, Nov. 3, 2017, pp. 3-14. [cited by applicant]
Carlini, N., et al., “Towards Evaluating the Robustness of Neural Networks”, arXiv:1608.04644v2, Mar. 22, 2017, 19 pages. [cited by applicant]
Cavallari, G.B., et al., “Unsupervised representation learning using convolutional and stacked auto-encoders: a domain and cross-domain feature space analysis”, arXiv:1811.00473v1, Nov. 1, 2018, 7 pages. [cited by applicant]
Chen, P.Y., et al., “ZOO: Zeroth Order Optimization Based Black-box Attacks to Deep Neural Networks Without Training Substitute Models”, AlSec'17, Nov. 3, 2017, arXiv:1708.03999v2, Nov. 2, 2017, 13 pages. [cited by applicant]
Donsker, M.D., et al., “Asymptotic Evaluation of Certain Markov Process Expectations for Large Time. IV”, Communications on Pure and Applied Mathematics, Received Aug. 1982, pp. 183-212, vol. 36 (1983). [cited by applicant]
Geurts, P., et al., “Extremely randomized trees”, Machine Learning 2006, Revised Oct. 29, 2005, Accepted Nov. 15, 2005, Published online Mar. 2, 2006, pp. 3-42, 63. [cited by applicant]
Goodfellow, I.J., et al., “Explaining and Harnessing Adversarial Examples”, Published as a conference paper at ICLR 2015, arXiv:1412.6572v3, Mar. 20, 2015, 11 pages. [cited by applicant]
Heusel, M., et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 2017, 12 pages. [cited by applicant]
Hjelm, R. D., et al., “Learning deep representations by mutual information estimation and maximization”, Published as a conference paper at ICLR 2019, arXiv:1808.06670v5, Feb. 22, 2019, 24 pages. [cited by applicant]
Liu, S., et al., “Min-max optimization without gradients: Convergence and applications to black-box evasion and poisoning attacks”, Proceedings of the 37 th International Conference on Machine Learning PMLR vol. 119, Ju… [cited by applicant]
Lu, S., “Hybrid Block Successive Approximation for One-Sided Non-Convex Min-Max Problems: Algorithms and Applications”, arXiv:1902.08294v1, Feb. 21, 2019, 31 pages. [cited by applicant]
Madry, A., et al., “Towards Deep Learning Models Resistant to Adversarial Attacks”, arXiv:1706.06083v4, Sep. 4, 2019, 28 pages. [cited by applicant]
Makhzani, A., et al., “Adversarial autoencoders”, arXiv:1511.05644v2, May 25, 2016, 16 pages. [cited by applicant]
Bhagoji, A.N., “Practical Black-box Attacks on Deep Neural Networks using Efficient Query Mechanisms”, In Proceedings of the European Conference on Computer Vision (ECCV 2018), First Online Oct. 6, 2018, 16 pages. [cited by applicant]
Papernot, N., et al., “Practical black-box attacks against machine learning”, Asia CCS '17, Apr. 2-6, 2017, arXiv:1602.02697v4, Mar. 19, 2017, 14 pages. [cited by applicant]
Ranzato, M.A., et al., “Unsupervised learning of invariant feature hierarchies with applications to object recognition”, 2007 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 17-22, 2007, 8 pages. [cited by applicant]
Razaviyayn, M., et al., “Nonconvex Min-Max Optimization: Applications, Challenges, and Recent Theoretical Advances”, arXiv:2006.08141v2, Aug. 18, 2020, 12 pages. [cited by applicant]
Su, D., et al., “Is Robustness the Cost of Accuracy?—A Comprehensive Study on the Robustness of 18 Deep Image Classification Models”, arXiv:1808.01688v2, Mar. 4, 2019, 24 pages. [cited by applicant]
Szegedy, C., et al., “Intriguing properties of neural networks”, arXiv:1312.6199v4, Feb. 19, 2014, 10 pages. [cited by applicant]
Tishby, N., et al., “The information bottleneck method”, Sep. 30, 1999, arXiv:physics/0004057v1, Apr. 24, 2000, 16 pages. [cited by applicant]
Tsipras, D., et al., “Robustness May Be at Odds with Accuracy”, Published as a conference paper at ICLR 2019, Sep. 27, 2018, Modified Feb. 22, 2019, 23 pages. [cited by applicant]
Weng, T.-W., et al., “Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach”, Published as a conference paper at ICLR 2018, arXiv:1801.10578v1, Jan. 31, 2018, 18 pages. [cited by applicant]
Zhai, X., et al., “S4I: Self-Supervised Semi-Supervised Learning”, arXiv:1905.03670v2, Jul. 23, 2019, 13 pages. [cited by applicant]
Zhu, S., et al., “Learning Adversarially Robust Representations via Worst-Case Mutual Information Maximization”, Proceedings of the 37th International Conference on Machine Learning, PMLR vol. 119, 2020, arXiv:2002.1179… [cited by applicant]
Zhu, X., et al., “Introduction to Semi-Supervised Learning”, Synthesis Lectures on Artificial Intelligence and Machine Learning, 130 pages. [cited by applicant]