IP Library › Granted Patent US 12,511,545
Granted Patent B2
US 12,511,545 · App. 17/337,812 · Granted Dec 30, 2025

Training robust neural networks via smooth activation functions

Inventors: Mingxing Tan (Newark, CA); Cihang Xie (Baltimore, MD); Boqing Gong (Bellevue, WA); Quoc V. Le (Sunnyvale, CA)
Assignee: GOOGLE LLC
G06N3/084G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,545
App. No.
17/337,812
Granted
Dec 30, 2025
Kind
B2
Abstract

Generally, the present disclosure is directed to the training of robust neural network models by using smooth activation functions. Systems and methods according to the present disclosure may generate and/or train neural network models with improved robustness without incurring a substantial accuracy penalty and/or increased computational cost, or without any such penalty at all. For instance, in some examples, the accuracy may improve. A smooth activation function may replace an original activation function in a machine-learned model when backpropagating a loss function through the model. Optionally, one activation function may be used in the model at inference time, and a replacement activation function may be used when backpropagating a loss function through the model. The replacement activation function may be used to update learnable parameters of the model and/or to generate adversarial examples for training the model.

Claims (53)

1 . A computer-implemented method for improved adversarial training to increase model robustness, the method comprising:

obtaining, by a computing system comprising one or more computing devices, a training example for a machine-learned model;

executing, by the computing system, a first forward pass of the machine-learned model using a non-smooth activation function in at least one layer to generate a first output based on the training example;

replacing, by the computing system and in the at least one layer for a first backward pass, the non-smooth activation function with a smooth replacement activation function for adversarial example generation;

backpropagating, by the computing system and in the first backward pass, a loss function through the at least one layer having the smooth replacement activation function for adversarial example generation;

perturbing, by the computing system and based on a gradient at the training example determined based on the backpropagating, the training example in a direction of the gradient corresponding to an increase in the loss function to obtain an adversarial example;

executing, by the computing system, a second forward pass of the machine-learned model using the non-smooth activation function in the at least one layer to generate a second output based on the adversarial example;

replacing, by the computing system and in the at least one layer for a second backward pass, the non-smooth activation function with the smooth replacement activation function for training the machine-learned model;

backpropagating, by the computing system and in the second backward pass, the loss function through the at least one layer having the smooth replacement activation function for training the machine-learned model; and

updating, by the computing system and based on a gradient at a model parameter determined based on the backpropagating, the model parameter in a direction of the gradient corresponding to a decrease in the loss function.

2 . The computer-implemented method of claim 1 , wherein, at least for said backpropagating, the smooth replacement activation function comprises a smooth approximation of a rectified linear unit.

3 . The computer-implemented method of claim 2 , wherein, at least for said backpropagating, the smooth replacement activation function of the machine-learned model comprises an activation function having a zero-value activation output for a plurality of negative activation inputs.

4 . The computer-implemented method of claim 1 , wherein the training output is generated based at least in part on a non-smooth activation function.

5 . The computer-implemented method of claim 1 , wherein the smooth replacement activation function comprises a learnable parameter, and wherein the method comprises:

updating, by the computing system, the learnable parameter based on the backpropagating in the second backward pass.

6 . The computer-implemented method of claim 1 , wherein the machine-learned model comprises an image processing model, and wherein the training example is an image, and wherein perturbing, by the computing system and based on the gradient at the training example determined based on the backpropagating, the training example in the direction of the gradient corresponding to the increase in the loss function to obtain the adversarial example comprises:

modifying, by the computing system, the image.

7 . The computer-implemented method of claim 1 , wherein the one or more computing devices consist of a user computing device, wherein obtaining, by the one or more computing devices, the training example comprises obtaining, by the user computing device, a personal training example that is stored at a local memory of the user computing device, and wherein the machine-learned model is also stored at the local memory of the user computing device.

8 . The computer-implemented method of claim 1 , wherein the computing system comprises a user computing device, and wherein the training example is image is an image captured by the user computing device.

9 . A computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining a training example for a machine-learned model;

executing a first forward pass of the machine-learned model using a non-smooth activation function in at least one layer to generate a first output based on the training example;

replacing, in the at least one layer for a first backward pass, the non-smooth activation function with a smooth replacement activation function for adversarial example generation;

backpropagating, in the first backward pass, a loss function through the at least one layer having the smooth replacement activation function for adversarial example generation;

perturbing, based on a gradient at the training example determined based on the backpropagating, the training example in a direction of the gradient corresponding to an increase in the loss function to obtain an adversarial example;

executing a second forward pass of the machine-learned model using the non-smooth activation function in the at least one layer to generate a second output based on the adversarial example;

replacing, in the at least one layer for a second backward pass, the non-smooth activation function with the smooth replacement activation function for training the machine-learned model;

backpropagating, in the second backward pass, the loss function through the at least one layer having the smooth replacement activation function for training the machine-learned model; and

updating, based on a gradient at a model parameter determined based on the backpropagating, the model parameter in a direction of the gradient corresponding to a decrease in the loss function.

10 . The computing system of claim 9 , wherein, at least for said backpropagating, the smooth replacement activation function comprises a smooth approximation of a rectified linear unit.

11 . The computing system of claim 10 , wherein, at least for said backpropagating, the smooth replacement activation function of the machine-learned model comprises an activation function having a zero-value activation output for a plurality of negative activation inputs.

12 . The computing system of claim 9 , wherein the training output is generated based at least in part on a non-smooth activation function.

13 . The computing system of claim 9 , wherein the smooth replacement activation function comprises a learnable parameter, and wherein the operations comprise:

updating the learnable parameter based on the backpropagating in the second backward pass.

14 . The computing system of claim 9 , wherein the computing system consists of a user computing device, wherein obtaining the training example comprises obtaining, by the user computing device, a personal training example that is stored at a local memory of the user computing device, and wherein the machine-learned model is also stored at the local memory of the user computing device.

15 . One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more processors, cause a computing system to perform operations, the operations comprising:

obtaining a training example for a machine-learned model;

executing a first forward pass of the machine-learned model using a non-smooth activation function in at least one layer to generate a first output based on the training example;

replacing, in the at least one layer for a first backward pass, the non-smooth activation function with a smooth replacement activation function for adversarial example generation;

backpropagating, in the first backward pass, a loss function through the at least one layer having the smooth replacement activation function for adversarial example generation;

perturbing, based on a gradient at the training example determined based on the backpropagating, the training example in a direction of the gradient corresponding to an increase in the loss function to obtain an adversarial example;

executing a second forward pass of the machine-learned model using the non-smooth activation function in the at least one layer to generate a second output based on the adversarial example;

replacing, in the at least one layer for a second backward pass, the non-smooth activation function with the smooth replacement activation function for training the machine-learned model;

backpropagating, in the second backward pass, the loss function through the at least one layer having the smooth replacement activation function for training the machine-learned model; and

updating, based on a gradient at a model parameter determined based on the backpropagating, the model parameter in a direction of the gradient corresponding to a decrease in the loss function.

16 . The one or more non-transitory computer-readable media of claim 15 , wherein, at least for said backpropagating, the smooth replacement activation function comprises a smooth approximation of a rectified linear unit.

17 . The one or more non-transitory computer-readable media of claim 16 , wherein, at least for said backpropagating, the smooth replacement activation function of the machine-learned model comprises an activation function having a zero-value activation output for a plurality of negative activation inputs.

18 . The one or more non-transitory computer-readable media of claim 15 , wherein the training output is generated based at least in part on a non-smooth activation function.

19 . The one or more non-transitory computer-readable media of claim 15 , wherein the smooth replacement activation function comprises a learnable parameter, and wherein the operations comprise:

updating the learnable parameter based on the backpropagating in the second backward pass.

20 . The one or more non-transitory computer-readable media of claim 15 , wherein the computing system consists of a user computing device, wherein obtaining the training example comprises obtaining, by the user computing device, a personal training example that is stored at a local memory of the user computing device, and wherein the machine-learned model is also stored at the local memory of the user computing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 12, 2021
From: TAN, MINGXING; LE, QUOC V.; XIE, CIHANG; GONG, BOQING
To: GOOGLE LLC
Reel/Frame 057164/0293 →
Continuity (2)
Provisional Application 63034173 · Jun 3, 2020
Related Publication 20210383237A1 · Dec 9, 2021
References Cited (95)
US 10783401B1 · Jiang · 2020 [cited by examiner]
US 11645499B2 · Guntoro · 2023 [cited by examiner]
US 20180137413A1 · Li · 2018 [cited by examiner]
US 20190114541A1 · Yang · 2019 [cited by examiner]
US 20200005143A1 · Zamora Esquivel · 2020 [cited by examiner]
US 20200097818A1 · Li · 2020 [cited by examiner]
US 20200285952A1 · Liu · 2020 [cited by examiner]
US 20200302234A1 · Walters · 2020 [cited by examiner]
US 20200364616A1 · Wong · 2020 [cited by examiner]
US 20200372305A1 · Streeter · 2020 [cited by examiner]
US 20210097343A1 · Goodsitt · 2021 [cited by examiner]
US 20210181754A1 · Cui · 2021 [cited by examiner]
US 20210279591A1 · Shamir · 2021 [cited by examiner]
Searching for activation function; Ramachandran et al. (Year: 2017). [cited by examiner]
Adversarial Machine learning at scale (Year: 2017). [cited by examiner]
Aprilpyone et al., “Block-wise Image Transformation with Secret Key for Adversarially Robust Defense”, arXiv:2010.00801v1, Oct. 2, 2020, 15 pages. [cited by applicant]
Athalye et al., “Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples”, arXiv:1802.00420v4, Jul. 31, 2018, 12 pages. [cited by applicant]
Barron, “Continuously Differentiable Exponential Linear Units”, arXiv:1704.07483v1, Apr. 24, 2017, 2 pages. [cited by applicant]
Bhagoji et al., “Enhancing Robustness of Machine Learning Systems via Data Transformations”, arXiv:1704.02654v4, Nov. 29, 2017, 15 pages. [cited by applicant]
Biswas et al., “Tanhsoft—A Family of Activation Functions Combining Tanh and Softplus”, arXiv:2009.03863v1, Sep. 8, 2020, 11 pages. [cited by applicant]
Buckman et al., “Thermometer Encoding: One Hot Way to Resist Adversarial Examples”, International Conference on Learning Representations, Apr. 30-May 3, 2018, Vancouver, Canada, 22 pages. [cited by applicant]
Carlini et al., “On Evaluating Adversarial Robustness”, arXiv:1902.06705v2, Feb. 20, 2019, 24 pages. [cited by applicant]
Chen et al., “Anti-Bandit Neural Architecture Search for Model Defense”, arXiv:2008.00698v2, Aug. 5, 2020, 18 pages. [cited by applicant]
Clevert et al., “Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)”, arXiv:1511.07289v1, Nov. 23, 2015, 14 pages. [cited by applicant]
Croce et al., “Reliable Evaluation of Adversarial Robustness with an Ensemble of Diverse Parameter-free Attacks”, Jul. 12-18, 2020, Virtual, 11 pages. [cited by applicant]
Cubuk et al., “AutoAugment: Learning Augmentation Strategies from Data”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, CA, 11 pages. [cited by applicant]
Dhillon et al., “Stochastic Activation Pruning for Robust Adversarial Defense”, arXiv:1803.01442v1, Mar. 5, 2018, 13 pages. [cited by applicant]
Ding et al., “MMA Training: Direct Input Space Margin Maximization through Adversarial Training”, arXiv:1812.02637v4, Mar. 4, 2020, 28 pages. [cited by applicant]
Dong et al., “Boosting Adversarial Attacks with Momentum”, Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2018, Salt Lake City, UT, pp. 9185-9193. [cited by applicant]
Dziugaite et al., “A study of the effect of JPG compression on adversarial images”, arXiv:1608.00853v1, Aug. 2, 2016, 8 pages. [cited by applicant]
Elfwing et al., “Sigmoid-weighted linear units for neural networks function approximation in reinforcement learning”, Neural Networks, vol. 107, 2018, pp. 3-11. [cited by applicant]
Galloway et al., “Batch Normalization is a Cause of Adversarial Vulnerability”, arXiv:1905.02161v2, May 29, 2019, 17 pages. [cited by applicant]
Gao et al., “Convergence of Adversarial Training in Overparametrized Neural Networks”, Conference on Neural Information Processing Systems, Dec. 8-14, 2019, Vancouver, Canada, 12 pages. [cited by applicant]
Goodfellow et al., “Explaining and Harnessing Adversarial Examples”, arXiv:1412.6572v3, Mar. 20, 2015, 11 pages. [cited by applicant]
Guo et al., “Countering Adversarial Images Using Input Transformations” arXiv:1711.00117v3, Jan. 25, 2018, 12 pages. [cited by applicant]
Guo et al., “When NAS Meets Robustness: In Search of Robust Architectures against Adversarial Attacks”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 631-640. [cited by applicant]
Hahnloser et al., “Digital selection and analogue ampli®cation coexist in a cortex-inspired silicon circuit”, Nature, vol. 405, Jun. 22, 2000, pp. 947-951. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, Conference on Computer Vision and Pattern Recognition, Jun. 26-Jul. 1, 2016, Las Vegas, NV, pp. 770-778. [cited by applicant]
Hendrycks et al., “Gaussian Error Linear Units (GELUs)”, arXiv:1606.08415, Nov. 11, 2018, 9 pages. [cited by applicant]
Hoffer et al., “Train longer, generalize better: closing the generalization gap in large batch training of neural networks”, Conference on Neural Information Processing Systems, Dec. 4-9, 2017, Long Beach, CA, 11 pages. [cited by applicant]
Huang et al., “Deep Networks with Stochastic Depth”, European Conference on Computer Vision, Oct. 8-16, 2016, Amsterdam, The Netherlands, 16 pages. [cited by applicant]
Ilyas et al., “Adversarial Examples are not Bugs, they are Features”, Conference on Neural Information Processing Systems, Dec. 8-14, 2019, Vancouver, Canada, 12 pages. [cited by applicant]
Izmailov et al., “Averaging Weights Leads to Wider Optima and Better Generalization”, arXiv:1803.05407v2, Aug. 8, 2018, 11 pages. [cited by applicant]
Kettunen et al., “E-LPIPS: Robust Perceptual Image Similarity via Random Transformation Ensembles”, arXiv:1906.03973v2, Jun. 11, 2019, 10 pages. [cited by applicant]
Kurakin et al., “Adversarial Machine Learning at Scale”, International Conference on Learning Representations, Apr. 24-26, 2017, Toulon, France, 17 pages. [cited by applicant]
Lee et al., “Robust Ensemble Model Training via Random Layer Sampling Against Adversarial Attack”, arXiv:2005.10757v1, May 21, 2020, 13 pages. [cited by applicant]
Li et al., “Visualizing the Loss Landscape of Neural Nets”, Conference on Neural Information Processing Systems, Dec. 3-8, 2018, Montreal, Canada, 11 pages. [cited by applicant]
Liao et al., “Defense against Adversarial Attacks Using High-Level Representation Guided Denoiser”, Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2018, Salt Lake City, UT, pp. 1778-1787. [cited by applicant]
Liu et al., “Towards Robust Neural Networks via Random Self-ensemble”, European Conference on Computer Vision, Sep. 8-14, 2018, Munich, Germany, 17 pages. [cited by applicant]
Lokhande et al., “Generating Accurate Pseudo-labels in Semi-Supervised Learning and Avoiding Overconfident Predictions via Hermite Polynomial Activations”, Conference on Computer Vision and Pattern Recognition, Jun. 14-… [cited by applicant]
Luo et al., “Foveation-based Mechanisms Alleviate Adversarial Examples”, arXiv:1511.06292v1, Nov. 19, 2015, 23 pages. [cited by applicant]
Luo et al., “Random Mask: Towards Robust Convolutional Neural Networks”, arXiv:2007.14249v1, Jul. 27, 2020, 26 pages. [cited by applicant]
Madry et al., “Towards Deep Learning Models Resistant to Adversarial Attacks”, International Conference on Learning Representations, Apr. 30-May 3, 2018, Vancouver, Canada, 23 pages. [cited by applicant]
Meng et al., “MagNet: a Two-Pronged Defense against Adversarial Examples”, ACM Conference on Computer and Communications Security, Oct. 30- Nov. 3rd, Dallas, TX, pp. 135-147. [cited by applicant]
Misra, “Mish: A Self Regularized Non-Monotonic Neural Activation Function”, arXiv:1908.08681v2, Oct. 2, 2019, 13 pages. [cited by applicant]
Nair et al., “Rectified Linear Units Improve Restricted Boltzmann Machines”, International Conference on Machine Learning, Jun. 21-24, 2010, Haifa, Israel, 8 pages. [cited by applicant]
Nakkiran, “Adversarial Robustness May Be at Odds With Simplicity”, arXiv:1901.00532v1, Jan. 2, 2019, 8 pages. [cited by applicant]
Pang et al., “Bag of Tricks for Adversarial Training”, International Conference on Learning Representations, May 3-7, 2021, Virtual, 21 pages. [cited by applicant]
Pang et al., “Improving Adversarial Robustness via Promoting Ensemble Diversity”, International Conference on Machine Learning, Jun. 10-15, 2019, Long Beach, California, 10 pages. [cited by applicant]
Pang et al., “Mixup Inference: Better Exploiting Mixup to Defend Adversarial Attacks”, International Conference on Learning Representations, Apr. 26-May 1, 2020 Virtual, 14 pages. [cited by applicant]
Papernot et al., “Practical Black-Box Attacks against Machine Learning”, Asia Conference on Computer and Communications Security, Apr. 2-6, 2017, Abu Dhabi, United Arab Emirates, pp. 506-519. [cited by applicant]
Prakash et al., “Deflecting Adversarial Attacks with Pixel Deflection”, Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2018, Salt Lake City, UT, pp. 8571-8580. [cited by applicant]
Qin et al., “Adversarial Robustness through Local Linearization”, Conference on Neural Information Processing Systems, Dec. 8-14, 2019, Vancouver, Canada, 10 pages. [cited by applicant]
Raff et al., “Barrage of Random Transforms for Adversarially Robust Defense”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, CA, pp. 6528-6537. [cited by applicant]
Ramachandran et al., “Searching for Activation Functions”, arXiv:1710.05941v2, Oct. 27, 2017, 13 pages. [cited by applicant]
Rozsa et al., “Improved Adversarial Robustness by Reducing Open Space Risk via Tent Activations”, arXiv:1908.02435, Aug. 7, 2019, 15 pages. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge”, arXiv:1409.0575v3, Jan. 30, 2015, 43 pages. [cited by applicant]
Samangouei et al., “Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models”, International Conference on Learning Representations, Apr. 30-May 3, 2018, Vancouver, Canada, 17 pages. [cited by applicant]
Schott et al., “Towards the First Adversarially Robust Neural Network Model on MNIST”, International Conference on Learning Representations, May 6-9, 2019, New Orleans, Louisiana, 17 pages. [cited by applicant]
Shafahi et al., “Adversarial Training for Free!”, Conference on Neural Information Processing Systems, Dec. 8-14, 2019, Vancouver, Canada, 12 pages. [cited by applicant]
Shamir et al., “Smooth Activations and Reproducibility in Deep Networks”, arXiv:2010.009931v2, Dec. 1, 2020, 23 pages. [cited by applicant]
Sinha et al., “Certifying Some Distributional Robustness with Principled Adversarial Training”, International Conference on Learning Representations, Apr. 30-May 3, 2018, Vancouver, Canada, 34 pages. [cited by applicant]
Song et al., “PixelDefend: Leveraging Generative Models to Understand and Defend Against Adversarial Examples”, International Conference on Learning Representations, Apr. 30-May 3, 2018, Vancouver, Canada, 20 pages. [cited by applicant]
Srivastava et al., “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”, Journal of Machine Learning Research, vol. 15, Jun. 2014, pp. 1929-1958. [cited by applicant]
Strauss et al., “Ensemble Methods as a Defense to Adversarial Perturbations Against Deep Neural Networks”, arXiv:1709.03423v1., Sep. 11, 2017, 11 pages. [cited by applicant]
Su et al., “Is Robustness the Cost of Accuracy?—A Comprehensive Study on the Robustness of 18 Deep Image Classification Models”, European Conference on Computer Vision, Sep. 8-14, 2018, Munich, Germany, 18 pages. [cited by applicant]
Szegedy et al., “Intriguing properties of neural networks”, arXiv:1312.6199v4, Feb. 19, 2014, 10 pages. [cited by applicant]
Tan et al., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks”, International Conference on Machine Learning, Jun. 10-15, 2019, Long Beach, CA, 10 pages. [cited by applicant]
Tsipras et al., “Robustness May Be at Odds with Accuracy”, arXiv:1805.12152v5, Sep. 9, 2019, 24 pages. [cited by applicant]
Wang et al., “Defensive Dropout for Hardening Deep Neural Networks under Adversarial Attacks”, International Conference on Computer-Aided Design, Nov. 5-8, 2018, San Diego, CA, 8 pages. [cited by applicant]
Wang et al., “Improving Adversarial Robustness Requires Revisiting Misclassified Examples”, International Conference on Learning Representations, Apr. 26-May 1, 2020, Virtual, 14 pages. [cited by applicant]
Wang et al., “On the Convergence and Robustness of Adversarial Training”, International Conference on Machine Learning, Jun. 10-15, 2019, Long Beach, CA, 10 pages. [cited by applicant]
Wang et al., “Protecting Neural Networks with Hierarchical Random Switching: Towards Better Robustness-Accuracy Trade-off for Stochastic Defenses”, International Joint Conference on Artificial Intelligence, Aug. 10-16, … [cited by applicant]
Wong et al., “Fast is Better than Free: Revisiting Adversarial Training”, International Conference on Learning Representations, Apr. 26-May 1, 2020, Virtual, 17 pages. [cited by applicant]
Xiao et al., “Enhancing Adversarial Defense by k-Winners-Take-All”, International Conference on Learning Representations, Apr. 26-May 1, 2020, Virtual, 30 pages. [cited by applicant]
Xiao et al., “One Man's Trash is Another Man's Treasure: Resisting Adversarial Examples by Adversarial Examples”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, pp. 412-421. [cited by applicant]
Xie et al., “Aggregated Residual Transformations for Deep Neural Networks”, Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, Honolulu, Hawaii, pp. 1492-1500. [cited by applicant]
Xie et al., “Feature Denoising for Improving Adversarial Robustness”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, CA, pp. 501-509. [cited by applicant]
Xie et al., “Intriguing Properties of Adversarial Training at Scale”, International Conference on Learning Representations, Apr. 26-May 1, 2020, Virtual, 14 pages. [cited by applicant]
Xie et al., “Mitigating Adversarial Effects Through Randomization”, International Conference on Learning Representations, Apr. 30-May 3, 2018, Vancouver, Canada, 16 pages. [cited by applicant]
Xie et al., “Self-training with Noisy Student improves ImageNet classification”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 10687-10698. [cited by applicant]
Xie et al., “Smooth Adversarial Training”, arXiv:2006.14536v2, Jul. 11, 2021, 11 pages. [cited by applicant]
Xu et al., “Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks”, Network and Distributed Systems Security Symposium, Feb. 18-21, 2018, San Diego, CA, 15 pages. [cited by applicant]
Zhang et al., “Theoretically Principled Trade-off between Robustness and Accuracy”, International Conference on Machine Learning, Jun. 10-Jun. 15, 2019, Long Beach, California, 11 pages. [cited by applicant]
Zoph et al., “Neural Architecture Search with Reinforcement Learning”, arXiv:1611.01578v1, Nov. 5, 2016, 15 pages. [cited by applicant]