IP Library › Granted Patent US 12,321,825
Granted Patent B2
US 12,321,825 · App. 17/210,934 · Granted Jun 3, 2025

Training neural networks with limited data using invertible augmentation operators

Inventors: Tero Tapani Karras (Helsinki, FI); Miika Samuli Aittala (Helsinki, FI); Janne Johannes Hellsten (Helsinki, FI); Samuli Matias Laine (Vantaa, FI); Jaakko T. Lehtinen (Helsinki, FI); Timo Oskari Aila (Tuusula, FI)
Assignee: NVIDIA Corporation
G06N20/00G06T1/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,825
App. No.
17/210,934
Granted
Jun 3, 2025
Kind
B2
Abstract

Embodiments of the present disclosure relate to a technique for training neural networks, such as a generative adversarial neural network (GAN), using a limited amount of data. Training GANs using too little example data typically leads to discriminator overfitting, causing training to diverge and produce poor results. An adaptive discriminator augmentation mechanism is used that significantly stabilizes training with limited data providing the ability to train high-quality GANs. An augmentation operator is applied to the distribution of inputs to a discriminator used to train a generator, representing a transformation that is invertible to ensure there is no leakage of the augmentations into the images generated by the generator. Reducing the amount of training data that is needed to achieve convergence has the potential to considerably help many applications and may the increase use of generative models in fields such as medicine.

Claims (39)

1. A computer-implemented method for training a neural network model comprising a generator and a discriminator, comprising:

receiving training data including example output data and ground truth outputs;

processing latent codes, according to parameters, by the generator to produce generated data;

applying at least one augmentation to the generated data to produce augmented generated data, wherein an augmentation operator is invertible and specifies the at least one augmentation;

processing only the augmented generated data by the discriminator to produce values; and

adjusting the parameters to reduce differences between the values and the ground truth outputs.

2. The computer-implemented method of claim 1 , wherein the example output data is associated with a first distribution and the augmentation operator transforms the first distribution into an augmented distribution that matches a second distribution associated with the augmented generated data.

3. The computer-implemented method of claim 1 , wherein the example output data comprises a first subset of output data produced by the generator and a second subset of real data, the ground truth outputs comprise values indicating either a first state or a second state, and adjusting the parameters causes a first portion of the generated data to match the first state more closely and a second portion of the generated data to match the second state more closely.

4. The computer-implemented method of claim 1 , wherein the generated data comprises images and the at least one augmentation is disabled for producing the augmented generated data.

5. The computer-implemented method of claim 1 , wherein applying further comprises randomly disabling application of the at least one augmentation for a portion of the generated data based on an augmentation strength value.

6. The computer-implemented method of claim 5 , further comprising dynamically adjusting the augmentation strength value based on an overfitting heuristic over the course of repeated processing of the latent codes and applying.

7. The computer-implemented method of claim 6 , wherein the augmentation strength value is increased or decreased based on a comparison between the overfitting heuristic and a pre-defined target value.

8. The computer-implemented method of claim 6 , wherein at least a portion of the generated data is compared to a reference to produce a result and the augmentation strength value is adjusted based on the result.

9. The computer-implemented method of claim 1 , wherein the at least one augmentation is differentiable.

10. The computer-implemented method of claim 1 , wherein the at least one augmentation is implemented as a sequence of different augmentations.

11. The computer-implemented method of claim 1 , wherein the example output data comprises a first subset of output data produced by the generator based on second parameters and a second subset of real data.

12. The computer-implemented method of claim 11 , further comprising, after adjusting the parameters, adjusting the second parameters to cause distributions of the first subset and the second subset to match more closely.

13. The computer-implemented method of claim 1 , wherein the discriminator is trained by:

applying at least one augmentation to the example output data to produce augmented example output data;

processing the augmented example output data by the discriminator according to second parameters to produce second values; and

adjusting the second parameters to reduce differences between the second values and the ground truth outputs.

14. The computer-implemented method of claim 1 , wherein the neural network model is used for accelerating applications including at least one of autonomous vehicle platforms, deep learning, high-accuracy speech, image, and text recognition, intelligent video analytics, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analytics, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimizations, or personalized user recommendations.

15. The computer-implemented method of claim 1 , wherein at least one of the steps of processing of the latent codes, applying, processing the augmented generated data, or adjusting are performed on at least a portion of a graphics processing unit.

16. A system, comprising:

a memory that stores training data including example output data and ground truth outputs; and

a processor that is coupled to the memory and implements a neural network model comprising a generator and a discriminator, wherein the neural network model is trained by:

processing latent codes, according to parameters, by the generator to produce generated data;

applying at least one augmentation to the generated data to produce augmented generated data, wherein an augmentation operator is invertible and specifies the at least one augmentation;

processing only the augmented generated data by the discriminator to produce values; and

adjusting the parameters to reduce differences between the values and the ground truth outputs.

17. The system of claim 16 , wherein the example output data is associated with a first distribution and the augmentation transforms the first distribution into an augmented distribution that matches a second distribution associated with the augmented generated data.

18. A non-transitory computer-readable media storing computer instructions that, when executed by one or more processors, cause the one or more processors to train a neural network model comprising a generator and a discriminator by performing the steps of:

receiving training data including example output data and ground truth outputs;

processing latent codes, according to parameters, by the generator to produce generated data;

applying at least one augmentation to the generated data to produce augmented generated data, wherein an augmentation operator is invertible and specifies the at least one augmentation;

processing only the augmented generated data by the discriminator to produce values; and

adjusting the parameters to reduce differences between the values and the ground truth outputs.

19. The non-transitory computer-readable media of claim 18 , wherein the example output data is associated with a first distribution and the augmentation operator transforms the first distribution into an augmented distribution that matches a second distribution associated with the augmented generated data.

20. The non-transitory computer-readable media of claim 18 , wherein applying further comprises randomly disabling application of the at least one augmentation for a portion of the generated data based on an augmentation strength value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2021
From: KARRAS, TERO TAPANI; AITTALA, MIIKA SAMULI; HELLSTEN, JANNE JOHANNES; LAINE, SAMULI MATIAS; LEHTINEN, JAAKKO T.; AILA, TIMO OSKARI
To: NVIDIA CORPORATION
Reel/Frame 055701/0227 →
Continuity (2)
Provisional Application 63035448 · Jun 5, 2020
Related Publication 20210383241A1 · Dec 9, 2021
References Cited (65)
US 20170351952A1 · Zhang et al. · 2017 [cited by applicant]
US 20180019388A1 · Fukami et al. · 2018 [cited by applicant]
US 20180101784A1 · Rolfe et al. · 2018 [cited by applicant]
US 20190114544A1 · Sundaram et al. · 2019 [cited by applicant]
US 20190130278A1 · Karras et al. · 2019 [cited by applicant]
US 20190373293A1 · Bortman · 2019 [cited by examiner]
US 20200090045A1 · Baker · 2020 [cited by applicant]
US 20210201078A1 · Yao · 2021 [cited by examiner]
US 20210216860A1 · Poghosyan · 2021 [cited by examiner]
CN 109313724A · 2019 [cited by applicant]
CN 110059793A · 2019 [cited by applicant]
CN 110800062A · 2020 [cited by applicant]
CN 110892417A · 2020 [cited by applicant]
CN 111145106A · 2020 [cited by applicant]
WO 2016159017A1 · 2016 [cited by applicant]
Zhao, S., Liu, Z., Lin, J., Zhu, J. Y., & Han, S. (2020). Differentiable augmentation for data-efficient gan training. Advances in neural information processing systems, 33, 7559-7570. (Year: 2020). [cited by examiner]
Shao, S., Wang, P., & Yan, R. (2019). Generative adversarial networks for data augmentation in machine fault diagnosis. Computers in Industry, 106, 85-93. (Year: 2019). [cited by examiner]
Inoue, H. (2018). Data augmentation by pairing samples for images classification. arXiv preprint arXiv:1801.02929. (Year: 2018). [cited by examiner]
Tran, N. T., Tran, V. H., Nguyen, N. B., Nguyen, T. K., & Cheung, N. M. (2020). Towards good practices for data augmentation in gan training. arXiv preprint arXiv:2006.05338, 2(3), 3. (Year: 2020). [cited by examiner]
Yadav, S. S., & Jadhav, S. M. (2019). Deep convolutional neural network based medical image classification for disease diagnosis. Journal of Big data, 6(1), 1-18. (Year: 2019). [cited by examiner]
Zhao, Z., Zhang, Z., Chen, T., Singh, S., & Zhang, H. (2020). Image augmentations for gan training. arXiv preprint arXiv:2006.02595. (Year: 2020). [cited by examiner]
Ching, C. W., Lin, T. C., Chang, K. H., Yao, C. C., & Kuo, J. J. (Dec. 2020). Model partition defense against gan attacks on collaborative learning via mobile edge computing. In GLOBECOM 2020—2020 IEEE Global Communicat… [cited by examiner]
Zhang, M. et al., “DeepRoad: GAN-Based metamorphic testing and input validation framework for autonomous driving systems,” 2018 33rd IEEE/ACM Intl. Conf. on Automated Software Engineering (Sep. 2018) pp. 132-142. (Year:… [cited by examiner]
Aksac, A., et al., “BreCaHAD: A dataset for breast cancer histopathological annotation and diagnosis,” BMC Research Notes, 12, 2019. [cited by applicant]
Arjovsky, M., et al., “Towards principled methods for training generative adversarial networks,” In Proc. ICLR, 2017. [cited by applicant]
Binkowski, M., et al., “Demystifying MMD GANs,” In Proc. ICLR, 2018. [cited by applicant]
Bora, A., et al., “AmbientGAN: Generative models from lossy measurements,” In Proc. ICLR, 2018. [cited by applicant]
Brock, A., et al., “Large Scale GAN training for high fidelity natural image synthesis,” In Proc. ICLR, 2019. [cited by applicant]
Chen, T., et al., “Self-supervised GANs via auxiliary rotation loss,” In Proc. CVPR, 2019. [cited by applicant]
Choi, Y., et al., “StarGAN v2: Diverse image synthesis for multiple domains,” In Proc CVPR, 2020. [cited by applicant]
Cubuk, E.D., et al., “AutoAugment: Learning augmentation strategies from data,” In Proc. CVPR, 2019. [cited by applicant]
Cubuk, E.D., et al., “Rand Augment: Practical automated data augmentation with a reduced search space,” CoRR, abs/1909.13719, 2019. [cited by applicant]
DeVries, T., et al., “Improved regularization of convolutional neural networks with cutout,” CoRR, abs/1708.04552, 2017. [cited by applicant]
Gong, X., et al., “AutoGAN: Neural architecture search for generative adversarial networks,” In Proc. ICCV, 2019. [cited by applicant]
Goodfellow, I., et al., “Generative adversarial nets,” In Proc. NIPS, 2014. [cited by applicant]
Gulrajani, I., et al., “Improved training of Wasserstein GANs,” In Proc. NIPS, pp. 5769-5779, 2017. [cited by applicant]
Gurumurthy, S., et al., “DeLiGAN: Generative adversarial networks for diverse and limited data,” In Proc. CVPR, 2017. [cited by applicant]
He, T., et al., “Bag of tricks for image classification with convolutional neural networks,” In Proc. CVPR, 2019. [cited by applicant]
Heusel, M., “GANs trained by a two time-scale update rule converge to a local Nash equilibrium,” In Proc. NIPS, 2017. [cited by applicant]
Karras, T., et al., “Progressive growing of GANs for improved quality, stability and variation,” In Proc. ICLR, 2018. [cited by applicant]
Karras, T., et al., “A style-based generator architecture for generative adversarial networks,” In Proc. CVPR, 2018. [cited by applicant]
Karras, T., et al., “Analyzing and improving the image quality of StyleGAN,” In Proc. CVPR, 2020. [cited by applicant]
Kavalerov, W., et al., “cGANs with multi-hinge loss,” CoRR, abs/1912.04216, 2019. [cited by applicant]
Krizhevsky, A., “Learning multiple layers of features from tiny images,” Technical report, University of Toronto, 2009. [cited by applicant]
Laine, S., et al., “Temporal ensembling for semi-supervised learning,” In Proc., ICLR, 2017. [cited by applicant]
Miyato, T., et al., “Spectral normalization for generative adversarial networks,” In Proc. ICLR, 2018. [cited by applicant]
Miyato, T., et al., “cGANs with projection discriminator,” In Proc. ICLR, 2018. [cited by applicant]
Mo, S., et al., “Freeze the discriminator: a simple baseline for fine-tuning GANs,” CoRR, abs/2002.10964, 2020. [cited by applicant]
Noguchi, A., et al., “Image generation from small datasets via batch statistics adaptation,” In Proc. ICCV, 2019. [cited by applicant]
Sajjadi, M., et al., “Regularization with stochastic transformations and perturbations for deep semi-supervised learning,” In Proc. NIPS, 2016. [cited by applicant]
Salimans, T., et al., “Improved techniques for training GANs,” In Proc., NIPS, 2016. [cited by applicant]
Schonfeld, E., et al., “A U-net based discriminator for generative adversarial networks,” CoRR, abs/2002.12655, 2020. [cited by applicant]
Sendik, O., et al., “Unsupervised multi-modal styled content generation,” CoRR, abs/2001.03640, 2020. [cited by applicant]
Shorten, C., et al., “A survey on image data augmentation for deep learning,” Journal of Big Data, 6, 2019. [cited by applicant]
Sonderby, C., et al., “Amortised MAP inference for image super-resolution,” In Proc. ICLR, 2017. [cited by applicant]
Srivastava, N., et al., “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research, 15:1929-1958, 2014. [cited by applicant]
Wang, Y., et al., “MineGAN: Effective knowledge transfer from GANs to target domains with few images,” In Proc. CVPR, 2020. [cited by applicant]
Yi, X., et al., “Generative adversarial network in medical imaging: A review.” Medical Image Analysis, 58, 2019. [cited by applicant]
Zhang, D., et al., “PA-GAN: Improving GAN training by progressive augmentation,” In Proc. NeurIPS, 2019. [cited by applicant]
Zhang, H., et al., “Consistency regularization for generative adversarial networks,” In Proc. ICLR, 2019. [cited by applicant]
Zhao, Y., et al., “Feature quantization improves GAN training,” CoRR abs/2004.02088, 2020. [cited by applicant]
Zhao, Z., et al., “Improved consistency regularization for GANs,” CoRR, abs/2002.04724, 2020. [cited by applicant]
Berthelot, D., et al., “BEGAN: Boundary Equilibrium Generative Adversarial Networks,” arXiv:1703.10717v4, May 31, 2017. [cited by applicant]
Zhao, S., et al., “Differentiable Augmentation for Data-Efficient GAN Training,” arXiv:2006.10738v4, Dec. 7, 2020. [cited by applicant]
Zhao, Z., et al., “Image Augmentations for GAN Training,” arXiv:2006.02595v1, Jun. 4, 2020. [cited by applicant]