IP Library › Granted Patent US 12,737,608
Granted Patent B2
US 12,737,608 · App. 17/089,492 · Granted Sep 15, 2026

Deep hierarchical variational autoencoder

Inventors: Arash Vahdat (Mountain View, CA); Jan Kautz (Lexington, MA)
Assignee: NVIDIA CORPORATION
G06N3/08G06N3/04G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,608
App. No.
17/089,492
Granted
Sep 15, 2026
Kind
B2
Abstract

One embodiment of the present invention sets forth a technique for performing machine learning. The technique includes inputting a training dataset into a variational autoencoder (VAE) comprising an encoder network, a prior network, and a decoder network. The technique also includes training the VAE by updating one or more parameters of the VAE based on a smoothness of one or more outputs produced by the VAE from the training dataset. The technique further includes producing generative output that reflects a first distribution of the training dataset by applying the decoder network to one or more values sampled from a second distribution of latent variables generated by the prior network.

Claims (51)

1 . A method for performing machine learning, comprising:

inputting a set of training images into a machine learning model that comprises an encoder portion, a prior, and a decoder portion;

training the machine learning model by updating one or more parameters of the machine learning model based on an objective function, the objective function controlling a smoothness of one or more outputs produced by the machine learning model when processing the set of training images, wherein training the machine learning model comprises:

applying, to the objective function, a plurality of balancing coefficients to generate a modified objective function, the plurality of balancing coefficients including a first balancing coefficient that has a first value during a first training period and a second balancing coefficient that has a second value, wherein the second value is higher than the first value during a second training period that occurs after the first training period, and

updating the one or more parameters of the machine learning model based on the modified objective function; and

producing a new image that reflects one or more visual attributes associated with the set of training images by applying the decoder portion to a value generated based on an output of the prior.

2 . The method of claim 1 , wherein the new image comprises a face that is not found in the set of training images.

3 . The method of claim 1 , wherein the new image comprises an animal or a vehicle that is not found in the set of training images.

4 . A method for performing machine learning, comprising:

inputting a training dataset into a variational autoencoder (VAE) comprising an encoder network, a prior network, and a decoder network;

training the VAE by updating one or more parameters of the VAE based on an objective function, the objective function controlling a smoothness of one or more outputs produced by the VAE when processing the training dataset, wherein training the VAE comprises:

applying, to the objective function, a plurality of balancing coefficients to generate a modified objective function, the plurality of balancing coefficients including a first balancing coefficient that has a first value during a first training period and a second balancing coefficient that has a second value, wherein the second value is higher than the first value during a second training period that occurs after the first training period, and

updating the one or more parameters of the VAE based on the objective function; and

producing generative output that reflects a first distribution of the training dataset by applying the decoder network to one or more values sampled from a second distribution of latent variables generated by the prior network.

5 . The method of claim 4 , wherein applying the decoder network to the one or more values comprises applying batch normalization to one or more layers of the decoder network based on a momentum parameter that increases a rate at which a running statistic associated with the batch normalization catches up to a batch statistic associated with the batch normalization.

6 . The method of claim 5 , wherein training the VAE comprises applying a regularization parameter to a scaling parameter associated with the batch normalization.

7 . The method of claim 5 , wherein applying the batch normalization to the one or more layers comprises combining the batch normalization with a Swish activation function.

8 . The method of claim 5 , wherein applying the batch normalization to the one or more layers of the decoder network comprises recalculating batch statistics associated with the batch normalization based on the one or more values sampled from the second distribution.

9 . The method of claim 4 , wherein the VAE comprises a hierarchy of groups of the latent variables, and wherein a first sample from a first group in the hierarchy is combined with a feature map and passed to a second group following the first group in the hierarchy for use in generating a second sample from the second group.

10 . The method of claim 4 , wherein the VAE comprises a residual cell and the residual cell comprises a first batch-normalization (BN) layer with a first Swish activation function, a first convolutional layer following the first BN layer with the first Swish activation function, a second BN layer with a second Swish activation function, a second convolutional layer following the second BN layer with the second Swish activation function, and a squeeze and excitation (SE) layer.

11 . The method of claim 4 , wherein the VAE comprises a residual cell and the residual cell comprises a first BN layer, a first convolutional layer following the first BN layer, a second BN layer with a first Swish activation function, and a depthwise separable convolution layer following the second BN layer.

12 . The method of claim 11 , wherein the residual cell further comprises a third BN layer with a second Swish activation function, a second convolutional layer following the third BN layer, a fourth BN layer following the second convolutional layer, and an SE layer following the fourth BN layer.

13 . The method of claim 1 , wherein the plurality of balancing coefficients includes a third balancing coefficient having a value that is determined based on a corresponding value of a Kullback-Leibler (KL) term included in the objective function.

14 . The method of claim 4 , wherein training the VAE comprises:

storing a first subset of activations generated by the VAE from the training dataset during a forward pass associated with training the VAE; and

recalculating a second subset of activations generated by the VAE based on the stored first subset of activations during a backward pass associated with training the VAE, wherein the second subset of activations is determined based on reducing a memory consumption associated with training of the VAE.

15 . The method of claim 4 , wherein training the VAE comprises storing a first portion of the one or more parameters using a first precision and storing a second portion of the one or more parameters using a second precision that is lower than the first precision.

16 . A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to perform the steps of:

inputting a training dataset into a variational autoencoder (VAE) comprising an encoder network, a prior network, and a decoder network;

training the VAE by updating one or more parameters of the VAE based on an objective function, the objective function controlling a smoothness of one or more outputs produced by the VAE when processing the set of training images, wherein training the VAE comprises:

applying, to the objective function, a plurality of balancing coefficients to generate a modified objective function, the plurality of balancing coefficients including a first balancing coefficient that has a first value during a first training period and a second balancing coefficient that has a second value, wherein the second value is higher than the first value during a second training period that occurs after the first training period, and

updating the one or more parameters of the VAE based on the modified objective function; and

producing generative output that reflects a first distribution of the training dataset by applying the decoder network to one or more values sampled from a second distribution of latent variables generated by the prior network.

17 . The non-transitory computer readable medium of claim 16 , wherein training the VAE further comprises updating the one or more parameters of the VAE based on regularization of a scaling parameter associated with batch normalization of one or more layers of the VAE via a regularization term that is added to the scaling parameter, the regularization term being based on a norm of a plurality of scaling parameters associated with a plurality of batch normalization layers of the VAE.

18 . The non-transitory computer readable medium of claim 16 , wherein applying the decoder network to the one or more values comprises applying a batch normalization to one or more layers of the VAE based on a momentum parameter that increases a rate at which a running statistic associated with the batch normalization catches up to a batch statistic associated with the batch normalization.

19 . The non-transitory computer readable medium of claim 16 , wherein the VAE comprises a hierarchy of groups of the latent variables, and wherein a first sample from a first group in the hierarchy is combined with a feature map and passed to a second group following the first group in the hierarchy for use in generating a second sample from the second group.

20 . The non-transitory computer readable medium of claim 19 , wherein the encoder network comprises a bottom-up model and a top-down model that perform bidirectional inference of the groups of the latent variables based on the training dataset.

21 . The non-transitory computer readable medium of claim 20 , wherein producing the generative output comprises:

executing the top-down model to sample the one or more values along the hierarchy of groups of the latent variables; and

inputting the sampled one or more values into the decoder network to produce the generative output.

22 . The non-transitory computer readable medium of claim 16 , wherein applying the decoder network to the one or more values comprises recalculating batch statistics associated with a batch normalization based on the one or more values sampled from the second distribution.

23 . A system, comprising:

a memory that stores instructions, and

a processor that is coupled to the memory and, when executing the instructions, is configured to:

sample one or more values from a first distribution of latent variables associated with an encoder network included in a variational autoencoder (VAE), wherein the encoder network comprises a first residual cell, wherein the first residual cell comprises a first batch-normalization (BN) layer fused with a first Swish activation function, a first convolutional layer following the first BN layer fused with the first Swish activation function, a second BN layer fused with a second Swish activation function, a second convolutional layer following the second BN layer fused with the second Swish activation function, and a first squeeze and excitation (SE) layer following the second convolutional layer;

train the VAE by updating one or more parameters of the VAE based on an objective function, the objective function controlling a smoothness of one or more outputs produced by the VAE when processing the set of training images, wherein training the VAE comprises:

applying, to the objective function, a plurality of balancing coefficients to generate a modified objective function, the plurality of balancing coefficients including a first balancing coefficient that has a first value during a first training period and a second balancing coefficient that has a second value, wherein the second value is higher than the first value during a second training period that occurs after the first training period, and

updating the one or more parameters of the machine learning model based on the modified objective function; and

sample from a second distribution to produce generative output associated with the data.

24 . The system of claim 23 , wherein the one or more values are sampled using a second residual cell comprising a third BN layer, a third convolutional layer following the third BN layer, a fourth BN layer fused with a third Swish activation function, and a depthwise separable convolution layer following the fourth BN layer.

25 . The system of claim 24 , wherein the second residual cell further comprises a fifth BN layer fused with a fourth Swish activation function, a fourth convolutional layer following the fifth BN layer, a sixth BN layer following the fourth convolutional layer, and a second SE layer following the sixth BN layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2020
From: VAHDAT, ARASH; KAUTZ, JAN
To: NVIDIA CORPORATION
Reel/Frame 054276/0466 →
Continuity (2)
Provisional Application 63041038 · Jun 18, 2020
Related Publication 20210397945A1 · Dec 23, 2021
References Cited (118)
US 20180101784A1 · Rolfe · 2018 [cited by examiner]
US 20190147339A1 · Nachum · 2019 [cited by examiner]
US 20190180732A1 · Ping et al. · 2019 [cited by applicant]
US 20190318244A1 · Alvarez et al. · 2019 [cited by applicant]
US 20190393903A1 · Mandt · 2019 [cited by examiner]
US 20200019859A1 · Kyriazopoulou Panagiotopoulou · 2020 [cited by examiner]
US 20200082534A1 · Nikolov et al. · 2020 [cited by applicant]
US 20200090050A1 · Rolfe et al. · 2020 [cited by applicant]
US 20200143240A1 · Baker · 2020 [cited by applicant]
US 20200380370A1 · Lie · 2020 [cited by examiner]
CN 109886388A · 2019 [cited by applicant]
CN 110506278A · 2019 [cited by applicant]
CN 110533620A · 2019 [cited by applicant]
CN 110895715A · 2020 [cited by applicant]
CN 111243045A · 2020 [cited by applicant]
CN 111258992A · 2020 [cited by applicant]
WO 2019067960A1 · 2019 [cited by applicant]
WO 2019079895A1 · 2019 [cited by applicant]
WO 2020064990A1 · 2020 [cited by applicant]
Lei Cai, Multi-Stage Variational Auto-Encoders for Coarse-to-Fine Image Generation, 2019, Proceedings of the 2019 SIAM International Conference on Data Mining (SDM), pp. 630-638 (Year: 2019). [cited by examiner]
Kingma et al., “Improved variational inference with inverse autoregressive flow”, 2017, arXiv, v2, pp. 1-16 (Year: 2017). [cited by examiner]
Yoshida et al., “Spectral norm regularization for improving the generalizability of deep learning”, 2017, arXiv, v1, pp. 1-12 (Year: 2017). [cited by examiner]
Hou et al., “Deep Feature Consistent Variational Autoencoder”, 2017, 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), vol. 2017, pp. 1133-1141 (Year: 2017). [cited by examiner]
Chollet, “Xception: Deep Learning With Depthwise Separable Convolutions”, 2017, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), vol. 2017, pp. 1251-1258 (Year: 2017). [cited by examiner]
Ioffe, “Batch Renormalization: Towards Reducing Minibatch Dependence in Batch-Normalized Models”, 2017, Advances in Neural Information Processing Systems, vol. 30 (2017), pp. 1-9 (Year: 2017). [cited by examiner]
Ramachandran et al., “Searching for Activation Functions”, 2017, arXiv, v2, pp. 1-13 (Year: 2017). [cited by examiner]
Hu et al., “Squeeze-and-Excitation Networks”, 2019, arXiv, v4, pp. 1-13 (Year: 2019). [cited by examiner]
Micikevicius et al., “Mixed Precision Training”, 2017, arXiv, v3, pp. 1-12 (Year: 2017). [cited by examiner]
Miyato et al., “Spectral normalization for generative adversarial networks”, 2018, arXiv, v1, pp. 1-26 (Year: 2018). [cited by examiner]
Li et al., “Leaf Classification Utilizing Densely Connected Convolutional Networks with a Self-gated Activation Function”, 2018, Intelligent Computing Methodologies, vol. 2018, pp. 370-375 (Year: 2018). [cited by examiner]
Liu et al., “Evolving Normalization-Activation Layers”, Jun. 11, 2020, arXiv, v4, pp. 1-16 (Year: 2020). [cited by examiner]
Gulati et al., “ResU-Net: Residual Convolutional Neural Network for Prostate MRI Segmentation”, May 16, 2020, arXiv, v1, pp. 1-5 (Year: 2020). [cited by examiner]
Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications”, 2017, arXiv, v1, pp. 1-9 (Year: 2019). [cited by examiner]
Sandler et al., “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, 2018, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), vol. 2018, pp. 4510-4520 (Year: 2018). [cited by examiner]
Razavi et al., “Generating Diverse High-Fidelity Images with VQ-VAE-2”, 2019, Advances in Neural Information Processing Systems 32, vol. 32, pp. 1-11 (Year: 2019). [cited by examiner]
Van Laarhoven, “L2 Regularization versus Batch and Weight Normalization”, 2017, arXiv, v1, pp. 1-9 (Year: 2017). [cited by examiner]
Luo et al., “Towards Understanding Regularization in Batch Normalization”, 2019, arXIv, v4, pp. 1-23 (Year: 2019). [cited by examiner]
Chen et al., “Training Deep Nets with Sublinear Memory Cost”, 2016, arXIv, v2, pp. 1-12 (Year: 2016). [cited by examiner]
Xie et al., “SRUN: Spectral Regularized Unsupervised Networks for Hyperspectral Target Detection”, 2019, IEEE Transactions on Geoscience and Remote Sensing, vol. 58, pp. 1463-1474 (Year: 2019). [cited by examiner]
Yu et al., “Survey of Multi-Task Learning”, Chinese Journal of Computers, DOI 10.11897/SP.J1016.2020.01340, vol. 43, No. 7, Jul. 2020, pp. 1340-1378. [cited by applicant]
He et al., “Identity Mappings in Deep Residual Networks”, In European conference on computer vision, DOI: 10.1007/978-3-319-46493-0 38, Springer, 2016, pp. 630-645. [cited by applicant]
Chen et al., “Training Deep Nets with Sublinear Memory Cost”, arXiv:1604.06174, 2016, pp. 1-12. [cited by applicant]
Martens et al., “Training Deep and Recurrent Networks with Hessian-Free Optimization”, In Neural networks: Tricks of the Trade, Springer, 2012, pp. 479-535. [cited by applicant]
Yoshida et al., “Spectral Norm Regularization for Improving the Generalizability of Deep Learning”, arXiv: 1705.10941, 2017, pp. 1-12. [cited by applicant]
Miyato et al., “Spectral normalization for generative adversarial networks”, In International Conference on Learning Representations, 2018, pp. 1-26. [cited by applicant]
Chen et al., “VFlow: More Expressive Generative Flows with Variational Data Augmentation”, arXiv:2002.09741, 2020, 16 pages. [cited by applicant]
Huang et al., “Augmented Normalizing Flows:Bridging the Gap Between Generative Flows and Latent Variable Models”, arXiv:2002.07101, 2020, 27 pages. [cited by applicant]
Ho et al., “Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design”, In Proceedings of the 36th International Conference on Machine Learning, arXiv: 1902.00275, 2019, 16 b… [cited by applicant]
Kingma et al., “Glow: Generative flow with Invertible 1x1 Convolutions”, In Advances in Neural Information Processing Systems, arXiv:1807.03039, 2018, pp. 10236-10245. [cited by applicant]
Dinh et al., “Density Estimation Using Real NVP”, arXiv:1605.08803, ICLR, 2016, pp. 1-32. [cited by applicant]
Tomczak et al., “VAE with a VampPrior”, In International Conference on Artificial Intelligence and Statistics, vol. 84, 2018, pp. 1214-1223. [cited by applicant]
Ma et al., “MAE: Mutual posterior-divergence regularization for variational autoencoders”, In The International Conference on Learning Representations (ICLR), arXiv:1901.01498, 2019, pp. 1-16. [cited by applicant]
Chen et al., “Variational Lossy Autoencoder”, arXiv:1611.02731, 2016, pp. 1-17. [cited by applicant]
Ma et al., “MaCow: Masked Convolutional Generative Flow”, In Advances in Neural Information Processing Systems, arXiv:1902.04208, 2019, pp. 1-18. [cited by applicant]
Menick et al., “Generating high fidelity images with subscale pixel networks and multidimensional upscaling”, In International Conference on Learning Representations, arXiv:1812.01608, 2019, pp. 1-15. [cited by applicant]
Parmar et al., “Image Transformer”, In Proceedings of the 35th International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Salimans et al., “PixelCNN++: Improving the pixelCNN with discretized logistic mixture likelihood and other modifications”, arXiv:1701.05517, 2017, pp. 1-10. [cited by applicant]
Oord et al., “Conditional Image Generation with PixelCNN Decoders”, In Advances in Neural Information Processing Systems, 2016, pp. 4790-4798. [cited by applicant]
Lecun, Yann, “The Mnist Database of handwritten digits”, Retrieved from http://yann.lecun.com/exdb/mnist/, 1998, 7 pages. [cited by applicant]
Krizhevsky, Alex, “Learning Multiple Layers of Features from Tiny Images”, Apr. 8, 2009, 60 pages. [cited by applicant]
Deng et al., “ImageNet: A Large-Scale Hierarchical Image Database”, In Computer Vision and Pattern Recognition (CVPR), 2009, pp. 248-255. [cited by applicant]
Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, In Proceedings of the EEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 4401-4410. [cited by applicant]
Brock et al., “Large scale GAN training for high fidelity natural image synthesis”, In International Conference on Learning Representations, arXiv:1809.11096, 2019, pp. 1-35. [cited by applicant]
Radford et al., “Unsupervised representation learning with deep convolutional generative adversarial networks”, arXiv:1511.06434, 2015, pp. 1-16. [cited by applicant]
Grover et al., “Bias Correction of Learned Generative Models using Likelihood-Free ImportanceWeighting”, In Advances in Neural Information Processing Systems, 2019, pp. 11056-11068. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization”, arXiv: 1412.6980, 2014, pp. 1-15. [cited by applicant]
Bjorck et al., “Understanding Batch Normalization”, 32nd Conference on Neural Information Processing Systems, 2018, pp. 1-12. [cited by applicant]
Oberman et al., “Lipschitz Regularized Deep Neural Networks Converge and Generalize”, Oct. 3, 2018, pp. 1-17. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes”, In the International Conference on Learning Representations (ICLR), arXiv:1312.6114, 2014, 14 pages. [cited by applicant]
Rezende et al., “Stochastic Backpropagation and Approximate Inference in Deep Generative Models”, In International Conference on Machine Learning, vol. 32, 2014, pp. 1278-1286. [cited by applicant]
Rezende et al., “Variational Inference with Normalizing Flows”, arXiv:1505.05770, vol. 37, 2015, 10 pages. [cited by applicant]
Kingma et al., “Improved Variational Inference with Inverse Autoregressive Flow”, In Advances in Neural Information Processing Systems, 2016, pp. 4743-4751. [cited by applicant]
Gregor et al., Draw: A Recurrent Neural Network for Image Generation, In 32nd International Conference on Machine Learning, vol. 37, 2015, pp. 1462-1471. [cited by applicant]
Cremer et al., “Inference Suboptimality in Variational Autoencoders”, In Proceedings of the 35th International Conference on Machine Learning, arXiv:1801.03558, 2018, 12 pages. [cited by applicant]
Marino et al., “Iterative Amortized Inference”, In Proceedings of the 35th International Conference on Machine earning, 2018, 10 pages. [cited by applicant]
MaalØe et al., “Auxiliary Deep Generative Models”, In Proceedings of the 33rd International Conference on Machine Learning, arXiv:1602.05473, 2016, 9 pages. [cited by applicant]
Ranganath et al., “Hierarchical Variational Models”, In Proceedings of the 33rd International Conference on Machine Learning, 2016, 10 pages. [cited by applicant]
Vahdat et al., “Undirected Graphical Models as Approximate Posteriors”, In International Conference on Machine Learning (ICML), arXiv:1901.03440, 2020, 13 pages. [cited by applicant]
Burda et al., “Importance Weighted Autoencoders”, ICLR 2016, arXiv:1509.00519, 2015, 14 pages. [cited by applicant]
Li et al., “Rényi Divergence Variational Inference”, In Advances in Neural Information Processing Systems, 2016, pp. 1073-1081. [cited by applicant]
Bornschein et al., “Bidirectional Helmholtz Machines”, In International Conference on Machine Learning, 2016, pp. 2511-2519. [cited by applicant]
Masrani et al., “The Thermodynamic Variational Objective”, In Advances in Neural Information Processing Systems, 2019, pp. 11521-11530. [cited by applicant]
Roeder et al., “Sticking the Landing: Simple, Lower-Variance Gradient Estimators for Variational Inference”, In Advances in Neural Information Processing Systems, arXiv:1703.09194, 2017, pp. 6925-6934. [cited by applicant]
Tucker et al., “Doubly Reparameterized Gradient Estimators for Monte Carlo Objectives”, arXiv:1810.04152, 2018, pp. 1-14. [cited by applicant]
Maddison et al., “The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”, ICLR, arXiv:1611.00712, 2016, pp. 1-20. [cited by applicant]
Jang et al., “Categorical Reparameterization with Gumbel-Softmax”, arXiv:1611.01144, ICLR, 2016, pp. 1-13. [cited by applicant]
Rolfe, Jason Tyler, “Discrete Variational Autoencoders”, arXiv:1609.02200, ICLR, 2016, pp. 1-33. [cited by applicant]
Vahdat et al., “DVAE++: Discrete Variational Autoencoders with Overlapping Transformations”, In International Conference on Machine Learning (ICML), arXiv:1802.04920, 2018, 16 pages. [cited by applicant]
Vahdat et al., “DVAE#: Discrete Variational Autoencoders with Relaxed Boltzmann Priors”, In Neural Information Processing Systems, arXiv:1805.07445, 2018, pp. 1-11. [cited by applicant]
Tucker et al., “REBAR: Low-variance, unbiased gradient estimates for discrete latent variable models”, In 31st Conference on Neural Information Processing Systems, 2017, pp. 2624-2633. [cited by applicant]
Grathwohl et al., “Backpropagation Through the Void: Optimizing Control Variates for Black-Box Gradient Estimation”, In International Conference on Learning Representations (ICLR), arXiv:1711.00123, 2018, pp. 1-17. [cited by applicant]
Bowman et al., “Generating Sentences from a Continuous Space”, In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, arXiv:1511.06349, 2016, pp. 10-21. [cited by applicant]
Razavi et al., “Preventing Posterior Collapse with delta-VAEs”, In The International Conference on Learning Representations (ICLR), arXiv:1901.03416, 2019, pp. 1-24. [cited by applicant]
Gulrajani et al., “PixeIVAE: A Latent Variable Model for Natural Images”, arXiv:1611.05013, 2016, pp. 1-9. [cited by applicant]
Lucas et al., “Don't Blame the Elbo! A Linear VAE Perspective on Posterior Collapse”, In Advances in Neural Information Processing Systems, arXiv:1911.02469, 2019, pp. 1-21. [cited by applicant]
Karras et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation”, In International Conference on Learning Representations, arXiv:1710.10196, 2018, pp. 1-26. [cited by applicant]
Barber et al., “Information Maximization in Noisy Channels : A Variational Approach”, In Advances in Neural Information Processing Systems, 2004, 8 pages. [cited by applicant]
Alemi et al., “Deep Variational Information Bottleneck”, In The International Conference on Learning Representations (ICLR), 2017, pp. 1-19. [cited by applicant]
Schwartz-Ziv et al., “Opening the black box of Deep Neural Networks via Information”, arXiv: 1703.00810, 2017, pp. 1-19. [cited by applicant]
Wu et al., “On the quantitative analysis of decoder-based generative models”, In The International Conference on Learning Representations (ICLR), arXiv:1611.04273, 2017, pp. 1-17. [cited by applicant]
Dieleman et al., “The challenge of realistic music generation:modelling raw audio at scale”, In Advances in Neural Information Processing Systems, 2018, pp. 7989-7999. [cited by applicant]
Chen et al., “PixelSNAIL: An Improved Autoregressive Generative Model”, In International Conference on Machine Learning, 2018, 9 pages. [cited by applicant]
Sadeghi et al., “PixelVAE++: Improved PixelVAE with Discrete Prior”, arXiv:1908.09948, 2019, pp. 1-12. [cited by applicant]
MaalØe et al., et al., “BIVA: A Very Deep Hierarchy of Latent Variables for Generative Modeling”, In 33rd Conference on Neural Information Processing Systems, 2019, pp. 6548-6558. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, In 32nd International Conference on Machine Learning, vol. 37, 2015, pp. 448-456. [cited by applicant]
Chollet, Francois, “Xception: Deep Learning with Depthwise Separable Convolutions”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, arXiv:1610.02357, 2017, pp. 1251-1258. [cited by applicant]
Razavi et al., “Generating Diverse High-Fidelity Images with VQ-VAE-2”, In Advances in Neural Information Processing Systems, 2019, pp. 14837-14847. [cited by applicant]
Oord et al., “Pixel Recurrent Neural Networks”, In Proceedings of the 33rd International Conference on International Conference on Machine Learning, JMLR. org, vol. 48, 2016, pp. 1747-1756. [cited by applicant]
SØnderby et al., “Ladder Variational Autoencoders”, In 30th Conference on Neural Information Processing Systems, 2016, pp. 3738-3746. [cited by applicant]
Klushyn et al., “Learning Hierarchical Priors in VAEs”, In 33rd Conference on Neural Information Processing Systems, arXiv:1905.04982, 2019, pp. 1-17. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, In Computer Vision and Pattern Recognition, DOI 10.1109/CVPR.2016.90, 2016, pp. 770-778. [cited by applicant]
Sandler et al., “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, In Proceedings of the IEEE conference on computer vision and pattern recognition, DOI 10.1109/CVPR.2018.00474, 2018, pp. 4510-4520. [cited by applicant]
Salimans et al., “Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks”, In 30th Conference on Neural Information Processing Systems, 2016, pp. 1-9. [cited by applicant]
Ramachandran et al., “Searching for Activation Functions”, arXiv:1710.05941, 2017, pp. 1-13. [cited by applicant]
Tan et al., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks”, In Proceedings of the 36th International Conference on Machine Learning, 2019, pp. 6105-6114. [cited by applicant]
Chen et al., “Residual Flows for Invertible Generative Modeling”, In Advances in Neural Information Processing Systems, arXiv:1906.02735, 2019, pp. 1-21. [cited by applicant]
Clevert et al., “Fast and accurate deep network learning by exponential linear units (ELUs)”, arXiv:1511.07289, 2015, pp. 1-14. [cited by applicant]
Hu et al., “Squeeze-and-Excitation Networks”, arXiv: 1709.01507, 2017, pp. 1-13. [cited by applicant]