Deep hierarchical variational autoencoder
One embodiment of the present invention sets forth a technique for performing machine learning. The technique includes inputting a training dataset into a variational autoencoder (VAE) comprising an encoder network, a prior network, and a decoder network. The technique also includes training the VAE by updating one or more parameters of the VAE based on a smoothness of one or more outputs produced by the VAE from the training dataset. The technique further includes producing generative output that reflects a first distribution of the training dataset by applying the decoder network to one or more values sampled from a second distribution of latent variables generated by the prior network.
1 . A method for performing machine learning, comprising:
inputting a set of training images into a machine learning model that comprises an encoder portion, a prior, and a decoder portion;
training the machine learning model by updating one or more parameters of the machine learning model based on an objective function, the objective function controlling a smoothness of one or more outputs produced by the machine learning model when processing the set of training images, wherein training the machine learning model comprises:
applying, to the objective function, a plurality of balancing coefficients to generate a modified objective function, the plurality of balancing coefficients including a first balancing coefficient that has a first value during a first training period and a second balancing coefficient that has a second value, wherein the second value is higher than the first value during a second training period that occurs after the first training period, and
updating the one or more parameters of the machine learning model based on the modified objective function; and
producing a new image that reflects one or more visual attributes associated with the set of training images by applying the decoder portion to a value generated based on an output of the prior.
2 . The method of claim 1 , wherein the new image comprises a face that is not found in the set of training images.
3 . The method of claim 1 , wherein the new image comprises an animal or a vehicle that is not found in the set of training images.
4 . A method for performing machine learning, comprising:
inputting a training dataset into a variational autoencoder (VAE) comprising an encoder network, a prior network, and a decoder network;
training the VAE by updating one or more parameters of the VAE based on an objective function, the objective function controlling a smoothness of one or more outputs produced by the VAE when processing the training dataset, wherein training the VAE comprises:
applying, to the objective function, a plurality of balancing coefficients to generate a modified objective function, the plurality of balancing coefficients including a first balancing coefficient that has a first value during a first training period and a second balancing coefficient that has a second value, wherein the second value is higher than the first value during a second training period that occurs after the first training period, and
updating the one or more parameters of the VAE based on the objective function; and
producing generative output that reflects a first distribution of the training dataset by applying the decoder network to one or more values sampled from a second distribution of latent variables generated by the prior network.
5 . The method of claim 4 , wherein applying the decoder network to the one or more values comprises applying batch normalization to one or more layers of the decoder network based on a momentum parameter that increases a rate at which a running statistic associated with the batch normalization catches up to a batch statistic associated with the batch normalization.
6 . The method of claim 5 , wherein training the VAE comprises applying a regularization parameter to a scaling parameter associated with the batch normalization.
7 . The method of claim 5 , wherein applying the batch normalization to the one or more layers comprises combining the batch normalization with a Swish activation function.
8 . The method of claim 5 , wherein applying the batch normalization to the one or more layers of the decoder network comprises recalculating batch statistics associated with the batch normalization based on the one or more values sampled from the second distribution.
9 . The method of claim 4 , wherein the VAE comprises a hierarchy of groups of the latent variables, and wherein a first sample from a first group in the hierarchy is combined with a feature map and passed to a second group following the first group in the hierarchy for use in generating a second sample from the second group.
10 . The method of claim 4 , wherein the VAE comprises a residual cell and the residual cell comprises a first batch-normalization (BN) layer with a first Swish activation function, a first convolutional layer following the first BN layer with the first Swish activation function, a second BN layer with a second Swish activation function, a second convolutional layer following the second BN layer with the second Swish activation function, and a squeeze and excitation (SE) layer.
11 . The method of claim 4 , wherein the VAE comprises a residual cell and the residual cell comprises a first BN layer, a first convolutional layer following the first BN layer, a second BN layer with a first Swish activation function, and a depthwise separable convolution layer following the second BN layer.
12 . The method of claim 11 , wherein the residual cell further comprises a third BN layer with a second Swish activation function, a second convolutional layer following the third BN layer, a fourth BN layer following the second convolutional layer, and an SE layer following the fourth BN layer.
13 . The method of claim 1 , wherein the plurality of balancing coefficients includes a third balancing coefficient having a value that is determined based on a corresponding value of a Kullback-Leibler (KL) term included in the objective function.
14 . The method of claim 4 , wherein training the VAE comprises:
storing a first subset of activations generated by the VAE from the training dataset during a forward pass associated with training the VAE; and
recalculating a second subset of activations generated by the VAE based on the stored first subset of activations during a backward pass associated with training the VAE, wherein the second subset of activations is determined based on reducing a memory consumption associated with training of the VAE.
15 . The method of claim 4 , wherein training the VAE comprises storing a first portion of the one or more parameters using a first precision and storing a second portion of the one or more parameters using a second precision that is lower than the first precision.
16 . A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to perform the steps of:
inputting a training dataset into a variational autoencoder (VAE) comprising an encoder network, a prior network, and a decoder network;
training the VAE by updating one or more parameters of the VAE based on an objective function, the objective function controlling a smoothness of one or more outputs produced by the VAE when processing the set of training images, wherein training the VAE comprises:
applying, to the objective function, a plurality of balancing coefficients to generate a modified objective function, the plurality of balancing coefficients including a first balancing coefficient that has a first value during a first training period and a second balancing coefficient that has a second value, wherein the second value is higher than the first value during a second training period that occurs after the first training period, and
updating the one or more parameters of the VAE based on the modified objective function; and
producing generative output that reflects a first distribution of the training dataset by applying the decoder network to one or more values sampled from a second distribution of latent variables generated by the prior network.
17 . The non-transitory computer readable medium of claim 16 , wherein training the VAE further comprises updating the one or more parameters of the VAE based on regularization of a scaling parameter associated with batch normalization of one or more layers of the VAE via a regularization term that is added to the scaling parameter, the regularization term being based on a norm of a plurality of scaling parameters associated with a plurality of batch normalization layers of the VAE.
18 . The non-transitory computer readable medium of claim 16 , wherein applying the decoder network to the one or more values comprises applying a batch normalization to one or more layers of the VAE based on a momentum parameter that increases a rate at which a running statistic associated with the batch normalization catches up to a batch statistic associated with the batch normalization.
19 . The non-transitory computer readable medium of claim 16 , wherein the VAE comprises a hierarchy of groups of the latent variables, and wherein a first sample from a first group in the hierarchy is combined with a feature map and passed to a second group following the first group in the hierarchy for use in generating a second sample from the second group.
20 . The non-transitory computer readable medium of claim 19 , wherein the encoder network comprises a bottom-up model and a top-down model that perform bidirectional inference of the groups of the latent variables based on the training dataset.
21 . The non-transitory computer readable medium of claim 20 , wherein producing the generative output comprises:
executing the top-down model to sample the one or more values along the hierarchy of groups of the latent variables; and
inputting the sampled one or more values into the decoder network to produce the generative output.
22 . The non-transitory computer readable medium of claim 16 , wherein applying the decoder network to the one or more values comprises recalculating batch statistics associated with a batch normalization based on the one or more values sampled from the second distribution.
23 . A system, comprising:
a memory that stores instructions, and
a processor that is coupled to the memory and, when executing the instructions, is configured to:
sample one or more values from a first distribution of latent variables associated with an encoder network included in a variational autoencoder (VAE), wherein the encoder network comprises a first residual cell, wherein the first residual cell comprises a first batch-normalization (BN) layer fused with a first Swish activation function, a first convolutional layer following the first BN layer fused with the first Swish activation function, a second BN layer fused with a second Swish activation function, a second convolutional layer following the second BN layer fused with the second Swish activation function, and a first squeeze and excitation (SE) layer following the second convolutional layer;
train the VAE by updating one or more parameters of the VAE based on an objective function, the objective function controlling a smoothness of one or more outputs produced by the VAE when processing the set of training images, wherein training the VAE comprises:
applying, to the objective function, a plurality of balancing coefficients to generate a modified objective function, the plurality of balancing coefficients including a first balancing coefficient that has a first value during a first training period and a second balancing coefficient that has a second value, wherein the second value is higher than the first value during a second training period that occurs after the first training period, and
updating the one or more parameters of the machine learning model based on the modified objective function; and
sample from a second distribution to produce generative output associated with the data.
24 . The system of claim 23 , wherein the one or more values are sampled using a second residual cell comprising a third BN layer, a third convolutional layer following the third BN layer, a fourth BN layer fused with a third Swish activation function, and a depthwise separable convolution layer following the fourth BN layer.
25 . The system of claim 24 , wherein the second residual cell further comprises a fifth BN layer fused with a fourth Swish activation function, a fourth convolutional layer following the fifth BN layer, a sixth BN layer following the fourth convolutional layer, and a second SE layer following the sixth BN layer.