IP Library Granted Patent US 10,373,055
Granted Patent B1
US 10,373,055 · App. 15/600,696 · Granted Aug 6, 2019

Training variational autoencoders to generate disentangled latent factors

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,373,055
App. No.
15/600,696
Granted
Aug 6, 2019
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a variational auto-encoder (VAE) to generate disentangled latent factors on unlabeled training images. In one aspect, a method includes receiving the plurality of unlabeled training images, and, for each unlabeled training image, processing the unlabeled training image using the VAE to determine the latent representation of the unlabeled training image and to generate a reconstruction of the unlabeled training image in accordance with current values of the parameters of the VAE, and adjusting current values of the parameters of the VAE by optimizing a loss function that depends on a quality of the reconstruction and also on a degree of independence between the latent factors in the latent representation of the unlabeled training image.

Claims (31)

1. A method performed by one or more computers for training a variational auto-encoder (VAE) to generate disentangled latent factors on a plurality of unlabeled training images,

wherein the VAE has a plurality of parameters and is configured to receive an input image, process the input image to determine a latent representation of the input image that includes a plurality of latent factors, and to process the latent representation to generate a reconstruction of the input image, and

wherein the method comprises:

receiving the plurality of unlabeled training images, and, for each unlabeled training image:

processing the unlabeled training image using the VAE to determine the latent representation of the unlabeled training image and to generate a reconstruction of the unlabeled training image in accordance with current values of the parameters of the VAE, and

adjusting current values of the parameters of the VAE by determining a gradient of a loss function with respect to the parameters of the VAE, wherein the loss function depends on a quality of the reconstruction of the unlabeled training image and also on a degree of independence between the latent factors in the latent representation of the unlabeled training image.

2. The method of claim 1 , wherein the loss function is of the form L=Q−B(KL), where Q is a term that depends on the quality of the reconstruction of the unlabeled training image, KL is a term that measures the degree of independence between the latent factors in the latent representation of the unlabeled training image and an effective capacity of the latent bottleneck, and B is a tunable parameter.

3. The method of claim 2 , wherein B is a value in the range between 2 exclusive and 250 inclusive.

4. The method of claim 3 , wherein B is four.

5. The method of claim 3 , wherein the value of B is dependent on a number of latent factors in the latent representation of the input image.

6. The method of claim 1 , wherein the generative factors are densely sampled from their respective continuous distributions among the plurality of unlabeled training images.

7. The method of claim 1 , wherein the degree of independence between the latent factors is computed within a latent bottleneck with restricted effective capacity.

8. The method of claim 1 , wherein the VAE includes a latent bottleneck layer and a capacity of the latent bottleneck layer is adjusted simultaneously with the degree of independence.

9. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations for training a variational auto-encoder (VAE) to generate disentangled latent factors on a plurality of unlabeled training images, wherein the VAE has a plurality of parameters and is configured to receive an input image, process the input image to determine a latent representation of the input image that includes a plurality of latent factors, and to process the latent representation to generate a reconstruction of the input image, and wherein the operations comprise:

receiving the plurality of unlabeled training images, and, for each unlabeled training image:

processing the unlabeled training image using the VAE to determine the latent representation of the unlabeled training image and to generate a reconstruction of the unlabeled training image in accordance with current values of the parameters of the VAE, and

adjusting current values of the parameters of the VAE by determining a gradient of a loss function with respect to the parameters of the VAE, wherein the loss function that-depends on a quality of the reconstruction of the unlabeled training image and also on a degree of independence between the latent factors in the latent representation of the unlabeled training image.

10. The system of claim 9 , wherein the loss function is of the form L=Q−B(KL), where Q is a term that depends on the quality of the reconstruction of the unlabeled training image, KL is a term that measures the degree of independence between the latent factors in the latent representation of the unlabeled training image and an effective capacity of the latent bottleneck, and B is a tunable parameter.

11. The system of claim 10 , wherein B is a value in the range between 2 exclusive and 250 inclusive.

12. The system of claim 11 , wherein B is four.

13. The system of claim 11 , wherein the value of B is dependent on a number of latent factors in the latent representation of the input image.

14. The system of claim 9 , wherein the generative factors are densely sampled from their respective continuous distributions among the plurality of unlabeled training images.

15. The system of claim 9 , wherein the degree of independence between the latent factors is computed within a latent bottleneck with restricted effective capacity.

16. The system of claim 9 , wherein the VAE includes a latent bottleneck layer and a capacity of the latent bottleneck layer is adjusted simultaneously with the degree of independence.

17. A computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for training a variational auto-encoder (VAE) to generate disentangled latent factors on a plurality of unlabeled training images, wherein the VAE is configured to receive an input image, process the input image to determine a latent representation of the input image that includes a plurality of latent factors, and to process the latent representation to generate a reconstruction of the input image, and wherein the operations comprise:

receiving the plurality of unlabeled training images, and, for each unlabeled training image:

processing the unlabeled training image using the VAE to determine the latent representation of the unlabeled training image and to generate a reconstruction of the unlabeled training image in accordance with current values of the parameters of the VAE, and

adjusting current values of the parameters of the VAE by determining a gradient of a loss function with respect to the parameters of the VAE, wherein the loss function depends on a quality of the reconstruction of the unlabeled training image and also on a degree of independence between the latent factors in the latent representation of the unlabeled training image.

18. The computer storage medium of claim 17 , wherein the loss function is of the form L=Q−B(KL), where Q is a term that depends on the quality of the reconstruction of the unlabeled training image, KL is a term that measures the degree of independence between the latent factors in the latent representation of the unlabeled training image and an effective capacity of the latent bottleneck, and B is a tunable parameter.

19. The computer storage medium of claim 18 , wherein B is a value in the range between 2 exclusive and 250 inclusive.

20. The computer storage medium of claim 18 , wherein B is four.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071109/0414 →
CORRECTIVE ASSIGNMENT TO CORRECT THE DECLARATION PREVIOUSLY RECORDED AT REEL: 044567 FRAME: 0001. ASSIGNOR(S) HEREBY CONFIRMS THE DECLARATION . Recorded Jan 13, 2022
From: DEEPMIND TECHNOLOGIES LIMITED
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 058721/0801 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2017
From: GOOGLE INC.
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 044242/0116 →
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2017
From: MATTHEY-DE-L'ENDROIT, LOIC; PAL, ARKA TILAK; MOHAMED, SHAKIR; GLOROT, XAVIER; HIGGINS, IRINA; LERCHNER, ALEXANDER
To: GOOGLE INC.
Reel/Frame 042672/0369 →
Cited By (6)
US 12,327,188 US 12,333,427 US 12,412,089 US 12,511,529 US 12,530,592 US 12,536,664