IP Library Granted Patent US 11,336,908
Granted Patent B2
US 11,336,908 · App. 16/586,837 · Granted May 17, 2022

Compressing images using neural networks

Inventors: Daniel Pieter Wierstra (London, GB); Karol Gregor (London, GB); Frederic Olivier Besse (London, GB)
Assignee: DeepMind Technologies Limited
H04N19/30G06N3/0454G06N3/0472G06N3/08G06T9/002H04N19/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,336,908
App. No.
16/586,837
Granted
May 17, 2022
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for compressing images using neural networks. One of the methods includes receiving an image; processing the image using an encoder neural network, wherein the encoder neural network is configured to receive the image and to process the image to generate an output defining values of a first number of latent variables that each represent a feature of the image; generating a compressed representation of the image using the output defining the values of the first number of latent variables; and providing the compressed representation of the image for use in generating a reconstruction of the image.

Claims (44)

1. A method comprising:

receiving a lossy compressed representation of an image, wherein the lossy compressed representation of the image defines values of a first number of latent variables that each represent a feature of the image; and

generating a reconstruction of the image from the lossy compressed representation of the image, comprising:

selecting a value of one or more additional latent variables that are not in the first number of latent variables randomly from a prior distribution; and

generating the reconstruction of the image by conditioning a generative neural network on the values of the first number of latent variables and the randomly selected values of the additional latent variables that are not in the first number of latent variables, wherein:

the generative neural network has been trained jointly with an encoder neural network as a variational auto-encoder network that includes the encoder neural network and the generative neural network, and the encoder neural network is configured to process the image to generate an encoder output that define parameters of statistical distributions of the first number of latent variables and the additional latent variables.

2. The method of claim 1 , wherein generating the reconstruction comprises:

decompressing the values of the first number of latent variables from the lossy compressed representation.

3. The method of claim 1 , wherein receiving the lossy compressed representation comprises:

retrieving the lossy compressed representation from memory.

4. The method of claim 1 , wherein receiving the lossy compressed representation comprises:

receiving the lossy compressed representation over a network from an encoder system.

5. The method of claim 1 , further comprising generating the lossy compressed representation of the image, comprising:

processing the image using the encoder neural network, wherein the encoder neural network is configured to receive the image and to process the image to generate an output defining values of a second number of latent variables that is greater than the first number; and

generating the lossy compressed representation of the image using the first number of the latent variables.

6. The method of claim 1 , wherein the features are arranged in a hierarchy from least abstract representation of the image to most abstract representation of the image, and wherein the first number of latent variables are the latent variables that represent features at a predetermined number of highest levels in the hierarchy.

7. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:

receiving a lossy compressed representation of an image, wherein the lossy compressed representation of the image defines values of a first number of latent variables that each represent a feature of the image; and

generating a reconstruction of the image from the lossy compressed representation of the image, comprising:

selecting a value of one or more additional latent variables that are not in the first number of latent variables randomly from a prior distribution; and

generating the reconstruction of the image by conditioning a generative neural network on the values of the first number of latent variables and the randomly selected values of the additional latent variables that are not in the first number of latent variables, wherein:

the generative neural network has been trained jointly with an encoder neural network as a variational auto-encoder network that includes the encoder neural network and the generative neural network, and the encoder neural network is configured to process the image to generate an encoder output that define parameters of statistical distributions of the first number of latent variables and the additional latent variables.

8. The system of claim 7 , wherein generating the reconstruction comprises:

decompressing the values of the first number of latent variables from the lossy compressed representation.

9. The system of claim 7 , wherein receiving the lossy compressed representation comprises:

retrieving the lossy compressed representation from memory.

10. The system of claim 7 , wherein receiving the lossy compressed representation comprises:

receiving the lossy compressed representation over a network from an encoder system.

11. The system of claim 7 , the operations further comprising generating the lossy compressed representation of the image, comprising:

processing the image using the encoder neural network, wherein the encoder neural network is configured to receive the image and to process the image to generate an output defining values of a second number of latent variables that is greater than the first number; and

generating the lossy compressed representation of the image using the first number of the latent variables.

12. The system of claim 7 , wherein the features are arranged in a hierarchy from least abstract representation of the image to most abstract representation of the image, and wherein the first number of latent variables are the latent variables that represent features at a predetermined number of highest levels in the hierarchy.

13. One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving a lossy compressed representation of an image, wherein the lossy compressed representation of the image defines values of a first number of latent variables that each represent a feature of the image; and

generating a reconstruction of the image from the lossy compressed representation of the image, comprising:

selecting a value of one or more additional latent variables that are not in the first number of latent variables randomly from a prior distribution; and

generating the reconstruction of the image by conditioning a generative neural network on the values of the first number of latent variables and the randomly selected values of the additional latent variables that are not in the first number of latent variables, wherein:

the generative neural network has been trained jointly with an encoder neural network as a variational auto-encoder network that includes the encoder neural network and the generative neural network, and the encoder neural network is configured to process the image to generate an encoder output that define parameters of statistical distributions of the first number of latent variables and the additional latent variables.

14. The computer-readable storage media of claim 13 , wherein generating the reconstruction comprises:

decompressing the values of the first number of latent variables from the lossy compressed representation.

15. The computer-readable storage media of claim 13 , wherein receiving the lossy compressed representation comprises:

retrieving the lossy compressed representation from memory.

16. The computer-readable storage media of claim 13 , wherein receiving the lossy compressed representation comprises:

receiving the lossy compressed representation over a network from an encoder system.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071109/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2020
From: WIERSTRA, DANIEL PIETER; GREGOR, KAROL; BESSE, FREDERIC OLIVIER
To: GOOGLE INC.
Reel/Frame 051670/0641 →
CHANGE OF NAME Recorded Jan 30, 2020
From: GOOGLE INC.
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 051675/0942 →
Continuity (3)
Continuation 15396332 · Dec 30, 2016
Provisional Application 62292167 · Feb 5, 2016
Related Publication 20200029084A1 · Jan 23, 2020