IP Library Granted Patent US 11,599,972
Granted Patent B1
US 11,599,972 · App. 17/748,604 · Granted Mar 7, 2023

Method and system for lossy image or video encoding, transmission and decoding

Inventors: Jan Xu (London, GB); Chri Besenbruch (London, GB); Arsalan Zafar (London, GB)
Assignee: DEEP RENDER LTD.
G06T5/002G06N3/0454G06N3/08G06T3/40G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,599,972
App. No.
17/748,604
Granted
Mar 7, 2023
Kind
B1
Abstract

There is provided a method for lossy image or video encoding and transmission, including the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to produce a latent representation, performing a quantization process on the latent representation to produce a quantized latent, and transmitting the quantized latent to a second computer system.

Claims (35)

1. A method for lossy image or video encoding, transmission and decoding, the method comprising the steps of:

receiving an input image at a first computer system;

encoding the input image using a first trained neural network to produce a latent representation;

performing a quantization process on the latent representation to produce a quantized latent;

transmitting the quantized latent to a second computer system;

decoding the quantized latent using a trained denoising model, wherein the initial input to the trained denoising model is a sample from a standard normal conditioned with the upsampled quantized latent, to produce an output image;

wherein the output image is an approximation of the input image.

2. The method of claim 1 , wherein the trained denoising model is a second trained neural network.

3. The method of claim 1 , wherein the trained denoising performs an iterative process and includes a denoising function configured to predict a noise vector;

wherein the denoising function receives as input an output of the previous iterative step, the data based on the latent representation and parameters describing a noise distribution; and

the noise vector is applied to the output of the previous iterative step to obtain the output of the current iterative step.

4. The method of claim 3 , wherein the parameters describing the noise distribution specify the variance of the noise distribution.

5. The method of claim 3 , wherein the noise distribution is a gaussian distribution.

6. The method of claim 1 , wherein the data based on the latent representation is upsampled prior to the application of the trained denoising model.

7. A method of training one or more models including neural networks, the one or more models being for use in lossy image or video encoding, transmission and decoding, the method comprising the steps of:

receiving a first input training image;

encoding the first input training image using a first neural network to produce a latent representation;

performing a quantization process on the latent representation to produce a quantized latent;

decoding the quantized latent using a trained denoising model, wherein the initial input to the trained denoising model is a sample from a standard normal conditioned with the upsampled quantized latent, to produce an output image, wherein the output image is an approximation of the first input training image;

evaluating a loss function based on the rate of the quantized latent;

evaluating a gradient of the loss function;

back-propagating the gradient of the loss function through the first neural network to update the parameters of the first neural network;

repeating the above steps using a first set of training images to produce a first trained neural network.

8. The method of claim 7 , wherein the loss function includes a denoising loss; and

the denoising process includes a denoising function configured to predict a noise vector;

wherein the denoising function receives as input the first input training image with added noise, the data based on the latent representation and parameters describing a noise distribution;

the denoising loss is evaluated based on a difference between the predicted noise vector and the noise added to the first training image;

back-propagation the gradient of the loss function is additionally performed through the denoising model to update the parameters of the denoising model to produce a trained denoising model.

9. A data processing system, comprising at least one computer system configured to perform a method for lossy image or video encoding, transmission and decoding, the method comprising the steps of:

receiving an input image at a first computer system;

encoding the input image using a first trained neural network to produce a latent representation;

performing a quantization process on the latent representation to produce a quantized latent;

transmitting the quantized latent to a second computer system;

decoding the quantized latent using a trained denoising model, wherein the initial input to the trained denoising model is a sample from a standard normal conditioned with the upsampled quantized latent, to produce an output image;

wherein the output image is an approximation of the input image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2026
From: DEEP RENDER LTD
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 073864/0596 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2022
From: XU, JAN; BESENBRUCH, CHRI; ZAFAR, ARSALAN
To: DEEP RENDER LTD.
Reel/Frame 061771/0138 →
Priority Claims (1)
GB 2118730 · Dec 22, 2021 · national
Cited By (7)
US 12,323,634 US 12,354,244 US 12,470,715 US 12,501,050 US 12,511,797 US 12,536,713 US 12,555,043