IP Library Granted Patent US 12,731,298
Granted Patent B2
US 12,731,298 · App. 18/723,595 · Granted Sep 8, 2026

Method and data processing system for lossy image or video encoding, transmission and decoding

Inventors: Arsalan Zafar (London, GB); Jan Xu (London, GB); Christian Besenbruch (London, GB); Bilal Abbasi (London, GB); Aleksandar Cherganski (London, GB); Chris Finlay (London, GB); Christian Etmann (London, GB)
Assignee: InterDigital VC Holdings, Inc.
G06T9/002
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,298
App. No.
18/723,595
Granted
Sep 8, 2026
Kind
B2
Abstract

A method for lossy image and video encoding, transmission and decoding, the method comprising the steps of: receiving an input image at a first computer system; encoding the input image using a first trained neural network to produce a latent representation; performing a quantization process on the latent representation to produce a quantized latent; transmitting the quantized latent to a second computer system; decoding the quantized latent using a denoising process to produce an output image, wherein the output image is an approximation of the input image.

Claims (30)

1 . A method of training one or more neural networks, the one or more neural networks being for use in lossy image or video encoding, transmission and decoding, the method comprising the steps of:

receiving a first input training image;

encoding the first input training image using a first neural network to produce a latent representation;

performing a quantization process on the latent representation to produce a quantized latent;

decoding the quantized latent using a second neural network to produce an output image, wherein the output image is an approximation of the input training image; evaluating a loss function based on differences between the output image and the input training image;

evaluating a gradient of the loss function;

back-propagating the gradient of the loss function through the first neural network and the second neural network to update the parameters of the first neural network and the second neural network; and

repeating the above steps using a first set of training images to produce a first trained neural network and a second trained neural network;

wherein the differences between the output image and the input training image is determined based on the output of a neural network acting as a discriminator;

the neural network acting as a discriminator receives the output image as an input and outputs one or more values associated with one or more sub-sections of the output image, wherein each value indicates the likelihood that the corresponding sub-section of the output image is a fake sub-section, and

back-propagation of the gradient of the loss function is additionally used to update the parameters of the neural network acting as a discriminator.

2 . The method of claim 1 , wherein the output of the neural network acting as a discriminator is converted to a probability distribution, wherein the value of the probability distribution is defined for each of the one or more sub-sections and is proportionate to the value indicating the likelihood that the corresponding sub-section of the output image is a fake sub-section.

3 . The method of claim 2 , wherein the conversion to a probability distribution is performed using a softmax function.

4 . The method of claim 1 , further including the step of providing the one or more sub-sections of the output image to a neural network acting as a sub-discriminator;

wherein the neural network acting as a sub-discriminator outputs one or more values associated with the one or more sub-sections of the output image, each value indicating the likelihood that the corresponding sub-section of the output image is a fake sub-section; and the differences between the output image and the input training image is additionally determined based on the output of the neural network acting as a sub-discriminator; and back-propagation of the gradient of the loss function is additionally used to update the parameters of the neural network acting as a sub-discriminator.

5 . The method of claim 4 , wherein the one or more sub-sections of the output image are determined by sampling the probability distribution.

6 . The method of claim 4 , wherein two to five sub-sections of the output image are provided to the neural network acting as a sub-discriminator, preferably wherein three sub-sections of the output image are provided.

7 . The method of claim 1 , wherein

the neural network acting as a discriminator additionally receives the quantized latent as an input.

8 . The method of claim 2 , further comprising the steps of, after the output of the neural network acting as a discriminator is converted to a probability distribution:

sampling the probability distribution to select a sub-section of the output image; encoding the corresponding sub-section of the input image to the selected sub-section of the output image using the first neural network to produce a sub-latent representation; performing a quantization process on the sub-latent representation to produce a quantized sub-latent;

decoding the quantized sub-latent using a second neural network to produce an output sub-image, wherein the output sub-image is an approximation of the sub-section of the input image;

wherein the evaluation of the loss function and back propagation of the gradient of the loss function to update the parameters of the neural networks is performed based on the output sub-image and the sub-section of the input image.

9 . A method for lossy image or video encoding, transmission and decoding, the method comprising the steps of:

receiving an input image at a first computer system;

encoding the first input training image using a first trained neural network to produce a latent representation;

performing a quantization process on the latent representation to produce a quantized latent;

transmitting the quantized latent to a second computer system; and

decoding the quantized latent using a second trained neural network to produce an output image, wherein the output image is an approximation of the input training image; wherein the first trained neural network and the second trained neural network have been trained according to the method of claim 1 .

10 . A data processing system configured to perform the method of claim 1 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2026
From: ZAFAR, ARSALAN; XU, JAN; BESENBRUCH, CHRISTIAN; ABBASI, BILAL; CHERGANSKI, ALEKSANDAR; FINLAY, CHRIS; ETMANN, CHRISTIAN
To: DEEP RENDER LTD.
Reel/Frame 074290/0947 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2026
From: DEEP RENDER LTD
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 073864/0596 →
Priority Claims (4)
GB 2118730 · Dec 22, 2021 · national
GB 2118863 · Dec 22, 2021 · national
GB 2200899 · Jan 25, 2022 · national
GB 2201471 · Feb 4, 2022 · national
Continuity (1)
Related Publication 20250086843A1 · Mar 13, 2025
References Cited (14)
US 11468265B2 · Chhabra · 2022 [cited by examiner]
US 11544880B2 · Park · 2023 [cited by examiner]
US 12481877B2 · Brehmer · 2025 [cited by examiner]
US 12505342B2 · Kim · 2025 [cited by examiner]
US 12505595B2 · Liu · 2025 [cited by examiner]
US 12541955B1 · Hatamizadeh · 2026 [cited by examiner]
US 20220215265A1 · Jiang · 2022 [cited by examiner]
WO 2021220008A1 · 2021 [cited by applicant]
Chitwan, Saharia , et al., “Image Super-Resolution via Iterative Refinement”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,Jun. 30, 2021 (Jun. 30, 2021). [cited by applicant]
Ho, J. , et al., Denoising diffusion probabilistic models. arXiv preprint arXiv:2006.11239, 2020. [cited by applicant]
Ma, S. , et al., “Image and Video Compression With Neural Networks: A Review”, Apr. 10, 2019 (Apr. 10, 2019), p. 1-16. [cited by applicant]
Mentzer, F. , et al., High-fidelity generative image compression. arXiv preprint arXiv:2006.09965, 2020. [cited by applicant]
Yang, R. , et al., “Lossy Image Compression with Conditional Diffusion Models”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,Dec. 9, 2022 (Dec. 9, 2022). [cited by applicant]
International Search Report, dated Apr. 3, 2023, and Written Opinion issued in International Application No. PCT/EP2022/087271. [cited by applicant]