IP Library Granted Patent US 12,739,383
Granted Patent B2
US 12,739,383 · App. 18/513,581 · Granted Sep 15, 2026

Image encoding and decoding, video encoding and decoding: methods, systems and training methods

Inventors: Chri Besenbruch (London, GB); Aleksandar Cherganski (London, GB); Christopher Finlay (London, GB); Alexander Lytchier (London, GB); Jonathan Rayner (London, GB); Tom Ryder (London, GB); Jan Xu (London, GB); Arsalan Zafar (London, GB)
Assignee: InterDigital VC Holdings, Inc.
H04N19/13G06V10/422H04N19/124H04N19/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,739,383
App. No.
18/513,581
Granted
Sep 15, 2026
Kind
B2
Abstract

Lossy or lossless compression and transmission, comprising the steps of: (i) receiving an input image; (ii) encoding it to produce a y latent representation; (iii) encoding the y latent representation to produce a z hyperlatent representation; (iv) quantizing the z hyperlatent representation to produce a quantized z hyperlatent representation; (v) entropy encoding the quantized z hyperlatent representation into a first bitstream, (vi) processing the quantized z hyperlatent representation to obtain a location entropy parameter μy, an entropy scale parameter σy, and a context matrix Ay of the y latent representation; (vii) processing the y latent representation, the location entropy parameter μy and the context matrix Ay, to obtain quantized latent residuals; (viii) entropy encoding the quantized latent residuals into a second bitstream; and (ix) transmitting the bitstreams.

Claims (38)

1 . A computer implemented method of training a neural network for use in lossy image or video compression, the method comprising:

(i) receiving an input image;

(ii) encoding the input image using a first neural network to produce a latent representation, and decoding the latent representation using a second neural network to produce a reconstruction of the input image;

(iii) producing a first feature map, associated with the input image;

(iv) producing a second feature map, associated with the reconstruction of the input image;

(v) evaluating a function based on differences between the first feature map and the second feature map;

(vi) evaluating a gradient of the function;

(vii) back-propagating the gradient of the function through the first neural network and the second neural network to update the weights of the first neural network and the second neural network;

(viii) repeating steps (i) to (vii) to produce a trained first neural network and a trained second neural network.

2 . The method of claim 1 , comprising producing the first feature map and the second feature map using a third neural network.

3 . The method of claim 2 , wherein the third neural network comprises a neural network trained for a task other than image or video compression.

4 . The method of claim 3 , wherein the task other than image or video compression comprises a classification task.

5 . The method of claim 2 , wherein the first feature map and second feature map comprise one or more outputs from one or more layers of the third neural network.

6 . The method of claim 5 , wherein the first feature map and/or the second feature map comprise one or more outputs from one or more intermediate layers of the third neural network.

7 . The method of claim 6 , wherein the third neural network comprises a neural network trained for a task other than for use in lossy image or video compression.

8 . The method of claim 1 , wherein evaluating the function comprises estimating a difference metric between the first feature map and the second feature map.

9 . The method of claim 8 , wherein the difference metric comprises a cosine distance metric.

10 . The method of claim 8 , wherein the difference metric comprises a mean-squared error metric.

11 . The method of claim 8 , wherein the estimating the difference metric comprises estimating differences between the first and second feature maps with the first and second feature maps in a first operand order, and with the first and second feature maps in a second operand order.

12 . The method of claim 1 , wherein the function is further based on differences between the input image and the reconstruction of the input image.

13 . The method of claim 1 , wherein the function defines a first discriminator network and wherein evaluating the function comprises, with the first discriminator network, estimating a probability that the first and/or second feature maps are respectively associated with the input image or the reconstruction of the input image.

14 . The method of claim 13 , comprising:

(ix) back-propagating the gradient of the function through the first discriminator network to update the weights of the discriminator network; and

(x) repeating steps (i) to (vi) and (ix) to produce a trained first discriminator network.

15 . The method of claim 13 , wherein the first and/or second feature maps each comprise first and second tensors of different dimensions, wherein the first discriminator network comprises a plurality of sub-networks, and wherein the method comprises:

with the plurality of sub-networks combining the first and second tensors of different dimensions before estimating said probability.

16 . The method of claim 13 , wherein the first and/or second feature maps comprise one or more outputs from one or more layers of a third neural network, wherein evaluating the function comprises:

with the first discriminator network, estimating a probability that the outputs from a first layer of the third neural network are respectively associated with the input image or the reconstruction of the input image; and

with the second discriminator network, estimating a probability that the outputs from a second layer of the third neural network are respectively associated with the input image or the reconstruction of the input image.

17 . A data processing system configured to perform the method of claim 1 .

18 . A non-transitory computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of claim 1 .

19 . A decoder method comprising:

receiving a bitstream encoded with a latent representation of an image; and

decoding the latent representation using a decoder neural network to reconstruct the image;

wherein the decoder neural network is trained according to the training method of claim 1 , and wherein the training loss function is based on differences between feature maps associated with training images and feature maps associated with corresponding reconstructions of the training images.

20 . A decoder apparatus, the apparatus comprising:

a processor configured to execute a decoder neural network to decode a latent representation of an image to produce a reconstruction of the image;

the decoder neural network trained according to the training method of claim 1 , wherein the training loss function is based on differences between feature maps associated with training images and feature maps associated with corresponding reconstructions of the training images.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 74702 FRAME: 313. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 21, 2026
From: BESENBRUCH, CHRI; CHERGANSKI, ALEKSANDAR; FINLAY, CHRISTOPHER; LYTCHIER, ALEXANDER; RAYNER, JONATHAN; RYDER, TOM; XU, JAN; ZAFAR, ARSALAN
To: DEEP RENDER LTD.
Reel/Frame 075441/0391 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2026
From: DEEP RENDER LTD
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 073864/0596 →
Priority Claims (1)
GB 2016824 · Oct 23, 2020 · national
Continuity (5)
Continuation 18105338 · Feb 3, 2023
Continuation 17748502 · May 19, 2022
Continuation 17740798 · May 10, 2022
Continuation PCTGB2021052770 · Oct 25, 2021
Related Publication 20240107022A1 · Mar 28, 2024
References Cited (21)
US 10373300B1 · Besenbruch et al. · 2019 [cited by applicant]
US 10489936B1 · Zafar et al. · 2019 [cited by applicant]
US 11475542B2 · Munkberg · 2022 [cited by examiner]
US 11606560B2 · Besenbruch et al. · 2023 [cited by applicant]
US 11677948B2 · Besenbruch · 2023 [cited by examiner]
US 20190171929A1 · Abadi · 2019 [cited by examiner]
US 20190180136A1 · Bousmalis et al. · 2019 [cited by applicant]
US 20200364574A1 · Kim · 2020 [cited by examiner]
US 20210004677A1 · Menick et al. · 2021 [cited by applicant]
US 20220272352A1 · Dinh et al. · 2022 [cited by applicant]
Agustsson, Eirikur , et al., Generative adversarial networks for extreme learned image compression. In Proceedings of the IEEE International Conference on Computer Vision, pp. 221-231, 2019. [cited by applicant]
Blau, Yochai , et al., The perception-distortion tradeoff. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. [cited by applicant]
Goodfellow, Ian , et al., Generative adversarial nets. Advances in neural information processing systems, 27, 2014. [cited by applicant]
Habibian, Amirhossein , et al., “Video Compression With Rate-Distortion Autoencoders”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 14, 2019 (Aug. 14, 2019). [cited by applicant]
Han, Jun , et al., “Deep Probabilistic Video Compression”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 5, 2018 (Oct. 5, 2018). [cited by applicant]
Jan, Xu , et al., “Efficient Context-Aware 1-21 Lossy Image Compression”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE, Jun. 14, 2020 (Jun. 14, 2020). [cited by applicant]
Kingma, Diederik , et al., Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. [cited by applicant]
Mechrez, Roey , et al., Maintaining natural image statistics with the contextual loss. In Asian Conference on Computer Vision, pp. 427-443. Springer, 2018. [cited by applicant]
Mechrez, Roey , et al., The contextual loss for image transformation with non-aligned data. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 768-783, 2018. [cited by applicant]
Ronneberger, Olaf , et al., U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-assisted Intervention, pp. 234-241. Springer, 2015. [cited by applicant]
Zhang, Richard , et al., The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. [cited by applicant]