IP Library Granted Patent US 12,646,226
Granted Patent B2
US 12,646,226 · App. 18/458,473 · Granted Jun 2, 2026

Methods and systems of training neural networks for lossy image or video encoding, transmission and decoding

Inventors: Chris Finlay (London, GB); Jonathan Rayner (London, GB); Jan Xu (London, GB); Christian Besenbruch (London, GB); Arsalan Zafar (London, GB); Sebastjan Cizel (London, GB); Vira Koshkina (London, GB)
Assignee: InterDigital VC Holdings, Inc.
G06T9/002G06F30/27G06N3/045G06N3/08G06N20/00G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,226
App. No.
18/458,473
Granted
Jun 2, 2026
Kind
B2
Abstract

A method of training one or more neural networks, the one or more neural networks being for use in lossy image or video encoding, transmission and decoding. The method comprises encoding an input image using a first neural network to produce a latent representation and decoding the latent representation using a second neural network to produce an output image. At least one of the plurality of layers of the first or second neural network comprises a transformation and a function based on an output of the transformation is use to update the parameters of the first neural network and the second neural network.

Claims (16)

1 . A method of training one or more neural networks comprising the steps of:

a) receiving an input image at a first computer system;

b) encoding the input image using a first neural network to produce a latent representation;

c) decoding the latent representation using a second neural network to produce an output image, wherein the output image is reconstruction of the input image;

wherein the first neural network and the second neural network each comprise a plurality of layers;

at least one of the plurality of layers of the first or second neural network comprises a transformation, wherein the output of the transformation is a pre-activation feature map;

d) evaluating a function comprising a term which is a function of a difference between the output image and the input image, a term which is the mean of the element-wise squares of the pre-activation latent representation and a term which is a function of the feature map;

e) updating the parameters of the first neural network and the second neural network based on the evaluated function; and

f) repeating the steps a) to e) a plurality of times until the evaluated function reaches a minima using a first set of input images to produce a first trained neural network and a second trained neural network.

2 . The method of claim 1 , wherein the term which is the mean of the element-wise squares of the pre-activation feature map is a sum of the mean of the element-wise squares of a plurality of pre-activation feature maps, wherein each feature map is in a different layer of the plurality of layers of the first or second neural network.

3 . The method of claim 2 , wherein the contribution of each of the sum of the mean of the element-wise squares of the plurality of pre-activation feature maps to the term is scaled by a predetermined value.

4 . The method of claim 1 , wherein the contribution of the term which is the mean of the element-wise squares of the pre-activation feature map to the update to the parameters of the first neural network and the second neural network is scaled by a predetermined value.

5 . The method of claim 1 , wherein the contribution of the term which is the mean of the element-wise squares of the pre-activation feature map to the update to the parameters of the first neural network and the second neural network is scaled by a value; and

the value is additionally updated in at least one of the plurality of repetitions of steps a) to e).

6 . The method of claim 1 , wherein the transformation of the at least one of the plurality of layers of the first or second neural network is an affine linear transformation.

7 . A non-transitory, computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to perform the method of claim 1 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2026
From: DEEP RENDER LTD
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 073864/0596 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2024
From: FINLAY, CHRIS; RAYNER, JONATHAN; XU, JAN; BESENBRUCH, CHRISTIAN; ZAFAR, ARSALAN; CIZEL, SEBASTJAN; KOSHKINA, VIRA
To: DEEP RENDER LTD.
Reel/Frame 069432/0095 →
Priority Claims (1)
GB 2115399 · Oct 26, 2021 · national
Continuity (2)
Continuation PCTEP2022080015 · Oct 26, 2022
Related Publication 20240070925A1 · Feb 29, 2024
References Cited (19)
US 11252417B2 · Andreopoulos · 2022 [cited by examiner]
US 12223426B2 · Sung · 2025 [cited by examiner]
US 20190046068A1 · Ceccaldi · 2019 [cited by examiner]
US 20230074979A1 · Brehmer · 2023 [cited by examiner]
CN 110868598A · 2020 [cited by examiner]
WO WO2021097421A1 · 2021 [cited by examiner]
WO WO2022106014A1 · 2022 [cited by examiner]
STIC Provided translation of CN-110868598 A relied upon in Rejection of the claims. (Year: 2020). [cited by examiner]
Changyue Ma et al, “A Cross Channel Context Model for Latents in Deep Image Compression”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,Mar. 4, 2021 (Mar. 4, 2021). [cited by applicant]
Caron, M., Bojanowski, P., Joulin, A. & Douze, M. (2018) Deep Clustering for Unsupervised Learning of Visual Learning. [cited by applicant]
Da Silva Renam C et al, “Distortion scalable learned image compression”, 2019 IEEE 21st International Workshop on Multimedia Signal Processing (MMSP), IEEE,Sep. 27, 2019 (Sep. 27, 2019), p. 1-6. [cited by applicant]
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J. & Houlsby, N. (2020) An Image is Worth 16x16 Words: Transformers … [cited by applicant]
Jang, E., Gu, S. & Poole, B. (2016) Categorical Reparameterization with Gumbel-Softmax. [cited by applicant]
Lin Chaoyi et al, “A Spatial RNN Codec for End-to-End Image Compression”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE,Jun. 13, 2020 (Jun. 13, 2020), p. 13266-13274. [cited by applicant]
Michele Benzi, Gene H. Golub, and Jörg Liesen. Numerical solution of saddle point problems. Acta Numerica, 14:1-137, May 2005. [cited by applicant]
Un Han et al, “Deep Probabilistic Video Compression”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,Oct. 5, 2018 (Oct. 5, 2018). [cited by applicant]
Wu, B., Xu, C., Dai, X., Wan, A., Zhang, P., Yan, Z., Tomizuka, M., Gonzalez, J., Keutzer, K. & Vajda, P. (2020) Visual Transformers: Token-based Image Representation and Processing for Computer Vision. [cited by applicant]
Balle et al., “End-to-end Optimized Image Compression,” Conf. paper at ICLR 2017; arXiv:1611.01704v3 (2017). [cited by applicant]
Balle et al., “Variational image compression with a scale hyperprior,” arXiv:1802.01436v2 (2018). [cited by applicant]