IP Library Granted Patent US 12,323,593
Granted Patent B2
US 12,323,593 · App. 18/230,361 · Granted Jun 3, 2025

Image compression and decoding, video compression and decoding: methods and systems

Inventors: Chri Besenbruch (London, GB); Ciro Cursio (London, GB); Christopher Finlay (London, GB); Vira Koshkina (London, GB); Alexander Lytchier (London, GB); Jan Xu (London, GB); Arsalan Zafar (London, GB)
Assignee: DEEP RENDER LTD.
H04N19/126G06N3/045G06N3/084G06T3/4046G06T9/002G06V10/774H04N19/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,323,593
App. No.
18/230,361
Filed
Aug 4, 2023
Granted
Jun 3, 2025
Kind
B2
Art Unit
2482
USPC
375/240.03
Abstract

A computer-implemented method for lossy image or video compression, transmission and decoding, the method including the steps of: (i) receiving an input image at a first computer system; (ii) encoding the input image using a first trained neural network, using the first computer system, to produce a latent representation; (iii) quantizing the latent representation using the first computer system to produce a quantized latent; (iv) entropy encoding the quantized latent into a bitstream, using the first computer system; (v) transmitting the bitstream to a second computer system; (vi) the second computer system entropy decoding the bitstream to produce the quantized latent; (vii) the second computer system using a second trained neural network to produce an output image from the quantized latent, wherein the output image is an approximation of the input image.

Claims (20)

1. A computer implemented method of training a first neural network and a second neural network based on training images in which each respective training image includes human scored data relating to a perceived level of distortion in the respective training image as evaluated by a group of humans, the first and second neural networks being for use in lossy image or video compression, transmission and decoding, the method including the steps of:

(i) generating a set of training images by forward passing a set of images through a trained AI-based compression pipeline, each training image having one or more artefacts representative of artefacts introduced by the trained AI-based compression pipeline, and by human scoring each generated training image based on a perceived level of distortion;

(ii) receiving an input training image from the generated set of training images;

(iii) encoding the input training image using the first neural network, to produce a latent representation;

(iv) quantizing the latent representation to produce a quantized latent;

(v) using the second neural network to produce an output image from the quantized latent, wherein the output image is an approximation of the input image;

(vi) evaluating a loss function based on differences between the output image and the input training image;

(vii) evaluating a gradient of the loss function;

(viii) back-propagating the gradient of the loss function through the second neural network and through the first neural network, to update weights of the second neural network and of the first neural network;

(ix) repeating steps (i) to (viii) using a set of training images, to produce a trained first neural network and a trained second neural network; and

(x) storing the weights of the trained first neural network and of the trained second neural network;

wherein the loss function includes a weighted sum of a rate term and a distortion term, and wherein the distortion term is a function based on human scored data of the respective training image from the generated set of training images.

2. The method of claim 1 , wherein the loss function is evaluated as a weighted sum of differences between the output image and the input training image, and the estimated bits of the quantized image latents.

3. The method of claim 1 , wherein at least one thousand training images are used.

4. The method of claim 1 , wherein the training images include a plurality of distortion types.

5. The method of claim 1 , wherein the training images include at least one distortion type corresponding to one or more distortion types introduced using AI-based compression encoder-decoder pipelines.

6. The method of claim 1 , wherein the human scored data is based on human labelled data.

7. The method of claim 1 , wherein in step (v) the loss function includes a component that represents the human visual system.

8. The method of claim 7 , wherein a distortion term of the loss function comprises a MSE and/or PSNR term.

9. The method of claim 1 , wherein said generating comprises generating the training set of images by forward passing the set of images through the trained AI-based compression pipeline at different time steps in the training of the trained AI-based compression pipeline.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2026
From: DEEP RENDER LTD
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 073864/0596 →
Priority Claims (12)
GB 2006275 · Apr 29, 2020 · national
GB 2008241 · Jun 2, 2020 · national
GB 2011176 · Jul 20, 2020 · national
GB 2012461 · Aug 11, 2020 · national
GB 2012462 · Aug 11, 2020 · national
GB 2012463 · Aug 11, 2020 · national
GB 2012465 · Aug 11, 2020 · national
GB 2012467 · Aug 11, 2020 · national
GB 2012468 · Aug 11, 2020 · national
GB 2012469 · Aug 11, 2020 · national
GB 2016824 · Oct 23, 2020 · national
GB 2019531 · Dec 10, 2020 · national
Continuity (6)
Continuation 18055666 · Nov 15, 2022
Continuation 17740716 · May 10, 2022
Continuation PCTGB2021051041 · Apr 29, 2021
Provisional Application 63017295 · Apr 29, 2020
Provisional Application 63053807 · Jul 20, 2020
Related Publication 20240195971A1 · Jun 13, 2024
References Cited (64)
US 5048095A · Bhanu et al. · 1991 [cited by applicant]
US 9990687B1 · Kaufhold et al. · 2018 [cited by applicant]
US 10373300B1 · Besenbruch et al. · 2019 [cited by applicant]
US 10489936B1 · Zafar et al. · 2019 [cited by applicant]
US 10880551B2 · Topiwala et al. · 2020 [cited by applicant]
US 10886943B2 · Choi et al. · 2021 [cited by applicant]
US 10930263B1 · Mahyar · 2021 [cited by applicant]
US 10965948B1 · Appalaraju et al. · 2021 [cited by applicant]
US 11310509B2 · Topiwala · 2022 [cited by examiner]
US 11330264B2 · Zhou et al. · 2022 [cited by applicant]
US 11375194B2 · Liu et al. · 2022 [cited by applicant]
US 11388416B2 · Habibian et al. · 2022 [cited by applicant]
US 11445222B1 · Andreopoulos · 2022 [cited by examiner]
US 11481633B2 · Krishnamoorthy · 2022 [cited by applicant]
US 11526734B2 · Yang et al. · 2022 [cited by applicant]
US 11544536B2 · Gesmundo · 2023 [cited by applicant]
US 11610154B1 · Teig et al. · 2023 [cited by applicant]
US 11748615B1 · Wu et al. · 2023 [cited by applicant]
US 20100332423A1 · Kapoor et al. · 2010 [cited by applicant]
US 20160292589A1 · Taylor et al. · 2016 [cited by applicant]
US 20170230675A1 · Wierstra et al. · 2017 [cited by applicant]
US 20180139450A1 · Gao et al. · 2018 [cited by applicant]
US 20180176578A1 · Rippel et al. · 2018 [cited by applicant]
US 20190188573A1 · Lehman et al. · 2019 [cited by applicant]
US 20190289296A1 · Kottke · 2019 [cited by examiner]
US 20200021865A1 · Topiwala · 2020 [cited by examiner]
US 20200027247A1 · Minnen et al. · 2020 [cited by applicant]
US 20200090069A1 · Mandt et al. · 2020 [cited by applicant]
US 20200097742A1 · Ratnesh Kumar et al. · 2020 [cited by applicant]
US 20200104640A1 · Poole et al. · 2020 [cited by applicant]
US 20200111501A1 · Sung et al. · 2020 [cited by applicant]
US 20200226421A1 · Almazan et al. · 2020 [cited by applicant]
US 20200304802A1 · Habibian et al. · 2020 [cited by applicant]
US 20200364574A1 · Kim · 2020 [cited by examiner]
US 20200372686A1 · Wen et al. · 2020 [cited by applicant]
US 20200401916A1 · Rolfe et al. · 2020 [cited by applicant]
US 20210004677A1 · Menick et al. · 2021 [cited by applicant]
US 20210042606A1 · Bai et al. · 2021 [cited by applicant]
US 20210067808A1 · Schroers et al. · 2021 [cited by applicant]
US 20210142534A1 · Liu et al. · 2021 [cited by applicant]
US 20210152831A1 · Liu et al. · 2021 [cited by applicant]
US 20210166151A1 · Kennel et al. · 2021 [cited by applicant]
US 20210211741A1 · Andreopoulos · 2021 [cited by examiner]
US 20210281867A1 · Golinski et al. · 2021 [cited by applicant]
US 20210286270A1 · Middlebrooks et al. · 2021 [cited by applicant]
US 20210360259A1 · Wang et al. · 2021 [cited by applicant]
US 20210366161A1 · Wong · 2021 [cited by examiner]
US 20210390335A1 · Du et al. · 2021 [cited by applicant]
US 20210397895A1 · Sun et al. · 2021 [cited by applicant]
US 20220101106A1 · Van Der Wilk et al. · 2022 [cited by applicant]
US 20220103839A1 · Van Rozendaal et al. · 2022 [cited by applicant]
US 20220327363A1 · Xu et al. · 2022 [cited by applicant]
US 20230093734A1 · Zheng et al. · 2023 [cited by applicant]
Leon-Garcia , “Probability and random processes for electrical engineering,” Pearson Education India (1994). [cited by applicant]
Chen , et al., “Neural ordinary differential equations,” Advances in neural information processing systems; 31 (2018). [cited by applicant]
Elsken , et al., “Neural architecture search: A survey,”The Journal of Machine Learning Research, 1997-2017 (2019). [cited by applicant]
Li , et al., “Sgas: Sequential Greedy Architecture Search,” In Proceedings of the IEEE/CVF Conf of Computer Vision and Pattern Recognition, pp. 1620-1630 (2020). [cited by applicant]
Molina , et al., “Pade Activation Units: End-to-end Learning of Flexible Activation Functions in Deep Networks,” arXiv preprint arXiv: 1907.06732 (2019). [cited by applicant]
Ziegler , et al., “Latent normalizing flows for discrete sequences,” Intl. Conf. on Machine Learning; PMLR (2019). [cited by applicant]
Balle et al. , “End-to-end optimized image compression,” arXiv preprint arXiv: 1611.01704 (2016). [cited by applicant]
Cheng et al. , “Energy compaction-based image compression using convolutional autoencoder,” IEEE Transactions on Multimedia 22.4, pp. 860-873 (2019). [cited by applicant]
Habibian, Amirhossein , et al., “Video Compression with Rate-Distortion Autoencoders,” arxiv.org, Cornell Univ. Library (Aug. 14, 2019) XP081531236. [cited by applicant]
Han, Jun , et al., “Deep Probabilistic Video Compression,” arxiv.org, Cornell Univ. Library, (Oct. 5, 2018) XP080930310. [cited by applicant]
Yan et al. , “Deep autoencoder-based lossy geometry compression for point clouds,” arXiv preprint arXiv: 1905.03691 (2019). [cited by applicant]
Cited By (3)
US 12,380,624 US 12,387,466 US 12,647,611