IP Library › Granted Patent US 12,382,068
Granted Patent B2
US 12,382,068 · App. 18/949,892 · Granted Aug 5, 2025

High-performance and low-complexity neural compression from a single image, video or audio data

Inventors: Emilien Dupont (London, GB); Hyun Jik Kim (London, GB); Matthias Stephan Bauer (London, GB); Lucas Marvin Theis (London, GB)
Assignee: DeepMind Technologies Limited
H04N19/189H04N19/119H04N19/124H04N19/167H04N19/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,382,068
App. No.
18/949,892
Granted
Aug 5, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium for encoding input data comprising input data values corresponding to respective input data grid points of an input data grid, such as image, video or audio data.

Claims (85)

1. A method of encoding input data performed by one or more data processing apparatus, the input data comprising input data values corresponding to respective input data grid points of an input data grid, the method comprising:

optimizing an objective function by jointly optimizing parameters of a synthesis neural network, parameters of a decoder neural network, and a set of respective latent values, each respective latent value corresponding to a respective latent grid point of a respective one of a plurality of latent grids having different respective resolutions, and wherein the optimizing comprises, at each of a plurality of optimization iterations:

generating a respective updated value for each of the respective latent values, comprising:

generating a respective noise distribution for the optimization iteration, wherein the respective noise distribution for the optimization iteration has a shape that is more uniform than the respective noise distribution for a preceding optimization iteration;

sampling a respective noise value for each respective latent value from the respective noise distribution for the optimization iteration;

updating each of the respective latent values by adding the respective noise value for the respective latent value to the respective latent value;

processing the respective updated values using the synthesis neural network to generate respective reconstructed data values, each corresponding to a respective input data value of the input data values;

determining a respective probability distribution for each respective updated value using the decoder neural network;

evaluating the objective function by determining (i) an error between the input data values and the respective reconstructed data values and (ii) a compressibility term based on the respective probability distributions;

determining gradients of the objective function with respect to the parameters of the synthesis neural network, the parameters of the decoder neural network, and the set of respective latent values; and

using the gradients to update one or more of: the parameters of the synthesis neural network, the parameters of the decoder neural network, and the respective latent values;

quantizing the optimized respective latent values; and

encoding the quantized respective latent values using a probability distribution for the respective latent values, the probability distribution being defined by the decoder neural network.

2. The method according to claim 1 , wherein generating a respective updated value for each of the respective latent values comprises:

prior to updating each of the respective latent values by adding the respective noise value for the respective latent value to the respective latent value, applying a soft-rounding function to each of the respective latent values to update each respective latent value.

3. The method according to claim 1 , wherein the respective noise distribution for the optimization iteration is non-uniform.

4. The method according to claim 3 , wherein generating the respective noise distribution for the optimization iteration comprises determining the shape of the respective noise distribution based on a shape parameter that controls the shape of the respective noise distribution.

5. The method according to claim 2 , wherein the soft-rounding function depends on a temperature parameter that controls a smoothness of the soft-rounding function, the temperature parameter being adjusted between the optimization iterations.

6. The method according to claim 1 , wherein:

using the gradients comprises multiplying the gradients by a learning rate before updating the one or more of: the parameters of the synthesis neural network, the parameters of the decoder neural network, and the respective latent values; and

the learning rate varies between the optimization iterations according to a cosine schedule.

7. The method according to claim 1 , wherein the optimizing further comprises, at each of a plurality of further optimization iterations:

quantizing the respective latent values using a further hard-rounding function;

using a soft-rounding estimator to determine further gradients of the objective function using the quantized respective latent values, wherein the soft-rounding estimator provides a smooth approximation to the gradient of the further hard-rounding function; and

using the further gradients determined using the soft-rounding estimator to update one or more of: the parameters of the synthesis neural network, the parameters of the decoder neural network, and the respective latent values.

8. The method according to claim 7 , wherein the soft-rounding estimator depends on a temperature parameter that controls the smoothness of the gradient of a soft-rounding function corresponding to the soft-rounding estimator, the temperature parameter being adjusted between the further optimization iterations such that the gradient of the soft-rounding function increasingly resembles the gradient of the further hard-rounding function.

9. The method according to claim 7 , wherein:

using the further gradients comprises multiplying the further gradients by a further learning rate before updating the one or more of: the parameters of the synthesis neural network, the parameters of the decoder neural network, and the respective latent values; and

the further learning rate is decreased between the further optimization iterations.

10. The method according to claim 7 , wherein the further hard-rounding function quantizes the respective latent values in steps smaller than the steps used by a hard-rounding function for quantizing the respective latent values after the optimizing.

11. The method according to claim 9 , wherein the further hard-rounding function quantizes the respective latent values in steps smaller than one.

12. The method according to claim 1 , wherein the compressibility term is based on a respective probability for each respective latent value, and wherein the respective probabilities are determined by, for each of the respective latent grid points of each of the latent grids:

determining a causal subset of the respective latent values comprising respective latent values corresponding to respective latent grid points that precede the respective latent grid point in the latent grid; and

using the decoder neural network to determine, from the respective latent values in the causal subset, one or more probability distribution parameters defining a conditional probability distribution for a respective latent value at the respective latent grid point; and

using the conditional probability distribution defined by the one or more probability distribution parameters to determine the respective probability of the respective latent value conditioned on the respective latent values in the causal subset.

13. The method according to claim 12 , wherein for one or more of the latent grids, the causal subsets of the respective latent values for each of the respective latent grid points of the latent grid comprise respective latent values of another of the latent grids that has a resolution that is less than a resolution of the latent grid.

14. The method according to claim 13 , wherein the other latent grid immediately precedes the latent grid when the latent grids are arranged in ascending order of resolution.

15. The method according to claim 12 , wherein the decoder neural network comprises a plurality of decoder subnetworks, each decoder subnetwork being configured to process the respective latent values of a respective one of the latent grids.

16. The method according to claim 15 , wherein each decoder subnetwork processes the respective latent values of the respective one of the latent grids independently of the respective latent values of the other latent grids.

17. The method according to claim 12 , wherein the decoder neural network comprises an input layer, an output layer, and intermediate layers between the input and output layers, the intermediate layers comprising activation functions for determining features of the respective latent values at different resolutions, wherein one or more of the intermediate layers comprises a modulation layer configured to apply a transformation to features of respective latent values at a first resolution conditioned on features of respective latent values at a second resolution.

18. The method according to claim 12 , wherein output values of the decoder neural network are exponentiated to determine the one or more probability distribution parameters and wherein a pre-determined shift value is added to the output values prior to exponentiation.

19. The method according to claim 1 , wherein one or more of the synthesis neural network or the decoder neural network comprises activation functions that have a higher computational complexity than Rectified Linear Units.

20. The method according to claim 1 , wherein determining (i) an error between the input data values and the respective reconstructed data values and (ii) a compressibility term based on the respective probability distributions comprises determining the respective reconstructed data values by:

for each of the latent grids, generating from the respective latent values of the latent grid, upsampled latent data comprising respective upsampled latent values for the input data grid points; and

using the synthesis neural network to determine, from the upsampled latent data for each of the latent grids, the respective reconstructed data values for the input data grid points.

21. The method according to claim 1 , wherein each of the input data grid points corresponds to a respective time interval or frequency component of an audio waveform and the input data values correspond to amplitudes of the audio waveform.

22. The method according to claim 1 , wherein each of the input data grid points corresponds to a respective pixel location in one or more images or part of one or more images.

23. The method according to claim 1 , wherein each of the input data grid points corresponds to a respective frame in a sequence of image frames of a video and a respective pixel location in the respective frame.

24. The method according to claim 23 , further comprising:

dividing the video into a plurality of video patches, each video patch corresponding to one or more of: a proper subset of the respective pixel locations of each image frame or a proper subset of the image frames of the video; and wherein the input data comprises one of the plurality of video patches.

25. The method according to claim 23 , wherein the compressibility term is based on a respective probability for each respective latent value, and wherein the respective probabilities are determined by, for each of the respective latent grid points of each of the latent grids:

determining a causal subset of the respective latent values comprising respective latent values corresponding to respective latent grid points that precede the respective latent grid point in the latent grid; and

using the decoder neural network to determine, from the respective latent values in the causal subset, one or more probability distribution parameters defining a conditional probability distribution for a respective latent value at the respective latent grid point; and

using the conditional probability distribution defined by the one or more probability distribution parameters to determine the respective probability of the respective latent value conditioned on the respective latent values in the causal subset, and wherein the causal subset of the respective latent values for each of the respective latent grid points comprises selected latent values of another latent grid corresponding to an image frame of the video that precedes the image frame corresponding to the respective latent grid point in the sequence of image frames.

26. The method according to claim 1 , wherein the optimizing is performed for input data corresponding to a plurality of training examples to generate a corresponding synthesis neural network, decoder neural network and set of respective latent values for each of the one or more training examples, wherein at least some of the parameters of one or more of the synthesis neural networks or the decoder neural networks are shared between the training examples.

27. The method according to claim 26 , wherein the input data of the plurality of training examples correspond to different respective images or videos, or to parts of one image or video.

28. The method according to claim 1 , further comprising providing the encoded latent values in a bitstream.

29. A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform a method of encoding input data, the input data comprising input data values corresponding to respective input data grid points of an input data grid, the method comprising:

optimizing an objective function by jointly optimizing parameters of a synthesis neural network, parameters of a decoder neural network, and a set of respective latent values, each respective latent values corresponding to a respective latent grid point of a respective one of a plurality of latent grids having different respective resolutions, and wherein the optimizing comprises, at each of a plurality of optimization iterations:

generating a respective updated value for each of the respective latent values, comprising:

generating a respective noise distribution for the optimization iteration, wherein the respective noise distribution for the optimization iteration has a shape that is more uniform than the respective noise distribution for a preceding optimization iteration;

sampling a respective noise value for each respective latent value from the respective noise distribution for the optimization iteration;

updating each of the respective latent values by adding the respective noise value for the respective latent value to the respective latent value;

processing the respective updated values using the synthesis neural network to generate respective reconstructed data values, each corresponding to a respective input data value of the input data values;

determining a respective probability distribution for each respective updated value using the decoder neural network;

evaluating the objective function by determining (i) an error between the input data values and the respective reconstructed data values and (ii) a compressibility term based on the respective probability distributions;

determining gradients of the objective function with respect to the parameters of the synthesis neural network, the parameters of the decoder neural network, and the set of respective latent values; and

using the gradients to update one or more of: the parameters of the synthesis neural network, the parameters of the decoder neural network, and the respective latent values;

quantizing the optimized respective latent values; and

encoding the quantized respective latent values using a probability distribution for the latent values, the probability distribution being defined by the decoder neural network.

30. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for encoding input data performed by one or more data processing apparatus, the input data comprising input data values corresponding to respective input data grid points of an input data grid, the operations comprising:

optimizing an objective function by jointly optimizing parameters of a synthesis neural network, parameters of a decoder neural network, and a set of respective latent values, each respective latent value corresponding to a respective latent grid point of a respective one of a plurality of latent grids having different respective resolutions, and wherein the optimizing comprises, at each of a plurality of optimization iterations:

generating a respective updated value for each of the respective latent values, comprising:

generating a respective noise distribution for the optimization iteration, wherein the respective noise distribution for the optimization iteration has a shape that is more uniform than the respective noise distribution for a preceding optimization iteration;

sampling a respective noise value for each respective latent value from the respective noise distribution for the optimization iteration;

updating each of the respective latent values by adding the respective noise value for the respective latent value to the respective latent value;

processing the respective updated values using the synthesis neural network to generate respective reconstructed data values, each corresponding to a respective input data value of the input data values;

determining a respective probability distribution for each respective updated value using the decoder neural network;

evaluating the objective function by determining (i) an error between the input data values and the respective reconstructed data values and (ii) a compressibility term based on the respective probability distributions;

determining gradients of the objective function with respect to the parameters of the synthesis neural network, the parameters of the decoder neural network, and the set of respective latent values; and

using the gradients to update one or more of: the parameters of the synthesis neural network, the parameters of the decoder neural network, and the respective latent values;

quantizing the optimized respective latent values; and

encoding the quantized respective latent values using a probability distribution for the respective latent values, the probability distribution being defined by the decoder neural network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071550/0092 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 20, 2025
From: DUPONT, EMILIEN; KIM, HYUN JIK; BAUER, MATTHIAS STEPHAN; THEIS, LUCAS MARVIN
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 071172/0591 →
Continuity (2)
Provisional Application 63600412 · Nov 17, 2023
Related Publication 20250168368A1 · May 22, 2025
References Cited (159)
US 6405341B1 · Maru · 2002 [cited by examiner]
US 6411236B1 · Kermani · 2002 [cited by examiner]
US 6493389B1 · Bailleul · 2002 [cited by examiner]
US 7742643B2 · Yang · 2010 [cited by examiner]
US 7793189B2 · Yano · 2010 [cited by examiner]
US 8190550B2 · Bouchard · 2012 [cited by examiner]
US 8244524B2 · Shirakawa · 2012 [cited by examiner]
US 10133933B1 · Fisher · 2018 [cited by examiner]
US 10395098B2 · Hwang · 2019 [cited by examiner]
US 10474988B2 · Fisher · 2019 [cited by examiner]
US 10839259B2 · Shazeer · 2020 [cited by examiner]
US 10970441B1 · Zhang · 2021 [cited by examiner]
US 11151447B1 · Chen · 2021 [cited by examiner]
US 11214268B2 · Gonzalez Aguirre · 2022 [cited by examiner]
US 11347995B2 · Bender · 2022 [cited by examiner]
US 11531852B2 · Vahdat · 2022 [cited by examiner]
US 11677948B2 · Besenbruch · 2023 [cited by examiner]
US 11854182B2 · Wang · 2023 [cited by examiner]
US 20050166129A1 · Yano · 2005 [cited by examiner]
US 20060013493A1 · Yang · 2006 [cited by examiner]
US 20070094023A1 · Gallino · 2007 [cited by examiner]
US 20100074330A1 · Fu · 2010 [cited by examiner]
US 20100106511A1 · Shirakawa · 2010 [cited by examiner]
US 20100318490A1 · Bouchard · 2010 [cited by examiner]
US 20120177234A1 · Rank · 2012 [cited by examiner]
US 20170055003A1 · Yu · 2017 [cited by examiner]
US 20170236000A1 · Hwang · 2017 [cited by examiner]
US 20190043003A1 · Fisher · 2019 [cited by examiner]
US 20190130213A1 · Shazeer · 2019 [cited by examiner]
US 20190135300A1 · Gonzalez Aguirre · 2019 [cited by examiner]
US 20200184375A1 · Ohwa · 2020 [cited by examiner]
US 20200379893A1 · Arechiga Gonzalez · 2020 [cited by examiner]
US 20210097401A1 · Ramalho · 2021 [cited by examiner]
US 20210133590A1 · Amroabadi · 2021 [cited by examiner]
US 20210158211A1 · Talwar · 2021 [cited by examiner]
US 20210232958A1 · Yamaguchi · 2021 [cited by examiner]
US 20210303967A1 · Bender · 2021 [cited by examiner]
US 20210374608A1 · El-Khamy · 2021 [cited by examiner]
US 20210383222A1 · Saxton · 2021 [cited by examiner]
US 20220156479A1 · Wang · 2022 [cited by examiner]
US 20220188636A1 · Pham · 2022 [cited by examiner]
US 20220293247A1 · Gibson · 2022 [cited by examiner]
US 20220382880A1 · Castiglione · 2022 [cited by examiner]
US 20220383036A1 · Nazi · 2022 [cited by examiner]
US 20220415453A1 · Ronneberger · 2022 [cited by examiner]
US 20230061004A1 · Wang · 2023 [cited by examiner]
US 20230073174A1 · Clark · 2023 [cited by examiner]
US 20230154055A1 · Besenbruch · 2023 [cited by examiner]
US 20230267216A1 · Santana De Oliveira · 2023 [cited by examiner]
US 20230299788A1 · Agustsson · 2023 [cited by examiner]
US 20230329646A1 · Zhou · 2023 [cited by examiner]
US 20230351042A1 · De · 2023 [cited by examiner]
US 20230368337A1 · Karras · 2023 [cited by examiner]
US 20230401424A1 · Jiang · 2023 [cited by examiner]
US 20240086760A1 · Fraboni · 2024 [cited by examiner]
US 20240354553A1 · Finlay · 2024 [cited by examiner]
US 20240386275A1 · Yang · 2024 [cited by examiner]
CN 106059640A · 2016 [cited by examiner]
CN 112866668A · 2021 [cited by examiner]
CN 113313238A · 2021 [cited by examiner]
CN 113807265A · 2021 [cited by examiner]
CN 114037051A · 2022 [cited by examiner]
CN 114088658A · 2022 [cited by examiner]
WO WO2022140583A1 · 2022 [cited by examiner]
WO WO2022248891A1 · 2022 [cited by examiner]
Schwarz et al., “Modality-agnostic variational compression of implicit neural representations.” arXiv preprint arXiv:2301.09479 (2023). (Year: 2023). [cited by examiner]
Gong et al., “Differentiable soft quantization: Bridging full-precision and low-bit neural networks.” In Proceedings of the IEEE/CVF international conference on computer vision, pp. 4852-4861. 2019. (Year: 2019). [cited by examiner]
Wikipedia, Simulated annealing, published on Aug. 26, 2023 (Year: 2023). [cited by examiner]
Wikipedia, autoencoder, published on Nov. 12, 2023 (Year: 2023). [cited by examiner]
Ballé et al., “End-to-end optimization of nonlinear transform codes for perceptual quality,” 2016 Picture Coding Symposium (PCS), Nuremberg, Germany, 2016, pp. 1-5 (Year: 2016). [cited by examiner]
Li, Zhiyuan, and Sanjeev Arora. “An exponential learning rate schedule for deep learning.” arXiv preprint arXiv:1910.07454 (2019). (Year: 2019). [cited by examiner]
Machine translation of WO 2022140583 A1 (Year: 2022). [cited by examiner]
Machine translation of WO 2022248891 A1 (Year: 2022). [cited by examiner]
Herglotz et al., “Beyond Bjøntegaard: Limits of Video Compression Performance Comparisons,” 2022 IEEE International Conference on Image Processing (ICIP), Bordeaux, France, 2022, pp. 46-50 (Year: 2022). [cited by examiner]
Lewkowycz, “How to decay your learning rate.” arXiv preprint arXiv:2103.12682 (2021). (Year: 2021). [cited by examiner]
Normal distribution—Wikipedia (Year: 2023). [cited by examiner]
Laplace distribution—Wikipedia (Year: 2023). [cited by examiner]
International Search Report and Written Opinion in International Appln. No. PCT/EP2024/082596, dated Feb. 13, 2025, 17 pages. [cited by applicant]
Agustsson et al., “Scale-space flow for end-to-end optimized video compression,” Presented at the 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, pp. 8503-8512. [cited by applicant]
Agustsson et al., “Universally Quantized Neural Compression,” Presented at the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Dec. 6-12, 2020, 10 pages. [cited by applicant]
Bai et al., “PS-NeRV: Patch-Wise Stylized Neural Representations for Videos,” 2023 IEEE International Conference on Image Processing (ICIP), Oct. 8-11, 2023, pp. 41-45. [cited by applicant]
Ballé et al., “End-to-end Optimized Image Compression,” CoRR, Submitted on Mar. 3, 2017, arXiv:1611.01704v3, 27 pages. [cited by applicant]
Ballé et al., “Variational image compression with a scale hyperprior,” Presented at the 6th International Conference on Learning Representations, Apr. 30-May 3, 2018, 47 pages. [cited by applicant]
Begaint et al., “CompressAI: a PyTorch library and evaluation platform for end-to-end compression research,” CoRR, Nov. 5, 2020, arxiv.org/abs/2011.03029, 19 pages. [cited by applicant]
bellard.org [online], “BPG Image Format,” available on or before Nov. 11, 2023, via Internet Archive: Wayback Machine URL<https://web.archive.org/web/20231111133027/https://bellard.org/bpg/>, retrieved on Nov. 25, 2024,… [cited by applicant]
Bjontegaard et al., “Calculation of average PSNR differences between RD-curves,” Paper, Presented at the 13th Meeting of the Video Coding Experts Group (VCEG), Austin Texas, USA, Apr. 2-4, 2001; ITU SG16 Doc, VCEG-M33, … [cited by applicant]
Bross et al., “Overview of the Versatile Video Coding (VVC) Standard and Its Applications,” IEEE Transactions on Circuits and Systems for Video Technology, Oct. 2021, 31(10):3736-3764. [cited by applicant]
Campos et al., “Content Adaptive Optimization for Neural Image Compression,” Presented at the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Jun. 15-20, 2019, 5 pages. [cited by applicant]
Catania et al., “Nif: A Fast Implicit Image Compression with Bottleneck Layers and Modulated Sinusoidal Activations,” Presented at the 31st ACM International Conference on Multimedia (MM'23), Oct. 29-Nov. 3, 2023, pp. 9… [cited by applicant]
Chen et al., “HNeRV: A Hybrid Neural Representation for Videos,” Presented at the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 17-24, 2023, pp. 10270-10279. [cited by applicant]
Chen et al., “NeRV: Neural Representations for Videos,” Presented at the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Dec. 6-12, 2020, 12 pages. [cited by applicant]
Cheng et al., “Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules,” Presented at the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 202… [cited by applicant]
Clare et al., “Wavefront Parallel Processing for HEVC Encoding and Decoding,” Presented at the 6th Meeting of the Joint Collaborative Team on Video Coding (JCT-VC), Jul. 14-22, 2011, 16 pages. [cited by applicant]
Damodaran et al., “Rqat-inr: Improved implicit neural image compression,” 2023 Data Compression Conference (DCC), Mar. 21-24, 2023, pp. 208-217. [cited by applicant]
Davies et al., “On the Effectiveness of Weight-Encoded Neural Implicit 3D Shapes,” CoRR, Submitted on Jan. 17, 2021, arXiv:2009.09808v3, 13 pages. [cited by applicant]
Dupont et al., “COIN: COmpression with Implicit Neural representations,” CoRR, Mar. 3, 2021, arXiv:2103.03123v2, 12 pages. [cited by applicant]
Dupont et al., “COIN++: Neural Compression Across Modalities,” Transactions on Machine Learning Research, Nov. 2022, 26 pages. [cited by applicant]
Gao et al., “SINCO: A Novel Structural Regularizer for Image Compression Using Implicit Neural Representations,” Presented at the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing… [cited by applicant]
Girish et al., “Scalable HAsh-grid Compression for Implicit Neural Representations,” Presented at the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 1-6, 2023, pp. 17513-17524. [cited by applicant]
Gomes et al., “Video Compression with Entropy-Constrained Neural Representations,” Presented at the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 1-6, 2023, pp. 18497-18506. [cited by applicant]
Gordon et al., “On Quantizing Implicit Neural Representations,” Presented at the 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Jan. 2-7, 2023, pp. 341-350. [cited by applicant]
Guo et al., “Compression with Bayesian Implicit Neural Representations,” CoRR, Submitted on Oct. 29, 2023, arXiv:2305.19185v5, 19 pages. [cited by applicant]
Guo et al., “Variable Rate Image Compression With Content Adaptive Optimization,” Presented at the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2020), Seattle, Jun. 14-19, 2020, pp. 122-123. [cited by applicant]
Guo-Hua et al., “EVC: Towards Real-Time Neural Image Compression with Mask Decay,” Presented at the Eleventh International Conference on Learning Representations, May 1-5, 2023, 23 pages. [cited by applicant]
Han et al., “Deep Generative Video Compression,” CoRR, Oct. 5, 2018, arxiv.org/abs/1810.02845, 15 pages. [cited by applicant]
He et al., “Checkerboard Context Model for Efficient Learned Image Compression,” Presented at the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 20-25, 2021, pp. 14771-14780. [cited by applicant]
He et al., “ELIC: Efficient Learned Image Compression With Unevenly Grouped Space-Channel Contextual Adaptive Coding,” Presented at the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 18… [cited by applicant]
He et al., “REcombiner: Robust and Enhanced Compression with Bayesian Implicit Neural Representations,” CoRR, Submitted on Sep. 29, 2023, arXiv:2309.17182v1, 24 pages. [cited by applicant]
Hendrycks et al. “Gaussian Error Linear Units (GELUs),” CoRR, Submitted on Jun. 6, 2023, arXiv:1606.08415v5, 10 pages. [cited by applicant]
Henry et al., “Wavefront Parallel Processing,” Presented at the 5th Meeting of the Collaborative Team on Video Coding (JCT-VC), Mar. 16-23, 2011, 9 pages. [cited by applicant]
Huang et al., “Compressing multidimensional weather and climate data into neural networks,” Presented at the Eleventh International Conference on Learning Representations, May 1-5, 2023, 17 pages. [cited by applicant]
[No Author Listed] “Information Technology—Digital Compression and Coding of Continuous-Tone Still Images—Requirements and Guidelines,” International Telecommunication Union (ITU-T), Recommendation T.81, Sep. 1992, 186 … [cited by applicant]
Isik et al., “LVAC: Learned volumetric attribute compression for point clouds using coordinate based networks,” Frontiers in Signal Processing, Oct. 11, 2022, 2:1008812, 18 pages. [cited by applicant]
Jiang et al., “MLIC: Multi-Reference Entropy Model for Learned Image Compression,” Presented at the 31st ACM International Conference on Multimedia, Oct. 29-Nov. 3, 2023, pp. 7618-7627. [cited by applicant]
Johnston et al., “Computationally Efficient Neural Image Compression,” CoRR, Dec. 18, 2019, arXiv:1912.08771v1, 9 pages. [cited by applicant]
Kumaraswamy, “A generalized probability density function for double-bounded random processes,” Journal of Hydrology, Mar. 1980, 46(1-2):79-88. [cited by applicant]
Kwan et al., “HiNeRV: Video Compression with Hierarchical Encoding-based Neural Representation,” Presented at the 37th Conference on Neural Information Processing Systems (NeurIPS 2023), Dec. 10-16, 2023, 13 pages. [cited by applicant]
Ladune et al., “Cool-Chic: Coordinate-based Low Complexity Hierarchical Image Codec,” Presented at the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 1-6, 2023, pp. 13515-13522. [cited by applicant]
Lanzendo et al., “Siamese SIREN: Audio Compression with Implicit Neural Representations,” CoRR, Submitted on Jun. 22, 2023, arXiv:2306.12957v1, 7 pages. [cited by applicant]
Le et al., “MobileCodec: Neural Inter-frame Video Compression on Mobile Devices,” Presented at the 13th ACM Multimedia Systems Conference (MMSys '22), Jun. 14- 17, 2022, pp. 324-330. [cited by applicant]
Lee et al., “FFNeRV: Flow-Guided Frame-Wise Neural Representations for Videos,” CoRR, Submitted on Aug. 7, 2023, arXiv:2212.12294v2, 15 pages. [cited by applicant]
Lee et al., “Meta-learning Sparse Implicit Neural Representations,” Presented at the 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Dec. 6-14, 2021, 12 pages. [cited by applicant]
Leguay et al., “Low-complexity Overfitted Neural Image Codec,” CoRR, Jul. 24, 2023, arXiv:2307.12706v1, 6 pages. [cited by applicant]
Li et al., “Compressing Volumetric Radiance Fields to 1 MB,” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 17-24, 2023, pp. 4222-4231. [cited by applicant]
Li et al., “E-NeRV: Expedite neural video representation with disentangled spatial-temporal context,” European Conference on Computer Vision, Nov. 4, 2022, pp. 267-284. [cited by applicant]
Lu et al., “Compressive neural representations of volumetric scalar fields, ” Computer Graphics Forum, Jun. 29, 2021, 40(3)135-146. [cited by applicant]
Lu et al., “Content adaptive and error propagation aware deep video compression,” European Conference on Computer Vision, Aug. 23-28, 2020, pp. 456-472. [cited by applicant]
Lv et al., “Dynamic low-rank instance adaptation for universal neural image compression,” Proceedings of the 31st ACM International Conference on Multimedia, Oct. 27, 2023, pp. 632-642. [cited by applicant]
Mancini et al., “Lossy compression of multidimensional medical images using sinusoidal activation networks: An evaluation study,” International Workshop on Computational Diffusion MRI, Dec. 1, 2022, pp. 26-37. [cited by applicant]
Martin, “Range encoding: An algorithm for removing redundancy from a digitized message,” Video & Data Recording Conference, Jul. 24-27, 1979, 9 pages. [cited by applicant]
Mentzer et al., “VCT: A video compression transformer,” Advances in Neural Information Processing Systems 35, 2022, 13 pages. [cited by applicant]
Mercat et al., “UVG dataset: 50/120fps 4K sequences for video codec analysis and development,” Proceedings of the 11th ACM Multimedia Systems Conference, May 27, 2020, pp. 297-302. [cited by applicant]
Mikami et al., “An efficient image compression method based on neural network: An overfitting approach,” 2021 IEEE International Conference on Image Processing (ICIP), Sep. 19-22, 2021, pp. 2084-2088. [cited by applicant]
Minnen et al., “Advancing the rate-distortion-computation frontier for neural image compression,” 2023 IEEE International Conference on Image Processing (ICIP), Oct. 8-11, 2023, pp. 2940-2944. [cited by applicant]
Minnen et al., “Joint autoregressive and hierarchical priors for learned image compression,” Advances in neural information processing systems 31, 2018, 10 pages. [cited by applicant]
netflixtechblog.com [online], “Per-Title Encode Optimization,” Dec. 14, 2015, retrieved on Nov. 25, 2024, retrieved from URL<https://netflixtechblog.com/per-title-encode-optimization-7e99442b62a2>, 23 pages. [cited by applicant]
Perez et al., “FiLM: Visual Reasoning with a General Conditioning Layer,” CoRR, Submitted on Dec. 18, 2017, arXiv:1709.07871v2, 13 pages. [cited by applicant]
Pham et al., “Autoencoding implicit neural representations for image compression,” ICML 2023 Workshop Neural Compression: From Information Theory to Applications, 2023, 13 pages. [cited by applicant]
r0K.us [online], “Kodak Lossless True Color Image Suite,” available on or before Oct. 18, 2023, via Internet Archive: Wayback Machine URL<https://web.archive.org/web/20231018203019/https://r0k.US/graphics/kodak/>, retri… [cited by applicant]
Ramirez et al., “L0onie: Compressing COINs with L0-constraints,” CoRR, Jul. 8, 2022, arxiv.org/abs/2207.04144, 14 pages. [cited by applicant]
Rippel et al., “ELF-VC: Efficient learned flexible-rate video coding,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 14479-14488. [cited by applicant]
Rozendaal et al., “Instance-Adaptive video compression: Improving neural codecs by training on the test set,” CoRR, Nov. 19, 2021, arxiv.org/abs/2111.10302. [cited by applicant]
Rozendaal et al., “Mobilenvc: Real-time 1080p neural video compression on a mobile device,” CoRR, Oct. 2, 2023, arxiv.org/abs/2310.01258, 15 pages. [cited by applicant]
Rozendaal et al., “Over-fitting for fun and profit: Instance-adaptive data compression,” International Conference on Learning Representations, Jan. 12, 2021, 18 pages. [cited by applicant]
Santurkar et al., “Generative Compression,” CoRR, Mar. 4, 2017, arxiv.org/abs/1703.01467, 10 pages. [cited by applicant]
Schwarz et al., “Meta-learning sparse compression networks,” Transactions on Machine Learning Research, Aug. 2022, 15 pages. [cited by applicant]
Schwarz et al., “Modality-agnostic variational compression of implicit neural representations,” Proceedings of the 40th International Conference on Machine Learning, Jul. 23, 2023, 1258:30342-30364. [cited by applicant]
Sheibanifard et al., “A novel implicit neural representation for vol. data,” Applied Sciences, Mar. 3, 2023, 13(5):3242. [cited by applicant]
Shi et al., “AlphaVC: High-performance and efficient learned video compression,” European Conference on Computer Vision, Nov. 9, 2022, pp. 616-631. [cited by applicant]
Strumpler et al., “Implicit neural representations for image compression,” European Conference on Computer Vision, Nov. 1, 2022, pp. 74-91. [cited by applicant]
Sullivan et al., “Overview of the high efficiency video coding (HEVC) standard,” IEEE Transactions on Circuits and Systems for Video Technology, Dec. 2012, 22(12):1649-1668. [cited by applicant]
Takikawa et al., “Variable bitrate neural fields,” CoRR, Jun. 15, 2022, arxiv.org/abs/2206.07707, 10 pages. [cited by applicant]
Theis et al., “Lossy image compression with compressive autoencoders,” CoRR, Mar. 1, 2017, arXiv:1703.00395, 19 pages. [cited by applicant]
Wu et al., “Video Compression through Image Interpolation,” CoRR, Apr. 18, 2018, arxiv.org/abs/1804.06919, 18 pages. [cited by applicant]
Xiang et al., “MIMT: Masked image modeling transformer for video compression,” The Eleventh International Conference on Learning Representations, Feb. 1, 2023, 17 pages. [cited by applicant]
Xie et al., “Neural fields in visual computing and beyond,” Computer Graphics Forum, May 24, 2022, pp. 641-676. [cited by applicant]
Yang et al., “An introduction to neural data compression,” Foundations and Trends® in Computer Graphics and Vision, Apr. 25, 2023, 15(2):113-200. [cited by applicant]
Yang et al., “Computationally-efficient neural image compression with shallow decoders,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 530-540. [cited by applicant]
Yang et al., “Improving inference for neural image compression,” Advances in Neural Information Processing Systems, 2020, 33:573-584. [cited by applicant]