IP Library Granted Patent US 12,634,464
Granted Patent B2
US 12,634,464 · App. 18/855,278 · Granted May 19, 2026

Deep-learning-based compression method using frequency decomposition

Inventors: Hyomin Choi (Los Altos, CA); Fabien Racape (Los Altos, CA); Shahab Hamidi-Rad (Los Altos, CA); Simon Feltman (Los Altos, CA)
Assignee: InterDigital VC Holdings, Inc.
H04N19/13H04N19/167H04N19/30H04N19/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,634,464
App. No.
18/855,278
Granted
May 19, 2026
Kind
B2
Abstract

In one implementation, we propose an end-to-end image video compression method that decomposes the spatial frequencies of the input content into a partitioned latent representation. Decomposed frequencies in the latent space are analyzed and grouped into separate latent representation or separate tensors, each tensor being jointly optimized to be decoded independently one from another. Therefore, the decoder can independently decode the tensors in a scalable manner to progressively reconstruct the input. This method enables quality scalability by progressively transmitting individual latent representations of decomposed frequency data, separated in the produced latent space. Furthermore, the quality scalability of region of interest (ROI) is enabled by which the decoder takes only corresponding latent representations in the enhancement tensors as input together with latent representations already delivered to the decoder.

Claims (44)

1 . A method of video encoding, comprising:

decomposing at least a part of an image into a plurality of frequency groups by a plurality of decomposition layers, wherein each decomposition layer performs an intra frequency process and an inter frequency process, and wherein a frequency group corresponds to a set of frequency bands;

generating a respective latent representation in a latent space for each frequency group of said plurality of frequency groups; and

entropy encoding one or more of said respective latent representations.

2 . The method of claim 1 , wherein said intra frequency process is performed by a set of convolutional layers.

3 . The method of claim 1 , wherein said inter frequency process comprises:

performing a frequency decomposition transform on an input data set associated with a frequency group, followed by one or more convolutional layers to form an output data set; and

adding said output data set to a result of an intra frequency process for another frequency group neighboring to said frequency group.

4 . The method of claim 1 , further comprising:

signaling location information of at least a region of interest of said image.

5 . A method of video decoding, comprising:

obtaining one or more latent representations in a latent space, wherein each of said one or more latent representations corresponds to a frequency group of one or more frequency groups, wherein a frequency group corresponds to a set of frequency bands;

obtaining said one or more frequency groups from said one or more latent representations; and

composing at least a part of an image from said one or more frequency groups by a plurality of frequency composition layers, wherein each frequency composition layer performs an intra frequency process and an inter frequency process.

6 . The method of claim 5 , wherein said intra frequency process is performed by a set of convolutional layers.

7 . The method of claim 5 , wherein said inter frequency process is performed by a set of convolutional layers.

8 . The method of claim 5 , further comprising:

obtaining another one or more latent representations in said latent space; and

obtaining another one or more frequency groups from said another one or more latent representations, wherein said at least a part of said image is composed further based on said another one or more frequency groups.

9 . The method of claim 8 , wherein said another one or more latent representations associated with another one or more frequency groups are set to a constant value.

10 . The method of claim 5 , further comprising:

receiving location information of at least a region of interest of said image.

11 . An apparatus, comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to:

decompose at least a part of an image into a plurality of frequency groups by a plurality of decomposition layers, wherein each decomposition layer performs an intra frequency process and an inter frequency process, and wherein a frequency group corresponds to a set of frequency bands;

generate a respective latent representation in a latent space for each frequency group of said plurality of frequency groups; and

entropy encode one or more of said respective latent representations.

12 . The apparatus of claim 11 , wherein said intra frequency process is performed by a set of convolutional layers.

13 . The apparatus of claim 11 , wherein said inter frequency process comprises:

performing a frequency decomposition transform on an input data set associated with a frequency group, followed by one or more convolutional layers to form an output data set; and

adding said output data set to a result of an intra frequency process for another frequency group neighboring to said frequency group.

14 . The apparatus of claim 11 , wherein said one or more processors are further configured to:

signal location information of at least a region of interest of said image.

15 . An apparatus, comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to:

obtain one or more latent representations in a latent space, wherein each of said one or more latent representations corresponds to a frequency group of one or more frequency groups, wherein a frequency group corresponds to a set of frequency bands;

obtain said one or more frequency groups from said one or more latent representations; and

compose at least a part of an image from said one or more frequency groups by a plurality of frequency composition layers, wherein each frequency composition layer performs an intra frequency process and an inter frequency process.

16 . The apparatus of claim 15 , wherein said intra frequency process is performed by a set of convolutional layers.

17 . The apparatus of claim 15 , wherein said inter frequency process is performed by a set of convolutional layers.

18 . The apparatus of claim 15 , wherein said one or more processors are further configured to:

obtain another one or more latent representations in said latent space; and

obtain another one or more frequency groups from said another one or more latent representations, wherein said at least a part of said image is composed further based on said another one or more frequency groups.

19 . The apparatus of claim 18 , wherein said another one or more latent representations associated with another one or more frequency groups are set to a constant value.

20 . The apparatus of claim 15 , wherein said one or more processors are further configured to:

receive location information of at least a region of interest of said image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2024
From: CHOI, HYOMIN; RACAPE, FABIEN; HAMIDI-RAD, SHAHAB; FELTMAN, SIMON
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068880/0239 →
Continuity (2)
Provisional Application 63338562 · May 5, 2022
Related Publication 20250247538A1 · Jul 31, 2025
References Cited (18)
US 5235420A · Gharavi · 1993 [cited by examiner]
US 5864780A · Kossentini · 1999 [cited by examiner]
US 12354402B2 · Wasnik · 2025 [cited by examiner]
US 20080310506A1 · Xu · 2008 [cited by examiner]
US 20160125599A1 · Stampanoni · 2016 [cited by examiner]
US 20170155924A1 · Gokhale · 2017 [cited by examiner]
Chen, et al., “Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution”, arXiv:1904.05049v3 (cs.CV), pp. 1-12, Aug. 18, 2019. [cited by applicant]
Choi et al., “Frequency-aware Learned Image Compression for Quality Scalability”, arxiv.org, Cornell University Libary, 201 Olin Library Cornell University Itacha, NY 14853, arXiv2301.01290v1, [eess.IV] Jan. 3, 2023. 6 … [cited by applicant]
Agarwal et al., “Deep Learning Based Image Compression in DWT Domain”, 2022 8th International Conference on Advanced Computing and Communication Systems (ICACCS), IEEE, vol. 1, Mar. 25, 2022, pp. 445-453. [cited by applicant]
Ma et al., “iWave: CNN-Based Wavelet-Like Transform for Image Compression”, IEEE Transactions on Multimedia, vol. 22, No. 7, pp. 1667-1679, Jul. 2020. [cited by applicant]
He et al., “An Enhanced Multi-frequency Learned Image Compression Method”, in Pattern Recognition and Computer Vision: 4th Chinese Conference, PRCV 2021, Beijing, China, Oct. 29-Nov. 1, 2021, Proceedings, Part III 4, pp… [cited by applicant]
Yang et al., “Deep Image Compression in the Wavelet Transform Domain Based on High Frequency Sub-Band Prediction”, IEEE Access, vol. 7, Apr. 16, 2019, 14 pages. [cited by applicant]
Ma et al., “End-to-End Optimized Versatile Image Compression With Wavelet-Like Transform”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 3, Mar. 2022, 17 pages. [cited by applicant]
Balle et al., “Variational Image Compression with a Scale Hyperprior”, Cornell University Library, Electrical Engineering and Systems Science, Image and Video Processing, arXiv: 1802.01436v2, May 1, 2018, 23 pages. [cited by applicant]
Cui et al., “G-VAE: A Continuously Variable Rate Deep Image Compression Framework”, arXiv preprint arXiv: 2003.02012 2.3 (2020), pp. 1-8. [cited by applicant]
Minnen et al., “Joint Autoregressive and Hierarchical Priors for Learned Image Compression”, Advances in Neural Information Processing Systems 31, arXiv:1809.02736v1 (cs.CV), pp. 1-22, Sep. 8, 2018. [cited by applicant]
Balle et al., “Density Modeling of Images Using a Generalized Normalization Transformation”, arXiv: 1511.06281v4 (cs, LG), pp. 1-14, Feb. 29, 2016. [cited by applicant]
Shi et al., “Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network”, arXiv:1609.05158v2 (cs.CV), pp. 1-10, Sep. 23, 2016. [cited by applicant]