IP Library Granted Patent US 12,481,877
Granted Patent B2
US 12,481,877 · App. 17/411,936 · Granted Nov 25, 2025

Instance-adaptive image and video compression in a network parameter subspace using machine learning systems

Inventors: Johann Hinrich Brehmer (Amsterdam, NL); Ties Jehan Van Rozendaal (Amsterdam, NL); Yunfan Zhang (Amsterdam, NL); Taco Sebastiaan Cohen (Amsterdam, NL)
Assignee: QUALCOMM Incorporated
G06N3/08G06F18/211
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,877
App. No.
17/411,936
Granted
Nov 25, 2025
Kind
B2
Abstract

Techniques are described for compressing data using machine learning systems. An example process can include receiving input data for compression by a neural network compression system. The process can include determining, based on the input data, a set of updated model parameters for the neural network compression system, wherein the set of updated model parameters is selected from a subspace of model parameters. The process can include generating at least one bitstream including a compressed version of the input data and a compressed version of one or more subspace coordinates that correspond to the set of updated model parameters. The process can include outputting the at least one bitstream for transmission to a receiver.

Claims (59)

1 . A method of processing image data, comprising:

receiving input data for compression by a neural network compression system;

determining, based on the input data, a set of updated model parameters for the neural network compression system, wherein the set of updated model parameters is selected from a subspace of model parameters based on minimizing, at inference time of a rate-distortion autoencoder (RD-AE) of the neural network compression system, a combined rate-distortion-model rate (RDM) loss corresponding to a rate-distortion loss of the RD-AE at the inference time and a number of model update bits used to represent the subspace of model parameters:

generating at least one bitstream including a compressed version of the input data and a compressed version of one or more subspace coordinates that correspond to the set of updated model parameters; and

outputting the at least one bitstream for transmission to a receiver.

2 . The method of claim 1 , wherein the subspace of model parameters includes a portion of a plurality of weight vectors.

3 . The method of claim 2 , wherein each of the plurality of weight vectors correspond to a weight vector used during training of the neural network compression system.

4 . The method of claim 2 , wherein the portion of the plurality of weight vectors is determined using at least one of principal component analysis (PCA), sparse principal component analysis (SPCA), and model-agnostic meta-learning (MAML).

5 . The method of claim 1 , further comprising:

generating a set of global model parameters based on a training dataset used to train the neural network compression system, wherein the one or more subspace coordinates that correspond to the set of updated model parameters are relative to the set of global model parameters.

6 . The method of claim 5 , wherein determining the set of updated model parameters from the subspace of model parameters comprises:

tuning the set of global model parameters using the input data, wherein the set of global model parameters are tuned based on a bit size of the compressed version of the input data and a distortion between the input data and reconstructed data generated from the compressed version of the input data.

7 . The method of claim 1 , further comprising:

quantizing the one or more subspace coordinates to yield one or more quantized subspace coordinates, wherein the at least one bitstream comprises a compressed version of the one or more quantized subspace coordinates.

8 . The method of claim 7 , wherein the at least one bitstream comprises a plurality of encoded quantization parameters used for quantizing the one or more subspace coordinates.

9 . The method of claim 1 , wherein generating the at least one bitstream comprises:

entropy encoding the one or more subspace coordinates using a model prior.

10 . The method of claim 1 , further comprising:

sending the subspace of model parameters to the receiver.

11 . An apparatus comprising:

at least one memory; and

at least one processor coupled to the at least one memory and configured to:

receive input data for compression by a neural network compression system;

determine, based on the input data, a set of updated model parameters for the neural network compression system, wherein the set of updated model parameters is selected from a subspace of model parameters based on minimizing, at inference time of a rate-distortion autoencoder (RD-AE) of the neural network compression system, a combined rate-distortion-model rate (RDM) loss corresponding to a rate-distortion loss of the RD-AE at the inference time and a number of model update bits used to represent the subspace of model parameters;

generate at least one bitstream including a compressed version of the input data and a compressed version of one or more subspace coordinates that correspond to the set of updated model parameters; and

output the at least one bitstream for transmission to a receiver.

12 . The apparatus of claim 11 , wherein the subspace of model parameters includes a portion of a plurality of weight vectors.

13 . The apparatus of claim 12 , wherein each of the plurality of weight vectors correspond to a weight vector used during training of the neural network compression system.

14 . The apparatus of claim 12 , wherein the portion of the plurality of weight vectors is determined using at least one of principal component analysis (PCA), sparse principal component analysis (SPCA), and model-agnostic meta-learning (MAML).

15 . The apparatus of claim 11 , where the at least one processor is further configured to:

generate a set of global model parameters based on a training dataset used to train the neural network compression system, wherein the one or more subspace coordinates that correspond to the set of updated model parameters are relative to the set of global model parameters.

16 . The apparatus of claim 15 , wherein to determine the set of updated model parameters from the subspace of model parameters the at least one processor is further configured to:

tune the set of global model parameters using the input data, wherein the set of global model parameters are tuned based on a bit size of the compressed version of the input data and a distortion between the input data and reconstructed data generated from the compressed version of the input data.

17 . The apparatus of claim 11 , where the at least one processor is further configured to:

quantize the one or more subspace coordinates to yield one or more quantized subspace coordinates, wherein the at least one bitstream comprises a compressed version of the one or more quantized subspace coordinates.

18 . The apparatus of claim 17 , wherein the at least one bitstream comprises a plurality of encoded quantization parameters used for quantizing the one or more subspace coordinates.

19 . The apparatus of claim 11 , wherein to generate the at least one bitstream the at least one processor is further configured to:

entropy encode the one or more subspace coordinates using a model prior.

20 . The apparatus of claim 11 , wherein the at least one processor is further configured to:

send the subspace of model parameters to the receiver.

21 . A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to:

receive input data for compression by a neural network compression system;

determine, based on the input data, a set of updated model parameters for the neural network compression system, wherein the set of updated model parameters is selected from a subspace of model parameters based on minimizing, at inference time of a rate-distortion autoencoder (RD-AE) of the neural network compression system, a combined rate-distortion-model rate (RDM) loss corresponding to a rate-distortion loss of the RD-AE at the inference time and a number of model update bits used to represent the subspace of model parameters;

generate at least one bitstream including a compressed version of the input data and a compressed version of one or more subspace coordinates that correspond to the set of updated model parameters; and

output the at least one bitstream for transmission to a receiver.

22 . The computer-readable storage medium of claim 21 , wherein the subspace of model parameters includes a portion of a plurality of weight vectors.

23 . The computer-readable storage medium of claim 22 , wherein each of the plurality of weight vectors correspond to a weight vector used during training of the neural network compression system.

24 . The computer-readable storage medium of claim 22 , wherein the portion of the plurality of weight vectors is determined using at least one of principal component analysis (PCA), sparse principal component analysis (SPCA), and model-agnostic meta-learning (MAML).

25 . The computer-readable storage medium of claim 21 , wherein the instructions further cause the one or more processors to:

generate a set of global model parameters based on a training dataset used to train the neural network compression system, wherein the one or more subspace coordinates that correspond to the set of updated model parameters are relative to the set of global model parameters.

26 . The computer-readable storage medium of claim 25 , wherein to determine the set of updated model parameters from the subspace of model parameters, the instructions further cause the one or more processors to:

tune the set of global model parameters using the input data, wherein the set of global model parameters are tuned based on a bit size of the compressed version of the input data and a distortion between the input data and reconstructed data generated from the compressed version of the input data.

27 . The computer-readable storage medium of claim 21 , where the instructions further cause the one or more processors to:

quantize the one or more subspace coordinates to yield one or more quantized subspace coordinates, wherein the at least one bitstream comprises a compressed version of the one or more quantized subspace coordinates.

28 . The computer-readable storage medium of claim 27 , wherein the at least one bitstream comprises a plurality of encoded quantization parameters used for quantizing the one or more subspace coordinates.

29 . The computer-readable storage medium of claim 21 , wherein to generate the at least one bitstream, the instructions cause the one or more processors to:

entropy encode the one or more subspace coordinates using a model prior.

30 . The computer-readable storage medium of claim 21 , wherein the instructions cause the one or more processors to:

send the subspace of model parameters to the receiver.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE FIRST NAME AND EXECUTION DATE FOR INVENTOR THREE PREVIOUSLY RECORDED ON REEL 058312 FRAME 0342. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 2, 2023
From: BREHMER, JOHANN HINRICH; VAN ROZENDAAL, TIES JEHAN; ZHANG, YUNFAN; COHEN, TACO SEBASTIAAN
To: QUALCOMM INCORPORATED
Reel/Frame 065547/0746 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2021
From: BREHMER, JOHANN HINRICH; VAN ROZENDAAL, TIES JEHAN; ZHANG, YUFAN; COHEN, TACO SEBASTIAAN
To: QUALCOMM INCORPORATED
Reel/Frame 058312/0342 →
Continuity (1)
Related Publication 20230074979A1 · Mar 9, 2023
References Cited (9)
US 20050265618A1 · Jebara · 2005 [cited by examiner]
US 20200311551A1 · Aytekin et al. · 2020 [cited by applicant]
EP 3716158A2 · 2020 [cited by examiner]
Yang, Fei, et al. “Slimmable compressive autoencoders for practical neural image compression.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021. (Year: 2021). [cited by examiner]
Yu, Jiahui, et al. “Slimmable neural networks.” arXiv preprint arXiv:1812.08928 (2018). (Year: 2018). [cited by examiner]
Habibian, Amirhossein, et al. “Video compression with rate-distortion autoencoders.” Proceedings of the IEEE/CVF international conference on computer vision. 2019.) (Year: 2019). [cited by examiner]
Gusak J., et al., “One Time is Not Enough: Iterative Tensor Decomposition for Neural Network Compression”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University, Ithaca, NY, 14853, Mar. 24, 2019 (Ma… [cited by applicant]
International Search Report and Written Opinion—PCT/US2022/074440—ISA/EPO—Nov. 25, 2022. [cited by applicant]
Zou N., et al., “Learning to Learn to Compress”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY, 14853, May 1, 2021 (May 1, 2021), 6 Pages, XP081950101, sections I, III, figure 1. [cited by applicant]
Cited By (1)
US 12,731,298