IP Library › Granted Patent US 12,265,898
Granted Patent B2
US 12,265,898 · App. 18/409,520 · Granted Apr 1, 2025

Compression of machine-learned models via entropy penalized weight reparameterization

Inventors: Deniz Oktay (Mountain View, CA); Saurabh Singh (Mountain View, CA); Johannes Balle (San Francisco, CA); Abhinav Shrivastava (Silver Springs, MD)
Assignee: GOOGLE LLC
G06N20/00G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,898
App. No.
18/409,520
Granted
Apr 1, 2025
Kind
B2
Abstract

Example aspects of the present disclosure are directed to systems and methods that learn a compressed representation of a machine-learned model (e.g., neural network) via representation of the model parameters within a reparameterization space during training of the model. In particular, the present disclosure describes an end-to-end model weight compression approach that employs a latent-variable data compression method. The model parameters (e.g., weights and biases) are represented in a “latent” or “reparameterization” space, amounting to a reparameterization. In some implementations, this space can be equipped with a learned probability model, which is used first to impose an entropy penalty on the parameter representation during training, and second to compress the representation using arithmetic coding after training. The proposed approach can thus maximize accuracy and model compressibility jointly, in an end-to-end fashion, with the rate-error trade-off specified by a hyperparameter.

Claims (56)

1. A computing system for using learned parameter decoder models to provide improved compressed representations of machine-learned models, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing:

a compressed machine-learned model that comprises one or more learned parameter decoding models configured to decode parameters of the compressed machine-learned model from underlying representations, wherein a respective learned parameter decoding model of the one or more learned parameter decoding models is associated with a respective group of parameters; and

instructions that are executable by the one or more processors to cause the computing system to perform operations, wherein the operations comprise:

decoding, using the respective learned parameter decoding model, a parameter of the respective group of parameters from an underlying representation;

processing an input to a layer of the compressed machine-learned model using the decoded parameter; and

generating, based on processing the input using the decoded parameter, a prediction from the compressed machine-learned model.

2. The computing system of claim 1 , wherein the respective learned parameter decoding model was trained based on a loss evaluated over outputs of the compressed machine-learned model.

3. The computing system of claim 1 , wherein the operations comprise:

evaluating the generated prediction to compute a loss; and

backpropagating the loss through the respective learned parameter decoding model to the underlying representation; and

updating the underlying representation based on the backpropagated loss.

4. The computing system of claim 3 , wherein the operations comprise:

updating the respective learned parameter decoding model based on the backpropagated loss.

5. The computing system of claim 1 , wherein the underlying representations are quantized.

6. The computing system of claim 1 , wherein the underlying representations are continuous during training of the compressed machine-learned model and quantized after training of the compressed machine-learned model.

7. The computing system of claim 1 , wherein the decoded parameter is decoded in real-time for generating the prediction.

8. The computing system of claim 1 , wherein the learned parameter decoding model comprises a neural network.

9. The computing system of claim 1 ,

wherein the compressed machine-learned model uses the underlying representations and the one or more machine-learned parameter decoding models to obtain decoded parameters for performing runtime operations of the compressed machine-learned model; and

wherein the operations comprise:

transmitting, to a recipient system, the underlying representations and the one or more machine-learned parameter decoding models;

wherein a first datasize of the underlying representations and the one or more machine-learned parameter decoding models is smaller than a second datasize of an uncompressed version of the machine-learned model that directly stores parameters for performing the runtime operations of the uncompressed machine-learned model.

10. The computing system of claim 1 , wherein the compressed machine-learned model comprises a second respective machine-learned parameter decoding model, and wherein:

the respective machine-learned parameter decoding model is configured to decode a weight value used to scale the input; and

the second respective machine-learned parameter decoding model is configured to decode a bias value added to the scaled input.

11. The computing system of claim 1 , wherein the respective group of parameters is one of a plurality of groups of parameters that are respectively associated with a plurality of machine-learned parameter decoding models, and wherein the plurality of groups of parameters comprise:

a group of parameters for convolutional layers of the compressed machine-learned model.

12. The computing system of claim 1 , wherein the respective group of parameters is one of a plurality of groups of parameters that are respectively associated with a plurality of machine-learned parameter decoding models, and wherein the plurality of groups of parameters comprise:

a group of parameters for a fully-connected layer of the compressed machine-learned model.

13. The computing system of claim 1 , wherein the respective group of parameters is one of a plurality of groups of parameters that are respectively associated with a plurality of machine-learned parameter decoding models, and wherein the plurality of groups of parameters comprise:

a first group of parameters for a first fully-connected layer of the compressed machine-learned model; and

a second group of parameters for a second fully-connected layer of the compressed machine-learned model.

14. The computing system of claim 1 , wherein:

the respective group of parameters is characterized by a distribution; and

the parameter comprises a sample from the distribution that is obtained by processing the underlying representation using the machine-learned parameter decoding model.

15. The computing system of claim 1 , wherein the underlying representations are represented using scalar quantization.

16. The computing system of claim 3 , wherein the loss is configured to penalize an entropy associated with the underlying representations.

17. The computing system of claim 16 , wherein the operations comprise:

storing the compressed machine-learned model with the underlying representations compressed according to an arithmetic code.

18. One or more non-transitory computer-readable media storing:

a compressed machine-learned model that comprises one or more learned parameter decoding models configured to decode parameters of the compressed machine-learned model from underlying representations, wherein a respective learned parameter decoding model of the one or more learned parameter decoding models is associated with a respective group of parameters; and

instructions that are executable by one or more processors to cause a computing system to perform operations, wherein the operations comprise:

decoding, using the respective learned parameter decoding model, a parameter of the respective group of parameters from an underlying representation;

processing an input to a layer of the compressed machine-learned model using the decoded parameter; and

generating, based on processing the input using the decoded parameter, a prediction from the compressed machine-learned model.

19. A computer-implemented method comprising:

accessing a compressed machine-learned model that comprises one or more learned parameter decoding models configured to decode parameters of the compressed machine-learned model from underlying representations, wherein a respective learned parameter decoding model of the one or more learned parameter decoding models is associated with a respective group of parameters;

decoding, using the respective learned parameter decoding model, a parameter of the respective group of parameters from an underlying representation;

processing an input to a layer of the compressed machine-learned model using the decoded parameter; and

generating, based on processing the input using the decoded parameter, a prediction from the compressed machine-learned model.

20. The computer-implemented method of claim 19 , comprising:

receiving, over a network connection, the compressed machine-learned model, wherein receiving the compressed machine-learned model comprises:

receiving parameters of the respective learned parameter decoding model; and

receiving one or more underlying representations corresponding to the respective group of parameters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2024
From: OKTAY, DENIZ; SINGH, SAURABH; SHRIVASTAVA, ABHINAV; BALLE, JOHANNES
To: GOOGLE LLC
Reel/Frame 066498/0813 →
Continuity (4)
Continuation 18165211 · Feb 6, 2023
Continuation 15931016 · May 13, 2020
Provisional Application 62848523 · May 15, 2019
Related Publication 20240220863A1 · Jul 4, 2024
References Cited (42)
US 11574232B2 · Oktay · 2023 [cited by examiner]
US 11907818B2 · Oktay · 2024 [cited by examiner]
US 20180174050A1 · Holt · 2018 [cited by examiner]
US 20180174275A1 · Bourdev · 2018 [cited by examiner]
US 20200304802A1 · Habibian · 2020 [cited by examiner]
Balle et al., “End-To-End Optimized Image Compression”, arXiv:1611.01704v1, Nov. 5, 2016, 24 pages. [cited by applicant]
Balle et al., “Variational Image Compression with A Scale Hyperprior”, arXiv:1802.01436v2, May 1, 2018, 23 pages. [cited by applicant]
Baskin et al., “UNIQ: Uniform Noise Injection for Non-Uniform Quantization of Neural Networks”, arXiv:1804.10969v3, Oct. 2, 2018, 10 pages. [cited by applicant]
Bengio et al., “Estimating or Propagating Gradients through Stochastic Neurons for Conditional Computation”, arXiv:1308.3432v1, Aug. 15, 2013, 12 pages. [cited by applicant]
Chen et al., “Compressing Convolutional Neural Networks in the Frequency Domain”, 22 [cited by applicant]
Chen et al., “Compressing Neural Networks with the Hashing Trick”, 32 [cited by applicant]
Courbariaux et al., “Binary Connect: Training Deep Neural Networks with Binary Weights During Propagations”, Advances in Neural Information Processing Systems 28, vol. 1, Dec. 2015, 9 pages. [cited by applicant]
Cun et al., “Optimal Brain Damage”, Advances in Neural Information Processing Systems, Feb. 1990, pp. 598-605. [cited by applicant]
Dubey et al., “Coreset-Based Neural Network Compression”, 15 [cited by applicant]
Github.com, “TensorFlow Compression”, https://github.com/tensorflow/tensorflow, retrieved on Mar. 25, 2021, 6 pages. [cited by applicant]
Gupta et al., “Deep Learning with Limited Numerical Precision”, 32 [cited by applicant]
Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding”, arXiv:1510.00149v3, Nov. 20, 2015, 13 pages. [cited by applicant]
Havasi et al., “Minimal Random Code Learning: Getting Bits Back from Compressed Model Parameters”, arXiv:1810.00440v1, Sep. 30, 2018, 11 pages. [cited by applicant]
He et al., Deep Residual Leaming for Image Recognition, 29 [cited by applicant]
Huffman, “A Method for the Construction of Minimum-Redundancy Codes”, Proceedings of the Institute of Radio Engineers, vol. 40, No. 9, Sep. 1952, pp. 1098-1101. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, arXiv:1502.03167v3, Mar. 2, 2015, 11 pages. [cited by applicant]
Kingma et al., “ADAM: A Method for Stochastic Optimization”, arXiv:1412.6980v1, Dec. 22, 2014, 9 pages. [cited by applicant]
Krizhevsky, “Learning Multiple Layers of Features from Tiny Images”, University of Toronto, Master's Thesis, Department of Computer Science, Apr. 8, 2009, 60 pages. [cited by applicant]
Lecun et al., “Gradient-Based Learning Applied to Document Recognition”, Proceedings of the Institute of Electrical and Electronics Engineers, Nov. 1998, 46 pages. [cited by applicant]
Lecun et al., Yann.lecun.com, “The MNIST Database”, http://yann.lecun.com/exdb/mnist/, retrieved on Jun. 24, 2020, 8 pages. [cited by applicant]
Li et al., “Pruning Filters for Efficient ConvNets”, arXiv:1608.08710v2, Sep. 15, 2016, 9 pages. [cited by applicant]
Li et al., “Ternary weight networks”, arXiv:1605.04711v2, Nov. 19, 2016, 5 pages. [cited by applicant]
Louizos et al., “Bayesian Compression for Deep Learning”, 31 [cited by applicant]
Louizos et al., “Relaxed Quantization for Discretized Neural Networks”, arXiv:1810.01875v1, Oct. 3, 2018, 14 pages. [cited by applicant]
Molchanov et al., “Variational Dropout Sparsifies Deep Neural Networks”, arXiv:1701.05369v3, Jun. 13, 2017, 10 pages. [cited by applicant]
Rissanen et al., “Universal Modeling and Coding”, Institute of Electrical and Electronics Engineers Transactions on Information Theory, vol. 27, No. 1, Jan. 1981, pp. 12-23. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge”, arXiv:1409.0575v3, Jan. 30, 2018, 43 pages. [cited by applicant]
Shannon, “A Mathematical Theory of Communication”, The Bell System Technical Journal, vol. 27, No. 3, 1948, 55 pages. [cited by applicant]
Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, arXiv:1409.1556v6, Apr. 10, 2015, 14 pages. [cited by applicant]
Srinivas et al., “Data-free Parameter Pruning for Deep Neural Networks”, arXiv:1507.06149, Jul. 22, 2015, 12 pages. [cited by applicant]
Theiss et al., “Lossy Image Compression with Compressive Autoencoders”, arXiv:1703.00395v1, Mar. 1, 2017, 19 pages. [cited by applicant]
Ullrich et al., “Soft Weight-Sharing for Neural Network Compression”, arXiv:1702.04008v2, May 9, 2017, 16 pages. [cited by applicant]
Wang et al., “CNNpack: Packing Convolutional Neural Networks in the Frequency Domain”, Annual Conference on Advances in Neural Information Processing Systems 2016, Dec. 5-10, 2016, Barcelona, Spain, 9 pages. [cited by applicant]
Wiedemann et al., “DeepCABAC: Context-adaptive Binary Arithmetic Coding for Deep Neural Network Compression”, arXiv:1905.08318v1, May 15, 2019, 4 pages. [cited by applicant]
Wiedemann et al., “Entropy-Constrained Training of Deep Neural Networks”, arXiv:1812.07520v1, Dec. 18, 2018, 8 pages. [cited by applicant]
Zagoruyko et al., “Wide Residual Networks”, 27 [cited by applicant]
Zhou et al., “Explicit Loss-Error-Aware Quantization for Low-Bit Deep Neural Networks”, Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 18-22, 2018, S… [cited by applicant]