IP Library Granted Patent US 12,450,476
Granted Patent B2
US 12,450,476 · App. 17/388,781 · Granted Oct 21, 2025

Compression of machine-learned models by vector quantization

Inventors: Ting Wei Liu (Dollard-des-Ormeaux, CA); Julieta Martinez Covarrubias (Toronto, CA); Jashan Sunil Shewakramani (Waterloo, CA); Raquel Urtasun (Toronto, CA); Wenyuan Zeng (Toronto, CA)
Assignee: AURORA OPERATIONS, INC.
G06N3/08G05B13/027G06N20/00H03M7/3082H03M7/70G05D1/0088G05D1/0274
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,476
App. No.
17/388,781
Granted
Oct 21, 2025
Kind
B2
Abstract

A computing system can include one or more processors and one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the computing system to perform operations including obtaining model structure data indicative of a plurality of parameters of a machine-learned model; determining a codebook comprising a plurality of centroids, the plurality of centroids having a respective index of a plurality of indices indicative of an ordering of the codebook; determining a plurality of codes respective to the plurality of parameters, the plurality of codes respectively comprising a code index of the plurality of indices corresponding to a closest centroid of the plurality of centroids to a respective parameter of the plurality of parameters; and providing encoded data as an encoded representation of the plurality of parameters of the machine-learned model, the encoded data comprising the codebook and the plurality of codes.

Claims (49)

1. A computing system comprising:

one or more processors; and

one or more computer-readable medium storing instructions that when executed by the one or more processors cause the computing system to perform operations, the operations comprising:

obtaining model structure data indicative of a plurality of parameters of a machine-learned model;

determining a codebook comprising a plurality of centroids, the plurality of centroids having a respective index of a plurality of indices indicative of an ordering of the codebook;

determining a plurality of codes respective to the plurality of parameters, the plurality of codes respectively comprising a code index of the plurality of indices corresponding to a closest centroid of the plurality of centroids to a respective parameter of the plurality of parameters; and

deploying the codebook and the plurality of codes as a compressed representation of the machine-learned model to an autonomous vehicle;

wherein determining the codebook comprising the plurality of centroids comprises learning the plurality of centroids concurrently with the plurality of codes to optimize a reconstruction error between the plurality of parameters and a reconstructed plurality of parameters that is reconstructed from the compressed representation of the machine-learned model.

2. The computing system of claim 1 , wherein the plurality of parameters comprises a plurality of weights of at least one layer of the machine-learned model, the plurality of weights comprising a weight matrix of the at least one layer, wherein a length of the codebook is less than a number of plurality of weights of the weight matrix.

3. The computing system of claim 2 , wherein the weight matrix comprises a plurality of subvectors, each subvector of the plurality of subvectors comprising a block of contiguous scalars in a column of the weight matrix, and wherein the plurality of codes are respective to the plurality of subvectors.

4. The computing system of claim 2 , wherein the at least one layer comprises a fully-connected (FC) layer, and wherein the plurality of weights comprises weights of connections from a prior layer to the fully-connected layer.

5. The computing system of claim 2 , wherein the at least one layer comprises a convolutional layer, wherein the plurality of weights comprises weights of a convolutional kernel, and wherein the weight matrix is reshaped into a two-dimensional matrix.

6. The computing system of claim 2 , wherein a set comprising the plurality of centroids of the codebook is smaller than a set comprising the plurality of parameters of the machine-learned model.

7. The computing system of claim 2 , wherein the weight matrix is permuted by a row permutation matrix, and wherein the operations comprise determining the row permutation matrix such that a determinant of a covariance of the plurality of weights is optimized.

8. The computing system of claim 7 , wherein determining the row permutation matrix comprises:

obtaining an initial row permutation matrix that optimizes a product of diagonal elements of the initial row permutational matrix, wherein obtaining the initial row permutation matrix comprises:

determining a plurality of buckets of row indices;

determining a variance of each row of the weight matrix;

assigning each row index of the plurality of buckets of row indices to a non-full bucket that results in a lowest variance of the plurality of buckets; and

interlacing rows from the plurality of buckets such that rows from a same bucket are placed a number of rows apart; and

iteratively searching a plurality of candidate permutations of the initial row permutation matrix to select the row permutation matrix as a selected candidate permutation of the plurality of candidate permutations based at least in part on a determinant of a covariance of the selected candidate permutation.

9. The computing system of claim 1 , wherein the reconstruction error is optimized by minimizing a covariance of the plurality of parameters.

10. The computing system of claim 1 , wherein the closest centroid to the respective parameter is closest to the respective parameter in Euclidean distance.

11. The computing system of claim 1 , wherein, subsequent to initialization of the plurality of codes and the codebook, the plurality of codes and the codebook are iteratively updated with random noise over one or more update iterations.

12. The computing system of claim 11 , wherein, subsequent to updating the plurality of codes and the codebook with random noise over the one or more update iterations, the plurality of centroids is fine-tuned by gradient-based learning.

13. The computing system of claim 1 , wherein the operations further comprise:

detecting one or more objects in an environment of the autonomous vehicle using the compressed representation of the machine-learned model; and

controlling the autonomous vehicle based on the one or more objects detected in the environment of the autonomous vehicle.

14. The computing system of claim 1 , wherein the codebook comprises a lookup table comprising the plurality of centroids, and wherein the code index for the respective parameter indexes the closest centroid in the lookup table.

15. A computer-implemented method for compressing a machine-learned model, the method comprising:

obtaining model structure data indicative of a plurality of parameters of a machine-learned model;

determining a codebook comprising a plurality of centroids, the plurality of centroids having a respective index of a plurality of indices indicative of an ordering of the codebook;

determining a plurality of codes respective to the plurality of parameters, the plurality of codes respectively comprising a code index of the plurality of indices corresponding to a closest centroid of the plurality of centroids to a respective parameter of the plurality of parameters; and

deploying the codebook and the plurality of codes as a compressed representation of the machine-learned model to an autonomous vehicle;

wherein the plurality of parameters comprises a plurality of weights of at least one layer of the machine-learned model, the plurality of weights comprising a weight matrix of the at least one layer; and

wherein determining the codebook comprising the plurality of centroids comprises learning the plurality of centroids concurrently with the plurality of codes to optimize a reconstruction error between the plurality of parameters and a reconstructed plurality of parameters that is reconstructed from the compressed representation of the machine-learned model.

16. The computer-implemented method of claim 15 , wherein the weight matrix is permuted by a row permutation matrix, and wherein the method comprises determining the row permutation matrix such that a determinant of a covariance of the plurality of weights is optimized;

wherein determining the row permutation matrix comprises:

obtaining an initial row permutation matrix that optimizes a product of diagonal elements of the initial row permutational matrix, wherein obtaining the initial row permutation matrix comprises:

determining a plurality of buckets of row indices;

determining a variance of each row of the weight matrix;

assigning each row index of the plurality of buckets of row indices to a non-full bucket that results in a lowest variance of the plurality of buckets; and

interlacing rows from the plurality of buckets such that rows from a same bucket are placed a number of rows apart; and

iteratively searching a plurality of candidate permutations of the initial row permutation matrix to select the row permutation matrix as a selected candidate permutation of the plurality of candidate permutations based at least in part on a determinant of a covariance of the selected candidate permutation.

17. The computer-implemented method of claim 15 , wherein the reconstruction error is optimized by minimizing a covariance of the plurality of parameters.

18. The computer-implemented method of claim 15 , wherein, subsequent to initialization of the plurality of codes and the codebook, the plurality of codes and the codebook are iteratively updated with random noise over one or more update iterations; and wherein, subsequent to updating the plurality of codes and the codebook with random noise over the one or more update iterations, the plurality of centroids is fine-tuned by gradient-based learning.

19. The computer-implemented method of claim 15 , further comprising:

detecting one or more objects in an environment of the autonomous vehicle using the compressed representation of the machine-learned model; and

controlling the autonomous vehicle based on the one or more objects detected in the environment of the autonomous vehicle.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 067733/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2023
From: URTASUN, RAQUEL
To: UATC, LLC
Reel/Frame 062642/0717 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2022
From: UBER TECHNOLOGIES, INC.
To: UATC, LLC
Reel/Frame 058962/0140 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2022
From: LIU, TING WEI; ZENG, WENYUAN; SHEWAKRAMANI, JASHAN SUNIL; COVARRUBIAS, JULIETA MARTINEZ
To: UATC, LLC
Reel/Frame 058795/0654 →
EMPLOYMENT AGREEMENT Recorded Jan 24, 2022
From: SOTIL, RAQUEL URTASUN
To: UBER TECHNOLOGIES, INC.
Reel/Frame 058826/0936 →
Continuity (2)
Provisional Application 63058041 · Jul 29, 2020
Related Publication 20220036184A1 · Feb 3, 2022
References Cited (46)
US 20220261616A1 · Li · 2022 [cited by examiner]
Cheng et al, “A Survey on Model Compression and Acceleration for Deep Neural Networks”, arXiv:1710.09282v9, Jun. 14, 2020, 10 pages. [cited by applicant]
Courbariaux et al., “Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or −1”, arXiv:1602.02830v3, Mar. 17, 2016, 11 pages. [cited by applicant]
Denil et al., “Predicting Parameters in Deep Learning”, Conference on Neural Information Processing Systems, Dec. 5-10, 2013, Lake Tahoe, Nevada, United States, 9 pages. [cited by applicant]
Denton et al., “Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation”, Conference on Neural Information Processing Systems, Dec. 8-13, 2014, Montreal, Canada, 9 pages. [cited by applicant]
Dong et al.“HAWQ-V2: Hessian Aware Trace-Weighted Quantization of Neural Networks”, arXiv:1911.03852v1, Nov. 10, 2019, 13 pages. [cited by applicant]
Dong et al., “HAWQ: Hessian-Aware Quantization of Neural Networks with Mixed Precision”, arXiv:1905.03696v1, Apr. 29, 2019, 12 pages. [cited by applicant]
Dong et al., “Learning to Prune Deep Neural Networks via Layer-wise Optimal Brain Surgeon”, Conference on Neural Information Processing Systems, Dec. 4-9, 2017, Long Beach, California, United States, 11 pages. [cited by applicant]
Ge et al., “Optimized Product Quantization for Approximate Nearest Neighbor Search”, Conference on Computer Vision and Pattern Recognition, Jun. 23-28, 2013, Portland, Oregon, pp. 2946-2953. [cited by applicant]
Gersho et al., “Vector Quantization and Signal Compression” Springer Science & Business Media, 1991, pages. [cited by applicant]
Gong et al., “Compressing Deep Convolutional Networks using Vector Quantization”, arXiv:1412.6115v1, Dec. 18, 2014, 10 pages. [cited by applicant]
Guo et al., “Dynamic Network Surgery for Efficient DNNs”, arXiv:1608.04493v2, Nov. 10, 2016, 9 pages. [cited by applicant]
Guo, “A Survey on Methods and Theories of Quantized Neural Networks”, arXiv:1808.04752v2, Dec. 16, 2018, 17 pages. [cited by applicant]
Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding”, arXiv:1510.00149v5, Feb. 15, 2016, 14 pages. [cited by applicant]
Han et al., “Learning both Weights and Connections for Efficient Neural Networks”, arXiv:1506.02626v3, Oct. 30, 2015, 9 pages. [cited by applicant]
Hassibi et al., “Second Order Derivatives for Network Pruning: Optimal Brain Surgeon”, Conference on Neural Information Processing Systems, 1993, Denver, Colorado, United States, pp. 164-171. [cited by applicant]
He et al., “Mask R-CNN”, arXiv:1703.06870v3, Jan. 24, 2018, 12 pages. [cited by applicant]
He et al., “AMC: AutoML for Model Compression and Acceleration on Mobile Devices”, arXiv:1802.03494v4, Jan. 16, 2019, 17 pages. [cited by applicant]
He et al., “Channel Pruning for Accelerating Very Deep Neural Networks”, arXiv:1707.06168v2, Aug. 21, 2017, 10 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, arXiv:1512.03385v1, Dec. 10, 2015, 12 pages. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network”, arXiv:1503.02531v1, Mar. 9, 2015, 9 pages. [cited by applicant]
Jaderberg et al., “Speeding up Convolutional Neural Networks with Low Rank Expansions”, arXiv:1405.3866v1, May 15, 2014, 12 pages. [cited by applicant]
Jegou et al., “Product Quantization for Nearest Neighbor Search”, Transactions on Pattern Analysis and Machine Intelligence, vol. 33 No. 1, 12 pages. [cited by applicant]
Kingma et al, “Adam: a Method for Stochastic Optimization”, arXiv:1412.6980v9, Jan. 30, 2017, 15 pages. [cited by applicant]
Lebedev et al., “Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition”, arXiv:1412.6553v3, Apr. 24, 2015, 11 pages. [cited by applicant]
Li et al, “Fully Quantized Network for Object Detection”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, California, United States, pp. 2810-2819. [cited by applicant]
Li et al., “Pruning Filters for Efficient ConvNets”, arXiv:1608.08710v3, Mar. 10, 2017, 13 pages. [cited by applicant]
Lin et al, “Focal Loss for Dense Object Detection”, International Conference on Computer Vision, Oct. 22-29, 2017, Venice, Italy, pp. 2980-2988. [cited by applicant]
Lin et al, “Microsoft COCO: Common Objects in Context”, arXiv:1405.0312v3, Feb. 21, 2015, 15 pages. [cited by applicant]
Lin et al, “Towards Accurate Binary Convolutional Neural Network”, arXiv1711.11294v1, Nov. 30, 2017, 14 pages. [cited by applicant]
Lourenco et al., “Iterated Local Search”, Handbook of MetaHeuristics, 49 pages. [cited by applicant]
Novikov et al., “Tensorizing Neural Networks”, Conference on Neural Information Processing Systems, Dec. 7-12, 2015, Montreal, Canada, 9 pages. [cited by applicant]
Rastegari et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks”, arXiv:1603.05279v4, Aug. 2, 2016, 17 pages. [cited by applicant]
Shayer et al, “Learning Discrete Weights Using the Local Reparameterization Trick”, arXiv:1710.07739v3, Feb. 2, 2018, 12 pages. [cited by applicant]
Sivic et al., “Video Google: a Text Retrieval Approach to Object Matching in Videos”, International Conference on Computer Vision, Oct. 14-17, 2003, Nice, France, 8 pages. [cited by applicant]
Son et al., “Clustering Convolutional Kernels to Compress Deep Neural Networks”, European Conference on Computer Vision, Sep. 8-14, 2018, Munich, Germany, 17 pages. [cited by applicant]
Stock et al., “And the Bit Goes Down: Revisiting the Quantization of Neural Networks”, arXiv:1907.05686v5, Nov. 9, 2020, 11 pages. [cited by applicant]
Thomee et al., “YFCC100M: The New Data in Multimedia Research”, arXiv:1503.01817v2, Apr. 25, 2016, 8 pages. [cited by applicant]
Tieleman et al, “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude”, COURSERA: Neural networks for machine learning, 4(2):26-31, 2012. [cited by applicant]
Tung et al., “Deep Neural Network Compression by In-Parallel Pruning-Quantization”, Pattern Analysis and Machine Intelligence, Dec. 2018, 12 pages. [cited by applicant]
Wang et al, “HAQ: Hardware-Aware Automated Quantization with Mixed Precision”, arXiv:1811.08886v3, Apr. 6, 2019, 10 pages. [cited by applicant]
Wu et al., “Quantized Convolutional Neural Networks for Mobile Devices”, arXiv:1512.06473v3, May 16, 2016, 11 pages. [cited by applicant]
Yalniz et al, “Billion-Scale Semi-Supervised Learning for Image Classification”, arXiv:1905.00546v1, May 2, 2019, 12 pages. [cited by applicant]
Zeger et al., “Globally Optimal Vector Quantizer Design by Stochastic Relaxation”, IEEE Transaction on Signal Processing, vol. 40, No. 2, Feb. 1992, pp. 310-322. [cited by applicant]
Zeger et al., “Stochastic Relaxation Algorithm for Improved Vector Quantiser Design”, Electronics Letters, vol. 25, No. 14, Jul. 6, 1989, pp. 896-898. [cited by applicant]
Zhu et al., “Trained Ternary Quantization”, arXiv:1612.01064v3, Feb. 23, 2017, 10 pages. [cited by applicant]
Cited By (1)
US 12,718,107