IP Library › Granted Patent US 12,362,764
Granted Patent B2
US 12,362,764 · App. 17/073,602 · Granted Jul 15, 2025

Neural network model compression with quantizability regularization

Inventors: Wei Jiang (San Jose, CA); Wei Wang (Palo Alto, CA); Shan Liu (San Jose, CA)
Assignee: TENCENT AMERICA LLC
H03M7/3059G06F18/214G06N3/04G06N3/063G06N3/084G06V10/771H03M7/702
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,362,764
App. No.
17/073,602
Granted
Jul 15, 2025
Kind
B2
Abstract

A method, computer program, and computer system is provided for compressing a neural network model. A multi-dimensional tensor corresponding to a set of weight coefficients associated with a neural network is reshaped. A subset of weight coefficients is identified from among the set of weight coefficients. A model of the neural network is compressed based on the identified subset of weight coefficients.

Claims (42)

1. A method for compressing a neural network model for deployment on a terminal, executable by a processor, comprising:

reshaping, for a layer in a deep neural network model, a weight tensor having a first dimension into a reshaped weight tensor having a second dimension, the second dimension being less than the first dimension,

wherein a size of the reshaped weight tensor is based on a number of input channels, a number of output channels, and an axis along which the weight tensor is reshaped;

partitioning, for the layer in the deep neural network model, the reshaped weight tensor into one or more blocks;

averaging, for the layer in the deep neural network model, weights within respective blocks of the one or more blocks;

ranking, for the layer in the deep neural network model, the one or more blocks of the reshaped weight tensor based on a loss associated with the respective blocks;

fixing, for the layer in the deep neural network model, the averaged weights within respective blocks of the one or more blocks for a predetermined number of ranked blocks and setting a respective item corresponding to a respective block in a quantization mask as a fixed value based on the average weight of the respective block;

training the deep neural network model based on updating un-fixed weights associated with a remaining number of ranked blocks;

compressing the deep neural network model, for each layer in the deep neural network model, based on the averaged weights for respective layers of the neural network.

2. The method of claim 1 , wherein updating the un-fixed weights is based on a gradient and a quantization mask associated with the averaged weights.

3. The method of claim 1 , wherein each layer of the deep neural network model is compressed separately.

4. A computer system for compressing a neural network model for deployment on a terminal, the computer system comprising:

one or more computer-readable non-transitory storage media configured to store computer program code; and

one or more computer processors configured to access said computer program code and operate as instructed by said computer program code, said computer program code including:

reshaping code configured to cause the one or more computer processors to reshape, for a layer in a deep neural network model, a weight tensor having a first dimension into a reshaped weight tensor having a second dimension, the second dimension being less than the first dimension,

wherein a size of the reshaped weight tensor is based on a number of input channels, a number of output channels, and an axis along which the weight tensor is reshaped;

partitioning code configured to cause the one or more computer processors to partition for the layer in the deep neural network model, the reshaped weight tensor into one or more blocks;

unifying code configured to cause the one or more computer processors to average, for the layer in the deep neural network model, weights within respective blocks of the one or more blocks;

ranking code configured to cause the one or more computer processors to rank, for the layer in the deep neural network model, the one or more blocks of the reshaped weight tensor based on a loss associated with the respective blocks;

fixing code configured to cause the one or more computer processors to fix, for the layer in the deep neural network model, the averaged weights within respective blocks of the one or more blocks for a predetermined number of ranked blocks and setting a respective item corresponding to a respective block in a quantization mask as a fixed value based on the average weight of the respective block;

training code configured to cause the one or more computer processors to train the deep neural network model based on updating un-fixed weights associated with a remaining number of ranked blocks; and

compressing code configured to cause the one or more computer processors to compress the deep neural network model, for each layer in the deep neural network model, based on the averaged weights for respective layers of the neural network.

5. The computer system of claim 4 , wherein the unifying code comprises:

quantizing code configured to cause the one or more computer processors to quantize weight coefficients; and

selecting code configured to cause the one or more computer processors to determine an average or a mean of weight coefficients based on minimizing a quantizability regularization loss value corresponding to a data loss value and a quantization loss value associated with the quantized weight coefficients.

6. The computer system of claim 5 , wherein the training code is further configured to cause the one or more computer processors to train the deep neural network model based on back-propagating the minimized quantizability regularization loss value.

7. The computer system of claim 4 , wherein updating the un-fixed weights is based on a gradient and a quantization mask associated with the averaged weights.

8. The computer system of claim 4 , wherein each layer of the deep neural network model is compressed separately.

9. A non-transitory computer readable medium having stored thereon a computer program for compressing a neural network model for deployment on a terminal, the computer program configured to cause one or more computer processors to:

reshape, for a layer in a deep neural network model, a weight tensor having a first dimension into a reshaped weight tensor having a second dimension, the second dimension being less than the first dimension,

wherein a size of the reshaped weight tensor is based on a number of input channels, a number of output channels, and an axis along which the weight tensor is reshaped;

partition, for the layer in the deep neural network model, the reshaped weight tensor into one or more blocks;

average, for the layer in the deep neural network model, weights within respective blocks of the one or more blocks;

rank, for the layer in the deep neural network model, the one or more blocks of the reshaped weight tensor based on a loss associated with the respective blocks;

fix, for the layer in the deep neural network model, the averaged weights within respective blocks of the one or more blocks for a predetermined number of ranked blocks and setting a respective item corresponding to a respective block in a quantization mask as a fixed value based on the average weight of the respective block;

train the deep neural network model based on updating un-fixed weights associated with a remaining number of ranked blocks;

compress the deep neural network model, for each layer in the deep neural network model, based on the averaged weights for respective layers of the neural network.

10. The computer readable medium of claim 9 , wherein the computer program is further configured to cause one or more computer processors average weights based on:

quantizing weight coefficients; and

determine an average or a mean of weight coefficients based on minimizing a quantizability regularization loss value corresponding to a data loss value and a quantization loss value associated with the quantized weight coefficients.

11. The computer readable medium of claim 10 , wherein training further comprises training the deep neural network model based on back-propagating the minimized quantizability regularization loss value.

12. The computer readable medium of claim 9 , wherein updating the un-fixed weights is based on a gradient and a quantization mask associated with the averaged weights.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2025
From: JIANG, WEI; WANG, WEI; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 070802/0083 →
Continuity (2)
Provisional Application 62954472 · Dec 28, 2019
Related Publication 20210201157A1 · Jul 1, 2021
References Cited (13)
US 10594338B1 · Lew · 2020 [cited by examiner]
US 20080207222A1 · Bhattacharya · 2008 [cited by examiner]
US 20190042743A1 · Chen · 2019 [cited by examiner]
US 20190294413A1 · Vantrease · 2019 [cited by examiner]
US 20190340499A1 · Burger · 2019 [cited by examiner]
US 20200151573A1 · Das · 2020 [cited by examiner]
US 20200380357A1 · Yao · 2020 [cited by examiner]
Chen, Shangyu, Wenya Wang, and Sinno Jialin Pan. “Deep neural network quantization via layer-wise optimization using limited training data.” Proceedings of the AAAI Conference on Artificial Intelligence. vol. 33. No. 01… [cited by examiner]
Farabet, Clément. Towards real-time image understanding with convolutional networks. Diss. Paris Est, 2013. (Year: 2013). [cited by examiner]
Carreira-Perpinán, Miguel A. “Model compression as constrained optimization, with application to neural nets. Part I: General framework.” arXiv preprint arXiv:1707.01209 (2017). (Year: 2017). [cited by examiner]
Shi, Zenglin, Yangdong Ye, and Yunpeng Wu. “Rank-based pooling for deep convolutional neural networks.” Neural Networks 83 (2016): 21-31. (Year: 2016). [cited by examiner]
Jianbo Ye et al., “Rethinking the Smaller-Norm-Less-Informative Assumption in Channel Pruning of Convolution Layers”, Published as a conference paper at ICLR 2018, pp. 1-11. [cited by applicant]
Song Han et al., “Learning both Weights and Connections for Efficient Neural Networks”, pp. 1-9. [cited by applicant]