IP Library › Granted Patent US 12,288,163
Granted Patent B2
US 12,288,163 · App. 17/031,817 · Granted Apr 29, 2025

Training method for quantizing the weights and inputs of a neural network

Inventors: Vahid Partovi Nia (Montreal, CA); Ryan Razani (Toronto, CA)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06N3/084G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,288,163
App. No.
17/031,817
Granted
Apr 29, 2025
Kind
B2
Abstract

A method and processing unit for training a neural network to selectively quantize weights of a filter of the neural network as either binary weights or ternary weights. A plurality of training iterations a performed that each comprise: quantizing a set of real-valued weights of a filter to generate a corresponding set of quantized weights; generating an output feature tensor based on matrix multiplication of an input feature tensor and the set of quantized weights; computing, based on the output feature tensor, a loss based on a regularization function that is configured to move the loss towards a minimum value when either: (i) the quantized weights move towards binary weights, or (ii) the quantized weights move towards a ternary weights; computing a gradient with an objective of minimizing the loss; updating the real-valued weights based on the computed gradient. When the training iterations are complete, a set of weights quantized from the updated real-valued weights is stored as either a set of binary weights or a set of ternary weights.

Claims (41)

1. A computer-implemented method for training a convolutional neural network (CNN) to selectively quantize weights of a filter of the CNN as either binary weights or ternary weights, the method comprising:

performing a plurality of training iterations for training the CNN that each comprise:

quantizing a set of real-valued weights of the filter to generate a corresponding set of quantized weights;

generating an output feature tensor based on matrix multiplication of an input feature tensor and the set of quantized weights;

computing, based on the output feature tensor, a loss based on a regularization function that is configured to move the loss towards a minimum value when either: (i) the quantized weights move towards binary weights, or (ii) the quantized weights move towards ternary weights, wherein the regularization function includes a learnable shape parameter, wherein changing a magnitude of the shape parameter in one direction causes the regularization function to approximate a binary regularization function and changing the magnitude of the shape parameter in an opposite direction causes the regularization function to approximate a ternary regularization function;

computing a gradient with an objective of minimizing the loss; and

backpropagating the loss through the CNN to update values of the real-valued weights based on the computed gradient;

when the training iterations are complete, storing a final set of weights quantized from the updated real-valued weights as either a set of binary weights or a set of ternary weights, the final set of weights representing a learned set of quantized weights; and

deploying the CNN model including the final set of weights to a computationally constrained hardware device, to cause the device to perform forward inference using the CNN model with respect to input data, the CNN model including the final set of weights representing a compressed CNN.

2. The method of claim 1 wherein the loss is further based on a difference between one or more values predicted by the CNN with respect to an original input feature tensor from a training set and corresponding one or more true values known for the original input feature tensor.

3. The method of claim 1 comprising sampling an initial set of real-valued weights from a bimodal distribution to use as the set of real-valued weights for a first iteration of the plurality of iterations.

4. The method of claim 1 wherein the matrix multiplication is part of a convolution operation, and the method comprises training a plurality of filters.

5. The method of claim 1 wherein the input feature tensor is a binarized input feature tensor.

6. The method of claim 5 comprising binarizing a real-valued input feature tensor to provide the binarized input feature tensor.

7. The method of claim 1 wherein generating the output feature tensor comprises applying an activation function that binarizes an output provided by the matrix multiplication of the input feature tensor and the set of quantized weights.

8. The method of claim 1 wherein each element in a set of binary weights has a value of either −1 or +1, and each element in a set of ternary weights has a value of either −1, or 0, or +1.

9. A processing unit for training a convolutional neural network (CNN) to selectively quantize weights of a filter of the CNN as either binary weights or ternary weights, the processing unit comprising a processor device and a persistent storage coupled to the processor device storing instructions that when executed by the processor device cause the processing unit to:

perform a plurality of training iterations for training the CNN that each comprise:

quantizing a set of real-valued weights of the filter to generate a corresponding set of quantized weights;

generating an output feature tensor based on matrix multiplication of an input feature tensor and the set of quantized weights;

computing, based on the output feature tensor, a loss based on a regularization function that is configured to move the loss towards a minimum value when either: (i) the quantized weights move towards binary weights, or (ii) the quantized weights move towards ternary weights, wherein the regularization function includes a learnable shape parameter, wherein changing a magnitude of the shape parameter in one direction causes the regularization function to approximate a binary regularization function and changing the magnitude of the shape parameter in an opposite direction causes the regularization function to approximate a ternary regularization function;

computing a gradient with an objective of minimizing the loss; and

backpropagating the loss through the CNN to update values of the real-valued weights based on the computed gradient;

when the training iterations are complete, store a final set of weights quantized from the updated real-valued weights as either a set of binary weights or a set of ternary weights, the final set of weights representing a learned set of quantized weights; and

deploy the CNN model including the final set of weights to a computationally constrained hardware device, to cause the device to perform forward inference using the CNN model with respect to input data, the CNN model including the final set of weights representing a compressed CNN.

10. The processing unit of claim 9 wherein the loss is further based on a difference between one or more values predicted by the CNN with respect to an original input feature tensor from a training set and corresponding one or more true values known for the original input feature tensor.

11. The processing unit of claim 9 wherein the processing unit is caused to sample an initial set of real-valued weights from a bimodal distribution to use as the set of real-valued weights for a first iteration of the plurality of iterations.

12. The processing unit of claim 9 wherein the matrix multiplication is part of a convolution operation, and the method comprises training a plurality of filters.

13. The processing unit of claim 9 wherein the input feature tensor is a binarized input feature tensor.

14. The processing unit of claim 13 wherein the processing unit is caused to binarize a real-valued input feature tensor to provide the binarized input feature tensor.

15. The processing unit of claim 9 wherein generating the output feature tensor comprises applying an activation function that binarizes an output provided by the matrix multiplication of the input feature tensor and the set of quantized weights.

16. The processing unit of claim 9 wherein each element in a set of binary weights has a value of either −1 or +1, and each element in a set of ternary weights has a value of either −1, or 0, or +1.

17. A non-transitory computer readable medium that persistently stores software instructions for training a convolutional neural network (CNN) to selectively quantize weights of a filter of the CNN as either binary weights or ternary weights, the software instructions including instructions for causing a processing unit to:

perform a plurality of training iterations for training the CNN that each comprise:

quantizing a set of real-valued weights of the filter to generate a corresponding set of quantized weights;

generating an output feature tensor based on matrix multiplication of an input feature tensor and the set of quantized weights;

computing, based on the output feature tensor, a loss based on a regularization function that is configured to move the loss towards a minimum value when either: (i) the quantized weights move towards binary weights, or (ii) the quantized weights move towards ternary weights, wherein the regularization function includes a learnable shape parameter, wherein changing a magnitude of the shape parameter in one direction causes the regularization function to approximate a binary regularization function and changing the magnitude of the shape parameter in an opposite direction causes the regularization function to approximate a ternary regularization function;

computing a gradient with an objective of minimizing the loss; and

backpropagating the loss through the CNN to update values of updating the real-valued weights based on the computed gradient;

when the training iterations are complete, store a final set of weights quantized from the updated real-valued weights as either a set of binary weights or a set of ternary weights, the final set of weights representing a learned set of quantized weights; and

deploy the CNN model including the final set of weights to a computationally constrained hardware device, to cause the device to perform forward inference using the CNN model with respect to input data, the CNN model including the final set of weights representing a compressed CNN.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2021
From: PARTOVI NIA, VAHID; RAZANI, RYAN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 054946/0374 →
Continuity (2)
Provisional Application 62905257 · Sep 24, 2019
Related Publication 20210089925A1 · Mar 25, 2021
References Cited (24)
US 20210295114A1 · Ye · 2021 [cited by examiner]
Ardakani et al., “Learning Recurrent Binary/Ternary Weights”, Jan. 24, 2019, arXiv:1809.11086v2, pp. 1-19. (Year: 2019). [cited by examiner]
Choi et al., “Learning Low Precision Deep Neural Networks through Regularization”, Sep. 1, 2018, arXiv:1809.00095v1, pp. 1-10. (Year: 2018). [cited by examiner]
Ian J. Goodfellow, Yoshua Bengio, and Aaron C. Courville. Deep Learning. Adaptive computation and machine learning. MIT Press 2016. [cited by applicant]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. Commun. ACM, 60(6) 2017. [cited by applicant]
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556 2014. [cited by applicant]
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Conference on Computer Vision and … [cited by applicant]
Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. CoRR, abs/1506.02640 2015. [cited by applicant]
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott E. Reed, Cheng-Yang Fu, and Alexander C. Berg. SSD: single shot multibox detector. CoRR, abs/1512.02325 2015. [cited by applicant]
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. Faster R-CNN: towards real-time object detection with region proposal networks. CoRR, abs/1506.01497 2015. [cited by applicant]
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Binaryconnect: Training deep neural networks with binary weights during propagations. In Advances in neural information processing systems (NIPS) 2015. [cited by applicant]
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Binarized neural networks. In Advances in neural information processing systems (NIPS) 2016. [cited by applicant]
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. Xnor-net: Imagenet classification using binary convolutional neural networks. CoRR, abs/1603.05279 2016. [cited by applicant]
Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou. Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients. CoRR, abs/1606.06160 2016. [cited by applicant]
Xiaofan Lin, Cong Zhao, and Wei Pan. Towards accurate binary convolutional neural network. In Advances in neural information processing systems (NIPS) 2017. [cited by applicant]
Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, and Yoshua Bengio. Neural networks with few multipli-cations. CoRR, abs/1510.03009 2015. [cited by applicant]
Fengfu Li and Bin Liu. Ternary weight networks. CoRR, abs/1605.04711 2016. [cited by applicant]
Chenzhuo Zhu, Song Han, Huizi Mao, and William J. Dally. Trained ternary quantization. CoRR, abs/1612.01064 2016. [cited by applicant]
Mouloud Belbahri, Eyyüb Sari, Sajad Darabi, and Vahid Partovi Nia. Foothill: A quasiconvex regularization for edge computing of deep neural networks. In Image Analysis and Recognition—16th International Conference (ICIA… [cited by applicant]
Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11) Nov. 1998. [cited by applicant]
Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Technical report, Citeseer 2009. [cited by applicant]
Forrest N. Iandola, Song Han , Matthew W. Moskewicz , Khalid Ashraf , William J. Dally, Kurt Keutzer; Squeezenet: Alexnet-Level Accuracy With 50X Fewer Parameters and <0.5MB Model Size, arXiv:1602.07360v4 Nov. 4, 2016. [cited by applicant]
Song Han, Huizi Mao, William J. Dally; Deep Compression: Compressing Deep Neural Networks With Pruning, Trained Quantization and Huffman Coding, arXiv:1510.00149v5 Feb. 15, 2016. [cited by applicant]
Morin et al:“Smart Ternary Quantization”,Sep. 25, 2019, total 8 pages. [cited by applicant]