IP Library Granted Patent US 12,387,101
Granted Patent B2
US 12,387,101 · App. 17/317,300 · Granted Aug 12, 2025

Systems and methods for pruning binary neural networks guided by weight flipping frequency

Inventors: Yixing Li (Tempe, AZ); Fengbo Ren (Tempe, AZ)
Assignee: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
G06N3/082G06N3/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,101
App. No.
17/317,300
Granted
Aug 12, 2025
Kind
B2
Abstract

Various embodiments of a system and method for pruning binary neural networks by analyzing weight flipping frequency and pruning the binary neural network based on the weight flipping frequency associated with each channel of the binary neural network are disclosed herein.

Claims (27)

1. A method, comprising:

providing a computer-implemented binary neural network, the neural network including a plurality of layers, each layer of the plurality of layers including a respective plurality of channels, each channel associated with a respective weight, wherein each weight of the plurality of weights is a binary value;

logging a flip count for each weight of the plurality of weights of the neural network during a pruning interval of a neural network training session, wherein a flip count for a respective weight comprises a count of a number of flips of the respective weight;

determining a portion of weights for each layer of the neural network that have a flip count greater than a predetermined value during the pruning interval; and

removing a portion of channels of the neural network corresponding to the portion of weights that have a flip count greater than the predetermined value during the pruning interval to generate a reduced neural network; and

training the reduced neural network.

2. The method of claim 1 , wherein the method is iteratively repeated until a maximum of the portion of insensitive weights is minimized.

3. The method of claim 1 , wherein the portion of channels which are removed from the neural network is proportional to the portion of insensitive weights.

4. The method of claim 3 , wherein a number of channels in an l.sup.th layer of the neural network is reduced by the portion of insensitive weights for the l.sup.th layer.

5. The method of claim 1 , wherein the pruning interval is defined as an interval during the training session that contributes to a final amount of accuracy gain towards an end of the training session.

6. The method of claim 1 , wherein the flip count for a weight of the plurality of weights is a number of times within the pruning interval of the training session that a value of the weight is altered.

7. The method of claim 1 , wherein the predetermined value is empirically selected.

8. The method of claim 1 , wherein the predetermined value is 2.

9. A computer-implemented neural network pruning system, comprising:

a computer-implemented binary neural network, the neural network including a plurality of layers, each layer of the plurality of layers including a respective plurality of channels, each channel associated with a respective weight, wherein each weight of the plurality of weights is a binary value;

a processor in communication with a tangible storage medium storing instructions that are executed by the processor to perform operations comprising:

log a flip count for each weight of the plurality of weights of the neural network during a pruning interval of a neural network training session, wherein a flip count for a respective weight comprises a count of a number of flips of the respective weight;

determine a portion of weights for each layer of the neural network that have a flip count greater than a predetermined value during the pruning interval; and

remove a portion of channels of the neural network corresponding to the portion of weights that have a flip count greater than the predetermined value during the pruning interval to generate a reduced neural network; and

train the reduced neural network.

10. The pruning system of claim 9 , wherein the operations are iteratively repeated until a maximum of the portion of insensitive weights is minimized.

11. The pruning system of claim 9 , wherein the portion of channels which are removed from the neural network is proportional to the portion of insensitive weights.

12. The pruning system of claim 11 , wherein a number of channels in an l.sup.th layer of the neural network is reduced by the portion of insensitive weights for the l.sup.th layer.

13. The pruning system of claim 9 , wherein the pruning interval is defined as an interval during the training session that contributes to a final amount of accuracy gain towards an end of the training session.

14. The pruning system of claim 9 , wherein the flip count for a weight of the plurality of weights is a number of times within the pruning interval of the training session that a value of the weight is altered.

15. The pruning system of claim 9 , wherein the predetermined value is empirically selected.

16. The method of claim 9 , wherein the predetermined value is 2.

Assignments (2)
CONFIRMATORY LICENSE Recorded Feb 25, 2022
From: ARIZONA STATE UNIVERSITY-TEMPE CAMPUS
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 059203/0512 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2021
From: LI, YIXING; REN, FENGBO
To: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
Reel/Frame 056244/0091 →
Continuity (2)
Provisional Application 63022913 · May 11, 2020
Related Publication 20210350242A1 · Nov 11, 2021
References Cited (21)
WO WO2020150678A1 · 2020 [cited by examiner]
Helwegen et al, “Latent Weights Do Not Exist: Rethinking Binarized Neural Network Optimization”, Jun. 5, 2019, arXiv, pp. 1-12 (Year: 2019). [cited by examiner]
Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. Tech. rep. Citeseer, 2009. [cited by applicant]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. “Imagenet classification with deep convolutional neural networks”. In: Advances in neural information processing systems. 2012, pp. 1097-1105. [cited by applicant]
Dongqing Zhang et al. “Lq-nets: Learned quantization for highly accurate and compact deep neural networks”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 365-382. [cited by applicant]
Itay Hubara et al. “Binarized neural networks”. In: Advances in neural information processing systems. 2016, pp. 4107-4115. [cited by applicant]
Jia Deng et al. “Imagenet: A large-scale hierarchical image database”. In: 2009 IEEE conference on computer vision and pattern recognition. Ieee. 2009, pp. 248-255. [cited by applicant]
Jian-Hao Luo, Jianxin Wu, and Weiyao Lin. “Thinet: A filter level pruning method for deep neural network compression”. In: Proceedings of the IEEE international conference on computer vision. 2017, pp. 5058-5066. [cited by applicant]
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. “Binaryconnect: Training deep neural networks with binary weights during propagations”. In: Advances in neural information processing systems. 2015, pp. 3123-3… [cited by applicant]
Min Lin, Qiang Chen, and Shuicheng Yan. “Network in network”. In: arXiv preprint arXiv:1312.4400 (2013). [cited by applicant]
Mohammad Rastegari et al. “Xnor-net: Imagenet classification using binary convolutional neural networks”. In: European Conference on Computer Vision. Springer. 2016, pp. 525-542. [cited by applicant]
Paulius Micikevicius et al. “Mixed precision training”. In: International Conference on Learning Representations (2018). [cited by applicant]
Pavlo Molchanov et al. “Importance estimation for neural network pruning”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2019, pp. 11264-11272. [cited by applicant]
Raghuraman Krishnamoorthi. “Quantizing deep convolutional networks for efficient inference: A whitepaper”. In: arXiv preprint arXiv:1806.08342 (2018). [cited by applicant]
Ruichi Yu et al. “Nisp: Pruning networks using neuron importance score propagation”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018, pp. 9194-9203. [cited by applicant]
Saman Biookaghazadeh, Ming Zhao, and Fengbo Ren. “Are FPGAs suitable for edge computing?” In: {USENIX} Workshop on Hot Topics in Edge Computing (HotEdge 18). 2018. [cited by applicant]
Shuang Liang et al. “FP-BNN: Binarized neural network on FPGA”. In: Neurocomputing 275 (2018), pp. 1072-1086. [cited by applicant]
Shuchang Zhou et al. “Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients”. In: arXiv preprint arXiv:1606.06160 (2016). [cited by applicant]
Song Han, Huizi Mao, and William J Dally. “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding”. In: International Conference on Learning Representations (2016). [cited by applicant]
Yuwei Hu et al. “Bitflow: Exploiting vector parallelism for binary neural networks on cpu”. In: 2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE. 2018, pp. 244-253. [cited by applicant]
Zhuang Liu et al. “Rethinking the value of network pruning”. In: arXiv preprint arXiv:1810.05270 (2018). [cited by applicant]