IP Library › Granted Patent US 12,248,877
Granted Patent B2
US 12,248,877 · App. 16/226,521 · Granted Mar 11, 2025

Hybrid neural network pruning

Inventors: Xiaofan Xu (Dublin, IE); Mi Sun Park (San Jose, CA); Cormac M. Brick (San Francisco, CA)
Assignee: Movidius Ltd.
G06N3/082G06N3/04G06N3/044G06T7/70G06N20/00G06T2207/10024G06T2207/10028G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,877
App. No.
16/226,521
Granted
Mar 11, 2025
Kind
B2
Abstract

A pruned version of a neural network is generated by determining pruned versions of each a plurality of layers of the network. The pruned version of each layer is determined by sorting a set of channels of the layer based on respective weight values of each channel in the set. A percentage of the set of channels are pruned based on the sorting to form a thinned version of the layer. Accuracy of a thinned version of the neural network is tested, where the thinned version of the neural network includes the thinned version of the layer. The thinned version of the layer is used to generate the pruned version of the layer based on the accuracy of the thinned version of the neural network exceeding a threshold accuracy value. A pruned version of the neural network is generated to include the pruned versions of the plurality of layers.

Claims (67)

1. At least one non-transitory machine accessible storage medium having instructions stored thereon, wherein the instructions when executed on a machine, cause the machine to prune a neural network by:

accessing data comprising a definition of the neural network, wherein the neural network comprises a layer having a plurality of channels;

generating a thinned version of the layer, wherein generating the thinned version of the layer comprises:

ranking the plurality of channels of the layer based on weight values of each of the plurality of channels,

selecting one or more channels from the plurality of channels based on the ranking,

pruning the layer by setting values of all weights in the selected one or more channels to zeros to form the thinned version of the layer, and

rounding a number of unpruned channels to a multiple corresponding to a hardware architecture;

providing input data to a thinned version of the neural network that includes the thinned version of the layer;

determining, based on an output of the neural network, that an accuracy of the thinned version of the neural network exceeds a threshold accuracy;

after determining that the accuracy of the thinned version of the neural network exceeds the threshold accuracy, further pruning the thinned version of the layer by modifying a weight in another channel of the layer to zero and keeping another weight in the another channel of the layer unmodified to form a further thinned version of the layer; and

generating a pruned version of the neural network that comprises the further thinned version of the layer.

2. The storage medium of claim 1 , wherein further pruning the thinned version of the layer comprises:

determining a descriptive statistic value from weights of the thinned version of the layer;

determining a weight threshold for the thinned version of the layer based on the descriptive statistic value; and

selecting the weight in the another channel based on a determination that the weight in the another channel is below the weight threshold,

wherein the another weight in the another channel is greater than or equal to the weight threshold.

3. The storage medium of claim 2 , wherein the weight threshold determined for the thinned version of the layer is different from a weight threshold determined for another layer in the neural network, and the neural network is pruned further by:

selecting a weight from weights of the another layer based on the weight threshold determined for the another layer; and

pruning the another layer by modifying the selected weight to zero to form a thinned version of the another layer,

wherein the pruned version of the neural network comprises the thinned version of the another layer.

4. The storage medium of claim 2 , wherein the descriptive statistic value comprises a mean of absolute values of the weights of the thinned version of the layer or a standard deviation of the absolute values of the weights of the thinned version of the layer.

5. The storage medium of claim 1 , wherein the neural network comprises an additional layer, and the additional layer is unpruned in the pruned version of the neural network.

6. The storage medium of claim 1 , wherein the neural network is pre-trained by using a particular data set, and the input data corresponds to the particular data set.

7. The storage medium of claim 1 , wherein the output of the neural network is generated through forward propagation of the input data through the thinned version of the neural network.

8. The storage medium of claim 1 , wherein the thinned version of the layer comprises a first iteration of the thinned version of the layer, the thinned version of the neural network comprises a first iteration of the thinned version of the neural network, and generating the thinned version of the layer further comprises:

determining that an accuracy of the first iteration of the thinned version of the neural network exceeds the threshold accuracy;

pruning one or more additional channels in the first iteration of the thinned version of the layer to form a second iteration of the thinned version of the layer; and

determining an accuracy of a second iteration of the thinned neural network, wherein the second iteration of the thinned neural network comprises the second iteration of the thinned layer.

9. The storage medium of claim 8 , wherein generating the thinned version of the layer further comprises:

determining that the accuracy of the second iteration of the thinned version of the neural network falls below the threshold accuracy; and

adopting the first iteration of the thinned version of the layer as the thinned version of the layer to be used to generate the pruned version of the neural network.

10. The storage medium of claim 1 , wherein the neural network comprises a convolutional neural network, and the layer is a hidden layer of the convolutional neural network.

11. The storage medium of claim 1 , wherein the hardware architecture comprises hardware architecture of a resource constrained computing device.

12. A system comprising:

a data processing apparatus; and

a non-transitory computer-readable memory storing computer program instructions executable by the data processing apparatus to perform operations for pruning a neural network by:

accessing data comprising a definition of a neural network, wherein the neural network comprises a layer having a plurality of channels,

generating a thinned version of the layer, wherein generating the thinned version of the layer comprises:

ranking the plurality of channels of the layer based on weight values of each of the plurality of channels,

selecting one or more channels from the plurality of channels based on the ranking,

pruning the layer by setting values of all weights in the selected one or more channels to zeros to form the thinned version of the layer, and

rounding a number of unpruned channels to a multiple corresponding to a hardware architecture,

providing input data to a thinned version of the neural network that includes the thinned version of the layer,

determining, based on an output of the neural network, that an accuracy of the thinned version of the neural network exceeds a threshold accuracy,

after determining that the accuracy of the thinned version of the neural network exceeds the threshold accuracy, further pruning the thinned version of the layer by modifying a weight in another channel of the layer to zero and keeping another weight in the another channel unmodified to form a further thinned version of the layer, and

generating a pruned version of the neural network that comprises the further thinned version of the layer.

13. The system of claim 12 , further comprising an interface to provide the pruned version of the neural network to a computing device.

14. The system of claim 13 , wherein the computing device comprises a resource-constrained computing device.

15. A method for pruning a neural network, the method comprising:

accessing data comprising a definition of a neural network, wherein the neural network comprises a layer having a plurality of channels;

generating a thinned version of the layer by:

ranking the plurality of channels of the layer based on weight values of each of the plurality of channels,

selecting one or more channels from the plurality of channels based on the ranking,

pruning the layer by setting values of all weights in the selected one or more channels to zeros to form the thinned version of the layer, and

rounding a number of unpruned channels to a multiple corresponding to a hardware architecture;

providing input data to a thinned version of the neural network that includes the thinned version of the layer;

determining, based on an output of the neural network, that an accuracy of the thinned version of the neural network exceeds a threshold accuracy;

after determining that the accuracy of the thinned version of the neural network exceeds the threshold accuracy, further pruning the thinned version of the layer by modifying a weight in another channel of the layer to zero and keeping another weight in the another channel unmodified to form a further thinned version of the layer; and

generating a pruned version of the neural network that comprises the further thinned version of the layer.

16. The method of claim 15 , wherein further pruning the thinned version of the layer comprises:

determining a weight threshold for the thinned version of the layer based on absolute values of weights in the thinned version of the layer; and

selecting the one or more weights in the another channel based on a determination that the one or more weights are below the weight threshold.

17. The method of claim 16 , wherein the weight threshold determined for the thinned version of the layer is different from a weight threshold determined for another layer in the neural network, and the method further comprises:

selecting a weight from a plurality of weights of the another layer based on the weight threshold determined for the another layer; and

pruning the another layer by modifying the selected weight to zero to form a thinned version of the another layer,

wherein the pruned version of the neural network comprises the thinned version of the another layer.

18. The method of claim 15 , wherein the neural network comprises one or more other layers in addition to the layer and the one or more other layers are unpruned in the pruned version of the neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2019
From: XU, XIAOFAN; PARK, MI SUN; BRICK, CORMAC M.
To: MOVIDIUS LTD.
Reel/Frame 050834/0834 →
Continuity (2)
Provisional Application 62675601 · May 23, 2018
Related Publication 20190362235A1 · Nov 28, 2019
References Cited (35)
US 10832136B2 · Kadav · 2020 [cited by examiner]
US 20160328643A1 · Liu · 2016 [cited by examiner]
US 20180365794A1 · Lee · 2018 [cited by examiner]
US 20190087729A1 · Byun · 2019 [cited by examiner]
US 20200234130A1 · Yan · 2020 [cited by examiner]
US 20200394520A1 · Kruglov · 2020 [cited by examiner]
WO 2019226686A2 · 2019 [cited by applicant]
Polyak, Adam, and Lior Wolf. “Channel-level acceleration of deep face representations.” IEEE Access 3 (2015): 2163-2175. (Year: 2015). [cited by examiner]
Li, Hao, et al. “Pruning filters for efficient convnets.” arXiv preprint arXiv:1608.08710 (2016). (Year: 2016). [cited by examiner]
Babaeizadeh, Mohammad, Paris Smaragdis, and Roy H. Campbell. “Noiseout: A simple way to prune neural networks.” arXiv preprint arXiv:1611.06211 (2016). (Year: 2016). [cited by examiner]
Guo, Yiwen, Anbang Yao, and Yurong Chen. “Dynamic network surgery for efficient dnns.” Advances in neural information processing systems 29 (2016). (Year: 2016). [cited by examiner]
Herv's, C., et al. “Optimization of computational neural network for its application in the prediction of microbial growth in foods.” Food science and technology international 7.2 (2001): 159-163. (Year: 2001). [cited by examiner]
Zhao, Wei, et al. “An FPGA implementation of a convolutional auto-encoder.” Applied Sciences 8.4 (2018): 504. (Year: 2018). [cited by examiner]
Liu, Zhuang, et al. “Learning efficient convolutional networks through network slimming.” Proceedings of the IEEE international conference on computer vision. 2017. (Year: 2017). [cited by examiner]
Collobert, Ronan, Koray Kavukcuoglu, and Clement Farabet. “Torch7: A matlab-like environment for machine learning.” BigLearn, NIPS workshop. No. CONF. 2011. (Year: 2011). [cited by examiner]
Gupta, Suyog, et al. “Deep learning with limited numerical precision.” International conference on machine learning. PMLR, 2015. (Year: 2015). [cited by examiner]
Anwar, Sajid, and Wonyong Sung. “Coarse pruning of convolutional neural networks with random masks.” (2016). (Year: 2016). [cited by examiner]
He, Yihui, Xiangyu Zhang, and Jian Sun. “Channel pruning for accelerating very deep neural networks.” Proceedings of the IEEE international conference on computer vision. 2017. (Year: 2017). [cited by examiner]
Chin, Ting-Wu, et al., “Layer-Compensated Pruning for Resource-Constrained Convolutional Neural Networks,” arXiv preprint, Oct. 18, 2012; 12 pages. [cited by applicant]
Guo, Yiwen, et al., “Dynamic Network Surgery for Efficient DNNs,” 30th Conference on Neural Information Processing Systems (NIPS); Barcelona, Spain, 2016; 9 pages. [cited by applicant]
Han, Song, et al.; “Learning both Weights and Connections for Efficient Neural Networks,” retrieved online at https://arxiv.org/abs/1506.02626; Oct. 30, 2015; 9 pages. [cited by applicant]
He, Kaiming, et al.; “Deep Residual Learning for Image Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016; 9 pages. [cited by applicant]
He, Yang, et al.; “Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks,” Proceedings of the 27th International Joint Conference on Artificial Intelligence, 2018; 7 pages. [cited by applicant]
He, Yihui, et al.; “AMC: AutoML for Model Compression and Acceleration on Mobile Devices,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018; 17 pages. [cited by applicant]
He, Yihui, et al.; “Channel Pruning for Accelerating Very Deep Neural Networks,” in The IEEE International Conference on Computer Vision (ICCV); Aug. 2017; 10 pages. [cited by applicant]
Huang, Gao, et al.; “Densely Connected Convolutional Networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017; 9 pages. [cited by applicant]
Li, Hao, et al.; “Pruning Filters for Efficient ConvNets,” published in the International Conference on Learning Representations (ICLR), 2017; 13 pages. [cited by applicant]
Lin, Ji, et al.; “Runtime Neural Pruning,” in the 31st Conference on Neural Information Processing Systems (NIPS), Long Beach, California; 2017; 11 pages. [cited by applicant]
Liu, Zhuang, et al.; “Learning Efficient Convolutional Networks through Network Slimming,” published in the International Conference on Computer Vision (ICCV), 2017; 9 pages. [cited by applicant]
Luo, Jian-Hao, et al.; “ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression,” retrieved online at https://arxiv.org/abs/1707.06342; Jul. 20, 2017; 9 pages. [cited by applicant]
Molchanov, Pavlo, et al.; “Pruning Convolutional Neural Networks for Resource Efficient Inference,” published in the International Conference on Learning Representations (ICLR), Jun. 2017; 17 pages. [cited by applicant]
Simonyan, Karen, et al.; “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in the International Conference on Learning Representations (ICLR), Apr. 2015; 14 pages. [cited by applicant]
Szegedy, Christian, et al.; “Going Deeper with Convolutions,” retrieved online at https://arxiv.org/abs/1409.4842; Sep. 17, 2014; 12 pages. [cited by applicant]
Wang, Huan, et al.; “Structured Probabilistic Pruning for Convolutional Neural Network Acceleration,” retrieved online at https://arxiv.org/abs/1709.06994; Sep. 10, 2018; 13 pages. [cited by applicant]
Yu, Ruichi, et al.; “NISP: Pruning Networks using Neuron Importance Score Propagation,” retrieved online at https://arxiv.org/abs/1711.05908; Mar. 21, 2018; 15 pages. [cited by applicant]
Cited By (2)
US 12,393,842 US 12,670,395