IP Library Granted Patent US 12,373,696
Granted Patent B2
US 12,373,696 · App. 16/449,009 · Granted Jul 29, 2025

Neural network hardware accelerator system with zero-skipping and hierarchical structured pruning methods

Inventors: Shai Litvak (Beit Hashmonai, IL); Eyal Hochberg (Ramat Gan, IL)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06N3/082G06F7/5443G06F17/15G06N3/02G06N3/063G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,696
App. No.
16/449,009
Granted
Jul 29, 2025
Kind
B2
Abstract

Systems and methods for structured-pruning, zero-skipping and accelerated processing of an artificial neural network (ANN) are described. The ANN may include one or more convolution layers. 2D channels in filters of the convolution layers comprise fully pruned channels (FPCs), each containing only zero weights, and mixed channels (MCs), each containing at least one non-zero weight. At least a portion of the MCs satisfy a limited zero sequence (LZS) condition limiting the number and location of zeroes in the MC. The LZS condition may be based on a number of weights that a zero-skipping circuit of a computing system for processing the ANN is configured to evaluate and skip in a single cycle. Thus, when processing the structurally-pruned ANN using the zero-skipping method, the computing system may avoid processing zero weights. This may allow speeding up the ANN processing and reducing the power required for processing.

Claims (33)

1. A computing system for processing an artificial neural network (ANN), comprising:

a processor comprising a zero-skipping circuit configured to process a maximal number of weights in a single cycle and to locate non-zero weights;

a memory; and

an ANN stored within the memory, wherein the ANN comprises a plurality of layers including one or more convolution layers, wherein each of the one or more convolution layers comprises a plurality of filters, each filter comprises a plurality of channels, each channel comprises a plurality of rows, and each row comprises a plurality of weights,

wherein at least one non-zero weight that is greater than zero and below a pruning threshold is added to at least one of the plurality of channels to satisfy a limited zero sequence (LZS) condition that limits a number or location of sequential zero weights in the one or more plurality of channels based on the maximal number of weights the zero-skipping circuit is configured to process in the single cycle, and

wherein the at least one non-zero weight corresponds to at least one weight of the one or more plurality of channels that was pruned based on the pruning threshold.

2. The computing system of claim 1 , wherein:

the at least one non-zero weight is pruned from the one or more plurality of channels based on the pruning threshold and re-inserted to satisfy the LZS condition.

3. The computing system of claim 1 , wherein:

the LZS condition comprises a maximum number of rows over which zero sequences extend.

4. The computing system of claim 3 , wherein:

the maximum number of rows is two.

5. The computing system of claim 1 , wherein:

the zero-skipping circuit is configured to skip zero weights based on the maximal number of weights the zero-skipping circuit is configured to process in a single cycle.

6. The computing system of claim 1 , wherein:

the one or more convolution layers comprises at least 33% zero weights.

7. The computing system of claim 1 , wherein:

the one or more fully pruned channels (FPCs) comprise at least 20% of channels in the one or more convolution layers.

8. The computing system of claim 1 , wherein:

the portion of the one or more plurality of channels comprise at least 95% of the one or more plurality of channels.

9. The computing system of claim 1 , wherein:

the LZS condition is based at least in part on a scanning order of the computing system.

10. The computing system of claim 9 , wherein:

the LZS condition applies to sequences of consecutive zero weights over more than one channel according to the scanning order of the computing system.

11. The computing system of claim 1 , wherein:

the ANN further comprises one or more fully connected layers that satisfy the LZS condition.

12. The computing system of claim 1 , further comprising:

a control unit configured to identify and skip fully pruned channel (FPCs) of the ANN.

13. The computing system of claim 1 , further comprising:

the zero-skipping circuit is further configured to evaluate a fixed number of lines in a look-ahead buffer and copy a line of weights to a multiplexing buffer based on the evaluating.

14. The computing system of claim 1 , further comprising:

the maximum number of weights depends on a size of a buffer which is configured to store weights to be processed.

15. The computing system of claim 1 , wherein the plurality of channels are iteratively pruned and unpruned based on the pruning threshold and a target pruning rate.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2019
From: LITVAK, SHAI; HOCHBERG, EYAL
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 049554/0624 →
Continuity (1)
Related Publication 20200401895A1 · Dec 24, 2020
References Cited (34)
US 11200495B2 · Wang · 2021 [cited by examiner]
US 20160358069A1 · Brothers et al. · 2016 [cited by applicant]
US 20180046898A1 · Lo · 2018 [cited by applicant]
US 20180046916A1 · Dally · 2018 [cited by examiner]
US 20180089562A1 · Jin et al. · 2018 [cited by applicant]
US 20190108436A1 · David · 2019 [cited by examiner]
US 20190147327A1 · Martin · 2019 [cited by examiner]
US 20190303757A1 · Wang · 2019 [cited by examiner]
US 20200285950A1 · Baum · 2020 [cited by examiner]
US 20200293876A1 · Phan · 2020 [cited by examiner]
US 20210027166A1 · Gorokhov · 2021 [cited by examiner]
US 20210097393A1 · Wang · 2021 [cited by examiner]
KR 1020180034853A · 2018 [cited by applicant]
KR 101989793B1 · 2019 [cited by applicant]
He, Y. et al., “Channel Pruning for Accelerating Very Deep Neural Networks” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 1389-1397 (Year: 2017). [cited by examiner]
Molchanov, P. et al., “Pruning Convolutional Neural Networks for Resource Efficient Inference” (Year: 2017). [cited by examiner]
Parashar, A. et al., “SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks” (Year: 2017). [cited by examiner]
Aghasi, A. et al., “Net-Trim: Convex Pruning of Deep Neural Networks with Performance Guarantee” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. (Year: 2017). [cited by examiner]
Anwar, S. et al., “Structured Pruning of Deep Convolutional Neural Networks” ACMJournal on Emerging Technologies in Computing Systems, vol. 13, No. 3, Article 32, Publication date: Feb. 2017. (Year: 2017). [cited by examiner]
Li, H. et al., “Pruning Filters for Efficient ConvNets” (Year: 2017). [cited by examiner]
Gao, X. et al., “Dynamic Channel Pruning: Feature Boosting and Suppression” (Year: 2019). [cited by examiner]
Hu, H. et al., “Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures”, (Year: 2016). [cited by examiner]
Hegde, K. et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition”, https://arxiv.org/abs/1804.06508 (Year: 2018). [cited by examiner]
Kim, D. et al., “ZeNA: Zero-Aware Neural Network Accelerator”, https://ieeexplore.ieee.org/document/8013151 (Year: 2017). [cited by examiner]
Gao, C. et al. “DeltaRNN: A Power-efficient Recurrent Neural Network Accelerator”, https://dl.acm.org/doi/10.1145/3174243.3174261 (Year: 2018). [cited by examiner]
Park, H. et al., “Zero and Data Reuse-aware Fast Convolution for Deep Neural Networks on GPU”, https://dl.acm.org/doi/abs/10.1145/2968456.2968476 (Year: 2016). [cited by examiner]
Han, S. et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization, and Huffman Coding”, https://arxiv.org/abs/1510.00149 (Year: 2016). [cited by examiner]
Frankle, J. et al., “The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks”, https://arxiv.org/abs/1803.03635 (Year: 2019). [cited by examiner]
Thomas Epelbaum, “Deep learning: Technical introduction”, Mediamobile, Sep. 12, 2017 (106 pages). [cited by applicant]
J. Cheng, et al. “Recent Advances in Efficient Computation of Deep Convolutional Neural Networks”, Frontiers of Information Technology & Electronic Engineering, vol. 19, No. 1, pp. 64-77, 2018. [cited by applicant]
S. Han, et al., “Deep Compression: Compressing Deep Neural Networks With Pruning, Trained Quantization and Huffman Coding”, ICLR 2016, pp. 1-14. [cited by applicant]
D. Kim et al., A Novel Zero Weight/Activation-Aware Hardware Architecture of Convolutional Neural Network, 2017 Design, Automation and Test in Europe, (Mar. 27, 2017), pp. 1462-1467. [cited by applicant]
Yang et al., “Thinning of convolutional neural network with mixed pruning”, IET Image Process., 2019, vol. 13 Iss. 5, pp. 779-784 (Jan. 17, 2019). [cited by applicant]
Office Action dated Aug. 23, 2024 in Korean Patent Application No. 10-2020-0016899, 13 pages (in Korean). [cited by applicant]