IP Library Granted Patent US 12,450,485
Granted Patent B2
US 12,450,485 · App. 16/197,986 · Granted Oct 21, 2025

Pruning neural networks that include element-wise operations

Inventors: Varun Praveen (Santa Clara, CA); Anil Ubale (Cupertino, CA); Parthasarathy Sriram (Los Altos, CA); Greg Heinrich (Nice, FR); Tayfun Gurel (Vantaa, FI)
Assignee: NVIDIA Corporation
G06N3/082G06N3/04H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,485
App. No.
16/197,986
Granted
Oct 21, 2025
Kind
B2
Abstract

Input layers of an element-wise operation in a neural network can be pruned such that the shape (e.g., the height, the width, and the depth) of the pruned layers matches. A pruning engine identifies all of the input layers into the element-wise operation. For each set of corresponding neurons in the input layers, the pruning engine equalizes the metrics associated with the neurons to generate an equalized metric associated with the set. The pruning engine prunes the input layers based on the equalized metrics generated for each unique set of corresponding neurons.

Claims (32)

1. A computer-implemented method comprising:

computing an average weight value of a plurality of neurons of one or more neural networks, wherein each of the plurality of neurons is located within at a matching location within a different layer included in a plurality of locations of different layers of the one or more neural networks, and wherein each of the plurality of neurons is associated with a matching feature type;

replacing a weight value of each of the plurality of neurons with the average weight value; and

deactivating each of the plurality of neurons based, at least in part, on whether the replaced weight value of the plurality of neurons exceeds a threshold value.

2. The method of claim 1 , wherein one or more outputs of the plurality of different layers are processed using one or more element-wise operations.

3. The method of claim 1 , wherein deactivating each of the plurality of neurons comprises performing one or more equalization operations on weights associated with the plurality of neurons.

4. The method of claim 3 , wherein performing the one or more equalization operations comprises applying an equalization operator to the weights associated with the plurality of neurons.

5. The method of claim 1 , wherein the average weight value is computed using an arithmetic mean or a geometric mean of weights associated with the plurality of neurons.

6. The method of claim 1 , wherein deactivating each of a plurality of neurons further comprises:

obtaining a desired dimensionality of the plurality of different layers to be deactivated.

7. The method of claim 1 , wherein the one or more neural networks comprise a residual network, and wherein the plurality of different layers includes a convolutional layer of the residual network and an identity layer of the residual network.

8. The method of claim 1 , wherein each of the plurality of neurons produces a different input into a given computational component of the one or more neural networks.

9. The method of claim 1 , wherein the weight value is computed using an arithmetic mean or a geometric mean of the weights associated with the plurality of neurons.

10. One or more processors comprising circuitry to:

computing compute an average weight value of a plurality of neurons of one or more neural networks, wherein each of the plurality of neurons is located within at a matching location within a different layer included in a plurality of locations of different layers of one or more neural networks, and wherein each of the plurality of neurons is associated with a matching feature type;

replacing replace a weight value of each of the plurality of neurons with the average weight value; and

deactivate each of the plurality of neurons based, at least in part, on whether the replaced weight value of the plurality of neurons exceeds a threshold value.

11. The one or more processors of claim 10 , wherein one or more outputs of the plurality of different layers are processed using one or more element-wise operations.

12. The one or more processors of claim 10 , wherein deactivation of each of the plurality of neurons comprise computing L2 norm of weights associated with the plurality of neurons.

13. The one or more processors of claim 10 , wherein deactivation of each of the plurality of neurons comprises performing one or more equalization operations on weights associated with the plurality of neurons.

14. The one or more processors of claim 13 , wherein performing the one or more equalization operations comprises applying an equalization operator to the weights associated with the plurality of neurons.

15. The one or more processors of claim 10 , wherein the average weight value is computed using an arithmetic mean or a geometric mean of the weights associated with the plurality of neurons.

16. The one or more processors of claim 10 , wherein deactivation each of the plurality of neurons comprises:

obtaining a desired dimensionality of the plurality of different layers to be deactivated.

17. The one or more processors of claim 10 , wherein the one or more neural networks comprise a residual network, and wherein the plurality of different layers include a convolutional layer of the residual network and an identity layer of the residual network.

18. A system comprising: one or more processors to:

compute an average weight value of a plurality of neurons of one or more neural networks, wherein each of the plurality of neurons is located at a within matching location within a different layer included in a plurality of locations of different layers of the one or more neural networks, and wherein each of the plurality of neurons is associated with a matching feature type;

replace a weight value of each of the plurality of neurons with the average weight value; and

deactivate each of the plurality of neurons based, at least in part, on whether the replaced weight value of the plurality of neurons exceeds a threshold value.

19. The system of claim 18 , wherein one or more outputs of the plurality of different layers are processed using one or more element-wise operations.

20. The system of claim 18 , wherein the deactivation of the plurality of neurons comprises performing an equalization operation on weights associated with the plurality of neurons.

21. The system of claim 18 , wherein the one or more neural networks comprise a residual network, and wherein the plurality of different layers includes a convolutional layer of the residual network and an identity layer of the residual network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2019
From: PRAVEEN, VARUN; UBALE, ANIL; SRIRAM, PARTHASARATHY; HEINRICH, GREG; GUREL, TAYFUN
To: NVIDIA CORPORATION
Reel/Frame 049471/0806 →
Continuity (1)
Related Publication 20200160185A1 · May 21, 2020
References Cited (44)
US 11423259B1 · Pratusevich · 2022 [cited by examiner]
US 20170102950A1 · Chamberlain · 2017 [cited by examiner]
US 20180046906A1 · Dally · 2018 [cited by examiner]
US 20180114114A1 · Molchanov · 2018 [cited by examiner]
US 20180137417A1 · Theodorakopoulos · 2018 [cited by examiner]
US 20190050710A1 · Wang · 2019 [cited by examiner]
US 20190050734A1 · Li · 2019 [cited by examiner]
US 20190122113A1 · Chen · 2019 [cited by examiner]
US 20190197406A1 · Darvish Rouhani · 2019 [cited by examiner]
US 20190378017A1 · Kung · 2019 [cited by examiner]
US 20200334537A1 · Yao · 2020 [cited by examiner]
US 20210027166A1 · Gorokhov · 2021 [cited by examiner]
US 20210182077A1 · Chen · 2021 [cited by examiner]
US 20230024840A1 · Chen · 2023 [cited by examiner]
CN 107977703A · 2018 [cited by applicant]
WO 2018119035A1 · 2018 [cited by applicant]
Chen, Tianshi, et al. “Diannao: A small-footprint high-throughput accelerator for ubiquitous machine-learning.” ACM SIGARCH Computer Architecture News 42.1 (2014): 269-284. (Year: 2014). [cited by examiner]
Wen, Wei, et al. “Learning structured sparsity in deep neural networks.” Advances in neural information processing systems 29 (2016): 1-9 (Year: 2016). [cited by examiner]
Molchanov, Pavlo, et al. “Pruning convolutional neural networks for resource efficient inference.” arXiv preprint arXiv:1611.06440 (2016): 1-17 (Year: 2016). [cited by examiner]
Li, Hao, et al. “Pruning filters for efficient convnets.” arXiv preprint arXiv:1608.08710v3 (2017): 1-13. (Year: 2017). [cited by examiner]
Zhou, Zhengguang, et al. “Online filter clustering and pruning for efficient convnets.” 2018 25th IEEE International Conference on Image Processing (ICIP). IEEE, Oct. 2018: 11-15. (Year: 2018). [cited by examiner]
Xie, Wenao, et al. “An energy-efficient FPGA-based embedded system for CNN application.” 2018 IEEE international conference on electron devices and solid state circuits (EDSSC). IEEE, Jun. 2018. (Year: 2018). [cited by examiner]
Yu, Ruichi, et al. “Nisp: Pruning networks using neuron importance score propagation.” Proceedings of the IEEE conference on computer vision and pattern recognition. Jun. 2018: 9194-9203 (Year: 2018). [cited by examiner]
Huang, Qiangui, et al. “Learning to prune filters in convolutional neural networks.” 2018 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, Mar. 2018: 709-718 (Year: 2018). [cited by examiner]
Anwar, Sajid, Kyuyeon Hwang, and Wonyong Sung. “Structured pruning of deep convolutional neural networks.” ACM Journal on Emerging Technologies in Computing Systems (JETC) 13.3 (2017): 1-18. (Year: 2017). [cited by examiner]
Russakovsky, Olga, et al. “Imagenet large scale visual recognition challenge.” International journal of computer vision 115 (2015): 211-252. (Year: 2015). [cited by examiner]
Ardakani, Arash, Carlo Condo, and Warren J. Gross. “Activation pruning of deep convolutional neural networks.” 2017 IEEE Global Conference on Signal and Information Processing (GlobalSIP). IEEE, 2017. (Year: 2017). [cited by examiner]
Ardakani, Arash, Carlo Condo, and Warren J. Gross. “A multi-mode accelerator for pruned deep neural networks.” 2018 16th IEEE International New Circuits and Systems Conference (NEWCAS). IEEE, Jun. 2018. (Year: 2018). [cited by examiner]
Chen et al., “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine Learning,” ASPLOS 14, Mar. 1, 2014, 16 pages. [cited by applicant]
Extended European Search Report mailed Apr. 22, 2020, for Application No. 19210256.4, 14 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Wen et al., “Learning Structured Sparsity in Deep Neural Networks,” Proceedings of the 30th International Conference on Neural Information Processing Systems, Dec. 5, 2016, 10 pages. [cited by applicant]
Xie et al., “An Energy-Efficient FPGA-Based Embedded System for CNN Application,” 2018 IEEE International Conference on Electron Devices and Solid State Circuits, Jun. 6, 2018, 2 pages. [cited by applicant]
Li et al., “Pruning Filters for Efficient Convnets”, arXiv:1608.08710v3 [cs.CV], Mar. 10, 2017, pp. 1-13. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, arXiv:1512.03385v1 [cs.CV], Dec. 10, 2015, pp. 1-12. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge”, arXiv:1409.0575v3 [cs.CV], Jan. 30, 2015, pp. 1-43. [cited by applicant]
Han et al., “Learning both Weights and Connections for Efficient Neural Networks”, arXiv:1506.02626v3 [cs.NE], Oct. 30, 2015, pp. 1135-1143. [cited by applicant]
Le Cun et al., “Optimal Brain Damage”, Jun. 1990, in D. Touretzky (ed.), Advances in Neural Information Processing Systems, vol. 2, pp. 598-605. [cited by applicant]
Decision of Refusal for European Application No. 19210256.4, mailed Oct. 28, 2024, 16 pages. [cited by applicant]
Office Action for Chinese Application No. 201911141802.3, mailed Dec. 6, 2024, 10 pages. [cited by applicant]
Office Action for Chinese Application No. 201911141802.3, mailed Jan. 8, 2024, 29 pages. [cited by applicant]
Extended European Search Report for Application No. 25150039.3, mailed Apr. 14, 2025, 10 pages. [cited by applicant]
Office Action for Chinese Application No. 201911141802.3, mailed Feb. 8, 2025, 10 pages. [cited by applicant]
Office Action for Chinese Application No. 201911141802.3, mailed Jul. 29, 2024, 8 pages. [cited by applicant]