IP Library › Granted Patent US 12,608,615
Granted Patent B2
US 12,608,615 · App. 17/943,176 · Granted Apr 21, 2026

Jointly pruning and quantizing deep neural networks

Inventors: Georgios Georgiadis (Porter Ranch, CA); Weiran Deng (Woodland Hills, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06N3/082G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,615
App. No.
17/943,176
Granted
Apr 21, 2026
Kind
B2
Abstract

A system and a method generate a neural network that includes at least one layer having weights and output feature maps that have been jointly pruned and quantized. The weights of the layer are pruned using an analytic threshold function. Each weight remaining after pruning is quantized based on a weighted average of a quantization and dequantization of the weight for all quantization levels to form quantized weights for the layer. Output feature maps of the layer are generated based on the quantized weights of the layer. Each output feature map of the layer is quantized based on a weighted average of a quantization and dequantization of the output feature map for all quantization levels. Parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the layer are updated using a cost function.

Claims (35)

1 . A neural network, comprising a plurality of layers, at least one layer comprising jointly pruned and quantized weights and output feature maps, the jointly pruned weights being pruned using an analytic threshold function, each weight remaining after being pruned further being quantized based on a weighted average of a quantization and dequantization of the weight for all quantization levels.

2 . The neural network of claim 1 , wherein the output feature maps are formed based on pruned and quantized weights of the layer, each output feature map being quantized based on a weighted average of a quantization and dequantization of the output feature map for all quantization levels and parameters of the analytic threshold function, and

wherein the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the layer are updated iteratively based on a cost function.

3 . The neural network of claim 2 , wherein the neural network is a full-precision trained neural network before weights and output feature maps of the at least one layer are jointly pruned and quantized.

4 . The neural network of claim 2 , wherein the cost function includes a pruning loss term, a weight quantization loss term and a feature map quantization loss term.

5 . The neural network of claim 2 , wherein the parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the layer are updated based on an optimization of the cost function.

6 . The neural network of claim 2 , wherein the parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the layer are iteratively updated based on an optimization of the cost function.

7 . The neural network of claim 2 , wherein the parameters of the analytic threshold function include a first parameter that controls a sharpness of the analytic threshold function, and a second parameter that controls a distance between a first edge and a second edge of the analytic threshold function.

8 . A method to prune weights and output feature maps of a layer of a neural network, the method comprising:

pruning weights of a layer of a neural network using an analytic threshold function, the neural network being a trained neural network;

quantizing each weight of the layer remaining after pruning based on a weighted average of a quantization and dequantization of the weight for all quantization levels to form quantized weights for the layer; and

determining output feature maps of the layer based on quantized weights of the layer.

9 . The method of claim 8 , further comprising:

quantizing each output feature map of the layer based on a weighted average of a quantization and dequantization of the output feature map for all quantization levels; and

iteratively updating parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the layer using a cost function.

10 . The method of claim 9 , wherein updating the parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the layer further comprises optimizing the cost function.

11 . The method of claim 10 , wherein the cost function includes a pruning loss term, a weight quantization loss term and a feature map quantization loss term.

12 . The method of claim 10 , further comprising iteratively pruning the weights, quantizing each weight of the layer, determining the output feature maps of the layer, quantizing each output feature map of the layer, and updating the parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the layer to optimize the cost function.

13 . The method of claim 12 , wherein the layer of the neural network is a first layer,

the method further comprising:

pruning weights of a second layer of the neural network using the analytic threshold function, the second layer being subsequent to the first layer in the neural network;

quantizing each weight of the second layer remaining after pruning based on a weighted average of a quantization and dequantization of the weight for all quantization levels to form quantized weights for the second layer;

determining the output feature maps of the second layer based on the quantized weights of the second layer;

quantizing each output feature map of the second layer based on a weighted average of a quantization and a dequantization of the output feature map for all quantization levels; and

updating parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the second layer by optimizing the cost function.

14 . The method of claim 9 , wherein the parameters of the analytic threshold function include a first parameter that controls a sharpness of the analytic threshold function, and second parameter that controls a distance between a first edge and a second edge of the analytic threshold function.

15 . A neural network analyzer, comprising:

an interface configured to receive a neural network that comprises a plurality of layers; and

a processing device configured to generate a neural network comprising at least one layer having weights and output feature maps that have been jointly pruned and quantized, to prune the weights of the at least one layer of the neural network using an analytic threshold function, to quantize each weight of the at least one layer remaining after pruning based on a weighted average of a quantization and dequantization of the weight for all quantization levels to form quantized weights for the at least one layer.

16 . The neural network analyzer of claim 15 , wherein the processing device is further configured to determine output feature maps of the at least one layer based on the quantized weights of the at least one layer, to quantize each output feature map of the at least one layer based on a weighted average of a quantization and dequantization of the output feature map for all quantization levels, and to update iteratively parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the at least one layer using a cost function.

17 . The neural network analyzer of claim 16 , wherein the interface is further configured to output the neural network comprising at least one layer having weights and output feature maps that have been jointly pruned and quantized.

18 . The neural network analyzer of claim 16 , wherein the neural network is a full-precision trained neural network before the weights and the output feature maps of the at least one layer are jointly pruned and quantized.

19 . The neural network analyzer of claim 16 , wherein the cost function includes a pruning loss term, a weight quantization loss term and a feature map quantization loss term.

20 . The neural network analyzer of claim 16 , wherein the parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the at least one layer are iteratively updated based on an optimization of the cost function, and

wherein the parameters of the analytic threshold function include a first parameter that controls a sharpness of the analytic threshold function, and second parameter that controls a distance between a first edge and a second edge of the analytic threshold function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2023
From: GEORGIADIS, GEORGIOS; DENG, WEIRAN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064866/0309 →
Continuity (3)
Continuation 16396619 · Apr 26, 2019
Provisional Application 62819484 · Mar 15, 2019
Related Publication 20230004813A1 · Jan 5, 2023
References Cited (24)
US 11475308B2 · Georgiadis · 2022 [cited by examiner]
US 20060195406A1 · Burges et al. · 2006 [cited by applicant]
US 20170286830A1 · El-Yaniv et al. · 2017 [cited by applicant]
US 20180046894A1 · Yao · 2018 [cited by applicant]
US 20180046915A1 · Sun et al. · 2018 [cited by applicant]
US 20180107926A1 · Choi et al. · 2018 [cited by applicant]
US 20180253401A1 · Cardinaux et al. · 2018 [cited by applicant]
US 20180285734A1 · Chen et al. · 2018 [cited by applicant]
US 20180285736A1 · Baum et al. · 2018 [cited by applicant]
US 20180300600A1 · Ma et al. · 2018 [cited by applicant]
US 20180365564A1 · Huang et al. · 2018 [cited by applicant]
CN 107832837A · 2018 [cited by applicant]
TW 201816669A · 2018 [cited by applicant]
TW 201842478A · 2018 [cited by applicant]
Wu et al., Compressing Complex Convolutional Neural Network Based on an Improved Deep Compression Algorithm; arXiv:1903.02358; Mar. 6, 2019; Total pp. 5 (Year: 2019). [cited by examiner]
Park et al., Weighted-Entropy-based Quantization for Deep Neural Networks; 2017 IEEE Conference on Computer Vision and Pattern Recognition; pp. 7197-7205 (Year: 2017). [cited by examiner]
Han, Song et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” arXiv:1510.00149v5 [cs.CV], 2016, 14 pages. [cited by applicant]
Manessi, Franco et al., “Automated Pruning for Deep Neural Network Compression,” 2018 24th International Conference on Pattern Recognition (ICPR), 2018, pp. 657-664. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/396,619, mailed Jun. 15, 2022. [cited by applicant]
Office Action for U.S. Appl. No. 16/396,619, mailed Mar. 22, 2022. [cited by applicant]
Rathi, Nitin et al., “STDP Based Pruning of Connections and Weight Quantization in Spiking Neural Networks for Energy-Efficient Recognition,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems… [cited by applicant]
Srivastava, Gaurav, “Joint Optimization of Quantization and Structured Sparsity for Compressed Deep Neural Networks,” ProQuest Dissertations & Theses Global: The Sciences and Engineering Collection, 2018, 64 pages. [cited by applicant]
Tung, Frederick et al., “CLIP-Q: Deep Network Compression Learning by In-Parallel Pruning-Quantization,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 7873-7882. [cited by applicant]
Manessi, Franco et al., “Automated Pruning for Deep Neural Network Compression,” https://arxiv.org/abs/1712.01721v1, Dec. 2017, 10 pages. [cited by applicant]