IP Library › Granted Patent US 12,314,857
Granted Patent B2
US 12,314,857 · App. 17/241,572 · Granted May 27, 2025

Method and device for deep neural network compression

Inventors: Juinn-Dar Huang (New Taipei, TW); Ya-Chu Chang (New Taipei, TW); Wei-Chen Lin (New Taipei, TW)
Assignee: ACER INCORPORATED
G06N3/082G06N3/04G06F7/49G06F7/491
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,857
App. No.
17/241,572
Granted
May 27, 2025
Kind
B2
Abstract

A method for deep neural network compression is provided. The method includes: using at least one weight of a deep neural network (DNN), setting a value of a P parameter, and combining every P weights in groups, and perform branch pruning and retraining, so that only one of each group has a non-zero weight, and the remaining weights are 0, wherein the remaining weights are evenly divided into branches to adjust a compression rate of the DNN and to adjust a reduction rate of the DNN.

Claims (37)

1. A method for deep neural network compression, comprising:

obtaining, by a processor, at least one weight of a deep neural network (DNN), setting a value of a P parameter, and combining every P weights in groups, wherein P is a number, the at least one weight is placed between two network layers corresponding to an input layer and an output layer adjacent to each other in the deep neural network, the processor comprising input layer storage units having nodes, wherein the processor further comprises a multiplexer, wherein input values of the nodes of the input layer are multiplied at the multiplexer by respective weights in one of the groups to generate first output values, the first output values accumulated into a final output value which is equal to an output value of one of plural nodes of the output layer;

setting, by the processor, a loop parameter L so that a branch pruning and retraining is performed on each group of P weights from the first loop to the Lth loop, setting a threshold T corresponding to each loop, wherein the threshold T increases gradually from a first threshold T1 to an L threshold TL, the first threshold T1 is used at the beginning of the first loop, and the L threshold TL is used in the Lth loop;

for each loop, setting, by the processor, the weights less than the threshold T to 0 and retaining the weights greater than the threshold T when obtaining the weights of the same group in the same loop according to the threshold T, wherein when all of the weights of the same group in the same loop are less than the threshold T, only the largest weight is retained, and remaining weights are set to 0; and

performing, by the processor, the branch pruning and retraining until the Lth loop, so that only one weight of each group is non-zero and the remaining weights are 0.

2. The method for deep neural network compression as claimed in claim 1 , wherein the DNN is a fully connected structure in which each node of the input layer is connected to each node of the output layer.

3. The method for deep neural network compression as claimed in claim 1 , wherein the DNN is a combination of multiple network layers, each group of network layers of the DNN is adjacently connected by the input layer of the same group of network layers and the output layer of the same group of network layers, and the output layer of the previous group of network layers is the input layer of the next group of network layers.

4. The method for deep neural network compression as claimed in claim 3 , wherein packet formats of the weight have 16-bit, 12-bit and 8-bit formats.

5. The method for deep neural network compression as claimed in claim 4 , wherein a binary fixed-point number extraction calculation is performed on the output value of the nodes of the output layer, wherein the number of digits before the decimal point of the output value of the nodes of the output layer plus the number of digits after the decimal point of the output value of the nodes of the output layer is 16 bits.

6. The method for deep neural network compression as claimed in claim 5 , wherein a binary fixed-point number extraction calculation is performed on the output value of the nodes of the output layer, wherein the number of digits before the decimal point of the output value plus at least 1 bit increased by adjustment, and plus the number of digits after the decimal point of the output value of the nodes of the output layer is 16 bits.

7. The method for deep neural network compression as claimed in claim 5 , wherein among the bits of the packet formats of the weight, at least one high bit counted from the highest bit is an address index, which is used to store an address corresponding to the input values of nodes of the input layer;

wherein the method further comprises:

inputting, by a select terminal of the multiplexer of the processor, the address index;

inputting, by an input of the multiplexer of the processor, the weights of the same group;

selecting, by the multiplexer of the processor, the input values of nodes of the input layer corresponding to the address index; and

outputting, by an output of the multiplexer of the processor, the input values of the nodes of the input layer corresponding to the address index.

8. The method for deep neural network compression as claimed in claim 7 , wherein among the bits of the packet formats of the weight, remaining bits are bits for storing the weights except for the address index represented by at least one high bit counted from the highest bit.

9. The method for deep neural network compression as claimed in claim 7 , wherein an address generator of the processor connected with a weight memory and the input layer storage units is used as an address counter, and the bits of the packet formats of the weight are stored in the weight memory.

10. A device for deep neural network compression, comprising:

a processor, configured to execute the following tasks:

obtaining at least one weight of a deep neural network (DNN), setting a value of a P parameter, and combining every P weights in groups, wherein P is a number, the at least one weight is placed between two network layers corresponding to an input layer and an output layer adjacent to each other in the deep neural network, the processor comprising input layer storage units having nodes, wherein the processor further comprises a multiplexer, wherein input values of the nodes of the input layer are multiplied at the multiplexer by respective weights in one of the groups to generate first output values; the first output values accumulated into a final output value which is equal to an output value of one of plural nodes of the output layer;

setting a loop parameter L so that a branch pruning and retraining is performed on each group of P weights from the first loop to the Lth loop, setting a threshold T corresponding to each loop, wherein the threshold T increases gradually from a first threshold T1 to an L threshold TL, the first threshold T1 is used at the beginning of the first loop, and the L threshold TL is used in the Lth loop;

for each loop, setting the weights less than the threshold T to 0 and retaining the weights greater than the threshold T when obtaining the weights of the same group in the same loop according to the threshold T, wherein when all of the weights of the same group in the same loop are less than the threshold T, only the largest weight is retained, and remaining weights are set to 0; and

performing the branch pruning and retraining until the Lth loop of the last loop, so that only one weight of each group is non-zero and the remaining weights are 0.

11. The device for deep neural network compression as claimed in claim 10 , wherein the DNN is a fully connected structure in which each node of the input layer is connected to each node of the output layer.

12. The device for deep neural network compression as claimed in claim 10 , wherein the DNN is a combination of multiple network layers, each group of network layers of the DNN is adjacently connected by the input layer of the same group of network layers and the output layer of the same group of network layers, and the output layer of the previous group of network layers is the input layer of the next group of network layers.

13. The device for deep neural network compression as claimed in claim 12 , wherein packet formats of the weight have 16-bit, 12-bit and 8-bit formats.

14. The device for deep neural network compression as claimed in claim 13 , wherein a binary fixed-point number extraction calculation is performed on the output value of the nodes of the output layer, wherein the number of digits before the decimal point of the output value of the nodes of the output layer plus the number of digits after the decimal point of the output value of the nodes of the output layer is 16 bits.

15. The device for deep neural network compression as claimed in claim 14 , wherein a binary fixed-point number extraction calculation is performed on the output value of the nodes of the output layer, wherein the number of digits before the decimal point of the output value plus at least 1 bit increased by adjustment, and plus the number of digits after the decimal point of the output value of the nodes of the output layer is 16 bits.

16. The device for deep neural network compression as claimed in claim 14 , wherein among the bits of the packet formats of the weight, at least one high bit counted from the highest bit is an address index, which is used to store an address corresponding to the input values of nodes of the input layer;

wherein the processor further executes the following tasks:

inputting, by a select terminal of the multiplexer of the processor, the address index;

inputting, by an input of the multiplexer of the processor, the weights of the same group;

selecting, by the multiplexer of the processor, the input values of nodes of the input layer corresponding to the address index; and

outputting, by an output of the multiplexer of the processor, the input values of the nodes of the input layer corresponding to the address index.

17. The device for deep neural network compression as claimed in claim 16 , wherein among the bits of the packet formats of the weight, remaining bits are bits for storing the weights except for the address index represented by at least one high bit counted from the highest bit.

18. The device for deep neural network compression as claimed in claim 16 , wherein an address generator of the processor connected with a weight memory and the input layer storage units is used as an address counter, and the bits of the packet formats of the weight are stored in the weight memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2021
From: HUANG, JUINN-DAR; CHANG, YA-CHU; LIN, WEI-CHEN
To: ACER INCORPORATED
Reel/Frame 056054/0335 →
Priority Claims (1)
TW 109116293 · May 15, 2020 · national
Continuity (1)
Related Publication 20210357758A1 · Nov 18, 2021
References Cited (14)
US 10223635B2 · Annapureddy et al. · 2019 [cited by applicant]
US 10621424B2 · Lin · 2020 [cited by applicant]
US 10832135B2 · Ji · 2020 [cited by examiner]
US 11568254B2 · Lee · 2023 [cited by examiner]
US 20200387782A1 · Hegde · 2020 [cited by examiner]
TW 201128542A · 2011 [cited by applicant]
TW 201627923A · 2016 [cited by applicant]
TW 201943263A · 2019 [cited by applicant]
Molchanov et al, “Pruning convolutional neural networks for resource efficient inference”, arXiv:1611.06440v2 [cs.LG] Jun. 8, 2017 (Year: 2017). [cited by examiner]
He et al, “Structured pruning for deep convolutional neural networks: a survey”, IEEE Transactions on pattern analysis and machine intelligence, vol. 46, No. 5, May 2024 (Year: 2024). [cited by examiner]
Extended European Search Report dated Mar. 15, 2022, issued in application No. EP 21172345.7. [cited by applicant]
Kang, H.J.; “Accelerator-Aware Pruning for Convolutional Neural Networks;” IEEE; Apr. 2018; pp. 1-11. [cited by applicant]
Anwar, S., et al.; “Structured Pruning of Deep Convolutional Neural Networks;” ACM Journal on Emerging Technologies in Computing Systems; vol. 13; No. 3; Article 32; Feb. 2017; pp. 32:1-32:18. [cited by applicant]
Chinese language office action dated Nov. 30, 2020, issued in application No. TW 109116293. [cited by applicant]