IP Library › Granted Patent US 12,488,246
Granted Patent B2
US 12,488,246 · App. 16/699,027 · Granted Dec 2, 2025

Processing method and accelerating device

Inventors: Zidong Du (Pudong New Area, CN); Xuda Zhou (Pudong New Area, CN); Shaoli Liu (Pudong New Area, CN); Tianshi Chen (Pudong New Area, CN)
Assignee: SHANGHAI CAMBRICON INFORMATION TECHNOLOGY CO., LTD.
G06N3/082G06F1/3296G06F9/3877G06F12/0875G06F13/16G06F16/285G06N3/04G06N3/044G06N3/048G06N3/063G06N3/084G06F2212/452G06F2213/0026
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,246
App. No.
16/699,027
Granted
Dec 2, 2025
Kind
B2
Abstract

The present disclosure provides a processing device including: a coarse-grained pruning unit configured to perform coarse-grained pruning on a weight of a neural network to obtain a pruned weight, an operation unit configured to train the neural network according to the pruned weight. The coarse-grained pruning unit is specifically configured to select M weights from the weights of the neural network through a sliding window, and when the M weights meet a preset condition, all or part of the M weights may be set to 0. The processing device can reduce the memory access while reducing the amount of computation, thereby obtaining an acceleration ratio and reducing energy consumption.

Claims (68)

1. A processing device, comprising:

a coarse-grained pruning circuit configured to perform coarse-grained pruning on weights of a neural network to obtain pruned weights, wherein the neural network includes a convolutional layer and a weight of the convolutional layer is a four-dimensional matrix (Nfin, Nfout, Kx, Ky), Nfin represents a count of input feature maps, Nfout represents a count of output feature maps, (Kx, Ky) is a size of a convolution kernel, and the convolutional layer has Nfin*Nfout*Kx*Ky weights;

and an operation circuit configured to train the neural network according to the pruned weight;

wherein the coarse-grained pruning circuit is configured to:

select M weights from the weights of the neural network through a sliding window;

determine that the M weights meet a preset condition, wherein the preset condition is that an information quantity of the M weights is less than a first given threshold;

and based on a determination that the M weights meet the preset condition,

set at least a portion of the selected M weights to 0 to obtain a portion of the pruned weights,

wherein the coarse-grained pruning circuit is further configured to:

perform coarse-grained pruning on the weight of the convolutional layer, where the sliding window is a four-dimensional sliding window with a size of Bfin*Bfout*Bx*By, Bfin is a positive integer greater than 0 and less than or equal to Nfin, Bfout is a positive integer greater than 0 and less than or equal to Nfout, Bx is a positive integer greater than 0 and less than or equal to Kx, and By is a positive integer greater than 0 and less than or equal to Ky,

make the sliding window slide Sfin stride in the direction of Bfin, or slide Sfout stride in the direction of Bfout, or slide S stride in the direction of Bx, or slide Sy stride in the direction of By, where Sfin is a positive integer greater than 0 and less than or equal to Bfin, Sfout is a positive integer greater than 0 and less than or equal to Bfout, Sx is a positive integer greater than 0 and less than or equal to Bx, and Sy is a positive integer greater than 0 and less than or equal to By,

select M weights from the Nfin*Nfout*Kx*Ky weights through the sliding window, and when the M weights meet the preset condition, all or part of the M weights may be set to 0, where M=Bfin*Bfout*Bx*Bv.

2. The processing device of claim 1 , wherein

the information quantity of the M weights is an arithmetic mean of an absolute value of the M weights, a geometric mean of the absolute value of the M weights or a maximum value of the absolute value of the M weights; the first given threshold is a first threshold, a second threshold or a third threshold; and the information quantity of the M weights being less than the first given threshold includes:

the arithmetic mean of the absolute value of the M weights being less than the first threshold, or the geometric mean of the absolute value of the M weights being less than the second threshold, or the maximum value of the M weights being less than the third threshold.

3. The processing device of claim 1 , wherein the coarse-grained pruning circuit and the operation circuit are configured to repeat performing coarse-grained pruning on the weights of the neural network and training the neural network according to the pruned weights until no weight meets the preset condition without losing a preset precision.

4. The processing device of claim 1 ,

wherein the neural network includes a fully connected layer and a Long Short Term Memory (LSTM) layer, where a weight of the fully connected layer is a two-dimensional matrix (Nin, Nout), Nin represents a count of input neurons and Nout represents a count of output neurons, and the fully connected layer has Nin*Nout weights;

a weight of LSTM layer is composed of m weights of the fully connected layer, m is a positive integer greater than 0, and an ith weight of the fully connected layer is (Nin_i, Nout_i), where i is a positive integer greater than 0 and less than or equal to m, Nin_i represents a count of input neurons of the ith weight of the fully connected layer and Nout_i represents a count of output neurons of the ith weight of the fully connected layer;

the coarse-grained pruning unit is specifically configured to: perform coarse-grained pruning on the weight of the fully connected layer, where the sliding window is a sliding window with a size of Bin*Bout, Bin is a positive integer greater than 0 and less than or equal to Nin, and Bout is a positive integer greater than 0 and less than or equal to Nout,

make the sliding window slide Sin stride in a direction of Bin, or slide Sout stride in a direction of Bout, where Sin is a positive integer greater than 0 and less than or equal to Bin, and Sout is a positive integer greater than 0 and less than or equal to Bout,

select M weights from the Nin*Nout weights through the sliding window, and when the M weights meet the preset condition, all or part of the M weights may be set to 0, where M=Bin*Bout,

perform coarse-grained pruning on the weight of the LSTM layer, where the size of the sliding window is Bin_i*Bout_i, Bin_i is a positive integer greater than 0 and less than or equal to Nini, and Bout_i is a positive integer greater than 0 and less than or equal to Nout_i,

make the sliding window slide Sini stride in the direction of Bin_i, or slide Sout_i stride in the direction of Bout_i, where Sin_i is a positive integer greater than 0 and less than or equal to Bin_i, and Sout_i is a positive integer greater than 0 and less than or equal to Bout_i, and

select M weights from the Bin_i*Bouti weights through the sliding window, and when the M weights meet the preset condition, all or part of the M weights may be set to 0, where M=Bin_i*Bouti.

5. The processing device of claim 1 , wherein the operation circuit is specifically configured to:

retrain the neural network by a back-propagation algorithm according to the pruned weight.

6. The processing device of claim 1 , further comprising

a quantization circuit configured to quantize the weight of the neural network and/or perform a first operation on the weight of the neural network after the coarse-grained pruning circuit performs coarse-grained pruning on the weight of the neural network, and before the operation circuit retrains the neural network according to the pruned weight to reduce a count of weight bits of the neural network.

7. A neural network operation device, comprising: a first processing device; and one or more second processing devices communicatively connected to the first processing device via a Peripheral Component Interconnect Express (PCIE) bus to transfer data to support operations of a neural network, wherein the neural network includes a convolutional layer and a weight of the convolutional layer is a four-dimensional matrix (Nfin, Nfout, Kx, Ky), Nfin represents a count of input feature maps, Nfout represents a count of output feature maps, (Kx, Ky) is a size of a convolution kernel, and the convolutional layer has Nfin*Nfout*Kx*Ky weights, wherein each of the first processing device and the one or more second processing devices includes:

a coarse-grained pruning circuit configured to perform coarse-grained pruning on weights of the neural network to obtain pruned weights;

and an operation circuit configured to train the neural network according to the pruned weight, wherein the coarse-grained pruning circuit is configured to:

select M weights from the weights of the neural network through a sliding window,

determine that the M weights meet a preset condition, wherein the preset condition is that an information quantity of the M weights is less than a first given threshold;

and set at least one of the selected M weights to zero to obtain a portion of the pruned weights based on a determination that the M weights meet the preset condition,

wherein the coarse-grained pruning circuit is further configured to:

perform coarse-grained pruning on the weight of the convolutional layer, where the sliding window is a four-dimensional sliding window with a size of Bfin*Bfout*Bx*By, Bfin is a positive integer greater than 0 and less than or equal to Nfin, Bfout is a positive integer greater than 0 and less than or equal to Nfout, Bx is a positive integer greater than 0 and less than or equal to Kx, and By is a positive integer greater than 0 and less than or equal to Ky,

make the sliding window slide Sfin stride in the direction of Bfin, or slide Sfout stride in the direction of Bfout, or slide S stride in the direction of Bx, or slide Sy stride in the direction of By, where Sfin is a positive integer greater than 0 and less than or equal to Bfin, Sfout is a positive integer greater than 0 and less than or equal to Bfout, Sx is a positive integer greater than 0 and less than or equal to Bx, and Sy is a positive integer greater than 0 and less than or equal to By,

select M weights from the Nfin*Nfout*Kx*Ky weights through the sliding window, and when the M weights meet the preset condition, all or part of the M weights set to 0, where M=Bfin*Bfout*Bx*Bv.

8. A processing method, comprising performing coarse-grained pruning on weights of a neural network to obtain pruned weights, wherein the neural network includes a convolutional layer and a weight of the convolutional layer is a four-dimensional matrix (Nfin, Nfout, Kx, Ky), Nfin represents a count of input feature maps, Nfout represents a count of output feature maps, (Kx, Ky) is a size of a convolution kernel, and the convolutional layer has Nfin*Nfout*Kx*Ky weights; and

training the neural network according to the pruned weights;

wherein the performing coarse-grained pruning on the weights of the neural network to obtain the pruned weights includes:

selecting M weights from the weights of the neural network through a sliding window,

determining that the M weights meet a preset condition, wherein the preset condition is that an information quantity of the M weights is less than a first given threshold;

based on a determination that the M weights meet the preset condition,

setting at least a portion of the selected M weights to zero to obtain a portion of the pruned weights, and

wherein the performing coarse-grained pruning further includes:

performing coarse-grained pruning on the weight of the convolutional layer, the sliding window is a four-dimensional sliding window with a size of Bfin*Bfout*Bx*By, where Bfin is a positive integer greater than 0 and less than or equal to Nfin, Bfout is a positive integer greater than 0 and less than or equal to Nfout, Bx is a positive integer greater than 0 and less than or equal to Kx, and By is a positive integer greater than 0 and less than or equal to Ky,

making the sliding window slide Sfin stride in the direction of Bfin, or slide Sfout stride in the direction of Bfout, or slide S stride in the direction of Bx, or slide Sy stride in the direction of By, where Sfin is a positive integer greater than 0 and less than or equal to Bfin, Sfout is a positive integer greater than 0 and less than or equal to Bfout, Sx is a positive integer greater than 0 and less than or equal to Bx, and Sy is a positive integer greater than 0 and less than or equal to By,

selecting M weights from the Nfin*Nfout*Kx*Ky weights through the sliding window, and when the M weights meet the preset condition, setting all or part of the M weights to 0, where M=Bfin*Bfout*Bx*By.

9. The processing method of claim 8 , wherein the information quantity of the M weights is an arithmetic mean of an absolute value of the M weights, a geometric mean of the absolute value of the M weights or a maximum value of the absolute value of the M weights; the first given threshold is a first threshold, a second threshold or a third threshold; and the information quantity of the M weights being less than the first given threshold includes:

the arithmetic mean of the absolute value of the M weights being less than the first threshold, or the geometric mean of the absolute value of the M weights being less than the second threshold, or the maximum value of the M weights being less than the third threshold.

10. The processing method of claim 8 , further comprising

repeating performing coarse-grained pruning on the weights of the neural network and training the neural network according to the pruned weights until no weight meets the preset condition without losing a preset precision.

11. The processing method of claim 8 ,

wherein the neural network includes a fully connected layer and a Long Short Term Memory (LSTM) layer, where a weight of the fully connected layer is a two-dimensional matrix (Nin, Nout), Nin represents a count of input neurons and Nout represents a count of output neurons, and the fully connected layer has Nin*Nout weights;

a weight of LSTM layer is composed of m weights of the fully connected layer, m is a positive integer greater than 0, and an ith weight of the fully connected layer is (Nin_i, Nout_i), where i is a positive integer greater than 0 and less than or equal to m, Nin_i represents a count of input neurons of the ith weight of the fully connected layer and Nout_i represents a count of output neurons of the ith weight of the fully connected layer;

the coarse-grained pruning unit is specifically configured to:

performing coarse-grained pruning on the weight of the fully connected layer, where the sliding window is a sliding window with a size of Bin*Bout, Bin is a positive integer greater than 0 and less than or equal to Nin, and Bout is a positive integer greater than 0 and less than or equal to Nout,

make the sliding window slide Sin stride in a direction of Bin, or slide Sout stride in a direction of Bout, where Sin is a positive integer greater than 0 and less than or equal to Bin, and Sout is a positive integer greater than 0 and less than or equal to Bout,

select M weights from the Nin*Nout weights through the sliding window, and when the M weights meet the preset condition, all or part of the M weights set to 0, where M=Bin*Bout,

perform coarse-grained pruning on the weight of the LSTM layer, where the size of the sliding window is Bin_i*Bout_i, Bin_i is a positive integer greater than 0 and less than or equal to Nini, and Bout_i is a positive integer greater than 0 and less than or equal to Nout_i,

make the sliding window slide Sini stride in the direction of Bin_i, or slide Sout_i stride in the direction of Bout_i, where Sin_i is a positive integer greater than 0 and less than or equal to Bin_i, and Sout_i is a positive integer greater than 0 and less than or equal to Bout_i, and

select M weights from the Bin_i*Bouti weights through the sliding window, and when the M weights meet the preset condition, all or part of the M weights may be set to 0, where M=Bin_i*Bouti.

12. The processing method of claim 8 , wherein the training the neural network according to the pruned weight is:

retraining the neural network by a back-propagation algorithm according to the pruned weight.

13. The processing method of claim 8 , wherein after performing coarse-grained pruning on the weight of the neural network, and before retraining the neural network, the method further includes:

quantizing the weight of the neural network and/or performing a first operation on the weight of the neural network to reduce a count of weight bits of the neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2019
From: DU, ZIDONG; ZHOU, XUDA; LIU, SHAOLI; CHEN, TIANSHI
To: SHANGHAI CAMBRICON INFORMATION TECHNOLOGY CO., LTD
Reel/Frame 051135/0680 →
Priority Claims (6)
CN 201710370905.1 · May 23, 2017 · national
CN 201710583336.9 · May 23, 2017 · national
CN 201710456759.4 · Jun 16, 2017 · national
CN 201710677987.4 · Aug 9, 2017 · national
CN 201710678038.8 · Aug 9, 2017 · national
CN 201710689666.6 · Aug 9, 2017 · national
Continuity (2)
Continuation In Part PCTCN2018088033 · May 23, 2018
Related Publication 20200097826A1 · Mar 26, 2020
References Cited (113)
US 4422141A · Shoji · 1983 [cited by applicant]
US 6360019B1 · Chaddha · 2002 [cited by applicant]
US 6772126B1 · Simpson et al. · 2004 [cited by applicant]
US 10127495B1 · Bopardikar et al. · 2018 [cited by applicant]
US 10657439B2 · Liu et al. · 2020 [cited by applicant]
US 11315018B2 · Molchanov et al. · 2022 [cited by applicant]
US 20030204311A1 · Bush · 2003 [cited by applicant]
US 20040180690A1 · Song et al. · 2004 [cited by applicant]
US 20050102301A1 · Flanagan · 2005 [cited by applicant]
US 20080249767A1 · Ertan · 2008 [cited by applicant]
US 20120189047A1 · Jiang et al. · 2012 [cited by applicant]
US 20140300758A1 · Tran · 2014 [cited by applicant]
US 20150332690A1 · Kim et al. · 2015 [cited by applicant]
US 20160358069A1 · Brothers · 2016 [cited by examiner]
US 20170061328A1 · Majumdar · 2017 [cited by examiner]
US 20170270408A1 · Shi et al. · 2017 [cited by applicant]
US 20180046900A1 · Dally · 2018 [cited by examiner]
US 20180075336A1 · Huang · 2018 [cited by examiner]
US 20180114114A1 · Molchanov et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180197081A1 · Ji · 2018 [cited by examiner]
US 20180285731A1 · Heifets et al. · 2018 [cited by applicant]
US 20180300603A1 · Ambardekar et al. · 2018 [cited by applicant]
US 20180314940A1 · Kundu et al. · 2018 [cited by applicant]
US 20190019311A1 · Hu et al. · 2019 [cited by applicant]
US 20190050709A1 · Yang et al. · 2019 [cited by applicant]
US 20190362235A1 · Xu et al. · 2019 [cited by applicant]
US 20200097806A1 · Chen et al. · 2020 [cited by applicant]
US 20200097826A1 · Du et al. · 2020 [cited by applicant]
US 20200097827A1 · Wang et al. · 2020 [cited by applicant]
US 20200097828A1 · Du et al. · 2020 [cited by applicant]
US 20200097831A1 · Wang et al. · 2020 [cited by applicant]
US 20200265301A1 · Burger · 2020 [cited by examiner]
US 20210182077A1 · Chen et al. · 2021 [cited by applicant]
US 20210224069A1 · Chen et al. · 2021 [cited by applicant]
CN 105512723A · 2016 [cited by applicant]
CN 106485316A · 2017 [cited by applicant]
CN 106548234A · 2017 [cited by applicant]
CN 106919942A · 2017 [cited by applicant]
CN 106991477A · 2017 [cited by applicant]
Kadetotad et al. (“Efficient Memory Compression in Deep Neural Networks Using Coarse-Grain Sparsification for Speech Applications”, ICCAD '16, Nov. 7-10, 2016) (Year: 2016). [cited by examiner]
Papandreou et al. (“Modeling Local and Global Deformations in Deep Learning: Epitomic Convolution, Multiple Instance Learning, and Sliding Window Detection”, IEEE, CVPR 2015). (Year: 2016). [cited by examiner]
Mao et al. (“Exploring the Regularity of Sparse Structure in Convolutional Neural Networks”, arXiv, 2017) (Year: 2017). [cited by examiner]
Anwar et al. (“Compact Deep Convolutional Neural Networks with Coarse Pruning”, ICLR 2016) (Year: 2016). [cited by examiner]
Zeng Dan et al.: “Compressing Deep Neural Network for Facial Landmarks Detection”, International Conference On Financial Cryptography and Data Security, Nov. 13, 2016, 11 Pages. [cited by applicant]
Song Han et al; “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman D2 Coding”, Internet: URL:https://arxiv.org/pdf/1510.00149v5.pdf; Feb. 15, 2016; 14 pages. [cited by applicant]
Fujii Tomoya et al; “An FPGA Realization of a Deep Convolutional Neural Network Using a Threshold Neuron Pruning”, International Conference On Financial Cryptography and Data Security; Mar. 31, 2017; 13 pages. [cited by applicant]
Sun Fangxuan et al.; “Intra-layer nonuniform quantization of convolutional neural network” 2016 8th International Conference On Wireless Communications & Signal Processing (WCSP), IEEE, Oct. 13, 2016, 5 pages. [cited by applicant]
Song Han et al.: “ESE: Efficient Speech Recognition Engine with Sparse LSTM on FPGA”, Proceedings of The 2017 ACM/SIGDA International Symposium On Field-Programmable Gate Arrays; Feb. 17, 2017; 10 pages. [cited by applicant]
Sajid Anwar et al; “Structured Pruning of Deep Convolutional Neural Networks”, ACM Journal On Emerging Technologies in Computing Systems; Feb. 9, 2017, 18 pages. [cited by applicant]
Yunchao Gong et al.; “Compressing Deep Convolutional Networks using Vector Quantization”, Internet: URL:https://arxiv.org/pdf/1412.6115.pdf ; Dec. 18, 2014, 10 pages. [cited by applicant]
Kadetotad Deepak et al.; “Efficient memory compression in deep neural networks using coarse-grain sparsification for speech applications”, 2016 IEEE/ACM International Conference On Computer-Aided Design, Nov. 7, 2016, 8… [cited by applicant]
Song Han et al.; “Leraning both weights and connections for efficient neural networks” Published as a conference paper at NIPS 2015; Internet URL: https://arxiv.org/abs/1506.02626; Oct. 30, 2015; 9 pages. [cited by applicant]
EP 18806558.5, European Search Report mailed Apr. 24, 2020, 13 pages. [cited by applicant]
EP 19214007.7, European Search Report mailed Apr. 15, 2020, 12 pages. [cited by applicant]
EP 19214010.1, European Search Report mailed Apr. 21, 2020, 12 pages. [cited by applicant]
EP 19214015.0, European Search Report mailed Apr. 21, 2020, 14 pages. [cited by applicant]
PCT/CN2018/088033—Search Report, mailed Aug. 21, 2018, 11 pages. (no English translation). [cited by applicant]
Moons, Bert, et al. “Energy-Efficient ConvNets Through Approximate Computing”, arXiv: 1603.06777v1, Mar. 22, 2016, 8 pages. [cited by applicant]
EP 19 214 010.1, Communication pursuant to Article 94(3), mailed Jan. 3, 2022, 11 pages. [cited by applicant]
Huang, Hongmei, et al, “Fault Prediction Method Based on RBF Network On-Line Learning”, Journal of Nanjing University of Aeronautics & Astronautics, vol. 39 No. 2, Apr. 2007, 4 pages. [cited by applicant]
CN 201710370905.1—Second Office Action, mailed Mar. 18, 2021,11 pages. (with English translation). [cited by applicant]
CN 201710583336.9—First Office Action, mailed Apr. 23, 2020, 15 pages. (with English translation). [cited by applicant]
CN 201710677987.4—Third Office Action, mailed Mar. 30, 2021, 16 pages. (with English translation). [cited by applicant]
CN 201710678038.8—First Office Action, mailed Oct. 10, 2020, 12 pages. (with English translation). [cited by applicant]
CN 201710689666.6—First Office Action, mailed Jun. 23, 2020, 19 pages. (with English translation). [cited by applicant]
CN 201710689666.6—Second Office Action, mailed Feb. 3, 2021, 18 pages. (with English translation). [cited by applicant]
CN 201710689595 X—First Office Action, mailed Sep. 27, 2020, 21 pages. (with English translation). [cited by applicant]
Liu, Shaoli, et al., “Cambricon: An Instruction Set Architecture for Neural Networks”, ACM/IEEE, 2016, 13 pages. [cited by applicant]
CN 201710689595.X—Second Office Action, mailed Jun. 6, 2021, 17 pages. (with English translation). [cited by applicant]
EP 18 806 558.5, Communication pursuant to Article 94(3), mailed Dec. 9, 2021, 11 pages. [cited by applicant]
EP 19 214 007.7, Communication pursuant to Article 94(3), mailed Dec. 8, 2021, 10 pages. [cited by applicant]
EP 19 214 015.0, Communication pursuant to Article 94(3), mailed Jan. 3, 2022, 12 pages. [cited by applicant]
PCT/CN2018/088033—Search Report, mailed Aug. 21, 2018, 19 pages. (with English translation). [cited by applicant]
U.S. Appl. No. 16/699,049—Final Office Action mailed on Apr. 11, 2024, 27 pages. [cited by applicant]
U.S. Appl. No. 16/699,055—Final Office Action mailed on Apr. 11, 2024, 31 pages. [cited by applicant]
Ahalt et al., “Competitive Learning Algorithms for Vector Quantization”, Neural Networks, vol. 3, Issue 3, pp. 277-290, 1990. [cited by applicant]
Choi et al., “Towards the Limit of Network Quantization”, ICLR 2017, Apr. 2017, pp. 1-14. [cited by applicant]
Chu et al., “Vector Quantization of Neural Networks”, IEEE Transactions on Neural Networks, vol. 9, No. 6, Nov. 1998, pp. 1235-1245. [cited by applicant]
Han et al., “EIE: Efficient Inference Engine On Compressed Deep Neural Network”, Available online at https://arxiv.org/pdf/1602.01528.pdf, May 3, 2016, 12 Pages. [cited by applicant]
He et al., “Effective Quantization Methods for Recurrent Neural Networks”, Available online at https://arxiv.org/pdf/1611.10176.pdf, Nov. 30, 2016, pp. 1-10. [cited by applicant]
Judd et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-free Deep Neural Network Computing”, Available online at https://arxiv.org/pdf/1705.00125.pdf, Apr. 29, 2017, pp. 1-6. [cited by applicant]
Lane et al., “Squeezing Deep Learning into Mobile and Embedded Devices”, IEEE Pervasive Computing, vol. 16, Issue 3, Jul. 27, 2017, pp. 82-88. [cited by applicant]
Mao et al., “Exploring the Granularity of Sparsity in Convolutional Neural Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Jul. 21-26, 2017, pp. 1927-1934. [cited by applicant]
Parashar et al., “SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks”, 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), Jun. 24-28, 2017, pp. 27-40. [cited by applicant]
U.S. Appl. No. 16/699,029—Non-Final Office Action mailed on Oct. 6, 2022, 15 pages. [cited by applicant]
U.S. Appl. No. 16/699,029—Notice of Allowance mailed on Mar. 14, 2023, 8 pages. [cited by applicant]
U.S. Appl. No. 16/699,032—Corrected Notice of Allowability mailed on Jan. 9, 2024, 2 pages. [cited by applicant]
U.S. Appl. No. 16/699,032—Non-Final Office Action mailed on Jul. 1, 2022, 15 pages. [cited by applicant]
U.S. Appl. No. 16/699,032—Notice of Allowance mailed on Oct. 25, 2023, 5 pages. [cited by applicant]
U.S. Appl. No. 16/699,046—Non-Final Office Action mailed on Sep. 22, 2022, 15 pages. [cited by applicant]
U.S. Appl. No. 16/699,046—Notice of Allowance mailed on Mar. 30, 2023, 9 pages. [cited by applicant]
U.S. Appl. No. 16/699,049—Non-Final Office Action mailed on Jun. 2, 2023, 25 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Final Office Action mailed on Feb. 15, 2024, 35 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Final Office Action mailed on Sep. 27, 2022, 36 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Non-Final Office Action mailed on Mar. 3, 2022, 32 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Non-Final Office Action mailed on May 8, 2023, 37 pages. [cited by applicant]
U.S. Appl. No. 16/699,055—Non-Final Office Action mailed on Jun. 8, 2023, 31 pages. [cited by applicant]
U.S. Appl. No. 62/486,432—Enhanced Neural Network Designs, filed on Apr. 17, 2017, 69 pages. [cited by applicant]
Yang et al., “Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21-26, 2017, pp. 6071-6079. [cited by applicant]
Yu et al., “Scalpel: Customizing DNN Pruning to the Underlying Hardware Parallelism”, 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), Jun. 24-28, 2017, pp. 548-560. [cited by applicant]
Zhou et al., “Cambricon-S: Addressing Irregularity in Sparse Neural Networks through A Cooperative Software/Hardware Approach”, 51st Annual IEEE/ACM International Symposium on Microarchitecture, Oct. 20, 2018, pp. 15-28. [cited by applicant]
EP19214015.0—Communication pursuant to Article 94(3) EPC mailed on Jun. 28, 2024, 7 pages. [cited by applicant]
U.S. Appl. No. 16/699,049—Non-Final Office Action mailed on Dec. 12, 2024, 29 pages. [cited by applicant]
U.S. Appl. No. 16/699,055—Non-Final Office Action mailed on Dec. 12, 2024, 33 pages,. [cited by applicant]
EP19214010.1—Communication pursuant to Article 94(3) EPC mailed on Dec. 9, 2024, 11 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Non-Final Office Action mailed on Sep. 9, 2024, 30 pages. [cited by applicant]
Lee et al., “Adaptive Vector Quantization Using a Self-development Neural Network”, IEEE Journal on Selected Areas in Communications, vol. 8, No. 8, Oct. 1990, pp. 1458-1471. [cited by applicant]
EP19214007.7—Communication pursuant to Article 94(3) EPC mailed on Apr. 23, 2024, 7 pages. [cited by applicant]
EP18806558.5—Summons to Attend Oral Proceedings mailed on Apr. 11, 2024, 13 pages. [cited by applicant]
EP19214010.1—Communication pursuant to Article 94(3) mailed on Jun. 28, 2024, 4 pages,. [cited by applicant]
U.S. Appl. No. 16/699,049—Advisory Action mailed on Jul. 30, 2024, 3 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Final Office Action mailed on Jun. 11, 2025, 37 pages. [cited by applicant]