IP Library › Granted Patent US 12,488,247
Granted Patent B2
US 12,488,247 · App. 16/699,049 · Granted Dec 2, 2025

Processing method and accelerating device

Inventors: Zidong Du (Pudong New Area, CN); Xuda Zhou (Pudong New Area, CN); Zai Wang (Pudong New Area, CN); Tianshi Chen (Pudong New Area, CN)
Assignee: SHANGHAI CAMBRICON INFORMATION TECHNOLOGY CO., LTD.
G06N3/082G06F1/3296G06F9/3877G06F12/0875G06F13/16G06F16/285G06N3/04G06N3/044G06N3/048G06N3/063G06N3/084G06F2212/452G06F2213/0026
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,247
App. No.
16/699,049
Granted
Dec 2, 2025
Kind
B2
Abstract

The present disclosure provides a processing device including: a coarse-grained pruning unit configured to perform coarse-grained pruning on a weight of a neural network to obtain a pruned weight, an operation unit configured to train the neural network according to the pruned weight. The coarse-grained pruning unit is specifically configured to select M weights from the weights of the neural network through a sliding window, and when the M weights meet a preset condition, all or part of the M weights may be set to 0. The processing device can reduce the memory access while reducing the amount of computation, thereby obtaining an acceleration ratio and reducing energy consumption.

Claims (59)

1. A data compression method, comprising: performing coarse-grained pruning on weights of a neural network, which includes: selecting one or more weights of the weights from the neural network through a sliding window, and setting all or part of the one or more weights to 0 when the one or more weights meet a preset condition, wherein the preset condition is a condition in which information quantity of the one or more weights is less than a first given threshold, and the information quantity is an arithmetic mean of an absolute value of the one or more weights, a geometric mean of the absolute value of the one or more weights, or a maximum value of the one or more weights; performing a first retraining on the neural network, where the one or more weights which have been set to 0 remain 0 in the first retraining; repeating selecting the one or more weights from the neural network through the sliding window; setting all or part of the one or more weights to 0 when the one or more weights meet the preset condition; and performing the first retraining on the neural network until no weight of the one or more weights is set to 0 without losing a preset precision; quantizing the weights of the neural network, which includes: grouping the weights of the neural network, performing a clustering operation on each group of the weights by using a clustering algorithm to generate multiple classes of the weights, computing a center weight of each class of the multiple classes, and replacing all the weights in each class of the multiple classes by the center weights; encoding the center weights and center weights codes to generate a weight codebook, wherein the weight codebook includes the center weights, the center weights codes, and a correspondence between the center weights and the respective center weight codes; and encoding the center weight codes to generate a weight dictionary.

2. The data compression method of claim 1 , wherein, after the encoding of the center weights, the method further includes:

performing a second retraining on the neural network, wherein,

only the weight codebook is trained during the second retraining of the neural network, and

the weight dictionary remains unchanged.

3. The data compression method of claim 1 , wherein

the first given threshold is a first threshold, a second threshold, or a third threshold; and

the information quantity of the one or more weights being less than the first given threshold includes:

the arithmetic mean of the absolute value of the one or more weights being less than the first threshold, or the geometric mean of the absolute value of the one or more weights being less than the second threshold, or the maximum value of the one or more weights being less than the third threshold.

4. The data compression method of claim 1 , wherein performing the coarse-grained pruning on the weights of a fully connected layer of the neural network includes: setting the weights of the fully connected layer being a two-dimensional matrix (Nin, Nout), where Nin represents a count of input neurons and Nout represents a count of output neurons, and the fully connected layer has Nin*Nout weights; setting a size of the sliding window being Bin*Bout, where Bin is a positive integer greater than 0 and less than or equal to Nin, and Bout is a positive integer greater than 0 and less than or equal to Nout; making the sliding window slide Sin stride in a direction of Bin, or slide Sout stride in a direction of Bout, where Sin is a positive integer greater than 0 and less than or equal to Bin, and Sout is a positive integer greater than 0 and less than or equal to Bout; and selecting the one or more weights from Nin*Nout weights through the sliding window, and when the one or more weights meet the preset condition, all or part of the one or more weights are set to 0, where M=Bin*Bout.

5. The data compression method of claim 1 , wherein performing the coarse-grained pruning on the weights of a convolutional layer of the neural network includes: setting the weights of the convolutional layer of the neural network being a four-dimensional matrix (Nfin, Nfout, Kx, Ky), where Nfin represents a count of input feature maps, Nfout represents a count of output feature maps, (Kx, Ky) is a size of a convolution kernel, and the convolutional layer has Nfin*Nfout*Kx*Ky weights; setting the sliding window being a four-dimensional sliding window with a size of Bfin*Bfout*Bx*By, where Bfin is a positive integer greater than 0 and less than or equal to Nfin, Bfout is a positive integer greater than 0 and less than or equal to Nfout, Bx is a positive integer greater than 0 and less than or equal to Kx, and By is a positive integer greater than 0 and less than or equal to Ky; making the sliding window slide Sfin stride in a direction of Bfin, or slide Sfout stride in a direction of Bfout, or slide S stride in a direction of Bx, or slide Sy stride in a direction of By, where Sfin is a positive integer greater than 0 and less than or equal to Bfin, Sfout is a positive integer greater than 0 and less than or equal to Bfout, Sx is a positive integer greater than 0 and less than or equal to Bx, and Sy is a positive integer greater than 0 and less than or equal to By; and selecting the one or more weights from Nfin*Nfout*Kx*Ky weights through the sliding window, and when the one or more weights meet the preset condition, all or part of the one or more weights are set to 0, where M=Bfin*Bfout*Bx*By.

6. The data compression method of claim 1 , wherein performing the coarse-grained pruning on the weights of a LSTM layer of the neural network includes: setting the weights of the LSTM layer being composed of a plurality weights of a fully connected layer, and an i th weight of the fully connected layer is a two-dimensional matrix (Nin_i, Nout_ 1 ), where 1 is a positive integer greater than 0 and less than or equal to m, Nin_i represents a count of input neurons of the i weight of the fully connected layer, and Nout_i represents a count of output neurons of the i weight of the fully connected layer; setting a size of the sliding window being Bin_i*Bout_ 1 , where Bin_ 1 is a positive integer greater than 0 and less than or equal to Nin_i, and Bout_i is a positive integer greater than 0 and less than or equal to Nout_ 1 ; making the sliding window slide Sin_i stride in a direction of Bin_ 1 , or slide Sout_ 1 stride in a direction of Bout_i 1 , where Sin_ 1 is a positive integer greater than 0 and less than or equal to Bin i, and Sout_i is a positive integer greater than 0 and less than or equal to Bout_i; and selecting one or more weights from Bin i*Bout_i weights through the sliding window, and when the one or more weights meet the preset condition, all or part of the one or more weights are set to 0, where M=Bin_i*Bout_ 1 .

7. The data compression method of claim 1 , wherein the first retraining adopts a back propagation algorithm, and the one or more weights which have been set to 0 remain 0 in the first retraining.

8. The data compression method of claim 1 ,

wherein the grouping of the weights of the neural network includes grouping into a group, layer-type-based grouping, inter-layer-based grouping, and/or intra-layer-based grouping,

wherein the grouping into a group includes grouping all the weights of the neural network into a group,

grouping the weights of the neural network according to the layer-type-based grouping method includes:

grouping the weights of all convolutional layers, the weights of all fully connected layers, and the weights of all LSTM layers in the neural network into one group respectively,

wherein the grouping of the weights of the neural network by the inter-layer-based grouping method includes:

grouping the weights of one or a plurality of convolutional layers, one or a plurality of fully connected layers and one or a plurality of LSTM layers in the neural network into one group respectively,

wherein the grouping of the weights of the neural network by the intra-layer-based grouping method includes:

segmenting the weights in one layer of the neural network, where each segmented part forms a group.

9. The data compression method of claim 1 , wherein the clustering algorithm includes K-means, K-medoids, Clara, and/or Clarans.

10. The data compression method of claim 1 , wherein a center weight selection method of a class is: minimizing a cost function J(w,w 0 ).

11. The data compression method of claim 10 , wherein the cost function meets a condition:

J

⁡

(

w

,

w

0

)

=

∑

i

=

1

n

⁢

⁢

(

w

i

-

w

0

)

2

where w is all weights of a class, w0 is a center weight of the class, n is a count of weights in the class, wi is an i th weight of the class, and i is a positive integer greater than 0 and less than or equal to n.

12. The data compression method of claim 2 , wherein the second retraining performed on the neural network after the clustering and the encoding includes:

performing the second retraining on the neural network after the clustering, and the encoding by using a back propagation algorithm, where the one or more weights that have been set to 0 in the second retraining remains 0 all time, and only the weight codebook is retrained, the weight dictionary is not retrained.

13. A data compression device, comprising: a memory configured to store an operation instruction; and a processor configured to: perform coarse-grained pruning on weights of a neural network, which includes: select one or more weights of the weights from the neural network through a sliding window, and set all or part of the one or more weights to 0 when the one or more weights meet a preset condition, wherein the preset condition is a condition in which information quantity of the one or more weights is less than a first given threshold, and the information quantity is an arithmetic mean of an absolute value of the one or more weights, a geometric mean of the absolute value of the one or more weights, or a maximum value of the one or more weights; perform a first retraining on the neural network, where the one or more weights which have been set to 0 remain 0 in the first retraining; repeat selection of the one or more weights from the neural network through the sliding window; set all or part of the one or more weights to 0 when the one or more weights meet the preset condition; and perform the first retraining on the neural network until no weight of the one or more weights is set to 0 without losing a preset precision; quantize the weights of the neural network, wherein the processor is further configured to: group the weights of the neural network, perform a clustering operation on each group of the weights by using a clustering algorithm to generate multiple classes of the weights, compute a center weight of each class of the multiple classes, and replace all the weights in each class of the multiple classes by the center weights; encode the center weights and center weights codes to generate a weight codebook, wherein the weight codebook includes the center weights, the center weights codes, and a correspondence between the center weights and the respective center weight codes; and encode the center weight codes to generate a weight dictionary.

14. The data compression device of claim 13 , the processor is further configured to perform a second retraining on the neural network, wherein, only the weight codebook is trained during the second retraining of the neural network, and the weight dictionary remains unchanged.

15. The data compression device of claim 13 , wherein

the first given threshold is a first threshold, a second threshold, or a third threshold; and

the information quantity of the one or more weights being less than the first given threshold includes:

the arithmetic mean of the absolute value of the one or more weights being less than the first threshold, or the geometric mean of the absolute value of the one or more weights being less than the second threshold, or the maximum value of the one or more weights being less than the third threshold.

16. An electronic device, comprising: a data compression device that includes: a memory configured to store an operation instruction; and a processor configured to: perform coarse-grained pruning on weights of a neural network, which includes: select one or more weights of the weights from the neural network through a sliding window, and set all or part of the one or more weights to 0 when the one or more weights meet a preset condition, wherein the preset condition is a condition in which information quantity of the one or more weights is less than a first given threshold, and the information quantity is an arithmetic mean of an absolute value of the one or more weights, a geometric mean of the absolute value of the one or more weights, or a maximum value of the one or more weights; perform a first retraining on the neural network, where the one or more weights which have been set to 0 remains 0 in the first retraining; repeat selection of the one or more weights from the neural network through the sliding window; set all or part of the one or more weights to 0 when the one or more weights meet the preset condition; and perform the first retraining on the neural network until no weight of the one or more weights is set to 0 without losing a preset precision; quantize the weights of the neural network, wherein the processor is further configured to: group the weights of the neural network, perform a clustering operation on each group of the weights by using a clustering algorithm to generate multiple classes of the weights, compute a center weight of each class of the multiple classes, and replace all the weights in each class of the multiple classes by the center weights; encode the center weights and center weights codes to generate a weight codebook, wherein the weight codebook includes the center weights, the center weights codes, and a correspondence between the center weights and the respective center weight codes; and encode the center weight codes to generate a weight dictionary.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2019
From: DU, ZIDONG; ZHOU, XUDA; WANG, ZAI; CHEN, TIANSHI
To: SHANGHAI CAMBRICON INFORMATION TECHNOLOGY CO., LTD
Reel/Frame 051136/0345 →
Priority Claims (1)
CN 201710677987.4 · Aug 9, 2017 · national
Continuity (3)
Continuation 16699027 · Nov 28, 2019
Continuation In Part PCTCN2018088033 · May 23, 2018
Related Publication 20200134460A1 · Apr 30, 2020
References Cited (111)
US 4422141A · Shoji · 1983 [cited by applicant]
US 6360019B1 · Chaddha · 2002 [cited by applicant]
US 6772126B1 · Simpson et al. · 2004 [cited by applicant]
US 10127495B1 · Bopardikar et al. · 2018 [cited by applicant]
US 10657439B2 · Liu et al. · 2020 [cited by applicant]
US 11315018B2 · Molchanov et al. · 2022 [cited by applicant]
US 20030204311A1 · Bush · 2003 [cited by applicant]
US 20040180690A1 · Song et al. · 2004 [cited by applicant]
US 20050102301A1 · Flanagan · 2005 [cited by applicant]
US 20080249767A1 · Ertan · 2008 [cited by applicant]
US 20120189047A1 · Jiang et al. · 2012 [cited by applicant]
US 20140300758A1 · Tran · 2014 [cited by applicant]
US 20150332690A1 · Kim · 2015 [cited by examiner]
US 20160358069A1 · Brothers · 2016 [cited by examiner]
US 20170061328A1 · Majumdar · 2017 [cited by examiner]
US 20170270408A1 · Shi et al. · 2017 [cited by applicant]
US 20180046900A1 · Dally · 2018 [cited by examiner]
US 20180075336A1 · Huang · 2018 [cited by examiner]
US 20180114114A1 · Molchanov et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180197081A1 · Ji · 2018 [cited by examiner]
US 20180285731A1 · Heifets et al. · 2018 [cited by applicant]
US 20180300603A1 · Ambardekar · 2018 [cited by examiner]
US 20180314940A1 · Kundu et al. · 2018 [cited by applicant]
US 20190019311A1 · Hu et al. · 2019 [cited by applicant]
US 20190050709A1 · Yang · 2019 [cited by examiner]
US 20190362235A1 · Xu et al. · 2019 [cited by applicant]
US 20200097806A1 · Chen et al. · 2020 [cited by applicant]
US 20200097826A1 · Du et al. · 2020 [cited by applicant]
US 20200097827A1 · Wang et al. · 2020 [cited by applicant]
US 20200097828A1 · Du et al. · 2020 [cited by applicant]
US 20200097831A1 · Wang et al. · 2020 [cited by applicant]
US 20200265301A1 · Burger · 2020 [cited by examiner]
US 20210182077A1 · Chen et al. · 2021 [cited by applicant]
US 20210224069A1 · Chen et al. · 2021 [cited by applicant]
CN 105512723A · 2016 [cited by applicant]
CN 106485316A · 2017 [cited by applicant]
CN 106548234A · 2017 [cited by applicant]
CN 106919942A · 2017 [cited by applicant]
CN 106991477A · 2017 [cited by applicant]
Kadetotad et al. (“Efficient Memory Compression in Deep Neural Networks Using Coarse-Grain Sparsification for Speech Applications”, ICCAD, 2016) (Year: 2016). [cited by examiner]
Papandreou et al. (“Modeling Local and Global Deformations in Deep Learning: Epitomic Convolution, Multiple Instance Learning, and Sliding Window Detection”, IEEE, 2015) (Year: 2015). [cited by examiner]
Han et al. (“Deep compression: compressing deep neural networks with pruning, trained quantization and Huffman coding”, ICLR 2016) (Year: 2016). [cited by examiner]
Zeng Dan et al.: “Compressing Deep Neural Network for Facial Landmarks Detection”, International Conference on Financial Cryptography and Data Security, Nov. 13, 2016, 11 Pages. [cited by applicant]
Song Han et al; “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman D2 Coding”, Internet: URL:https://arxiv.org/pdf/1510.00149v5.pdf; Feb. 15, 2016; 14 pages. [cited by applicant]
Fujii Tomoya et al; “An FPGA Realization of a Deep Convolutional Neural Network Using a Threshold Neuron Pruning”, International Conference on Financial Cryptography and Data Security; Mar. 31, 2017; 13 pages. [cited by applicant]
Sun Fangxuan et al.; “Intra-layer nonuniform quantization of convolutional neural network” 2016 8th International Conference on Wireless Communications & Signal Processing (WCSP), IEEE, Oct. 13, 2016, 5 pages. [cited by applicant]
Song Han et al.: “ESE: Efficient Speech Recognition Engine with Sparse LSTM on FPGA”, Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays; Feb. 17, 2017; 10 pages. [cited by applicant]
Sajid Anwar et al; “Structured Pruning of Deep Convolutional Neural Networks”, ACM Journal on Emerging Technologies in Computing Systems; Feb. 9, 2017, 18 pages. [cited by applicant]
Yunchao Gong et al.; “Compressing Deep Convolutional Networks using Vector Quantization”, Internet: URL:https://arxiv.org/pdf/1412.6115.pdf ; Dec. 18, 2014, 10 pages. [cited by applicant]
Kadetotad Deepak et al.; “Efficient memory compression in deep neural networks using coarse-grain sparsification for speech applications”, 2016 IEEE/ACM International Conference on Computer-Aided Design, Nov. 7, 2016, 8… [cited by applicant]
Song Han et al.; “Leraning both weights and connections for efficient neural networks” Published as a conference paper at NIPS 2015; Internet URL: https://arxiv.org/abs/1506.02626; Oct. 30, 2015; 9 pages. [cited by applicant]
EP 18806558.5, European Search Report mailed Apr. 24, 2020, 13 pages. [cited by applicant]
EP 19214007.7, European Search Report mailed Apr. 15, 2020, 12 pages. [cited by applicant]
EP 19214010.1, European Search Report mailed Apr. 21, 2020, 12 pages. [cited by applicant]
EP 19214015.0, European Search Report mailed Apr. 21, 2020, 14 pages. [cited by applicant]
Moons, Bert, et. al. “Energy-Efficient ConvNets Through Approximate Computing”, arXiv:1603.06777v1, Mar. 22, 2016, 8 pages. [cited by applicant]
EP 19 214 010.1, Communication pursuant to Article 94(3), mailed Jan. 3, 2022, 11 pages. [cited by applicant]
Huang, Hongmei, et al, “Fault Prediction Method Based on RBF Network On-Line Learning”, Journal of Nanjing University of Aeronautics & Astronautics, vol. 39 No. 2, Apr. 2007, 4 pages. [cited by applicant]
CN 201710370905.1—Second Office Action, mailed Mar. 18, 2021,11 pages. (with English translation). [cited by applicant]
CN 201710583336.9—First Office Action, mailed Apr. 23, 2020, 15 pages. (with English translation). [cited by applicant]
CN 201710677987.4—Third Office Action, mailed Mar. 30, 2021, 16 pages. (with English translation). [cited by applicant]
CN 201710678038.8—First Office Action, mailed Oct. 10, 2020, 12 pages. (with English translation). [cited by applicant]
CN 201710689666.6—First Office Action, mailed Jun. 23, 2020, 19 pages. (with English translation). [cited by applicant]
CN 201710689666.6—Second Office Action, mailed Feb. 3, 2021, 18 pages. (with English translation). [cited by applicant]
CN 201710689595.X—First Office Action, mailed Sep. 27, 2020, 21 pages. (with English translation). [cited by applicant]
Liu, Shaoli, et al., “Cambricon: An Instruction Set Architecture for Neural Networks”, ACM/IEEE, 2016, 13 pages. [cited by applicant]
CN 201710689595.X—Second Office Action, mailed Jun. 6, 2021, 17 pages. (with English translation). [cited by applicant]
EP 18 806 558.5, Communication pursuant to Article 94(3), mailed Dec. 9, 2021, 11 pages. [cited by applicant]
EP 19 214 007.7, Communication pursuant to Article 94(3), mailed Dec. 8, 2021, 10 pages. [cited by applicant]
EP 19 214 015.0, Communication pursuant to Article 94(3), mailed Jan. 3, 2022, 12 pages. [cited by applicant]
PCT/CN2018/088033—Search Report, mailed Aug. 21, 2018, 19 pages. (with English translation). [cited by applicant]
Ahalt et al., “Competitive Learning Algorithms for Vector Quantization”, Neural Networks, vol. 3, Issue 3, pp. 277-290, 1990. [cited by applicant]
Anwar et al. “Compact Deep Convolutional Neural Networks with Coarse Pruning”, Department of Electrical Engineering and Computer Science Seoul National University, ICLR, Oct. 30, 2016, pp. 1-10. [cited by applicant]
Choi et al., “Towards the Limit of Network Quantization”, ICLR 2017, Apr. 2017, pp. 1-14. [cited by applicant]
Chu et al., “Vector Quantization of Neural Networks”, IEEE Transactions on Neural Networks, vol. 9, No. 6, Nov. 1998, pp. 1235-1245. [cited by applicant]
Han et al., “EIE: Efficient Inference Engine on Compressed Deep Neural Network”, Available online at https://arxiv.org/pdf/1602.01528.pdf, May 3, 2016, 12 Pages. [cited by applicant]
He et al., “Effective Quantization Methods for Recurrent Neural Networks”, Available online at https://arxiv.org/pdf/1611.10176.pdf, Nov. 30, 2016, pp. 1-10. [cited by applicant]
Judd et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-free Deep Neural Network Computing”, Available online at https://arxiv.org/pdf/1705.00125.pdf, Apr. 29, 2017, pp. 1-6. [cited by applicant]
Lane et al., “Squeezing Deep Learning into Mobile and Embedded Devices”, IEEE Pervasive Computing, vol. 16, Issue 3, Jul. 27, 2017, pp. 82-88. [cited by applicant]
Mao et al., “Exploring the Granularity of Sparsity in Convolutional Neural Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Jul. 21-26, 2017, pp. 1927-1934. [cited by applicant]
Mao et al., “Exploring the Regularity of Sparse Structure in Convolutional Neural Networks”, Available online at https://arxiv.org/pdf/1705.08922.pdf, May 24, 2017, pp. 1-10. [cited by applicant]
Parashar et al., “SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks”, 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), Jun. 24-28, 2017, pp. 27-40. [cited by applicant]
U.S. Appl. No. 16/699,027—Non-Final Office Action mailed on Apr. 3, 2023, 22 pages. [cited by applicant]
U.S. Appl. No. 16/699,029—Non-Final Office Action mailed on Oct. 6, 2022, 15 pages. [cited by applicant]
U.S. Appl. No. 16/699,029—Notice of Allowance mailed on Mar. 14, 2023, 8 pages. [cited by applicant]
U.S. Appl. No. 16/699,032—Corrected Notice of Allowability mailed on Jan. 9, 2024, 2 pages. [cited by applicant]
U.S. Appl. No. 16/699,032—Non-Final Office Action mailed on Jul. 1, 2022, 15 pages. [cited by applicant]
U.S. Appl. No. 16/699,032—Notice of Allowance mailed on Oct. 25, 2023, 5 pages. [cited by applicant]
U.S. Appl. No. 16/699,046—Non-Final Office Action mailed on Sep. 22, 2022, 15 pages. [cited by applicant]
U.S. Appl. No. 16/699,046—Notice of Allowance mailed on Mar. 30, 2023, 9 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Final Office Action mailed on Feb. 15, 2024, 35 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Final Office Action mailed on Sep. 27, 2022, 36 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Non-Final Office Action mailed on Mar. 3, 2022, 32 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Non-Final Office Action mailed on May 8, 2023, 37 pages. [cited by applicant]
U.S. Appl. No. 16/699,055—Non-Final Office Action mailed on Jun. 8, 2023, 31 pages. [cited by applicant]
U.S. Appl. No. 62/486,432—Enhanced Neural Network Designs, filed Apr. 17, 2017, 69 pages. [cited by applicant]
Yang et al., “Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21-26, 2017, pp. 6071-6079. [cited by applicant]
Yu et al., “Scalpel: Customizing DNN Pruning to the Underlying Hardware Parallelism”, 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), Jun. 24-28, 2017, pp. 548-560. [cited by applicant]
Zhou et al., “Cambricon-S: Addressing Irregularity in Sparse Neural Networks through a Cooperative Software/Hardware Approach”, 51st Annual IEEE/ACM International Symposium on Microarchitecture, Oct. 20, 2018, pp. 15-28. [cited by applicant]
U.S. Appl. No. 16/699,055—Final Office Action mailed on Apr. 11, 2024, 31 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Non-Final Office Action mailed on Sep. 9, 2024, 30 pages. [cited by applicant]
Lee et al., “Adaptive Vector Quantization Using a Self-development Neural Network”, IEEE Journal on Selected Areas in Communications, vol. 8, No. 8, Oct. 1990, pp. 1458-1471. [cited by applicant]
U.S. Appl. No. 16/699,027—Final Office Action mailed on May 21, 2024, 17 pages. [cited by applicant]
EP19214007.7—Communication pursuant to Article 94(3) EPC mailed on Apr. 23, 2024, 7 pages. [cited by applicant]
EP18806558.5—Summons to attend oral proceedings mailed on Apr. 11, 2024, 13 pages. [cited by applicant]
EP19214015.0—Communication pursuant to Article 94(3) mailed on Jun. 28, 2024, 7 pages. [cited by applicant]
EP19214010.1—Communication pursuant to Article 94(3) mailed on Jun. 28, 2024, 4 pages. [cited by applicant]
U.S. Appl. No. 16/699,055—Non-Final Office Action mailed on Dec. 12, 2024, 33 pages. [cited by applicant]
U.S. Appl. No. 16/699,051—Final Office Action mailed on Jun. 11, 2025, 37 pages. [cited by applicant]
EP19214010.1—Communication pursuant to Article 94(3) EPC mailed on Dec. 9, 2024, 11 pages. [cited by applicant]