IP Library Granted Patent US 12,572,798
Granted Patent B1
US 12,572,798 · App. 17/696,819 · Granted Mar 10, 2026

Accounting for compute time in training of network

Inventors: Eric A. Sather (Palo Alto, CA); Steven L. Teig (Menlo Park, CA)
Assignee: Amazon Technologies, Inc.
G06N3/08G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,798
App. No.
17/696,819
Granted
Mar 10, 2026
Kind
B1
Abstract

Some embodiments provide a method for training a machine-trained (MT) network. The method receives a network having multiple layers. Each layer of a set of the layers includes multiple weight values. The method trains the network by alternately ( 1 ) propagating inputs through the network to generate outputs and adjusting the weight values based on differences between the generated outputs and expected outputs and ( 2 ) identifying sets of the weight values for removal according to a set of constraints that accounts for (i) a total number of weight values and (ii) an amount of time required to execute the network on a particular type of integrated circuit.

Claims (43)

1 . A method for training a machine-trained (MT) network, the method comprising:

receiving, for execution on a particular type of integrated circuit, a network comprising a plurality of layers, the plurality of layers comprising a set of prunable layers, wherein each layer of the set of prunable layers comprises a respective plurality of weight values; and

training the network by iteratively:

propagating inputs through the network to generate outputs and adjusting the respective plurality of weight values for each layer of the plurality of layers based on differences between the outputs and expected outputs;

determining, using the respective plurality of weight values for each layer of the plurality of layers, an estimated compute time for the network;

identifying, in response to the estimated compute time exceeding a compute time constraint and based on a parallelism computation capability of the particular type of integrated circuit, sets of weight values of the respective plurality of weight values for each layer of the set of prunable layers for removal; and

removing the sets of weight values from consideration during a subsequent training iteration of the network, wherein removing the sets of weight values reduces the estimated compute time.

2 . The method of claim 1 , wherein the respective plurality of weight values for reach of the set of prunable layers are arranged as filters with corresponding scales, wherein identifying the sets of weight values for removal comprises identifying filters for removal.

3 . The method of claim 2 , wherein identifying the sets of weight values for removal comprises:

projecting each corresponding scale of the corresponding scales to a respective set of states, wherein each corresponding scale is projected to one of (i) a first state in which the corresponding scale is set to zero and (ii) a second state in which the corresponding scale is not set to zero; and

identifying for removal the filters corresponding to the corresponding scales projected to the first state.

4 . The method of claim 3 , wherein a first scale in a first layer is projected to the first state and a second scale in the first layer is projected to the second state when the first scale is smaller than the second scale.

5 . The method of claim 3 , wherein a first scale having a first value and corresponding to a first filter in a first layer is projected to the first state and a second scale having a second value smaller than the first value and corresponding to a second filter in a second layer is projected to the second state.

6 . The method of claim 5 , wherein the first filter has more weight values than the second filter.

7 . The method of claim 5 , wherein (i) the first layer has a first number of filters such that removing the first filter reduces an amount of time required to execute the network on the particular type of integrated circuit and (ii) the second layer has a second number of filters such that removing the first filter does not reduce the amount of time required to execute the network on the particular type of integrated circuit.

8 . The method of claim 2 , wherein adjusting the respective plurality of weight values comprises adjusting (i) weight values of the filters and (ii) the corresponding scales.

9 . The method of claim 2 , wherein:

the particular type of integrated circuit comprises circuits for computing output values for a particular number of filters simultaneously.

10 . The method of claim 1 , wherein the compute time constraint is a constraint in a set of constraints for executing the network using the particular type of integrated circuit, the set of constraints comprising a weight value constraint, and wherein the weight value constraint that accounts for the total a total number of weight values for a layer of the set of prunable layers is based on a maximum amount of memory available on the particular type of integrated circuit.

11 . The method of claim 10 , wherein the weight value constraint that accounts for a total number of weight values for a layer of the set of prunable layers specifies a maximum number of weight values for a layer of the set of prunable layers after removal of the sets of weight values.

12 . The method of claim 10 , wherein identifying the sets of weight values for removal is based on the weight value constraint.

13 . A non-transitory machine-readable medium storing a program which when executed by a processor trains a machine-trained (MT) network, the program comprising sets of instructions for:

receiving, for execution on a particular type of integrated circuit, a network comprising a plurality of layers, the plurality of layers comprising a set of prunable layers, wherein each layer of the set of prunable layers comprises a respective plurality of weight values; and

training the network by iteratively:

propagating inputs through the network to generate outputs and adjusting the respective plurality of weight values for each layer of the plurality of layers based on differences between the outputs and expected outputs;

determining, using the respective plurality of weight values for each layer of the plurality of layers, an estimated compute time for the network;

identifying, in response to the estimated compute time exceeding a compute time constraint and based on a parallelism computation capability of the particular type of integrated circuit, sets of weight values of the respective plurality of weight values for each layer of the set of prunable layers removal; and

removing the sets of weight values from consideration during a subsequent training iteration of the network, wherein removing the sets of weight values reduces the estimated compute time.

14 . The non-transitory machine-readable medium of claim 13 , wherein the respective plurality of weight values for reach of the set of prunable layers are arranged as filters with corresponding scales, wherein identifying the sets of weight values for removal comprises identifying filters for removal.

15 . The non-transitory machine-readable medium of claim 14 , wherein the set of instructions for identifying the sets of weight values for removal comprises sets of instructions for:

projecting each corresponding scale of the corresponding scales to a respective set of states, wherein each corresponding scale is projected to one of (i) a first state in which the corresponding scale is set to zero and (ii) a second state in which the corresponding scale is not set to zero; and

identifying for removal the filters corresponding to the corresponding scales projected to the first state.

16 . The non-transitory machine-readable medium of claim 15 , wherein:

a first scale having a first value and corresponding to a first filter in a first layer is projected to the first state and a second scale having a second value smaller than the first value and corresponding to a second filter in a second layer is projected to the second state; and

the first filter has more weight values than the second filter.

17 . The non-transitory machine-readable medium of claim 15 , wherein:

a first scale having a first value and corresponding to a first filter in a first layer is projected to the first state and a second scale having a second value smaller than the first value and corresponding to a second filter in a second layer is projected to the second state;

the first layer has a first number of filters such that removing the first filter reduces an amount of time required to execute the network on the particular type of integrated circuit; and

the second layer has a second number of filters such that removing the first filter does not reduce the amount of time required to execute the network on the particular type of integrated circuit.

18 . The non-transitory machine-readable medium of claim 14 , wherein:

the particular type of integrated circuit comprises circuits for computing output values for a particular number of filters simultaneously.

19 . The non-transitory machine-readable medium of claim 13 , wherein the compute time constraint is a constraint in a set of constraints for executing the network using the particular type of integrated circuit, the set of constraints comprising a weight value constraint, and wherein the weight value constraint that accounts for a total number of weight values for a layer of the set of prunable layers is based on a maximum amount of memory available on the particular type of integrated circuit.

20 . The non-transitory machine-readable medium of claim 19 , wherein the weight value constraint that accounts for a total number of weight values for a layer of the set of prunable layers specifies a maximum number of weight values for a layer of the set of prunable layers after removal of the sets of weight values.

Assignments (3)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2022
From: SATHER, ERIC A.; TEIG, STEVEN L.
To: PERCEIVE CORPORATION
Reel/Frame 059696/0609 →
Continuity (4)
Continuation In Part 17089653 · Nov 4, 2020
Provisional Application 63189516 · May 17, 2021
Provisional Application 63178889 · Apr 23, 2021
Provisional Application 63065472 · Aug 13, 2020
References Cited (101)
US 9904874B2 · Shoaib et al. · 2018 [cited by applicant]
US 20130138589A1 · Yu · 2013 [cited by examiner]
US 20160086078A1 · Ji et al. · 2016 [cited by applicant]
US 20160174902A1 · Georgescu et al. · 2016 [cited by applicant]
US 20170286830A1 · El-Yaniv et al. · 2017 [cited by applicant]
US 20180107925A1 · Choi et al. · 2018 [cited by applicant]
US 20180197049A1 · Tran et al. · 2018 [cited by applicant]
US 20190012594A1 · Fukuda · 2019 [cited by examiner]
US 20190042948A1 · Lee et al. · 2019 [cited by applicant]
US 20190065896A1 · Lee · 2019 [cited by examiner]
US 20190138882A1 · Choi et al. · 2019 [cited by applicant]
US 20190138896A1 · Deng · 2019 [cited by applicant]
US 20190147323A1 · Li · 2019 [cited by examiner]
US 20190171927A1 · Diril et al. · 2019 [cited by applicant]
US 20190180184A1 · Deng · 2019 [cited by examiner]
US 20190188557A1 · Lowell et al. · 2019 [cited by applicant]
US 20190228274A1 · Georgiadis et al. · 2019 [cited by applicant]
US 20190286970A1 · Karaletsos · 2019 [cited by examiner]
US 20190340492A1 · Burger et al. · 2019 [cited by applicant]
US 20190354842A1 · Louizos et al. · 2019 [cited by applicant]
US 20190362235A1 · Xu · 2019 [cited by examiner]
US 20200005143A1 · Zamora Esquivel · 2020 [cited by examiner]
US 20200104692A1 · Hill · 2020 [cited by examiner]
US 20200134461A1 · Chai et al. · 2020 [cited by applicant]
US 20200202213A1 · Rouhani et al. · 2020 [cited by applicant]
US 20200202218A1 · Csefalvay · 2020 [cited by applicant]
US 20200210838A1 · Lo et al. · 2020 [cited by applicant]
US 20200264876A1 · Lo et al. · 2020 [cited by applicant]
US 20200302269A1 · Ovtcharov et al. · 2020 [cited by applicant]
US 20210019630A1 · Yao · 2021 [cited by examiner]
US 20210042626A1 · Krishnamoorthy · 2021 [cited by examiner]
US 20210073644A1 · Lin · 2021 [cited by examiner]
US 20210224642A1 · Moriya · 2021 [cited by examiner]
US 20210248459A1 · Li · 2021 [cited by examiner]
US 20210264271A1 · Gebre · 2021 [cited by examiner]
US 20210406672A1 · Hoang et al. · 2021 [cited by applicant]
Wang, “Fixed-point Factorized Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition. (Previously supplied). (Year: 2017). [cited by examiner]
Sze, “Efficient Processing of Deep Neural Networks: A Tutorial and Survey”, IEEE, 2017. (Previously supplied). (Year: 2017). [cited by examiner]
Zhang, “A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers”, 2018. (Year: 2018). [cited by examiner]
Bhattacharya, Sourav, et al., “Sparsification and Separation of Deep Learning Layers for Constrained Resource Inference on Wearables,” SenSys '16, Nov. 14-16, 2016, 15 pages, ACM, Stanford, CA, USA. [cited by applicant]
Han, Song, et al., “Learning Both Weights and Connections for Efficient Neural Networks,” Oct. 30, 2015, 9 pages, retrieved from https://arxiv.org/abs/1506.02626. [cited by applicant]
Hu, Hengyuan, et al., “Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures,” Jul. 12, 2016, 9 pages, retrieved from https://arxiv.org/abs/1607.03250v1. [cited by applicant]
Neklyudov, Kirill, et al., “Structured Bayesian Pruning via Log-Normal Multiplicative Noise,” Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 4-9, 2017, 10 pages, ACM, Long … [cited by applicant]
Wang, Peiqi, et al., “SNrram: An Efficient Sparse Neural Network Computation Architecture Based on Resistive Random-Access Memory,” DAC '18, Jun. 24-29, 2018, 7 pages, ACM, San Francisco, CA, USA. [cited by applicant]
Agostinelli, Forest, et al., “Learning Activation Functions to Improve Deep Neural Networks,” Apr. 21, 2015, 9 pages, retrieved from https://arxiv.org/abs/1412.6830. [cited by applicant]
Aizenberg, Igor, “Periodic Activation Function and a Modified Learning Algorithm for the Multivalued Neuron,” IEEE Transactions on Neural Networks, Dec. 2010, 11 pages, vol. 21, No. 12, IEEE. [cited by applicant]
Martens, James, “New Insights and Perspectives on the Natural Gradient Method,” Nov. 21, 2017, 59 pages, retrieved from https://arxiv.org/abs/1412.1193v9. [cited by applicant]
Withagen, Heini, “Reducing the Effect of Quantization by Weight Scaling,” Proceedings of 1994 IEEE International Conference on Neural Networks (ICNN '94), Jun. 28-Jul. 2, 1994, 3 pages, IEEE, Orlando, Florida, USA. [cited by applicant]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Andri, Renzo, et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Mar. 14, 2017, 14 pages, IEEE, New York,… [cited by applicant]
Bagherinezhad, Hessam, et al., “LCNN: Look-up Based Convolutional Neural Network,” Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 10 pages, IEEE, Honolulu, … [cited by applicant]
Bong, Kyeongryeol, et al., “A 0.62mW Ultra-Low-Power Convolutional-Neural-Network Face-Recognition Processor and a CIS Integrated with Always-On Haar-Like Face Detector,” Proceedings of 2017 IEEE International Solid-Sta… [cited by applicant]
Boo, Yoonho, et al., “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations,” 2017 IEEE Workshop on Signal Processing Systems (SiPS), Oct. 3-5, 2017, 6 pages, IEEE, Lorie… [cited by applicant]
Chen, Yu-Hsin, et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA 2… [cited by applicant]
Chen, Yu-Hsin, et al., “Using Dataflow to Optimize Energy Efficiency of Deep Neural Network Accelerators,” IEEE Micro, Jun. 14, 2017, 10 pages, vol. 37, Issue 3, IEEE, New York, NY, USA. [cited by applicant]
Courbariaux, Matthieu, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or -1,” Mar. 17, 2016, 11 pages, arXiv: 1602.02830v3, Computing Research Repository (CoR… [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations,” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 15),… [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Guo, Yiwen, et al., “Network Sketching: Exploring Binary Structure in Deep CNNs,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 9 pages, IEEE, Honolulu, HI. [cited by applicant]
He, Zhezhi, et al., “Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy,” Jul. 20, 2018, 8 pages, arXiv:1807.07948v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Hegde, Kartik, et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition,” Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA '18), Jun. 2-6, 2018, 14… [cited by applicant]
Huan, Yuxiang, et al., “A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity,” May 22, 2017, 5 pages, arXiv:1705.08009v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, … [cited by applicant]
Jain, Anil K., et al., “Artificial Neural Networks: A Tutorial,” Computer, Mar. 1996, 14 pages, vol. 29, Issue 3, IEEE. [cited by applicant]
Jouppi, Norman, P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17), Jun. 24-28, 2017, 17 pages, ACM, … [cited by applicant]
Judd, Patrick, et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing,” Apr. 29, 2017, 6 pages, arXiv:1705.00125v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, U… [cited by applicant]
Kingma, Diederik P., et al., “Auto-Encoding Variational Bayes,” May 1, 2014, 14 pages, arXiv:1312.6114v10, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Kong, Chen, et al., “Take it in your stride: Do we need striding in CNNs?,” Dec. 7, 2017, 9 pages, arXiv:1712.02502v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Leng, Cong, et al., “Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM,” Proceedings of 32nd AAAI Conference on Artificial Intelligence (AAAI-18), Feb. 2-7, 2018, 16 pages, Association for the Advance… [cited by applicant]
Li, Fengfu, et al., “Ternary Weight Networks,” May 16, 2016, 9 pages, arXiv:1605.04711v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Merolla, Paul, et al., “Deep Neural Networks are Robust to Weight Binarization and Other Non-linear Distortions,” Jun. 7, 2016, 10 pages, arXiv:1606.01981v1, Computing Research Repository (CoRR)—Cornell University, Itha… [cited by applicant]
Moshovos, Andreas, et al., “Exploiting Typical Values to Accelerate Deep Learning,” Computer, May 24, 2018, 13 pages, vol. 51-Issue 5, IEEE Computer Society, Washington, D.C. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 17/089,648, filed Nov. 4, 2020, 118 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 17/089,653, filed Nov. 4, 2020, 118 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 17/089,660, filed Nov. 4, 2020, 118 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned Related U.S. Appl. No. 17/696,809 with similar specification, filed Mar. 16, 2022, 109 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned Related U.S. Appl. No. 17/696,810 with similar specification, filed Mar. 16, 2022, 110 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned Related U.S. Appl. No. 17/696,812 with similar specification, filed Mar. 16, 2022, 110 pages, Perceive Corporation. [cited by applicant]
Park, Jongsoo, et al., “Faster CNNs with Direct Sparse Convolutions and Guided Pruning,” Jul. 28, 2017, 12 pages, arXiv:1608.01409v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Proceedings of 2016 European Conference on Computer Vision (ECCV '16), Oct. 8-16, 2016, 17 pages, Lecture Note… [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 12 pages, ICLR, Vanco… [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Vaswani, Sharan, “Exploiting Sparsity in Supervised Learning,” Month Unknown 2014, 9 pages, retrieved from https://vaswanis.github.io > optimization_report. [cited by applicant]
Wang, Min, et al., “Factorized Convolutional Neural Networks,” 2017 IEEE International Conference on Computer Vision Workshops (ICCVW '17), Oct. 22-29, 2017, 9 pages, IEEE, Venice, Italy. [cited by applicant]
Wen, Wei, et al., “Learning Structured Sparsity in Deep Neural Networks,” Oct. 18, 2016, 10 pages, arXiv:1608.03665v4, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Yan, Shi, “L1 Norm Regularization and Sparsity Explained for Dummies,” Aug. 27, 2016, 13 pages, retrieved from https://blog.mlreview.com/11-norm-regularization-and-sparsity-explained-for-dummies-5b0e4be3938a. [cited by applicant]
Yang, Tien-Ju, et al., “Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning,” Apr. 18, 2017, 9 pages, arXiv:1611.05128v4, Computer Research Repository (CoRR)—Comell University, Ithaca, NY… [cited by applicant]
Yang, Xuan, et al., “DNN Dataflow Choice Is Overrated,” Sep. 10, 2018, 13 pages, arXiv:1809.04070v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Zhang, Dongqing, et al., “LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks,” Jul. 26, 2018, 21 pages, arXiv:1807.10029v1, Computer Research Repository (CoRR)—Cornell University, Ithaca,… [cited by applicant]
Zhou, Shuchang, et al., “DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients,” Jul. 17, 2016, 14 pages, arXiv:1606.06160v2, Computer Research Repository (CoRR)—Cornell University,… [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv:1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Chen, Shangyu, et al., “Deep Neural Network Quantization via Layer-Wise Optimization Using Limited Training Data,” Proceeding of the 33rd AAAI Conference on Artificial Intelligence (AAAI-19), Jul. 2019, 8 pages, AAAI. [cited by applicant]
He, Yang, et al., “Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks,” Aug. 21, 2018, 8 pages, retrieved from https://arxiv.org/abs/1808.06866. [cited by applicant]
Hu, Yiming, et al., “A Novel Channel Pruning Method for Deep Neural Network Compression,” May 29, 2018, 10 pages, retrieved from https://arxiv.org/abs/1805.11394. [cited by applicant]
Jaderberg, Max, et al., “Speeding Up Convolutional Neural Networks with Low Rank Expansions,” May 15, 2014, 12 pages, retrieved from https://arxiv.org/abs/1405.3866. [cited by applicant]
Kozyrskiy, Nikolay, et al., “CNN Acceleration by Low-Rank Approximation with Quantized Factors,” Jun. 16, 2020, 14 pages, retrieved from https://arxiv.org/abs/2006.08878. [cited by applicant]
Molchanov, Pavlo, et al., “Pruning Convention Neural Networks for Resource Efficient Inference,” Jun. 8, 2017, 17 pages, retrieved from https://arxiv.org/abs/1611.06440. [cited by applicant]
Yang, Huanrui, et al., “Learning Low-rank Deep Neural Networks via Singular Vector Orthogonality Regularization and Singular Value Sparsification,” Apr. 20, 2020, 14 pages, retrieved from https://arxiv.org/abs/2004.0903… [cited by applicant]
Zhang, Xiangyu, et al., “Accelerating Very Deep Convolutional Networks for Classification and Detection,” Nov. 18, 2015, 14 pages, retrieved from https://arxiv.org/abs/1505.06798. [cited by applicant]
Jin, Canran,, et al., “Sparse Ternary Connect: Convolutional Neural Networks Using Ternarized Weights with Enhanced Sparsity,” 2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC), Jan. 22-25, 2018, 6… [cited by applicant]
Li, Yawei, et al., “Group Sparsity: The Hinge Between Filter Pruning and Decomposition for Network Compression,” Mar. 20, 2020, 14 pages, arXiv:2003.08935, Computer Research Repository (CoRR)—Cornell University, Ithaca,… [cited by applicant]
Wang, Peisong, et al., “Fixed-Point Factorized Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21-26, 2017, 9 pages, IEEE, Honolulu, Hawaii, USA. [cited by applicant]