IP Library Granted Patent US 12,299,068
Granted Patent B2
US 12,299,068 · App. 18/384,529 · Granted May 13, 2025

Reduced dot product computation circuit

Inventors: Kenneth Duong (San Jose, CA); Jung Ko (San Jose, CA); Steven L. Teig (Menlo Park, CA)
Assignee: Amazon Technologies, Inc.
G06F17/16G06N3/04G06N3/063G06N3/08G06T1/20G06T2207/20081G06T2207/20084G06T2207/20224G06T2207/30201G06V40/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,068
App. No.
18/384,529
Granted
May 13, 2025
Kind
B2
Abstract

Some embodiments provide an IC for implementing a machine-trained network with multiple layers. The IC includes a set of circuits to compute a dot product of (i) a first number of input values computed by other circuits of the IC and (ii) a set of predefined weight values, several of which are zero, with a weight value for each of the input values. The set of circuits includes (i) a dot product computation circuit to compute the dot product based on a second number of inputs and (ii) for each input value, at least two sets of wires for providing the input value to at least two of the dot product computation circuit inputs. The second number is less than the first number. Each input value with a corresponding weight value that is not equal to zero is provided to a different one of the dot product computation circuit inputs.

Claims (36)

1. A method for implementing a machine-trained (MT) network that comprises a plurality of processing nodes, the method comprising:

at a dot product circuit performing a dot product computation for a particular node of the MT network:

receiving (i) a first plurality of input values that are output values of a set of previous nodes of the MT network and (ii) data representing a set of weight values associated with the first plurality of input values;

from the first plurality of input values, selecting a second, smaller plurality of input values, said second plurality of input values comprising all input values from the first plurality of input values that are associated with non-zero weight values; and

computing a dot product of (i) the selected second plurality of input values and (ii) the weight values associated with the second plurality of input values.

2. The method of claim 1 , wherein the second plurality of input values comprises at least one input value associated with a weight value equal to zero.

3. The method of claim 1 , wherein each of the weight values is one of zero, a positive value, and a negation of the positive value.

4. The method of claim 1 , wherein each of the weight values is one of zero, one, and negative one.

5. The method of claim 1 , wherein selecting the second plurality of input values further comprises using a plurality of multiplexers (i) to receive the first plurality of input values and a set of selection signals, (ii) to select the second plurality of input values based on the set of selection signals, and (iii) to provide the selected second plurality of input values to a set of computation circuits.

6. The method of claim 1 , wherein the weight values are trained to ensure that a number of non-zero weight values associated with the first plurality of input values is equal to or less than a number of input values in the second plurality of input values.

7. The method of claim 1 , wherein the particular dot product circuit comprises (i) a set of input selection circuits and (ii) a set of dot product computation circuits.

8. The method of claim 7 , wherein the set of input selection circuits receive the first plurality of input values and select the second plurality of input values from the first plurality of input values.

9. The method of claim 8 , wherein the set of dot product computation circuits receives (i) the second plurality of input values and (ii) at least a subset of the data representing the set of weight values and performs the dot product computation.

10. The method of claim 8 , wherein receiving the first plurality of input values comprises receiving each input value of the first plurality of input values at two or more of the input selection circuits.

11. The method of claim 10 , wherein each input value is selected by at most one of the input selection circuits.

12. The method of claim 10 , wherein a cuckoo hashing algorithm is used to map the first plurality of input values to the input selection circuits.

13. The method of claim 1 , wherein the particular dot product circuit is a first dot product circuit and the dot product computation is a first partial dot product computation that is part of a complete dot product computation for the particular node, the method further comprising:

at one or more additional dot product circuits, performing additional partial dot product computations for the particular node of the MT network; and

combining results of the first partial dot product computation and the additional partial dot product computations to compute the complete dot product computation for the particular node of the MT network.

14. The method of claim 13 , wherein the set of previous nodes is a first set of previous nodes and the set of weight values is a first set of weight values, the method further comprising:

at a second one of the dot product circuits:

receiving (i) a third plurality of input values that are output values of a second set of previous nodes of the MT network and (ii) data representing a second set of weight values associated with the third plurality of input values;

from the third plurality of input values, selecting a fourth plurality of input values, said fourth plurality of input values comprising all input values from the third plurality of input values that are associated with non-zero weight values; and

computing a dot product of (i) the selected fourth plurality of input values and (ii) the weight values associated with the fourth plurality of input values.

15. The method of claim 14 , wherein:

the third plurality of input values comprises a same number of input values as the first plurality of input values;

the fourth plurality of input values comprises a same number of input values as the second plurality of input values; and

a number of input values in the third plurality of input values that are associated with non-zero weight values is different than a number of input values in the first plurality of input values that are associated with non-zero weight values.

16. The method of claim 13 further comprising computing an output for the particular node by applying a non-linear activation function to the complete dot product for the particular node.

17. The method of claim 16 further comprising, prior to computing the output for the particular node:

adding a bias value to the complete dot product to compute a first result; and

multiplying the first result by a scaling value to compute a second result,

wherein the non-linear activation function is applied to the second result to compute the output for the particular node.

18. The method of claim 17 , wherein the scaling value is based on a weight scaling value, wherein each of the weight values is one of zero, the weight scaling value, and a negation of the weight scaling value.

19. The method of claim 1 , wherein the set of previous nodes that output the first plurality of input values are nodes of a previous layer of the MT network.

20. The method of claim 1 , wherein receiving the first plurality of input values and the data representing the set of weight values comprises reading the first plurality of input values and the data representing the set of weight values from a memory.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2025
From: DUONG, KENNETH; KO, JUNG; TEIG, STEVEN L.
To: XCELSIS CORPORATION
Reel/Frame 070797/0042 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2025
From: XCELSIS CORPORATION
To: PERCEIVE CORPORATION
Reel/Frame 070797/0275 →
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
Continuity (6)
Continuation 17316639 · May 10, 2021
Continuation 16924360 · Jul 9, 2020
Continuation 16120387 · Sep 3, 2018
Provisional Application 62724589 · Aug 29, 2018
Provisional Application 62660914 · Apr 20, 2018
Related Publication 20240070225A1 · Feb 29, 2024
References Cited (164)
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 6038583A · Oberman et al. · 2000 [cited by applicant]
US 9710265B1 · Temam et al. · 2017 [cited by applicant]
US 9858636B1 · Lim et al. · 2018 [cited by applicant]
US 9904874B2 · Shoaib et al. · 2018 [cited by applicant]
US 10409604B2 · Kennedy et al. · 2019 [cited by applicant]
US 10445638B1 · Amirineni et al. · 2019 [cited by applicant]
US 10489478B2 · Lim et al. · 2019 [cited by applicant]
US 10515303B2 · Lie et al. · 2019 [cited by applicant]
US 10657438B2 · Lie et al. · 2020 [cited by applicant]
US 10740434B1 · Duong et al. · 2020 [cited by applicant]
US 10768856B1 · Diamant et al. · 2020 [cited by applicant]
US 10796198B2 · Franca-Neto · 2020 [cited by applicant]
US 10817042B2 · Desai et al. · 2020 [cited by applicant]
US 10853738B1 · Dockendorf et al. · 2020 [cited by applicant]
US 10977338B1 · Duong et al. · 2021 [cited by applicant]
US 11003736B2 · Duong et al. · 2021 [cited by applicant]
US 11138292B1 · Nair et al. · 2021 [cited by applicant]
US 11423289B2 · Judd et al. · 2022 [cited by applicant]
US 11809515B2 · Duong et al. · 2023 [cited by applicant]
US 20040078403A1 · Scheuermann et al. · 2004 [cited by applicant]
US 20110055308A1 · Mantor et al. · 2011 [cited by applicant]
US 20110307685A1 · Song · 2011 [cited by applicant]
US 20150339570A1 · Scheffler et al. · 2015 [cited by applicant]
US 20160086078A1 · Ji et al. · 2016 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342893A1 · Ross et al. · 2016 [cited by applicant]
US 20170011006A1 · Saber et al. · 2017 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170168775A1 · Tseng et al. · 2017 [cited by applicant]
US 20170243110A1 · Esquivel et al. · 2017 [cited by applicant]
US 20170323196A1 · Gibson et al. · 2017 [cited by applicant]
US 20180018559A1 · Yakopcic et al. · 2018 [cited by applicant]
US 20180046458A1 · Kuramoto · 2018 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180046905A1 · Li et al. · 2018 [cited by applicant]
US 20180046916A1 · Dally et al. · 2018 [cited by applicant]
US 20180101763A1 · Barnard et al. · 2018 [cited by applicant]
US 20180114569A1 · Strachan et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180164866A1 · Turakhia et al. · 2018 [cited by applicant]
US 20180181406A1 · Kuramoto · 2018 [cited by applicant]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180246855A1 · Redfern et al. · 2018 [cited by applicant]
US 20180285719A1 · Baum et al. · 2018 [cited by applicant]
US 20180285726A1 · Baum et al. · 2018 [cited by applicant]
US 20180285727A1 · Baum et al. · 2018 [cited by applicant]
US 20180285736A1 · Baum et al. · 2018 [cited by applicant]
US 20180293490A1 · Ma et al. · 2018 [cited by applicant]
US 20180293493A1 · Kalamkar et al. · 2018 [cited by applicant]
US 20180293691A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180300600A1 · Ma et al. · 2018 [cited by applicant]
US 20180307494A1 · Ould-Ahmed-Vall et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180307985A1 · Appu et al. · 2018 [cited by applicant]
US 20180308202A1 · Appu et al. · 2018 [cited by applicant]
US 20180314492A1 · Fais et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180322095A1 · Longley et al. · 2018 [cited by applicant]
US 20180322386A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180322387A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180329868A1 · Chen et al. · 2018 [cited by applicant]
US 20180365794A1 · Lee et al. · 2018 [cited by applicant]
US 20180373975A1 · Yu et al. · 2018 [cited by applicant]
US 20190012296A1 · Hsieh et al. · 2019 [cited by applicant]
US 20190026078A1 · Bannon et al. · 2019 [cited by applicant]
US 20190026237A1 · Talpes et al. · 2019 [cited by applicant]
US 20190026249A1 · Talpes et al. · 2019 [cited by applicant]
US 20190041961A1 · Desai et al. · 2019 [cited by applicant]
US 20190057036A1 · Mathuriya et al. · 2019 [cited by applicant]
US 20190073585A1 · Pu et al. · 2019 [cited by applicant]
US 20190087713A1 · Lamb et al. · 2019 [cited by applicant]
US 20190095776A1 · Kfir et al. · 2019 [cited by applicant]
US 20190114499A1 · Delaye et al. · 2019 [cited by applicant]
US 20190138891A1 · Kim et al. · 2019 [cited by applicant]
US 20190147338A1 · Pau et al. · 2019 [cited by applicant]
US 20190156180A1 · Nomura et al. · 2019 [cited by applicant]
US 20190171927A1 · Diril et al. · 2019 [cited by applicant]
US 20190179635A1 · Jiao et al. · 2019 [cited by applicant]
US 20190180167A1 · Huang et al. · 2019 [cited by applicant]
US 20190187983A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20190196970A1 · Han et al. · 2019 [cited by applicant]
US 20190205358A1 · Diril et al. · 2019 [cited by applicant]
US 20190205736A1 · Bleiweiss et al. · 2019 [cited by applicant]
US 20190205739A1 · Liu et al. · 2019 [cited by applicant]
US 20190205740A1 · Judd et al. · 2019 [cited by applicant]
US 20190205780A1 · Sakaguchi · 2019 [cited by applicant]
US 20190236437A1 · Shin et al. · 2019 [cited by applicant]
US 20190236445A1 · Das et al. · 2019 [cited by applicant]
US 20190266217A1 · Arakawa et al. · 2019 [cited by applicant]
US 20190266479A1 · Singh et al. · 2019 [cited by applicant]
US 20190294413A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294959A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294968A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190303741A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303749A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303750A1 · Kumar et al. · 2019 [cited by applicant]
US 20190325296A1 · Fowers et al. · 2019 [cited by applicant]
US 20190332925A1 · Modha · 2019 [cited by applicant]
US 20190347559A1 · Kang et al. · 2019 [cited by applicant]
US 20190385046A1 · Cassidy et al. · 2019 [cited by applicant]
US 20200005131A1 · Nakahara et al. · 2020 [cited by applicant]
US 20200042856A1 · Datta et al. · 2020 [cited by applicant]
US 20200042859A1 · Mappouras et al. · 2020 [cited by applicant]
US 20200089506A1 · Power et al. · 2020 [cited by applicant]
US 20200134461A1 · Chai et al. · 2020 [cited by applicant]
US 20200257930A1 · Nahr et al. · 2020 [cited by applicant]
US 20200342046A1 · Duong et al. · 2020 [cited by applicant]
US 20200364545A1 · Shattil · 2020 [cited by applicant]
US 20200380344A1 · Lie et al. · 2020 [cited by applicant]
US 20200380363A1 · Kwon et al. · 2020 [cited by applicant]
US 20210110236A1 · Shibata · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210241082A1 · Nagy et al. · 2021 [cited by applicant]
US 20210263995A1 · Duong et al. · 2021 [cited by applicant]
US 20220121914A1 · Huang et al. · 2022 [cited by applicant]
US 20220335562A1 · Surti et al. · 2022 [cited by applicant]
CN 108876698A · 2018 [cited by applicant]
CN 108280514B · 2020 [cited by applicant]
GB 2568086A · 2019 [cited by applicant]
WO 2020044527A1 · 2020 [cited by applicant]
Carbon, A., et al., “Pleura: A Scalable Energy-Efficient Programmable Hardware Accelerator for Neural Networks,” 2018 Design, Automation & Test in Europe Conference & Exhibition (Date 2018), Mar. 19-23, 2018, 6 pages, I… [cited by applicant]
Gokhale, Vinayak, et al., “Snowflake: A Model Agnostic Accelerator for Deep Convolutional Neural Networks,” Aug. 8, 2017, 11 pages, arXiv:1708.02579v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Jin, Canran,, et al., “Sparse Ternary Connect: Convolutional Neural Networks Using Terrorized Weights with Enhanced Sparsity,” 2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC), Jan. 22-25, 2018, 6… [cited by applicant]
Abtahi, Tahmid, et al., “Accelerating Convolutional Neural Network With FFT on Embedded Hardware,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Sep. 2018, 14 pages, vol. 26, No. 9, IEEE. [cited by applicant]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Andri, Renzo, et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Mar. 14, 2017, 14 pages, IEEE, New York,… [cited by applicant]
Ardakani, Arash, et al., “An Architecture to Accelerate Convolution in Deep Neural Networks,” IEEE Transactions on Circuits and Systems I: Regular Papers, Oct. 17, 2017, 14 pages, vol. 65, No. 4, IEEE. [cited by applicant]
Ardakani, Arash, et al., “Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks,” Proceedings of the 5th International Conference on Learning Representations (ICLR 2017), Apr.… [cited by applicant]
Bagherinezhad, Hessam, et al., “LCNN: Look-up Based Convolutional Neural Network,” Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 10 pages, IEEE, Honolulu, … [cited by applicant]
Bong, Kyeongryeol, et al., “A 0.62mW Ultra-Low-Power Convolutional-Neural-Network Face-Recognition Processor and a CIS Integrated with Always-On Haar-Like Face Detector,” Proceedings of 2017 IEEE International Solid-Sta… [cited by applicant]
Boo, Yoonho, et al., “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations,” 2017 IEEE Workshop on Signal Processing Systems (SiPS), Oct. 3-5, 2017, 6 pages, IEEE, Lorie… [cited by applicant]
Chen, Yu-Hsin, et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA 2… [cited by applicant]
Chen, Yu-Hsin, et al., “Using Dataflow to Optimize Energy Efficiency of Deep Neural Network Accelerators,” IEEE Micro, Jun. 14, 2017, 10 pages, vol. 37, Issue 3, IEEE, New York, NY, USA. [cited by applicant]
Courbariaux, Matthieu, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” Mar. 17, 2016, 11 pages, arXiv: 1602.02830v3, Computing Research Repository (CoR… [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations,” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS '15)… [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Fu, Yao, et al., “Embedded Vision with INT8 Optimization on Xilinx Devices,” WP490 (v1.0.1), Apr. 19, 2017, 15 pages, Xilinx, Inc., San Jose, CA, USA. [cited by applicant]
Guo, Yiwen, et al., “Network Sketching: Exploring Binary Structure in Deep CNNs,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 9 pages, IEEE, Honolulu, HI. [cited by applicant]
He, Zhezhi, et al., “Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy,” Jul. 20, 2018, 8 pages, arXiv:1807.07948v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Hegde, Kartik, et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition,” Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA '18), Jun. 2-6, 2018, 14… [cited by applicant]
Huan, Yuxiang, et al., “A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity,” May 22, 2017, 5 pages, arXiv:1705.08009v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, … [cited by applicant]
Jouppi, Norman, P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17), Jun. 24-28, 2017, 17 pages, ACM, … [cited by applicant]
Judd, Patrick, et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing,” Apr. 29, 2017, 6 pages, arXiv:1705.00125v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, U… [cited by applicant]
Leng, Cong, et al., “Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM,” Proceedings of 32nd AAAI Conference on Artificial Intelligence (AAAI-18), Feb. 2-7, 2018, 16 pages, Association for the Advance… [cited by applicant]
Li, Fengfu, et al., “Ternary Weight Networks,” May 16, 2016, 9 pages, arXiv:1605.04711v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Merolla, Paul, et al., “Deep Neural Networks are Robust to Weight Binarization and Other Non-linear Distortions,” Jun. 7, 2016, 10 pages, arXiv:1606.01981v1, Computing Research Repository (CoRR)—Cornell University, Itha… [cited by applicant]
Moons, Bert, et al., “ENVISION: A 0.26-to-10TOPS/W Subword-Parallel Dynamic-Voltage-Accuracy-Frequency-Scalable Convolutional Neural Network Processor in 28nm FDSOI,” Proceedings of 2017 IEEE International Solid-State C… [cited by applicant]
Moshovos, Andreas, et al., “Exploiting Typical Values to Accelerate Deep Learning,” Computer, May 24, 2018, 13 pages, vol. 51—Issue 5, IEEE Computer Society, Washington, D.C. [cited by applicant]
Park, Jongsoo, et al., “Faster CNNs with Direct Sparse Convolutions and Guided Pruning,” Jul. 28, 2017, 12 pages, arXiv:1608.01409v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Proceedings of 2016 European Conference on Computer Vision (ECCV '16), Oct. 8-16, 2016, 17 pages, Lecture Note… [cited by applicant]
Ren, Mengye, et al., “SBNet: Sparse Blocks Network for Fast Inference,” Jan. 7, 2018, 10 pages, arXiv:1801.02108v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 12 pages, ICLR, Vanco… [cited by applicant]
Shin, Dongjoo, et al., “DNPU: An 8.1TOPS/W Reconfigurable CNN-RNN Processor for General-Purpose Deep Neural Networks,” Proceedings of 2017 IEEE International Solid-State Circuits Conference (ISSCC 2017), Feb. 5-7, 2017,… [cited by applicant]
Sim, Jaehyeong, et al., “A 1.42TOPS/W Deep Convolutional Neural Network Recognition Processor for Intelligent IoE Systems,” Proceedings of 2016 IEEE International Solid-State Circuits Conference (ISSCC 2016), Jan. 31-Fe… [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Wang, Min, et al., “Factorized Convolutional Neural Networks,” 2017 IEEE International Conference on Computer Vision Workshops (ICCVW '17), Oct. 22-29, 2017, 9 pages, IEEE, Venice, Italy. [cited by applicant]
Wen, Wei, et al., “Learning Structured Sparsity in Deep Neural Networks,” Oct. 18, 2016, 10 pages, arXiv:1608.03665v4, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Yang, Xuan, et al., “DNN Dataflow Choice Is Overrated,” Sep. 10, 2018, 13 pages, arXiv:1809.04070v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Zhang, Shijin, et al., “Cambricon-X: An Accelerator for Sparse Neural Networks,” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO '16), Oct. 15-19, 2016, 12 pages, IEEE, Taipei, Taiwan. [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv:1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]