IP Library Granted Patent US 12,190,230
Granted Patent B2
US 12,190,230 · App. 17/982,474 · Granted Jan 7, 2025

Computation of neural network node by neural network inference circuit

Inventors: Kenneth Duong (San Jose, CA); Jung Ko (San Jose, CA); Steven L. Teig (Menlo Park, CA)
Assignee: Amazon Technologies, Inc.
G06N3/063G06F1/03G06F5/01G06F7/5443G06F9/30098G06F9/30145G06F17/10G06F17/16G06N3/048G06N3/06G06N3/08G06N3/084G06N5/04G06N5/046G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,230
App. No.
17/982,474
Granted
Jan 7, 2025
Kind
B2
Abstract

Some embodiments provide a neural network inference circuit (NNIC) for executing a neural network that includes multiple computation nodes at multiple layers. The NNIC includes a set of clusters of core computation circuits and a channel, connecting the core computation circuits, that includes separate segments corresponding to each of the clusters. The NNIC includes a fabric controller circuit, a cluster controller circuit for each of the clusters, and a core controller circuit for each of the core computation circuits. The fabric controller circuit receives high-level neural network instructions from a microprocessor and parses the high-level neural network instructions.

Claims (30)

1. For a neural network inference circuit, a method for executing a neural network that comprises a plurality of computation nodes, each of a set of the computation nodes comprising a dot product of input values and weight values, the method comprising:

to compute a particular computation node:

at each respective dot product core circuit of a plurality of dot product core circuits of the neural network inference circuit, computing a respective partial dot product using a respective set of input values and a respective set of weight values stored in a respective set of memories of the respective dot product core circuit; and

at a bus of the neural network inference circuit that comprises a plurality of aggregation circuits, combining the partial dot products computed by the plurality of dot product core circuits to compute the dot product for the particular computation node.

2. The method of claim 1 further comprising:

to compute each computation node of a plurality of additional computation nodes:

at each respective dot product core circuit of the plurality of dot product core circuits, computing a respective partial dot product for the computation node using a same respective set of input values and a different respective set of weight values stored in the respective set of memories of the respective dot product core circuit; and

at the bus, combining the partial dot products computed by the at least two dot product core circuits to compute the dot product for the particular computation node.

3. The method of claim 2 , wherein for a particular dot product core circuit, a same set of input values is used for the particular computation node and each of the additional computation nodes.

4. The method of claim 3 , wherein for the particular dot product core circuit, different sets of weight values are used for the particular computation node and each of the additional computation nodes.

5. The method of claim 2 , wherein the neural network inference circuits computes the particular computation node and each of the additional computation nodes simultaneously.

6. The method of claim 5 , wherein the particular computation node and each of the additional computation nodes belongs to a same layer of the neural network.

7. The method of claim 2 , wherein the partial dot products for each of the computation nodes are combined by a different aggregation circuit of the bus.

8. The method of claim 2 further comprising, prior to computing the particular computation node and the plurality of additional computation nodes:

loading the respective sets of input values into respective input buffers of the respective dot product core circuits from the respective sets of memories; and

loading the respective sets of weight values into respective pluralities of weight buffers of the respective dot product core circuits from the respective sets of memories.

9. The method of claim 8 , wherein the input values are loaded into a respective single input buffer in each respective dot product core circuit and the weight values are loaded into a respective plurality of weight buffers in each respective dot product core circuit.

10. The method of claim 9 , wherein a number of weight buffers into which the weight values are loaded in each dot product core circuit is the same as a number of computation nodes comprised by the particular computation node and the plurality of additional computation nodes.

11. The method of claim 9 , wherein each respective dot product core circuit comprises a respective plurality of adder tree circuits, wherein each respective adder tree circuit computes a respective partial dot product for one of the computation nodes by using the input values from the respective input buffer of the respective dot product core circuit and the weight values from one of the weight buffers of the respective dot product core circuit.

12. The method of claim 1 , wherein each of the dot product core circuits comprises a same number of adder tree circuits, each adder tree circuit having an index.

13. The method of claim 12 , wherein the partial dot products for the particular computation node are computed in the plurality of dot product core circuits by adder tree circuits having a same index.

14. The method of claim 13 , wherein the adder tree circuits having the same index in different dot product core circuits provide their computed partial dot products to a same aggregation circuit of the bus.

15. The method of claim 12 , wherein each respective aggregation circuit of the bus corresponds to respective adder tree circuits of each of the dot product core circuits having a same index.

16. The method of claim 1 , wherein each respective set of memories comprises one or more banks of random access memory (RAM), wherein a first block of the RAM for a particular dot product core circuit is allocated to the set of input values used by the particular dot product core circuit and a second block of the RAM for the particular dot product core circuit is allocated to the set of weight values used by the particular dot product core circuit.

17. The method of claim 16 , wherein the RAM for the particular dot product core circuit stores (i) all of the weight values used by the particular dot product core circuit for the entire neural network and (ii) input values for up to two layers of the neural network.

18. The method of claim 17 , wherein the input values for the particular computation node are overwritten by input values for subsequent nodes of the neural network.

19. The method of claim 1 further comprising, to compute the particular computation node, performing a set of post-processing operations on the dot product for the particular computation node at a set of post-processing circuits of the neural network inference circuit.

20. The method of claim 19 , wherein performing the set of post-processing operations comprises:

performing a bias operation and a scaling operation to the dot product at a first subset of the post-processing circuits; and

applying a non-linear activation function to an output of the first subset of the post-processing circuits at a second subset of the post-processing circuits.

Assignments (2)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
Continuity (8)
Continuation 16212616 · Dec 6, 2018
Provisional Application 62773164 · Nov 29, 2018
Provisional Application 62773162 · Nov 29, 2018
Provisional Application 62753878 · Oct 31, 2018
Provisional Application 62742802 · Oct 8, 2018
Provisional Application 62724589 · Aug 29, 2018
Provisional Application 62660914 · Apr 20, 2018
Related Publication 20230063274A1 · Mar 2, 2023
References Cited (167)
US 5621863A · Boulet et al. · 1997 [cited by applicant]
US 5717832A · Steimle et al. · 1998 [cited by applicant]
US 5740326A · Boulet et al. · 1998 [cited by applicant]
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 9710265B1 · Temam · 2017 [cited by examiner]
US 9858636B1 · Lim et al. · 2018 [cited by applicant]
US 9904874B2 · Shoaib et al. · 2018 [cited by applicant]
US 10445638B1 · Amirineni et al. · 2019 [cited by applicant]
US 10489478B2 · Lim et al. · 2019 [cited by applicant]
US 10664310B2 · Bokhari et al. · 2020 [cited by applicant]
US 10768856B1 · Diamant et al. · 2020 [cited by applicant]
US 10796198B2 · Franca-Neto · 2020 [cited by applicant]
US 10817042B2 · Desai et al. · 2020 [cited by applicant]
US 10853738B1 · Dockendorf et al. · 2020 [cited by applicant]
US 11138292B1 · Nair et al. · 2021 [cited by applicant]
US 11170289B1 · Duong et al. · 2021 [cited by applicant]
US 11205115B1 · Duong et al. · 2021 [cited by applicant]
US 11341397B1 · Duong et al. · 2022 [cited by applicant]
US 11361213B1 · Duong et al. · 2022 [cited by applicant]
US 11423289B2 · Judd et al. · 2022 [cited by applicant]
US 11501138B1 · Duong et al. · 2022 [cited by applicant]
US 11537853B1 · Afzal et al. · 2022 [cited by applicant]
US 11868867B1 · Afzal et al. · 2024 [cited by applicant]
US 20040078403A1 · Scheuermann et al. · 2004 [cited by applicant]
US 20110307685A1 · Song · 2011 [cited by applicant]
US 20150339570A1 · Scheffler et al. · 2015 [cited by applicant]
US 20160086078A1 · Ji et al. · 2016 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342893A1 · Ross et al. · 2016 [cited by applicant]
US 20170011006A1 · Saber et al. · 2017 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170243110A1 · Esquivel et al. · 2017 [cited by applicant]
US 20170300828A1 · Feng et al. · 2017 [cited by applicant]
US 20170323196A1 · Gibson et al. · 2017 [cited by applicant]
US 20170344882A1 · Ambrose et al. · 2017 [cited by applicant]
US 20180018559A1 · Yakopcic et al. · 2018 [cited by applicant]
US 20180025268A1 · Teig et al. · 2018 [cited by applicant]
US 20180046458A1 · Kuramoto · 2018 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180046905A1 · Li et al. · 2018 [cited by applicant]
US 20180046916A1 · Dally et al. · 2018 [cited by applicant]
US 20180101763A1 · Barnard et al. · 2018 [cited by applicant]
US 20180114569A1 · Strachan et al. · 2018 [cited by applicant]
US 20180121196A1 · Temam et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180164866A1 · Turakhia et al. · 2018 [cited by applicant]
US 20180181406A1 · Kuramoto · 2018 [cited by applicant]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180197068A1 · Narayanaswami et al. · 2018 [cited by applicant]
US 20180246855A1 · Redfern et al. · 2018 [cited by applicant]
US 20180285719A1 · Baum et al. · 2018 [cited by applicant]
US 20180285726A1 · Baum et al. · 2018 [cited by applicant]
US 20180285727A1 · Baum et al. · 2018 [cited by applicant]
US 20180285736A1 · Baum et al. · 2018 [cited by applicant]
US 20180293490A1 · Ma et al. · 2018 [cited by applicant]
US 20180293493A1 · Kalamkar et al. · 2018 [cited by applicant]
US 20180293691A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180300600A1 · Ma et al. · 2018 [cited by applicant]
US 20180307494A1 · Ould-Ahmed-Vall et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180308202A1 · Appu et al. · 2018 [cited by applicant]
US 20180314492A1 · Fais et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180322386A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180322387A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180329868A1 · Chen et al. · 2018 [cited by applicant]
US 20180365794A1 · Lee et al. · 2018 [cited by applicant]
US 20180373975A1 · Yu et al. · 2018 [cited by applicant]
US 20190012296A1 · Hsieh et al. · 2019 [cited by applicant]
US 20190026078A1 · Bannon et al. · 2019 [cited by applicant]
US 20190026237A1 · Talpes et al. · 2019 [cited by applicant]
US 20190026249A1 · Talpes et al. · 2019 [cited by applicant]
US 20190041961A1 · Desai et al. · 2019 [cited by applicant]
US 20190057036A1 · Mathuriya et al. · 2019 [cited by applicant]
US 20190073585A1 · Pu et al. · 2019 [cited by applicant]
US 20190087713A1 · Lamb et al. · 2019 [cited by applicant]
US 20190095776A1 · Kfir et al. · 2019 [cited by applicant]
US 20190114499A1 · Delaye et al. · 2019 [cited by applicant]
US 20190130265A1 · Ling et al. · 2019 [cited by applicant]
US 20190138891A1 · Kim et al. · 2019 [cited by applicant]
US 20190156180A1 · Nomura et al. · 2019 [cited by applicant]
US 20190171927A1 · Diril et al. · 2019 [cited by applicant]
US 20190179635A1 · Jiao et al. · 2019 [cited by applicant]
US 20190180167A1 · Huang et al. · 2019 [cited by applicant]
US 20190187983A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20190196970A1 · Han et al. · 2019 [cited by applicant]
US 20190205094A1 · Diril et al. · 2019 [cited by applicant]
US 20190205358A1 · Diril et al. · 2019 [cited by applicant]
US 20190205736A1 · Bleiweiss et al. · 2019 [cited by applicant]
US 20190205739A1 · Liu et al. · 2019 [cited by applicant]
US 20190205740A1 · Judd et al. · 2019 [cited by applicant]
US 20190205780A1 · Sakaguchi · 2019 [cited by applicant]
US 20190236437A1 · Shin et al. · 2019 [cited by applicant]
US 20190236445A1 · Das et al. · 2019 [cited by applicant]
US 20190266217A1 · Arakawa et al. · 2019 [cited by applicant]
US 20190266479A1 · Singh et al. · 2019 [cited by applicant]
US 20190294413A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294959A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294968A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190303741A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303749A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303750A1 · Kumar et al. · 2019 [cited by applicant]
US 20190325296A1 · Fowers et al. · 2019 [cited by applicant]
US 20190332925A1 · Modha · 2019 [cited by applicant]
US 20190347559A1 · Kang et al. · 2019 [cited by applicant]
US 20190385046A1 · Cassidy et al. · 2019 [cited by applicant]
US 20200005131A1 · Nakahara et al. · 2020 [cited by applicant]
US 20200042856A1 · Datta et al. · 2020 [cited by applicant]
US 20200042859A1 · Mappouras et al. · 2020 [cited by applicant]
US 20200089506A1 · Power et al. · 2020 [cited by applicant]
US 20200134461A1 · Chai et al. · 2020 [cited by applicant]
US 20200210838A1 · Lo et al. · 2020 [cited by applicant]
US 20200257930A1 · Nahr et al. · 2020 [cited by applicant]
US 20200364545A1 · Shattil · 2020 [cited by applicant]
US 20210110236A1 · Shibata · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210232897A1 · Bichler et al. · 2021 [cited by applicant]
US 20210241082A1 · Nagy et al. · 2021 [cited by applicant]
US 20220121914A1 · Huang et al. · 2022 [cited by applicant]
US 20220335562A1 · Surti et al. · 2022 [cited by applicant]
GB 2568086A · 2019 [cited by applicant]
Carbon, Alexandre, et al. “PNeuro: A scalable energy-efficient programmable hardware accelerator for neural networks.” 2018 Design, Automation & Test in Europe Conference & Exhibition. IEEE, Mar. 2018. (Year: 2018). [cited by examiner]
Gokhale, Vinayak, et al. “Snowflake: A model agnostic accelerator for deep convolutional neural networks.” arXiv preprint arXiv: 1708.02579. Aug. 2017. (Year: 2017). [cited by examiner]
Abtahi, Tahmid, et al., “Accelerating Convolutional Neural Network With FFT on Embedded Hardware,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Sep. 2018, 14 pages, vol. 26, No. 9, IEEE. [cited by applicant]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Andri, Renzo, et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Mar. 14, 2017, 14 pages, IEEE, New York,… [cited by applicant]
Ardakani, Arash, et al., “An Architecture to Accelerate Convolution in Deep Neural Networks,” IEEE Transactions on Circuits and Systems I: Regular Papers, Oct. 17, 2017, 14 pages, vol. 65, No. 4, IEEE. [cited by applicant]
Ardakani, Arash, et al., “Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks,” Proceedings of the 5th International Conference on Learning Representations (ICLR 2017), Apr.… [cited by applicant]
Bagherinezhad, Hessam, et al., “LCNN: Look-up Based Convolutional Neural Network,” Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 10 pages, IEEE, Honolulu, … [cited by applicant]
Bang, Suyoung, et al., “A 288pW Programmable Deep-Learning Processor with 270KB On-Chip Weight Storage Using Non-Uniform Memory Hierarchy for Mobile Intelligence,” Proceedings of 2017 IEEE International Solid-State Circ… [cited by applicant]
Bong, Kyeongryeol, et al., “A 0.62mW Ultra-Low-Power Convolutional-Neural-Network Face-Recognition Processor and a CIS Integrated with Always-On Haar-Like Face Detector,” Proceedings of 2017 IEEE International Solid-Sta… [cited by applicant]
Boo, Yoonho, et al., “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations,” 2017 IEEE Workshop on Signal Processing Systems (SiPS), Oct. 3-5, 2017, 6 pages, IEEE, Lorie… [cited by applicant]
Chen, Tianqi, et al., “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning,” Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI '18), Oct. 8-10, 2018, 17 pages, … [cited by applicant]
Chen, Yu-Hsin, et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA 2… [cited by applicant]
Chen, Yu-Hsin, et al., “Using Dataflow to Optimize Energy Efficiency of Deep Neural Network Accelerators,” IEEE Micro, Jun. 14, 2017, 10 pages, vol. 37, Issue 3, IEEE, New York, NY, USA. [cited by applicant]
Courbariaux, Matthieu, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” Mar. 17, 2016, 11 pages, arXiv:1602.02830v3, Computing Research Repository (CoRR… [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations,” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 15),… [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Fu, Yao, et al., “Embedded Vision with INT8 Optimization on Xilinx Devices,” WP490 (v1.0.1), Apr. 19, 2017, 15 pages, Xilinx, Inc., San Jose, CA, USA. [cited by applicant]
Guo, Yiwen, et al., “Network Sketching: Exploring Binary Structure in Deep CNNs,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 9 pages, IEEE, Honolulu, HI. [cited by applicant]
He, Zhezhi, et al., “Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy,” Jul. 20, 2018, 8 pages, arXiv:1807.07948v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Hegde, Kartik, et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition,” Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA '18), Jun. 2-6, 2018, 14… [cited by applicant]
Huan, Yuxiang, et al., “A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity,” May 22, 2017, 5 pages, arXiv:1705.08009v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, … [cited by applicant]
Jouppi, Norman, P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17), Jun. 24-28, 2017, 17 pages, ACM, … [cited by applicant]
Judd, Patrick, et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing,” Apr. 29, 2017, 6 pages, arXiv:1705.00125v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, U… [cited by applicant]
Leng, Cong, et al., “Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM,” Proceedings of 32nd AAAI Conference on Artificial Intelligence (AAAI-18), Feb. 2-7, 2018, 16 pages, Association for the Advance… [cited by applicant]
Li, Fengfu, et al., “Ternary Weight Networks,” May 16, 2016, 9 pages, arXiv:1605.04711v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Liu, Shaoli, et al., “Cambricon: An Instruction Set Architecture for Neural Networks,” 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Jun. 18-22, 2016, 13 pages, IEEE, Seoul, South Korea. [cited by applicant]
Merolla, Paul, et al., “Deep Neural Networks are Robust to Weight Binarization and Other Non-linear Distortions,” Jun. 7, 2016, 10 pages, arXiv:1606.01981v1, Computing Research Repository (CoRR)—Cornell University, Itha… [cited by applicant]
Moons, Bert, et al., “Envision: A 0.26-to-10TOPS/W Subword-Parallel Dynamic-Voltage-Accuracy-Frequency-Scalable Convolutional Neural Network Processor in 28nm FDSOI,” Proceedings of 2017 IEEE International Solid-State C… [cited by applicant]
Moshovos, Andreas, et al., “Exploiting Typical Values to Accelerate Deep Learning,” Computer, May 24, 2018, 13 pages, vol. 51-Issue 5, IEEE Computer Society, Washington, D.C. [cited by applicant]
Park, Jongsoo, et al., “Faster CNNs with Direct Sparse Convolutions and Guided Pruning,” Jul. 28, 2017, 12 pages, arXiv:1608.01409v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Proceedings of 2016 European Conference on Computer Vision (ECCV '16), Oct. 8-16, 2016, 17 pages, Lecture Note… [cited by applicant]
Ren, Mengye, et al., “SBNet: Sparse Blocks Network for Fast Inference,” Jan. 7, 2018, 10 pages, arXiv:1801.02108v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 12 pages, ICLR, Vanco… [cited by applicant]
Shin, Dongjoo, et al., “DNPU: An 8.1TOPS/W Reconfigurable CNN-RNN Processor for General-Purpose Deep Neural Networks,” Proceedings of 2017 IEEE International Solid-State Circuits Conference (ISSCC 2017), Feb. 5-7, 2017,… [cited by applicant]
Sim, Jaehyeong, et al., “A 1.42TOPS/W Deep Convolutional Neural Network Recognition Processor for Intelligent IoE Systems,” Proceedings of 2016 IEEE International Solid-State Circuits Conference (ISSCC 2016), Jan. 31-Fe… [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Wang, Min, et al., “Factorized Convolutional Neural Networks,” 2017 IEEE International Conference on Computer Vision Workshops (ICCVW '17), Oct. 22-29, 2017, 9 pages, IEEE, Venice, Italy. [cited by applicant]
Wen, Wei, et al., “Learning Structured Sparsity in Deep Neural Networks,” Oct. 18, 2016, 10 pages, arXiv:1608.03665v4, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Yang, Xuan, et al., “DNN Dataflow Choice Is Overrated,” Sep. 10, 2018, 13 pages, arXiv:1809.04070v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Zhang, Shijin, et al., “Cambricon-X: An Accelerator for Sparse Neural Networks,” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO '16), Oct. 15-19, 2016, 12 pages, IEEE, Taipei, Taiwan. [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv:1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Jin, Canran,, et al., “Sparse Ternary Connect: Convolutional Neural Networks Using Ternarized Weights with Enhanced Sparsity,” 2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC), Jan. 22-25, 2018, 6… [cited by applicant]
Han, Song, “Efficient Methods and Hardware for Deep Learning,” September 217, 125 pages, Stanford University, Palo Alto, CA, USA. [cited by applicant]