IP Library Granted Patent US 12,265,905
Granted Patent B2
US 12,265,905 · App. 17/984,228 · Granted Apr 1, 2025

Computation of neural network node with large input values

Inventors: Jung Ko (San Jose, CA); Kenneth Duong (San Jose, CA); Steven L. Teig (Menlo Park, CA)
Assignee: Amazon Technologies, Inc.
G06N3/063G06F1/03G06F5/01G06F7/5443G06F9/30098G06F9/30145G06F17/10G06F17/16G06N3/048G06N3/06G06N3/08G06N3/084G06N5/04G06N5/046G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,905
App. No.
17/984,228
Granted
Apr 1, 2025
Kind
B2
Abstract

Some embodiments provide a method for a circuit that executes a neural network including multiple nodes. The method loads a set of weight values for a node into a set of weight value buffers, a first set of bits of each input value of a set of input values for the node into a first set of input value buffers, and a second set of bits of each of the input values into a second set of input value buffers. The method computes a first dot product of the weight values and the first set of bits of each input value and a second dot product of the weight values and the second set of bits of each input value. The method shifts the second dot product by a particular number of bits and adds the first dot product with the bit-shifted second dot product to compute a dot product for the node.

Claims (32)

1. For a neural network inference circuit that executes a neural network comprising a plurality of computation nodes, the neural network inference circuit comprising a plurality of dot product cores for computing partial dot products, each of a set of the plurality of computation nodes comprising a dot product of input values and ternary weight values, a method for computing an output value for a particular computation node, the method comprising:

at each core of a set of the plurality of dot product cores of the neural network inference circuit, loading (i) data for a set of ternary weight values for the particular computation node into a weight value buffer of the core, (ii) a first portion of each input value of a set of input values for the particular computation node into a first input value buffer of the core, and (iii) a second portion of each of the input values into a second input value buffer of the core;

at a set of dot product computation circuits of the neural network inference circuit, the set of dot product computation circuits comprising a partial dot product computation circuit of each core of the set of dot product cores:

computing a first dot product between the set of ternary weight values from the weight value buffers of the set of cores and the first portion of each of the input values from the first input value buffers of the set of cores;

computing a second dot product between the set of ternary weight values and the second portion of each of the input values from the second input value buffers of the set of cores;

bit-shifting the second dot product to generate a bit-shifted second dot product; and

adding the first dot product with the bit-shifted second dot product to generate a computed dot product for the particular computation node; and

at a set of post-processing circuits of the neural network inference circuit, performing a set of post-processing operations to compute the output value for the particular computation node from the computed dot product for the particular computation node.

2. The method of claim 1 , wherein the first portion of each of the input values comprises a particular number of bits by which the second dot product is bit-shifted.

3. The method of claim 1 , wherein the first portion of each input value is a set of least significant bits of the input value and the second portion of each input value is a set of most significant bits of the input value.

4. The method of claim 1 , wherein (i) each input value is an 8-bit value, (ii) the first portion of each input value comprises a least significant 4 bits of the input value, and (iii) the second portion of each input value comprises a most significant 4 bits of the input value.

5. The method of claim 4 , wherein each of the first input value buffer and each second input value buffer comprises a set of slots for storing input values, wherein each slot stores 4 bits.

6. The method of claim 1 , wherein the partial dot product computation circuits of the set of dot product cores perform partial dot product computations of the first dot product and the second dot product.

7. The method of claim 6 , wherein a dot product bus of the neural network inference circuit aggregates the partial dot product computations of the first dot product and aggregates the partial dot product computations of the second dot product.

8. The method of claim 7 , wherein the dot product bus provides the first dot product and the second dot product to a dot product processing circuit of the neural network inference circuit that shifts the second dot product and adds the first dot product with the bit-shifted second dot product.

9. The method of claim 8 , wherein the dot product processing circuit provides the bit-shifted second dot product to the set of post-processing circuits.

10. The method of claim 9 , wherein:

each dot product core of the neural network inference circuit comprises a plurality of sets of weight value buffers and a plurality of partial dot product computation circuits for simultaneously computing partial dot products for different computation nodes;

the dot product bus comprises a plurality of independent aggregation circuits for aggregating partial dot products for different simultaneously-computed computation nodes; and

the neural network inference circuit comprises (i) a plurality of dot product processing circuits for simultaneously adding first dot products with bit-shifted second dot products for the different simultaneously-computed computation nodes and (ii) a plurality of sets of post-processing circuits for simultaneously performing sets of post-processing operations to compute output values for the different simultaneously-computed computation nodes.

11. The method of claim 1 , wherein the first dot product and the second dot product are computed in different clock cycles of the neural network inference circuit.

12. The method of claim 11 , wherein the set of dot product computation circuits computes the first dot product in a first clock cycle, the method further comprising storing the first dot product in a register of the set of dot product computation circuits for at least one clock cycle.

13. The method of claim 12 , wherein the set of dot product computation circuits computes the second dot product in a second clock cycle that is after the first clock cycle.

14. The method of claim 13 , wherein the set of dot product computation circuits bit-shifts the second dot product and adds the first dot product from the register with the bit-shifted second dot product in the second clock cycle.

15. The method of claim 11 , wherein the set of dot product computation circuits computes the second dot product and bit-shifts the second dot product in a first clock cycle, the method further comprising storing the bit-shifted second dot product in a register for at least one clock cycle.

16. The method of claim 15 , wherein the set of dot product computation circuits computes the first dot product in a second clock cycle that is after the first clock cycle.

17. The method of claim 16 , wherein the set of dot product computation circuits adds the bit-shifted second dot product from the register with the first dot product in the second clock cycle.

18. The method of claim 1 , wherein each ternary weight value is one of a positive value, a negation of the positive value, and zero.

19. The method of claim 1 , wherein performing the set of post-processing operations comprises:

at an adder circuit, adding a bias value to the computed dot product to compute a first intermediate result;

at a multiplier circuit, multiplying the first intermediate result by a scaling value to compute a second intermediate result; and

at a non-linear activation function circuit, applying a non-linear activation function to the second intermediate result to compute the output value for the particular computation node.

Assignments (2)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
Continuity (8)
Continuation 16212645 · Dec 6, 2018
Provisional Application 62773162 · Nov 29, 2018
Provisional Application 62773164 · Nov 29, 2018
Provisional Application 62753878 · Oct 31, 2018
Provisional Application 62742802 · Oct 8, 2018
Provisional Application 62724589 · Aug 29, 2018
Provisional Application 62660914 · Apr 20, 2018
Related Publication 20230076850A1 · Mar 9, 2023
References Cited (192)
US 5463573A · Yoshida · 1995 [cited by applicant]
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 6038583A · Oberman et al. · 2000 [cited by applicant]
US 8577951B1 · Langhammer · 2013 [cited by applicant]
US 9710265B1 · Temam et al. · 2017 [cited by applicant]
US 9858636B1 · Lim et al. · 2018 [cited by applicant]
US 9904874B2 · Shoaib et al. · 2018 [cited by applicant]
US 10409604B2 · Kennedy et al. · 2019 [cited by applicant]
US 10445638B1 · Amirineni et al. · 2019 [cited by applicant]
US 10489478B2 · Lim et al. · 2019 [cited by applicant]
US 10515303B2 · Lie et al. · 2019 [cited by applicant]
US 10614357B2 · Lie · 2020 [cited by applicant]
US 10657438B2 · Lie et al. · 2020 [cited by applicant]
US 10664310B2 · Bokhari et al. · 2020 [cited by applicant]
US 10699189B2 · Lie · 2020 [cited by applicant]
US 10768856B1 · Diamant et al. · 2020 [cited by applicant]
US 10796198B2 · Franca-Neto · 2020 [cited by applicant]
US 10817042B2 · Desai et al. · 2020 [cited by applicant]
US 10853738B1 · Dockendorf et al. · 2020 [cited by applicant]
US 10970630B1 · Aimone · 2021 [cited by applicant]
US 11138292B1 · Nair et al. · 2021 [cited by applicant]
US 11232347B2 · Lie · 2022 [cited by applicant]
US 11250326B1 · Ko et al. · 2022 [cited by applicant]
US 11295200B1 · Ko et al. · 2022 [cited by applicant]
US 11403530B1 · Ko et al. · 2022 [cited by applicant]
US 11423289B2 · Judd et al. · 2022 [cited by applicant]
US 11488004B2 · Lie · 2022 [cited by applicant]
US 11531727B1 · Ko et al. · 2022 [cited by applicant]
US 11537853B1 · Afzal et al. · 2022 [cited by applicant]
US 11868867B1 · Afzal et al. · 2024 [cited by applicant]
US 11960565B2 · Shibata · 2024 [cited by applicant]
US 20040078403A1 · Scheuermann et al. · 2004 [cited by applicant]
US 20110055308A1 · Mantor et al. · 2011 [cited by applicant]
US 20110307685A1 · Song · 2011 [cited by applicant]
US 20150339570A1 · Scheffler et al. · 2015 [cited by applicant]
US 20160086078A1 · Ji et al. · 2016 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160328643A1 · Liu et al. · 2016 [cited by applicant]
US 20160342893A1 · Ross et al. · 2016 [cited by applicant]
US 20170011006A1 · Saber et al. · 2017 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170168775A1 · Tseng et al. · 2017 [cited by applicant]
US 20170243110A1 · Esquivel et al. · 2017 [cited by applicant]
US 20170300828A1 · Feng et al. · 2017 [cited by applicant]
US 20170323196A1 · Gibson et al. · 2017 [cited by applicant]
US 20170344882A1 · Ambrose et al. · 2017 [cited by applicant]
US 20180018559A1 · Yakopcic et al. · 2018 [cited by applicant]
US 20180046458A1 · Kuramoto · 2018 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180046905A1 · Li et al. · 2018 [cited by applicant]
US 20180046916A1 · Dally et al. · 2018 [cited by applicant]
US 20180101763A1 · Barnard et al. · 2018 [cited by applicant]
US 20180114569A1 · Strachan et al. · 2018 [cited by applicant]
US 20180121196A1 · Temam et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180164866A1 · Turakhia et al. · 2018 [cited by applicant]
US 20180181406A1 · Kuramoto · 2018 [cited by applicant]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180197068A1 · Narayanaswami et al. · 2018 [cited by applicant]
US 20180246855A1 · Redfern et al. · 2018 [cited by applicant]
US 20180285719A1 · Baum et al. · 2018 [cited by applicant]
US 20180285726A1 · Baum et al. · 2018 [cited by applicant]
US 20180285727A1 · Baum et al. · 2018 [cited by applicant]
US 20180285736A1 · Baum et al. · 2018 [cited by applicant]
US 20180293490A1 · Ma et al. · 2018 [cited by applicant]
US 20180293493A1 · Kalamkar et al. · 2018 [cited by applicant]
US 20180293691A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180300600A1 · Ma et al. · 2018 [cited by applicant]
US 20180307494A1 · Ould-Ahmed-Vall et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180307985A1 · Appu et al. · 2018 [cited by applicant]
US 20180308202A1 · Appu et al. · 2018 [cited by applicant]
US 20180314492A1 · Fais et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180322095A1 · Longley et al. · 2018 [cited by applicant]
US 20180322386A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180322387A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180365794A1 · Lee et al. · 2018 [cited by applicant]
US 20180373975A1 · Yu et al. · 2018 [cited by applicant]
US 20190012296A1 · Hsieh et al. · 2019 [cited by applicant]
US 20190026078A1 · Bannon et al. · 2019 [cited by applicant]
US 20190026237A1 · Talpes et al. · 2019 [cited by applicant]
US 20190026249A1 · Talpes et al. · 2019 [cited by applicant]
US 20190041961A1 · Desai et al. · 2019 [cited by applicant]
US 20190057036A1 · Mathuriya et al. · 2019 [cited by applicant]
US 20190073585A1 · Pu et al. · 2019 [cited by applicant]
US 20190087713A1 · Lamb et al. · 2019 [cited by applicant]
US 20190095776A1 · Kfir et al. · 2019 [cited by applicant]
US 20190114499A1 · Delaye et al. · 2019 [cited by applicant]
US 20190114534A1 · Teng · 2019 [cited by applicant]
US 20190130265A1 · Ling et al. · 2019 [cited by applicant]
US 20190138882A1 · Choi · 2019 [cited by examiner]
US 20190138891A1 · Kim et al. · 2019 [cited by applicant]
US 20190147338A1 · Pau et al. · 2019 [cited by applicant]
US 20190156180A1 · Nomura et al. · 2019 [cited by applicant]
US 20190171927A1 · Diril et al. · 2019 [cited by applicant]
US 20190179635A1 · Jiao et al. · 2019 [cited by applicant]
US 20190180167A1 · Huang et al. · 2019 [cited by applicant]
US 20190187983A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20190196970A1 · Han et al. · 2019 [cited by applicant]
US 20190205094A1 · Diril et al. · 2019 [cited by applicant]
US 20190205358A1 · Diril et al. · 2019 [cited by applicant]
US 20190205736A1 · Bleiweiss et al. · 2019 [cited by applicant]
US 20190205739A1 · Liu et al. · 2019 [cited by applicant]
US 20190205740A1 · Judd et al. · 2019 [cited by applicant]
US 20190205780A1 · Sakaguchi · 2019 [cited by applicant]
US 20190236437A1 · Shin et al. · 2019 [cited by applicant]
US 20190236445A1 · Das et al. · 2019 [cited by applicant]
US 20190266217A1 · Arakawa et al. · 2019 [cited by applicant]
US 20190266479A1 · Singh et al. · 2019 [cited by applicant]
US 20190294413A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294959A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294968A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190303741A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303749A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303750A1 · Kumar et al. · 2019 [cited by applicant]
US 20190325296A1 · Fowers et al. · 2019 [cited by applicant]
US 20190332925A1 · Modha · 2019 [cited by applicant]
US 20190347559A1 · Kang et al. · 2019 [cited by applicant]
US 20190385046A1 · Cassidy et al. · 2019 [cited by applicant]
US 20200005131A1 · Nakahara et al. · 2020 [cited by applicant]
US 20200042856A1 · Datta et al. · 2020 [cited by applicant]
US 20200042859A1 · Mappouras et al. · 2020 [cited by applicant]
US 20200089506A1 · Power et al. · 2020 [cited by applicant]
US 20200134461A1 · Chai et al. · 2020 [cited by applicant]
US 20200210838A1 · Lo et al. · 2020 [cited by applicant]
US 20200257930A1 · Nahr · 2020 [cited by examiner]
US 20200364545A1 · Shattil · 2020 [cited by applicant]
US 20200380344A1 · Lie et al. · 2020 [cited by applicant]
US 20210110236A1 · Shibata · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210232897A1 · Bichler et al. · 2021 [cited by applicant]
US 20210241082A1 · Nagy et al. · 2021 [cited by applicant]
US 20220121914A1 · Huang et al. · 2022 [cited by applicant]
US 20220335562A1 · Surti et al. · 2022 [cited by applicant]
CN 108876698A · 2018 [cited by applicant]
CN 108280514B · 2020 [cited by applicant]
GB 2568086A · 2019 [cited by applicant]
WO 2020044527A1 · 2020 [cited by applicant]
Abtahi, Tahmid, et al., “Accelerating Convolutional Neural Network With FFT on Embedded Hardware,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Sep. 2018, 14 pages, vol. 26, No. 9, IEEE. [cited by applicant]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Andri, Renzo, et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Mar. 14, 2017, 14 pages, IEEE, New York,… [cited by applicant]
Ardakani, Arash, et al., “An Architecture to Accelerate Convolution in Deep Neural Networks,” IEEE Transactions on Circuits and Systems I: Regular Papers, Oct. 17, 2017, 14 pages, vol. 65, No. 4, IEEE. [cited by applicant]
Ardakani, Arash, et al., “Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks,” Proceedings of the 5th International Conference on Learning Representations (ICLR 2017), Apr.… [cited by applicant]
Bagherinezhad, Hessam, et al., “LCNN: Look-up Based Convolutional Neural Network,” Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 10 pages, IEEE, Honolulu, … [cited by applicant]
Boo, Yoonho, et al., “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations,” 2017 IEEE Workshop on Signal Processing Systems (SiPS), Oct. 3-5, 2017, 6 pages, IEEE, Lorie… [cited by applicant]
Chen, Tianqi, et al., “TVM: End-to-End Optimization Stack for Deep Learning,” Feb. 12, 2018, 19 pages, arXiv:1802.04799v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Chen, Yu-Hsin, et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA 2… [cited by applicant]
Chen, Yu-Hsin, et al., “Using Dataflow to Optimize Energy Efficiency of Deep Neural Network Accelerators,” IEEE Micro, Jun. 14, 2017, 10 pages, vol. 37, Issue 3, IEEE, New York, NY, USA. [cited by applicant]
Courbariaux, Matthieu, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” Mar. 17, 2016, 11 pages, arXiv:1602.02830v3, Computing Research Repository (CoRR… [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations,” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 15),… [cited by applicant]
Deng, Lei, et al., “GXNOR-Net: Training Deep Neural Networks with Ternary Weights and Activations without Full-Precision Memory under a Unified Discretization Framework,” Neural Networks 100, Feb. 2018, 10 pages, Elsevi… [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Fu, Yao, et al., “Embedded Vision with INT8 Optimization on Xilinx Devices,” WP490 (v1.0.1), Apr. 19, 2017, 15 pages, Xilinx, Inc., San Jose, CA, USA. [cited by applicant]
Gao, Mingyu, et al., “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,” Proceedings of the 22nd International Conference on Architectural Support for Programming Languages and Operating Systems… [cited by applicant]
Ghanekar, Sachin P., et al., “Signed-Digit-Based Multiplier-Free Realizations for Multirate Converters,” IEEE Transactions on Signal Processing, Mar. 1995, 12 pages, vol. 43, No. 3, IEEE. [cited by applicant]
Giri, Sweta, et al., “Implementation of Combinational Circuits Using Ternary Multiplexer,” International Journal of Computational Engineering Research, Mar.-Apr. 2012, 7 pages, vol. 2, No. 2, IJCER. [cited by applicant]
Guo, Yiwen, et al., “Network Sketching: Exploring Binary Structure in Deep CNNs,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 9 pages, IEEE, Honolulu, HI. [cited by applicant]
He, Zhezhi, et al., “Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy,” Jul. 20, 2018, 8 pages, arXiv:1807.07948v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Hegde, Kartik, et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition,” Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA '18), Jun. 2-6, 2018, 14… [cited by applicant]
Huan, Yuxiang, et al., “A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity,” May 22, 2017, 5 pages, arXiv:1705.08009v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, … [cited by applicant]
Jouppi, Norman, P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17), Jun. 24-28, 2017, 17 pages, ACM, … [cited by applicant]
Judd, Patrick, et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing,” Apr. 29, 2017, 6 pages, arXiv:1705.00125v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, U… [cited by applicant]
Leng, Cong, et al., “Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM,” Proceedings of 32nd AAAI Conference on Artificial Intelligence (AAAI-18), Feb. 2-7, 2018, 16 pages, Association for the Advance… [cited by applicant]
Li, Fengfu, et al., “Ternary Weight Networks,” May 16, 2016, 9 pages, arXiv:1605.04711v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Merolla, Paul, et al., “Deep Neural Networks are Robust to Weight Binarization and Other Non-linear Distortions,” Jun. 7, 2016, 10 pages, arXiv:1606.01981v1, Computing Research Repository (CoRR)—Cornell University, Itha… [cited by applicant]
Moons, Bert, et al., “Envision: A 0.26-to-10TOPS/W Subword-Parallel Dynamic-Voltage-Accuracy-Frequency-Scalable Convolutional Neural Network Processor in 28nm FDSOI,” Proceedings of 2017 IEEE International Solid-State C… [cited by applicant]
Moshovos, Andreas, et al., “Exploiting Typical Values to Accelerate Deep Learning,” Computer, May 24, 2018, 13 pages, vol. 51—Issue 5, IEEE Computer Society, Washington, D.C. [cited by applicant]
Non-published commonly owned related U.S. Appl. No. 16/212,645 with similar specification, filed Dec. 6, 2018, 110 pages, Perceive Corporation. [cited by applicant]
Park, Jongsoo, et al., “Faster CNNs with Direct Sparse Convolutions and Guided Pruning,” Jul. 28, 2017, 12 pages, arXiv:1608.01409v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Patterson, David, et al., “Computer Organization and Design: The Hardware/Software Interface—Chapter 4—The Processor,” Fifth Edition, Sep. 26, 2013, 124 pages, Elsevier, Inc. [cited by applicant]
Pawar, A. B., “Radix-2 Vs Radix-4 High Speed Multiplier,” International Journal of Advanced Research in Computer Science and Software Engineering, Mar. 2015, 5 pages, vol. 5, No. 3, IJARCSSE. [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Proceedings of 2016 European Conference on Computer Vision (ECCV '16), Oct. 8-16, 2016, 17 pages, Lecture Note… [cited by applicant]
Ren, Mengye, et al., “SBNet: Sparse Blocks Network for Fast Inference,” Jan. 7, 2018, 10 pages, arXiv:1801.02108v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 12 pages, ICLR, Vanco… [cited by applicant]
Shin, Dongjoo, et al., “DNPU: An 8.1TOPS/W Reconfigurable CNN-RNN Processor for General-Purpose Deep Neural Networks,” Proceedings of 2017 IEEE International Solid-State Circuits Conference (ISSCC 2017), Feb. 5-7, 2017,… [cited by applicant]
Sim, Jaehyeong, et al., “A 1.42TOPS/W Deep Convolutional Neural Network Recognition Processor for Intelligent IoE Systems,” Proceedings of 2016 IEEE International Solid-State Circuits Conference (ISSCC 2016), Jan. 31-Fe… [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Wang, Min, et al., “Factorized Convolutional Neural Networks,” 2017 IEEE International Conference on Computer Vision Workshops (ICCVW '17), Oct. 22-29, 2017, 9 pages, IEEE, Venice, Italy. [cited by applicant]
Wang, Peiqi, et al., “HitNet: Hybrid Ternary Recurrent Neural Network,” 32nd Conference on Neural Information Processing Systems (NeurIPS '18), Dec. 2018, 11 pages, Montreal, Canada. [cited by applicant]
Wen, Wei, et al., “Learning Structured Sparsity in Deep Neural Networks,” Oct. 18, 2016, 10 pages, arXiv:1608.03665v4, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Yang, Xuan, et al., “DNN Dataflow Choice Is Overrated,” Sep. 10, 2018, 13 pages, arXiv:1809.04070v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Zhang, Shijin, et al., “Cambricon-X: An Accelerator for Sparse Neural Networks,” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO '16), Oct. 15-19, 2016, 12 pages, IEEE, Taipei, Taiwan. [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv:1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Liu, Shaoli, et al., “Cambricon: An Instruction Set Architecture for Neural Networks,” 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Jun. 18-22, 2016, 13 pages, IEEE, Seoul, South Korea. [cited by applicant]
Carbon, A., et al., “Pleura: A Scalable Energy-Efficient Programmable Hardware Accelerator for Neural Networks,” 2018 Design, Automation & Test in Europe Conference & Exhibition (Date 2018), Mar. 19-23, 2018, 6 pages, I… [cited by applicant]
Gokhale, Vinayak, et al., “Snowflake: A Model Agnostic Accelerator for Deep Convolutional Neural Networks,” Aug. 8, 2017, 11 pages, arXiv:1708.02579v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Jin, Canran,, et al., “Sparse Ternary Connect: Convolutional Neural Networks Using Ternarized Weights with Enhanced Sparsity,” 2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC), Jan. 22-25, 2018, 6… [cited by applicant]
2017 Han, Song, “Efficient Methods and Hardware for Deep Learning,” Sep. 2017, 125 pages, Stanford University, Palo Alto, CA, USA. [cited by applicant]