IP Library Granted Patent US 12,639,557
Granted Patent B1
US 12,639,557 · App. 17/543,474 · Granted May 26, 2026

Neural network inference circuit performing matrix multiplication

Inventors: Kenneth Duong (San Jose, CA); Jung Ko (San Jose, CA); Steven L. Teig (Menlo Park, CA); Brian Thomas (Vancouver, CA)
Assignee: Amazon Technologies, Inc.
G06N3/063G06F17/16G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,557
App. No.
17/543,474
Granted
May 26, 2026
Kind
B1
Abstract

Some embodiments provide a neural network inference circuit (NNIC) for executing a network having multiple layers. The NNIC includes multiple circuit sets. Each circuit set includes a dot product circuit to compute dot products between weight values and activation values for at least a subset of a first set of the layers, a math function circuit to compute values based on computations using activation values for at least a subset of a second set of the layers, and a post-processing circuit to receive (i) values output by the dot product circuit and (ii) values output by the math function circuit and to perform post-processing operations on the received values. The NNIC includes a set of accumulation circuits. Each accumulation circuit is to accumulate outputs of math function circuits for layers of the second set of layers that perform matrix multiplication of sets of activation values output by previous layers.

Claims (35)

1 . A neural network inference circuit for executing a neural network comprising a plurality of layers, the neural network inference circuit comprising:

a plurality of clusters, each cluster comprising a respective plurality of circuit sets, each circuit set of the respective plurality of circuit sets comprising:

a dot product circuit to compute and output dot products between weight values and activation values for at least a subset of a first set of the plurality of layers;

a math function circuit to compute and output values based on computations using activation values for at least a subset of a second set of the plurality of layers; and

a post-processing circuit to receive (i) values output by the dot product circuit for the subset of the first set of the plurality of layers and (ii) values output by the math function circuit for the subset of the second set of the plurality of layers and to perform post-processing operations on the received values; and

a set of accumulation circuits, each accumulation circuit to accumulate outputs of a plurality of math function circuits included in the respective plurality of circuit sets for layers of the second set of the plurality of layers that perform matrix multiplication of a first set of activation values output by a previous neural network layer with a second set of activation values output by another previous neural network layer.

2 . The neural network inference circuit of claim 1 further comprising a plurality of cores to store activation values.

3 . The neural network inference circuit of claim 2 , wherein each cluster of the plurality of clusters comprises a respective set of the plurality of cores.

4 . The neural network inference circuit of claim 3 , wherein the respective set of the plurality of cores (i) perform dot product computations between weight values stored in the respective set of the plurality of cores and the activation values stored in the respective set of the plurality of cores and (ii) provide results of the dot product computations to a dot product bus that accumulates the results from the respective set of the plurality of cores in a first and the respective set of the plurality of cores in a second cluster.

5 . The neural network inference circuit of claim 4 , wherein the dot product bus comprises a plurality of lanes, each respective lane connecting to a respective dot product circuit of a respective circuit set in each of the plurality of clusters and providing accumulated dot product computation results to respective circuit sets.

6 . The neural network inference circuit of claim 4 , wherein the math function circuit receives activation values directly from the respective set of the plurality of cores that the math function circuit is associated with.

7 . The neural network inference circuit of claim 1 , wherein a particular layer of the second set of the plurality of layers comprises a matrix multiplication of activation values of a first layer by activation values of a second layer.

8 . The neural network inference circuit of claim 7 , wherein during execution of the particular layer, each math function circuit of a set of math function circuits receives a respective first activation value from the first layer and a respective second activation value from the second layer.

9 . The neural network inference circuit of claim 8 , wherein:

the respective first activation value is received by each math function circuit in the set of math function circuits during a first clock cycle of the neural network inference circuit;

the respective second activation value is received by each math function circuit in the set of math function circuits during a second clock cycle of the neural network inference circuit; and

each math function circuit of the set of math function circuits multiplies the respective first activation value by the respective second activation value during the second clock cycle.

10 . The neural network inference circuit of claim 8 , wherein a particular one of the set of accumulation circuits adds outputs of the set of math function circuits.

11 . The neural network inference circuit of claim 10 , wherein:

each output value of the particular layer is based on a particular number of multiplications between activation values from the first layer and activation values from the second layer; and

when the particular number of multiplications is larger than a number of math function circuits that provide their outputs to the particular one of the set of accumulation circuits, each math function circuit in the set of math function circuits performs a plurality of multiplications between activation values from the first layer and activation values from the second layer that are accumulated together by the particular one of the set of accumulation circuits to generate an output value of the particular layer.

12 . The neural network inference circuit of claim 11 , wherein the particular one of the set of accumulation circuits comprises a register for storing intermediate accumulated values.

13 . The neural network inference circuit of claim 10 , wherein each accumulation circuit in the set of accumulation circuits receives outputs from a same number of math function circuits during execution of the particular layer.

14 . The neural network inference circuit of claim 8 , wherein:

execution of the particular layer generates a plurality of channels of output activation values;

first output activation channel is based on multiplication of a first channel of the first layer by a first channel of the second layer; and

a second output activation channel is based on multiplication of a second channel of the first layer by a second channel of the second layer.

15 . The neural network inference circuit of claim 14 , wherein:

a first subset of the set of math function circuits receives activation values from the first channel of the first layer and the first channel of the second layer;

a second subset of the set of math function circuits receives activation values from the second channel of the first layer and the second channel of the second layer;

a first accumulation circuit accumulates outputs of each math function circuit in the first subset of the set of math function circuits; and

a second accumulation circuit accumulates outputs of each math function circuit in the second subset of the set of math function circuits.

16 . The neural network inference circuit of claim 14 , wherein the activation values of the first channel of the first layer and the first channel of the second layer are stored in a first core of the neural network inference circuit and the activation values of the second channel of the first layer and the second channel of the second layer are stored in a second core of the neural network inference circuit.

17 . The neural network inference circuit of claim 1 , wherein the neural network comprises an attention mechanism.

18 . The neural network inference circuit of claim 17 , wherein at least one of the layers that perform matrix multiplication are part of the attention mechanism of the neural network.

Assignments (2)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
Continuity (1)
Provisional Application 63243686 · Sep 13, 2021
References Cited (218)
US 5621863A · Boulet et al. · 1997 [cited by applicant]
US 5717832A · Steimle et al. · 1998 [cited by applicant]
US 5740326A · Boulet et al. · 1998 [cited by applicant]
US 5761442A · Barr et al. · 1998 [cited by applicant]
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 6038583A · Oberman et al. · 2000 [cited by applicant]
US 6453206B1 · Soraghan et al. · 2002 [cited by applicant]
US 6463438B1 · Veltri et al. · 2002 [cited by applicant]
US 6601052B1 · Lee et al. · 2003 [cited by applicant]
US 7788196B2 · Buscema · 2010 [cited by applicant]
US 9710265B1 · Temam et al. · 2017 [cited by applicant]
US 9858636B1 · Lim et al. · 2018 [cited by applicant]
US 9904874B2 · Shoaib et al. · 2018 [cited by applicant]
US 10409604B2 · Kennedy et al. · 2019 [cited by applicant]
US 10445638B1 · Amirineni et al. · 2019 [cited by applicant]
US 10489478B2 · Lim et al. · 2019 [cited by applicant]
US 10515303B2 · Lie et al. · 2019 [cited by applicant]
US 10657438B2 · Lie et al. · 2020 [cited by applicant]
US 10740434B1 · Duong et al. · 2020 [cited by applicant]
US 10768856B1 · Diamant et al. · 2020 [cited by applicant]
US 10796198B2 · Franca-Neto · 2020 [cited by applicant]
US 10817042B2 · Desai et al. · 2020 [cited by applicant]
US 10853738B1 · Dockendorf et al. · 2020 [cited by applicant]
US 10867247B1 · Teig · 2020 [cited by applicant]
US 10936951B1 · Teig · 2021 [cited by applicant]
US 11049013B1 · Duong et al. · 2021 [cited by applicant]
US 11138292B1 · Nair et al. · 2021 [cited by applicant]
US 11170289B1 · Duong et al. · 2021 [cited by applicant]
US 11205115B1 · Duong et al. · 2021 [cited by applicant]
US 11222257B1 · Ko et al. · 2022 [cited by applicant]
US 11250326B1 · Ko et al. · 2022 [cited by applicant]
US 11347297B1 · Ko et al. · 2022 [cited by applicant]
US 11423289B2 · Judd et al. · 2022 [cited by applicant]
US 11531868B1 · Duong et al. · 2022 [cited by applicant]
US 11568227B1 · Ko et al. · 2023 [cited by applicant]
US 11586910B1 · Duong et al. · 2023 [cited by applicant]
US 11868898B2 · Teig · 2024 [cited by applicant]
US 20040078403A1 · Scheuermann et al. · 2004 [cited by applicant]
US 20110055308A1 · Mantor et al. · 2011 [cited by applicant]
US 20110307685A1 · Song · 2011 [cited by applicant]
US 20160086078A1 · Ji et al. · 2016 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342889A1 · Thorson · 2016 [cited by examiner]
US 20160342891A1 · Ross · 2016 [cited by examiner]
US 20160342892A1 · Ross · 2016 [cited by examiner]
US 20160342893A1 · Ross et al. · 2016 [cited by applicant]
US 20170011006A1 · Saber et al. · 2017 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170168775A1 · Tseng et al. · 2017 [cited by applicant]
US 20170243110A1 · Esquivel et al. · 2017 [cited by applicant]
US 20170300828A1 · Feng et al. · 2017 [cited by applicant]
US 20170323196A1 · Gibson et al. · 2017 [cited by applicant]
US 20180018559A1 · Yakopcic et al. · 2018 [cited by applicant]
US 20180025268A1 · Teig et al. · 2018 [cited by applicant]
US 20180046458A1 · Kuramoto · 2018 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180046905A1 · Li et al. · 2018 [cited by applicant]
US 20180046916A1 · Dally et al. · 2018 [cited by applicant]
US 20180101763A1 · Barnard et al. · 2018 [cited by applicant]
US 20180114569A1 · Strachan et al. · 2018 [cited by applicant]
US 20180121196A1 · Temam et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180164866A1 · Turakhia et al. · 2018 [cited by applicant]
US 20180181406A1 · Kuramoto · 2018 [cited by applicant]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180197049A1 · Tran et al. · 2018 [cited by applicant]
US 20180197068A1 · Narayanaswami et al. · 2018 [cited by applicant]
US 20180246855A1 · Redfern et al. · 2018 [cited by applicant]
US 20180285719A1 · Baum et al. · 2018 [cited by applicant]
US 20180285726A1 · Baum et al. · 2018 [cited by applicant]
US 20180285727A1 · Baum et al. · 2018 [cited by applicant]
US 20180285736A1 · Baum et al. · 2018 [cited by applicant]
US 20180293490A1 · Ma et al. · 2018 [cited by applicant]
US 20180293493A1 · Kalamkar et al. · 2018 [cited by applicant]
US 20180293691A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180300600A1 · Ma et al. · 2018 [cited by applicant]
US 20180307494A1 · Ould-Ahmed-Vall et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180307985A1 · Appu et al. · 2018 [cited by applicant]
US 20180308202A1 · Appu et al. · 2018 [cited by applicant]
US 20180314492A1 · Fais et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180322095A1 · Longley et al. · 2018 [cited by applicant]
US 20180322386A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180322387A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180329868A1 · Chen et al. · 2018 [cited by applicant]
US 20180365794A1 · Lee et al. · 2018 [cited by applicant]
US 20180373975A1 · Yu et al. · 2018 [cited by applicant]
US 20190012296A1 · Hsieh et al. · 2019 [cited by applicant]
US 20190026078A1 · Bannon et al. · 2019 [cited by applicant]
US 20190026237A1 · Talpes et al. · 2019 [cited by applicant]
US 20190026249A1 · Talpes et al. · 2019 [cited by applicant]
US 20190041961A1 · Desai et al. · 2019 [cited by applicant]
US 20190057036A1 · Mathuriya et al. · 2019 [cited by applicant]
US 20190065453A1 · Bulgakov et al. · 2019 [cited by applicant]
US 20190073585A1 · Pu et al. · 2019 [cited by applicant]
US 20190087713A1 · Lamb et al. · 2019 [cited by applicant]
US 20190095776A1 · Kfir et al. · 2019 [cited by applicant]
US 20190114499A1 · Delaye et al. · 2019 [cited by applicant]
US 20190138891A1 · Kim et al. · 2019 [cited by applicant]
US 20190147338A1 · Pau et al. · 2019 [cited by applicant]
US 20190156180A1 · Nomura et al. · 2019 [cited by applicant]
US 20190171927A1 · Diril et al. · 2019 [cited by applicant]
US 20190179635A1 · Jiao et al. · 2019 [cited by applicant]
US 20190180167A1 · Huang et al. · 2019 [cited by applicant]
US 20190187983A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20190196970A1 · Han et al. · 2019 [cited by applicant]
US 20190205094A1 · Diril et al. · 2019 [cited by applicant]
US 20190205358A1 · Diril et al. · 2019 [cited by applicant]
US 20190205736A1 · Bleiweiss et al. · 2019 [cited by applicant]
US 20190205739A1 · Liu et al. · 2019 [cited by applicant]
US 20190205740A1 · Judd et al. · 2019 [cited by applicant]
US 20190205780A1 · Sakaguchi · 2019 [cited by applicant]
US 20190236437A1 · Shin et al. · 2019 [cited by applicant]
US 20190236445A1 · Das et al. · 2019 [cited by applicant]
US 20190266217A1 · Arakawa et al. · 2019 [cited by applicant]
US 20190266479A1 · Singh et al. · 2019 [cited by applicant]
US 20190272317A1 · Wroczynski et al. · 2019 [cited by applicant]
US 20190294413A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294959A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294968A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190303741A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303749A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303750A1 · Kumar et al. · 2019 [cited by applicant]
US 20190325296A1 · Fowers et al. · 2019 [cited by applicant]
US 20190332925A1 · Modha · 2019 [cited by applicant]
US 20190347559A1 · Kang et al. · 2019 [cited by applicant]
US 20190385046A1 · Cassidy et al. · 2019 [cited by applicant]
US 20200005131A1 · Nakahara et al. · 2020 [cited by applicant]
US 20200042856A1 · Datta et al. · 2020 [cited by applicant]
US 20200042859A1 · Mappouras et al. · 2020 [cited by applicant]
US 20200050941A1 · Zhuang et al. · 2020 [cited by applicant]
US 20200089506A1 · Power et al. · 2020 [cited by applicant]
US 20200134461A1 · Chai et al. · 2020 [cited by applicant]
US 20200234114A1 · Rakshit et al. · 2020 [cited by applicant]
US 20200249996A1 · Addepalli et al. · 2020 [cited by applicant]
US 20200257930A1 · Nahr et al. · 2020 [cited by applicant]
US 20200301668A1 · Li · 2020 [cited by applicant]
US 20200311207A1 · Kim et al. · 2020 [cited by applicant]
US 20200364545A1 · Shattil · 2020 [cited by applicant]
US 20200380344A1 · Lie et al. · 2020 [cited by applicant]
US 20210110236A1 · Shibata · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210241082A1 · Nagy et al. · 2021 [cited by applicant]
US 20220121914A1 · Huang et al. · 2022 [cited by applicant]
US 20220335562A1 · Surti et al. · 2022 [cited by applicant]
CN 108876698A · 2018 [cited by applicant]
CN 108280514B · 2020 [cited by applicant]
GB 2568086A · 2019 [cited by applicant]
WO 2020044527A1 · 2020 [cited by applicant]
Ardakani, Arash, et al., “An Architecture to Accelerate Convolution in Deep Neural Networks,” IEEE Transactions on Circuits and Systems I: Regular Papers, Oct. 17, 2017, 14 pages, vol. 65, No. 4, IEEE. [cited by applicant]
Abtahi, Tahmid, et al., “Accelerating Convolutional Neural Network With FFT on Embedded Hardware,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Sep. 2018, 14 pages, vol. 26, No. 9, IEEE. [cited by applicant]
Bilgili, Erdem, et al., “Applications of CNN with Trapezoidal Activation Function,” Springer Proceedings in Physics: Complex Computing-Networks, Jan. 2006, 9 pages, vol. 104, Springer, Berlin, Germany. [cited by applicant]
Carbon, A., et al., “Pleura: A Scalable Energy-Efficient Programmable Hardware Accelerator for Neural Networks,” 2018 Design, Automation & Test in Europe Conference & Exhibition (Date 2018), Mar. 19-23, 2018, 6 pages, I… [cited by applicant]
Chen, Guanrong, “Chaotification via Feedback Control: Theories, Methods, and Applications,” 2003 IEEE International Workshop on Workload Characterization, Aug. 20-22, 2003, 7 pages, IEEE, Saint Petersburg, Russia. [cited by applicant]
Gokhale, Vinayak, et al., “Snowflake: A Model Agnostic Accelerator for Deep Convolutional Neural Networks,” Aug. 8, 2017, 11 pages, arXiv: 1708.02579v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, N… [cited by applicant]
Jin, Canran,, et al., “Sparse Ternary Connect: Convolutional Neural Networks Using Ternarized Weights with Enhanced Sparsity,” 2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC), Jan. 22-25, 2018, 6… [cited by applicant]
Karan, Oguz, et al., “Diagnosing Diabetes using Neural Networks on Small Mobile Devices,” Expert Systems with Applications, Jan. 2012, 7 pages, vol. 39, Issue 1, Elsevier, Ltd. [cited by applicant]
Koehn, Philipp, “Combining Genetic Algorithms and Neural Networks: The Encoding Problem,” Dec. 1994, 2 pages, University of Tennessee, Knoxville, Tennessee, USA. [cited by applicant]
Kubosawa, Shunpei, “Neural Network and Computer Program Therefor,” May 4, 2015, 31 pages, National Institute of Information & Communications Technology. [cited by applicant]
Sopena, Josep M., et al., “Neural Networks with Periodic and Monotonic Activation Functions: A Comparative Study in Classification Problems,” 1999 Ninth International Conference on Artificial Neural Networks ICANN 99 (C… [cited by applicant]
Zeiler, M. D., et al., “On Rectified Linear Units for Speech Processing,” 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2013, 5 pages, IEEE. [cited by applicant]
Zhu, C., et al., “A Fourier Series Neural Network and Its Application to System Identification,” Journal of Dynamic Systems, Measurement, and Control, Sep. 1995, 9 pages, vol. 117, ASME. [cited by applicant]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Andri, Renzo, et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Mar. 14, 2017, 14 pages, IEEE, New York,… [cited by applicant]
Ardakani, Arash, et al., “Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks,” Proceedings of the 5th International Conference on Learning Representations (ICLR 2017), Apr.… [cited by applicant]
Bagherinezhad, Hessam, et al., “LCNN: Look-up Based Convolutional Neural Network,” Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 10 pages, IEEE, Honolulu, … [cited by applicant]
Bang, Suyoung, et al., “A 288μW Programmable Deep-Learning Processor with 270KB On-Chip Weight Storage Using Non-Uniform Memory Hierarchy for Mobile Intelligence,” Proceedings of 2017 IEEE International Solid-State Circ… [cited by applicant]
Bong, Kyeongryeol, et al., “A 0.62mW Ultra-Low-Power Convolutional-Neural-Network Face-Recognition Processor and a CIS Integrated with Always-On Haar-Like Face Detector,” Proceedings of 2017 IEEE International Solid-Sta… [cited by applicant]
Boo, Yoonho, et al., “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations,” 2017 IEEE Workshop on Signal Processing Systems (SiPS), Oct. 3-5, 2017, 6 pages, IEEE, Lorie… [cited by applicant]
Bruns, Erich, et al., “Mobile Phone-Enabled Museum Guidance with Adaptive Classification,” IEEE Computer Graphics and Applications, Jul. 9, 2008, 5 pages, vol. 28, Issue 4, IEEE. [cited by applicant]
Chakradhar, Srimat T., et al., “Toward Massively Parallel Automatic Test Generation,” IEEE Transactions on Computer-Aided Design, Sep. 1990, 14 pages, vol. 9, Issue 9, IEEE. [cited by applicant]
Chandra, Pravin, et al., “An Activation Function Adapting Training Algorithm for Sigmoidal Feedforward Networks,” Neurocomputing, Jun. 25, 2004, 9 pages, vol. 61, Elsevier. [cited by applicant]
Chen, Yu-Hsin, et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA 2… [cited by applicant]
Chen, Yu-Hsin, et al., “Using Dataflow to Optimize Energy Efficiency of Deep Neural Network Accelerators,” IEEE Micro, Jun. 14, 2017, 10 pages, vol. 37, Issue 3, IEEE, New York, NY, USA. [cited by applicant]
Courbariaux, Matthieu, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” Mar. 17, 2016, 11 pages, arXiv:1602.02830v3, Computing Research Repository (CoRR… [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations, ” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 15)… [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Fu, Yao, et al., “Embedded Vision with INT8 Optimization on Xilinx Devices,” WP490 (v1.0.1), Apr. 19, 2017, 15 pages, Xilinx, Inc., San Jose, CA, USA. [cited by applicant]
Guo, Yiwen, et al., “Network Sketching: Exploring Binary Structure in Deep CNNs,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 9 pages, IEEE, Honolulu, HI. [cited by applicant]
Hamadneh, Nawaf, et al., “Learning Logic Programming in Radial Basis Function Network via Genetic Algorithm,” Journal of Applied Sciences, Sep. 2012, 9 pages, vol. 12, Issue 9, Asian Network for Scientific Information. [cited by applicant]
He, Zhezhi, et al., “Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy,” Jul. 20, 2018, 8 pages, arXiv:1807.07948v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Hegde, Kartik, et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition,” Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA '18), Jun. 2-6, 2018, 14… [cited by applicant]
Huan, Yuxiang, et al., “A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity,” May 22, 2017, 5 pages, arXiv:1705.08009v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, … [cited by applicant]
Jain, Anil K., et al., “Artificial Neural Networks: A Tutorial,” Computer, Mar. 1996, 14 pages, vol. 29, Issue 3, IEEE. [cited by applicant]
Jouppi, Norman, P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17), Jun. 24-28, 2017, 17 pages, ACM, … [cited by applicant]
Judd, Patrick, et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing,” Apr. 29, 2017, 6 pages, arXiv:1705.00125v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, U… [cited by applicant]
Kang, Miao, et al., “Snap-drift ADaptive FUnction Neural Network (SADFUNN) for Optical and Pen-Based Handwritten Digit Recognition,” Proceedings of 10th International Conference on Engineering Applications of Neural Net… [cited by applicant]
Leng, Cong, et al., “Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM,” Proceedings of 32nd AAAI Conference on Artificial Intelligence (AAAI-18), Feb. 2-7, 2018, 16 pages, Association for the Advance… [cited by applicant]
Li, Fengfu, et al., “Ternary Weight Networks,” May 16, 2016, 9 pages, arXiv:1605.04711v1, Computing Research Repository (CoRR) —Cornell University, Ithaca, NY, USA. [cited by applicant]
Li, Hong-Xing, et al., “Interpolation Functions of Feedforward Neural Networks,” Computers & Mathematics with Applications, Dec. 2003, 14 pages, vol. 46, Issue 12, Elsevier Ltd. [cited by applicant]
Merolla, Paul, et al., “Deep Neural Networks are Robust to Weight Binarization and Other Non-linear Distortions,” Jun. 7, 2016, 10 pages, arXiv:1606.01981v1, Computing Research Repository (CoRR)—Cornell University, Itha… [cited by applicant]
Moons, Bert, et al., “Envision: A 0.26-to-10TOPS/W Subword-Parallel Dynamic-Voltage-Accuracy-Frequency-Scalable Convolutional Neural Network Processor in 28nm FDSOI,” Proceedings of 2017 IEEE International Solid-State C… [cited by applicant]
Moshovos, Andreas, et al., “Exploiting Typical Values to Accelerate Deep Learning,” Computer, May 24, 2018, 13 pages, vol. 51—Issue 5, IEEE Computer Society, Washington, D.C. [cited by applicant]
Non-Published Commonly Owned Related U.S. Appl. No. 17/543,446 with similar specification, filed Dec. 6, 2021, 129 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned Related U.S. Appl. No. 17/543,471 with similar specification, filed Dec. 6, 2021, 129 pages, Perceive Corporation. [cited by applicant]
Park, Jongsoo, et al., “Faster CNNs with Direct Sparse Convolutions and Guided Pruning,” Jul. 28, 2017, 12 pages, arXiv:1608.01409v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Pedrycz, Witold, et al., “fXOR Fuzzy Logic Networks,” Soft Computing, Dec. 2002, 15 pages, vol. 7, Issue 2, Springer-Verlag. [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Proceedings of 2016 European Conference on Computer Vision (ECCV '16), Oct. 8-16, 2016, 17 pages, Lecture Note… [cited by applicant]
Ren, Mengye, et al., “SBNet: Sparse Blocks Network for Fast Inference,” Jan. 7, 2018, 10 pages, arXiv:1801.02108v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 12 pages, ICLR, Vanco… [cited by applicant]
Shin, Dongjoo, et al., “DNPU: An 8.1TOPS/W Reconfigurable CNN-RNN Processor for General-Purpose Deep Neural Networks,” Proceedings of 2017 IEEE International Solid-State Circuits Conference (ISSCC 2017), Feb. 5-7, 2017,… [cited by applicant]
Sim, Jaehyeong, et al., “A 1.42TOPS/W Deep Convolutional Neural Network Recognition Processor for Intelligent IoE Systems,” Proceedings of 2016 IEEE International Solid-State Circuits Conference (ISSCC 2016), Jan. 31-Fe… [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Tan, Chew Lim, et al., “An Artificial Neural Network that Models Human Decision Making,” IEEE Computer, Mar. 1996, 7 pages, vol. 29, Issue 3, IEEE. [cited by applicant]
Varvak, Mark S., “Pattern Classification Using Radial Basis Function Neural Networks Enhanced with the Rvachev Function Method,” Proceedings of the 16th Iberoamerican Congress Conference on Progress in Pattern Recogniti… [cited by applicant]
Wang, Min, et al., “Factorized Convolutional Neural Networks,” 2017 IEEE International Conference on Computer Vision Workshops (ICCVW '17), Oct. 22-29, 2017, 9 pages, IEEE, Venice, Italy. [cited by applicant]
Wen, Bo, “Formulation and Modeling Approaches for Piecewise Linear Membership Functions in Fuzzy Nonlinear Programming,” Information Technology Journal, Mar. 21, 2014, 13 pages, vol. 13, Issue 9, SPARC. [cited by applicant]
Wen, Wei, et al., “Learning Structured Sparsity in Deep Neural Networks,” Oct. 18, 2016, 10 pages, arXiv:1608.03665v4, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Yang, Xuan, et al., “DNN Dataflow Choice Is Overrated,” Sep. 10, 2018, 13 pages, arXiv:1809.04070v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Zhang, Shijin, et al., “Cambricon-X: An Accelerator for Sparse Neural Networks,” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO '16), Oct. 15-19, 2016, 12 pages, IEEE, Taipei, Taiwan. [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv:1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Agostinelli, Forest, et al., “Learning Activation Functions to Improve Deep Neural Networks,” Apr. 21, 2015, 9 pages, retrieved from https://arxiv.org/abs/1412.6830. [cited by applicant]
Aizenberg, Igor, “Periodic Activation Function and a Modified Learning Algorithm for the Multivalued Neuron,” IEEE Transactions on Neural Networks, Dec. 2010, 11 pages, vol. 21, No. 12, IEEE. [cited by applicant]
Liu, Shaoli, et al., “Cambricon: An Instruction Set Architecture for Neural Networks,” 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Jun. 18-22, 2016, 13 pages, IEEE, Seoul, South Korea. [cited by applicant]