IP Library Granted Patent US 12,518,146
Granted Patent B1
US 12,518,146 · App. 16/717,925 · Granted Jan 6, 2026

Address decoding by neural network inference circuit read controller

Inventors: Jung Ko (San Jose, CA); Kenneth Duong (San Jose, CA); Steven L. Teig (Menlo Park, CA)
Assignee: Amazon Technologies, Inc.
G06N3/06G06F9/4401G06F12/0246G06F17/16G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,146
App. No.
16/717,925
Granted
Jan 6, 2026
Kind
B1
Abstract

Some embodiments provide a neural network inference circuit (NNIC) for executing a neural network (NN) that includes computation nodes at multiple layers. The NNIC includes a set of processing circuits for executing the computation nodes of the NN, a set of memories for storing data used by the processing circuits to execute the NN layers, and a read controller for retrieving the data from the memories for use by the processing circuits. The data is stored in the memories as multiple varying-size blocks. The read controller receives read instructions for a requested block of data to be used by the processing circuits for one or more computation nodes. The read instructions include a base memory address for multiple blocks of data, a size of the requested block of data, and a location of the requested block of data within the multiple blocks of data.

Claims (29)

1 . A neural network inference circuit for executing a neural network, the neural network comprising plurality of computation nodes at a plurality of layers, each of a set of the computation nodes comprising a dot product of input values and weight values, the neural network inference circuit comprising:

a set of processing circuits for executing the computation nodes of the neural network;

a set of memories for storing a plurality of varying-size blocks of input values and encoded weight data used by the set of processing circuits to execute the neural network layers;

a read controller for retrieving the data from the set of memories for use by the set of processing circuits; and

a first computing component for (i) providing instructions to the set of processing circuits and the read controller and (ii) providing data retrieved from the set of memories by the read controller to the set of processing circuits for the set of processing circuits to execute the computation nodes of the neural network,

wherein the read controller receives sets of read instructions for requested blocks of data to be used by the set of processing circuits for one or more computation nodes, each set of read instructions comprising (i) a base memory address for a plurality of blocks of data, (ii) a size of the requested block of data, and (iii) a location of the requested block of data within the plurality of blocks of data,

wherein the weight values for a particular layer are organized as a plurality of filters, each respective computation node of the particular layer computing a dot product of a respective group of input values to the particular layer and the weight values of a respective one of the filters, each filter stored in the set of memories as a separate block of encoded weight data, at least two different filters of the particular layer stored in the set of memories as two differently-sized blocks using different amounts of memory,

wherein a first set of read instructions for retrieving encoded weight data for a particular filter of the particular layer specifies a first block size, and wherein data retrieved using the first set of read instructions is used by the read controller to determine a second block size, wherein the second block size represents an amount by which the particular filter extends beyond a word boundary in the set of memories, and wherein a second set of read instructions for retrieving the encoded weight data for the particular filter specifies the second block size.

2 . The neural network inference circuit of claim 1 , wherein (i) the encoded weight data is stored in the set of memories at boot-up of the neural network inference circuit and (ii) the input values for each layer are stored in the set of memories during the execution of a previous layer.

3 . The neural network inference circuit of claim 2 , wherein for the particular layer, the size of the blocks of data is constant for blocks of the input values for the particular layer.

4 . The neural network inference circuit of claim 1 , wherein the plurality of filters are for convolutional operations.

5 . The neural network inference circuit of claim 1 , wherein the second block size is determined based on:

determining a first difference between a length of a first word and a length of other filters of the first word apart from a portion of the particular filter in the first word; and

subtracting the first difference from the first block size.

6 . The neural network inference circuit of claim 1 , wherein the input values for the particular layer are arranged in a plurality of two-dimensional grids, wherein each requested block of data comprising input values for the particular layer comprises at most one input value from each of the two-dimensional grids.

7 . The neural network inference circuit of claim 6 , wherein a particular requested block of data comprises one input value from each of a set of the two-dimensional grids, said input values in the particular requested block of data having a same set of coordinates in each respective two-dimensional grid of the set of two-dimensional grids.

8 . The neural network inference circuit of claim 7 , wherein the size of the particular requested block of data is based on a number of two-dimensional grids into which the input values for the particular layer are arranged.

9 . The neural network inference circuit of claim 1 , wherein the set of memories comprises random access memory (RAM), wherein the base memory address specifies a beginning of a particular fixed-length RAM word in the set of memories.

10 . The neural network inference circuit of claim 9 , wherein the location of the requested block of data specifies a number of blocks of data from the beginning of the particular fixed-length RAM word at which the requested block of data begins.

11 . The neural network inference circuit of claim 9 , wherein the read controller outputs to the first computing component, in response to each respective set of read instructions, a respective set of data having a fixed length of a RAM word, wherein the respective requested block of data is aligned to an edge of the respective output set of data.

12 . The neural network inference circuit of claim 11 , wherein when the size of the requested block of data is less than the fixed length of a RAM word, the output set of data comprises (i) the requested block of data and (ii) random data that is not used by the set of processing circuits.

13 . The neural network inference circuit of claim 1 , wherein the neural network inference circuit comprises a plurality of cores, each respective core comprising:

a respective subset of the processing circuits;

a respective subset of the memories; and

a respective read controller for retrieving data from the respective subset of memories for use by the respective subset of the processing circuits.

14 . The neural network inference circuit of claim 13 , wherein a subset of the cores are active for the particular layer, wherein each set of read instructions during execution of the particular layer is received by the respective read controller of each active core.

15 . The neural network inference circuit of claim 13 , wherein all of the memories belong to the plurality of cores, wherein an additional subset of the processing circuits used to execute the computation nodes of the neural network are separate from the plurality of cores.

16 . The neural network inference circuit of claim 9 , wherein the first block size is a fixed length of the RAM words.

17 . The neural network inference circuit of claim 1 , wherein the at least two different filters of the particular layer are stored as differently-sized blocks based on the different filters having different numbers of non-zero weight values.

Assignments (3)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2020
From: KO, JUNG; DUONG, KENNETH; TEIG, STEVEN L.
To: PERCEIVE CORPORATION
Reel/Frame 051851/0113 →
Continuity (15)
Continuation In Part 16212643 · Dec 6, 2018
Continuation In Part 16212617 · Dec 6, 2018
Continuation In Part 16120387 · Sep 3, 2018
Provisional Application 62946188 · Dec 10, 2019
Provisional Application 62886888 · Aug 14, 2019
Provisional Application 62873804 · Jul 12, 2019
Provisional Application 62853128 · May 27, 2019
Provisional Application 62797910 · Jan 28, 2019
Provisional Application 62792123 · Jan 14, 2019
Provisional Application 62773162 · Nov 29, 2018
Provisional Application 62773164 · Nov 29, 2018
Provisional Application 62753878 · Oct 31, 2018
Provisional Application 62742802 · Oct 8, 2018
Provisional Application 62724589 · Aug 29, 2018
Provisional Application 62660914 · Apr 20, 2018
References Cited (183)
US 5386531A · Blaner · 1995 [cited by examiner]
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 9710265B1 · Temam et al. · 2017 [cited by applicant]
US 9858636B1 · Lim et al. · 2018 [cited by applicant]
US 9904874B2 · Shoaib et al. · 2018 [cited by applicant]
US 10445638B1 · Amirineni et al. · 2019 [cited by applicant]
US 10489478B2 · Lim et al. · 2019 [cited by applicant]
US 10515303B2 · Lie et al. · 2019 [cited by applicant]
US 10657438B2 · Lie et al. · 2020 [cited by applicant]
US 10664310B2 · Bokhari et al. · 2020 [cited by applicant]
US 10768856B1 · Diamant et al. · 2020 [cited by applicant]
US 10796198B2 · Franca-Neto · 2020 [cited by applicant]
US 10817042B2 · Desai et al. · 2020 [cited by applicant]
US 10853738B1 · Dockendorf et al. · 2020 [cited by applicant]
US 10970630B1 · Aimone · 2021 [cited by applicant]
US 11138292B1 · Nair et al. · 2021 [cited by applicant]
US 11423289B2 · Judd et al. · 2022 [cited by applicant]
US 11537853B1 · Afzal et al. · 2022 [cited by applicant]
US 11568227B1 · Ko et al. · 2023 [cited by applicant]
US 11868867B1 · Afzal et al. · 2024 [cited by applicant]
US 20040078403A1 · Scheuermann et al. · 2004 [cited by applicant]
US 20110307685A1 · Song · 2011 [cited by applicant]
US 20150339570A1 · Scheffler et al. · 2015 [cited by applicant]
US 20160086078A1 · Ji et al. · 2016 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342893A1 · Ross et al. · 2016 [cited by applicant]
US 20170011006A1 · Saber et al. · 2017 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170109165A1 · Batley · 2017 [cited by examiner]
US 20170243110A1 · Esquivel et al. · 2017 [cited by applicant]
US 20170323196A1 · Gibson et al. · 2017 [cited by applicant]
US 20170344882A1 · Ambrose et al. · 2017 [cited by applicant]
US 20180018559A1 · Yakopcic et al. · 2018 [cited by applicant]
US 20180046458A1 · Kuramoto · 2018 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180046905A1 · Li et al. · 2018 [cited by applicant]
US 20180046916A1 · Dally et al. · 2018 [cited by applicant]
US 20180101763A1 · Barnard et al. · 2018 [cited by applicant]
US 20180114569A1 · Strachan et al. · 2018 [cited by applicant]
US 20180121196A1 · Temam · 2018 [cited by examiner]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180164866A1 · Turakhia et al. · 2018 [cited by applicant]
US 20180181406A1 · Kuramoto · 2018 [cited by applicant]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180197068A1 · Narayanaswami · 2018 [cited by examiner]
US 20180246855A1 · Redfern et al. · 2018 [cited by applicant]
US 20180285719A1 · Baum et al. · 2018 [cited by applicant]
US 20180285726A1 · Baum et al. · 2018 [cited by applicant]
US 20180285727A1 · Baum et al. · 2018 [cited by applicant]
US 20180285736A1 · Baum et al. · 2018 [cited by applicant]
US 20180293490A1 · Ma et al. · 2018 [cited by applicant]
US 20180293493A1 · Kalamkar et al. · 2018 [cited by applicant]
US 20180293691A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180300600A1 · Ma et al. · 2018 [cited by applicant]
US 20180307494A1 · Ould-Ahmed-Vall et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180307985A1 · Appu et al. · 2018 [cited by applicant]
US 20180308202A1 · Appu et al. · 2018 [cited by applicant]
US 20180314492A1 · Fais et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180322386A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180322387A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180329868A1 · Chen et al. · 2018 [cited by applicant]
US 20180365794A1 · Lee et al. · 2018 [cited by applicant]
US 20180373975A1 · Yu et al. · 2018 [cited by applicant]
US 20190012296A1 · Hsieh et al. · 2019 [cited by applicant]
US 20190026078A1 · Bannon et al. · 2019 [cited by applicant]
US 20190026237A1 · Talpes et al. · 2019 [cited by applicant]
US 20190026249A1 · Talpes et al. · 2019 [cited by applicant]
US 20190041961A1 · Desai et al. · 2019 [cited by applicant]
US 20190057036A1 · Mathuriya et al. · 2019 [cited by applicant]
US 20190065937A1 · Nowatzyk · 2019 [cited by examiner]
US 20190073585A1 · Pu et al. · 2019 [cited by applicant]
US 20190087713A1 · Lamb et al. · 2019 [cited by applicant]
US 20190095776A1 · Kfir et al. · 2019 [cited by applicant]
US 20190108436A1 · David · 2019 [cited by examiner]
US 20190114499A1 · Delaye et al. · 2019 [cited by applicant]
US 20190114534A1 · Teng · 2019 [cited by applicant]
US 20190138891A1 · Kim et al. · 2019 [cited by applicant]
US 20190147338A1 · Pau et al. · 2019 [cited by applicant]
US 20190156180A1 · Nomura et al. · 2019 [cited by applicant]
US 20190171927A1 · Diril et al. · 2019 [cited by applicant]
US 20190179635A1 · Jiao et al. · 2019 [cited by applicant]
US 20190180167A1 · Huang et al. · 2019 [cited by applicant]
US 20190187983A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20190196970A1 · Han et al. · 2019 [cited by applicant]
US 20190205094A1 · Diril · 2019 [cited by examiner]
US 20190205358A1 · Diril et al. · 2019 [cited by applicant]
US 20190205736A1 · Bleiweiss et al. · 2019 [cited by applicant]
US 20190205739A1 · Liu et al. · 2019 [cited by applicant]
US 20190205740A1 · Judd et al. · 2019 [cited by applicant]
US 20190205780A1 · Sakaguchi · 2019 [cited by applicant]
US 20190236437A1 · Shin et al. · 2019 [cited by applicant]
US 20190236445A1 · Das et al. · 2019 [cited by applicant]
US 20190266217A1 · Arakawa et al. · 2019 [cited by applicant]
US 20190266479A1 · Singh et al. · 2019 [cited by applicant]
US 20190294413A1 · Vantrease · 2019 [cited by examiner]
US 20190294959A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294968A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190303741A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303749A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303750A1 · Kumar et al. · 2019 [cited by applicant]
US 20190325296A1 · Fowers et al. · 2019 [cited by applicant]
US 20190332925A1 · Modha · 2019 [cited by applicant]
US 20190340493A1 · Coenen et al. · 2019 [cited by applicant]
US 20190347559A1 · Kang et al. · 2019 [cited by applicant]
US 20190385046A1 · Cassidy et al. · 2019 [cited by applicant]
US 20200005131A1 · Nakahara et al. · 2020 [cited by applicant]
US 20200042856A1 · Datta et al. · 2020 [cited by applicant]
US 20200042859A1 · Mappouras et al. · 2020 [cited by applicant]
US 20200089506A1 · Power et al. · 2020 [cited by applicant]
US 20200134461A1 · Chai et al. · 2020 [cited by applicant]
US 20200234114A1 · Rakshit et al. · 2020 [cited by applicant]
US 20200257930A1 · Nahr et al. · 2020 [cited by applicant]
US 20200301668A1 · Li · 2020 [cited by applicant]
US 20200364545A1 · Shattil · 2020 [cited by applicant]
US 20200380344A1 · Lie et al. · 2020 [cited by applicant]
US 20210110236A1 · Shibata · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210232897A1 · Bichler et al. · 2021 [cited by applicant]
US 20210241082A1 · Nagy et al. · 2021 [cited by applicant]
US 20220004854A1 · Lee et al. · 2022 [cited by applicant]
US 20220121914A1 · Huang et al. · 2022 [cited by applicant]
US 20220335562A1 · Surti et al. · 2022 [cited by applicant]
CN 108876698A · 2018 [cited by applicant]
CN 108280514B · 2020 [cited by applicant]
GB 2568086A · 2019 [cited by applicant]
WO 2020044527A1 · 2020 [cited by applicant]
Liu et al., Cambricon: An Instruction Set Architecture for Neural Networks, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, IEEE 2016, pp. 393-405 (Year: 2016). [cited by examiner]
Drepper, Ulrich, What Every Programmer Should Know About Memory, Nov. 21, 2007, 114 pages (Year: 2007). [cited by examiner]
Boo, Yoonho, et al., “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations,” 2017 IEEE Workshop on Signal Processing Systems (SiPS), Oct. 3-5, 2017, 6 pages, IEEE, Lorie… [cited by applicant]
Abtahi, Tahmid, et al., “Accelerating Convolutional Neural Network With FFT on Embedded Hardware,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Sep. 2018, 14 pages, vol. 26, No. 9, IEEE. [cited by applicant]
Ardakani, Arash, et al., “An Architecture to Accelerate Convolution in Deep Neural Networks,” IEEE Transactions on Circuits and Systems I: Regular Papers, Apr. 2018, 14 pages, vol. 65, No. 4, IEEE. [cited by applicant]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Andri, Renzo, et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Mar. 14, 2017, 14 pages, IEEE, New York,… [cited by applicant]
Bang, Suyoung, et al., “A 288pW Programmable Deep-Learning Processor with 270KB On-Chip Weight Storage Using Non-Uniform Memory Hierarchy for Mobile Intelligence,” Proceedings of 2017 IEEE International Solid-State Circ… [cited by applicant]
Bong, Kyeongryeol, et al., “A 0.62mW Ultra-Low-Power Convolutional-Neural-Network Face-Recognition Processor and a CIS Integrated with Always-On Haar-Like Face Detector,” Proceedings of 2017 IEEE International Solid-Sta… [cited by applicant]
Chen, Yu-Hsin, et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA 2… [cited by applicant]
Chen, Yu-Hsin, et al., “Using Dataflow to Optimize Energy Efficiency of Deep Neural Network Accelerators,” IEEE Micro, Jun. 14, 2017, 10 pages, vol. 37, Issue 3, IEEE, New York, NY, USA. [cited by applicant]
Cho, Minsik, et al., “MEC: Memory-Efficient Convolution for Deep Neural Network,” Jun. 21, 2017, 10 pages, arXiv:1706.06873v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Courbariaux, Matthieu, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” Mar. 17, 2016, 11 pages, arXiv:1602.02830v3, Computing Research Repository (CoRR… [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations, ” 9 pages, MIT Press, Montreal, Canada. arXiv:1511.00363v3, Apr. 18, 2016. [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Gao, Mingyu, et al., “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,” Proceedings of the 22nd International Conference on Architectural Support for Programming Languages and Operating Systems… [cited by applicant]
Guo, Yiwen, et al., “Network Sketching: Exploring Binary Structure in Deep CNNs,” 9 pages, arXiv:1706.02021v1, Jun. 7, 2017. [cited by applicant]
He, Zhezhi, et al., “Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy,” Jul. 20, 2018, 8 pages, arXiv:1807.07948v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Hegde, Kartik, et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition,” 14 pages, arXiv:1804.06508v1, Apr. 18, 2018. [cited by applicant]
Horowitz, Mark, “Computing's Energy Problem (and what can we do about it),” 2014 IEEE International Solid-State Circuits Conference (ISSCC 2014), Feb. 9-13, 2014, 5 pages, IEEE, San Francisco, CA, USA. [cited by applicant]
Huan, Yuxiang, et al., “A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity,” May 22, 2017, 5 pages, arXiv:1705.08009v1, Computer Research Repository (CoRR)—Comell University, Ithaca, NY, U… [cited by applicant]
Jeon, Dongsuk, et al., “A 23-mW Face Recognition Processor with Mostly-Read 5T Memory in 40-nm CMOS,” IEEE Journal on Solid-State Circuits, Jun. 2017, 15 pages, vol. 52, No. 6, IEEE, New York, NY, USA. [cited by applicant]
Jouppi, Norman, P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17), Jun. 24-28, 2017, 17 pages, ACM, … [cited by applicant]
Judd, Patrick, et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing,” Apr. 29, 2017, 6 pages, arXiv:1705.00125v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, U… [cited by applicant]
Leng, Cong, et al., “Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM,” arXiv:1707.09870v2, Sep. 13, 2017. [cited by applicant]
Li, Fengfu, et al., “Ternary Weight Networks,” May 16, 2016, 9 pages, arXiv:1605.04711v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Merolla, Paul, et al., “Deep Neural Networks are Robust to Weight Binarization and Other Non-linear Distortions,” Jun. 7, 2016, 10 pages, arXiv:1606.01981v1, Computing Research Repository (CoRR)—Cornell University, Itha… [cited by applicant]
Moons, Bert, et al., “Envision: A 0.26-to-10TOPS/W Subword-Parallel Dynamic-Voltage-Accuracy-Frequency-Scalable Convolutional Neural Network Processor in 28nm FDSOI,” Proceedings of 2017 IEEE International Solid-State C… [cited by applicant]
Moshovos, Andreas, et al., “Exploiting Typical Values to Accelerate Deep Learning,” Computer, May 24, 2018, 13 pages, vol. 51—Issue 5, IEEE Computer Society, Washington, D.C. [cited by applicant]
Non-published commonly owned U.S. Appl. No. 16/717,926, filed Dec. 17, 2019, 99 pages, Perceive Corporation. [cited by applicant]
Park, Jongsoo, et al., “Faster CNNs with Direct Sparse Convolutions and Guided Pruning,” Jul. 28, 2017, 12 pages, arXiv:1608.01409v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Pedram, Ardavan, et al., “Dark Memory and Accelerator-Rich System Optimization in the Dark Silicon Era,” Apr. 27, 2016, 8 pages, arXiv:1602.04183v3, Computer Research Repository (CoRR)—Comell University, Ithaca, NY, USA. [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” arXiv:1603.05279v4, Aug. 2, 2016. [cited by applicant]
Ren, Mengye, et al., “SBNet: Sparse Blocks Network for Fast Inference,” Jan. 7, 2018, 10 pages, arXiv:1801.02108v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Rutenbar, Rob A., et al., “Hardware Inference Accelerators for Machine Learning,” 2016 IEEE International Test Conference (ITC), Nov. 15-17, 2016, 39 pages, IEEE, Fort Worth, TX, USA. [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” arXiv:1710.07739v3, Feb. 2, 2018. [cited by applicant]
Shin, Dongjoo, et al., “DNPU: An 8.1TOPS/W Reconfigurable CNN-RNN Processor for General-Purpose Deep Neural Networks,” Proceedings of 2017 IEEE International Solid-State Circuits Conference (ISSCC 2017), Feb. 5-7, 2017,… [cited by applicant]
Sim, Jaehyeong, et al., “A 1.42TOPS/W Deep Convolutional Neural Network Recognition Processor for Intelligent IoE Systems,” Proceedings of 2016 IEEE International Solid-State Circuits Conference (ISSCC 2016), Jan. 31-Fe… [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Wang, Min, et al., “Factorized Convolutional Neural Networks,” 2017 IEEE International Conference on Computer Vision Workshops (ICCVW '17), Oct. 22-29, 2017, 9 pages, IEEE, Venice, Italy. [cited by applicant]
Yang, Tien-Ju, et al., “Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning,” Apr. 18, 2017, 9 pages, arXiv:1611.05128v4, Computer Research Repository (CoRR)—Comell University, Ithaca, NY… [cited by applicant]
Yang, Xuan, et al., “DNN Dataflow Choice Is Overrated,” Sep. 10, 2018, 13 pages, arXiv:1809.04070v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Zhang, Dongqing, et al., “LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks,” Jul. 26, 2018, 21 pages, arXiv:1807.10029v1, Computer Research Repository (CoRR)—Cornell University, Ithaca,… [cited by applicant]
Zhang, Shijin, et al., “Cambricon-X: An Accelerator for Sparse Neural Networks,” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO '16), Oct. 15-19, 2016, 12 pages, IEEE, Taipei, Taiwan. [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv:1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Hanlon, Jamie, “Why is So Much Memory Needed for Deep Neural Networks?,” Jan. 31, 2017, 6 pages, Graphcore, Bristol, United Kingdom, retrieved from https://www.graphcore.ai/posts/why-is-so-much-memory-needed-for-deep-ne… [cited by applicant]
Han, Song, “Efficient Methods and Hardware for Deep Learning,” Sep. 2017, 125 pages, Stanford University, Palo Alto, CA, USA. [cited by applicant]
Carbon, A., et al., “PNeuro: A Scalable Energy-Efficient Programmable Hardware Accelerator for Neural Networks,” 2018 Design, Automation & Test in Europe Conference & Exhibition (Date 2018), Mar. 19-23, 2018, 6 pages, I… [cited by applicant]
Gokhale, Vinayak, et al., “Snowflake: A Model Agnostic Accelerator for Deep Convolutional Neural Networks,” Aug. 8, 2017, 11 pages, arXiv:1708.02579v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
He, Kaiming, et al., “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification,” Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Dec. 7-13, 2015, pp. 1… [cited by applicant]
Jin, Canran, et al., “Sparse Ternary Connect Convolutional Neural Networks Using Ternarized Weights with Enhanced Sparsity,” 2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC), Jan. 22-25, 2018, 6 p… [cited by applicant]
Nair, Vinod, et al., “Rectified Linear Units Improve Restricted Boltzmann Machines,” Proceedings of the 27th International Conference on Machine Learning, Jun. 21-24, 2010, 8 pages, Omnipress, Haifa, Israel. [cited by applicant]