IP Library Granted Patent US 12,675,678
Granted Patent B1
US 12,675,678 · App. 17/306,744 · Granted Jul 7, 2026

Unified memory for integrated circuit executing neural network

Inventors: Jung Ko (San Jose, CA); Kenneth Duong (San Jose, CA); Steven L. Teig (Menlo Park, CA); Won Rhee (Los Altos, CA)
Assignee: Amazon Technologies, Inc.
G06N3/063G06F9/223G06F12/0246G06F12/06G06F12/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,678
App. No.
17/306,744
Filed
May 3, 2021
Granted
Jul 7, 2026
Kind
B1
Art Unit
2137
USPC
711/206
Abstract

Some embodiments provide an integrated circuit (IC). The IC includes a microprocessor circuit for loading configuration for a neural network and generating instructions for executing the neural network based on the configuration. The IC includes a neural network inference circuit for executing a neural network for input data according to instructions received from the central processing circuit. The IC includes an input processing circuit for receiving data and preparing input data for the neural network inference circuit. The IC includes a unified memory accessible by the microprocessor circuit, neural network inference circuit, and input processing circuit.

Claims (31)

1 . An integrated circuit (IC) comprising:

a microprocessor circuit for loading configuration for a neural network and generating instructions for executing the neural network based on the configuration;

a neural network inference circuit for executing the neural network for input data according to the instructions received from the microprocessor circuit, the neural network inference circuit comprising a plurality of cores;

an input processing circuit for receiving data and preparing the input data for the neural network inference circuit;

a unified memory accessible by the microprocessor circuit, the neural network inference circuit, and the input processing circuit, wherein the unified memory comprises a plurality of memory banks, and wherein a first core of the plurality of cores of the neural network inference circuit is configured to have exclusive access from among the plurality of cores to only a respective first subset of dedicated memory banks of the plurality of memory banks, and wherein a second core of the plurality of cores of the neural network inference circuit is configured to have exclusive access from among the plurality of cores to only a second subset of dedicated memory banks of the plurality of memory banks; and

at least a first direct connection from the first core of the plurality of cores to a first dedicated memory bank of a first subset of dedicated memory banks,

wherein the input processing circuit writes first input data to the first subset of dedicated memory banks and second input data to the second subset of dedicated memory banks, and

wherein a logical address of the first subset of dedicated memory banks is swapped with a logical address of the second subset of dedicated memory banks, after executing, by the neural network inference circuit, the first input data.

2 . The IC of claim 1 , wherein each memory bank of the plurality of memory banks is accessible by the input processing circuit and the microprocessor circuit.

3 . The IC of claim 1 , wherein each respective memory bank of the plurality of memory banks is associated with a respective core of the plurality of cores of the neural network inference circuit, and wherein each respective memory bank has a same memory configuration.

4 . The IC of claim 3 , wherein two memory banks have the same memory configuration when a number of physical memories for each of the two memory banks is the same and physical memories of each of the two memory banks have a same size.

5 . The IC of claim 1 , further comprising at least one direct connection from the first core of the plurality of cores to a first memory bank of a respective subset of dedicated memory banks a second direct connection from the second core of the plurality of cores to a second dedicated memory bank of the second subset of dedicated memory banks.

6 . The IC of claim 5 , further comprising a crossbar for providing access to each memory bank of the plurality of memory banks of the unified memory to the microprocessor circuit and the input processing circuit.

7 . The IC of claim 1 , wherein the neural network inference circuit writes weight values to the unified memory based on instructions provided from the microprocessor circuit at bootup of the IC.

8 . The IC of claim 7 , wherein the neural network inference circuit:

writes a set of input data to the unified memory based on instructions provided by the input processing circuit; and

executes the neural network for the input data to generate output data.

9 . The IC of claim 8 , wherein while executing the neural network for the set of input data the neural network inference circuit reads the set of input data along with a set of weight values from the unified memory, writes sets of intermediate activation values to the unified memory, and reads sets of intermediate activation values along with sets of weight values from the unified memory.

10 . The IC of claim 8 , wherein the microprocessor circuit stores in the unified memory a program comprising the instructions for the neural network inference circuit to execute the neural network.

11 . The IC of claim 10 , wherein the program is loaded from off-chip non-volatile storage at bootup of the IC.

12 . The IC of claim 1 , wherein the neural network inference circuit uses a first address space for addressing the unified memory while the microprocessor circuit and the input processing circuit use the first address space and a second address space for addressing the unified memory.

13 . The IC of claim 12 , wherein addresses in the first address space decode to specific physical memory locations.

14 . The IC of claim 13 , wherein:

a first address in the first address space is formatted to include (i) a first address prefix indicating that the first address is formatted in the first address space and (ii) a first physical address portion that refers to a first physical memory location and the first core of the plurality of cores, wherein the first core has exclusive access from among the plurality of cores to the first physical memory location; and

a second address in the second address space is formatted to include (i) a second address prefix indicating that the second address is formatted in the second address space and (ii) a first virtual address portion that refers to a first virtual memory address, wherein the first virtual memory address is used to identify a second physical memory location.

15 . The IC of claim 14 , wherein a mapping table is used to map the first virtual memory address to the second physical memory location.

16 . The IC of claim 15 , wherein the second address space allows a set of unused physical memory locations associated with two or more different cores of the plurality of cores to be treated as a contiguous block of virtual addresses used to identify the set of unused physical memory locations for access by the microprocessor circuit and the input processing circuit, wherein the set of unused physical memory locations are associated with static random access memory (SRAM).

17 . The IC of claim 12 , wherein the microprocessor circuit and the input processing circuit use the first address space for addressing the unified memory when reading and writing data that is also accessed by the neural network inference circuit.

18 . The IC of claim 17 , wherein the microprocessor circuit and the input processing circuit use the second address space for addressing the unified memory when reading and writing data that is not accessed by the neural network inference circuit.

19 . The IC of claim 1 , wherein the neural network is a first neural network, wherein the plurality of cores of the neural network inference circuit are separated into at least a first cluster of cores and a second cluster of cores.

20 . The IC of claim 19 , wherein the first cluster of cores is configured to execute the first neural network, and wherein the second cluster of cores is configured to execute a second neural network at a same time as the first neural network is executed.

Assignments (2)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
Continuity (1)
Provisional Application 63178933 · Apr 23, 2021
References Cited (198)
US 5621863A · Boulet et al. · 1997 [cited by applicant]
US 5717832A · Steimle et al. · 1998 [cited by applicant]
US 5740326A · Boulet et al. · 1998 [cited by applicant]
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 9201608B2 · Hendry · 2015 [cited by examiner]
US 9710265B1 · Temam et al. · 2017 [cited by applicant]
US 9858636B1 · Lim et al. · 2018 [cited by applicant]
US 9904874B2 · Shoaib et al. · 2018 [cited by applicant]
US 10445638B1 · Amirineni et al. · 2019 [cited by applicant]
US 10489478B2 · Lim et al. · 2019 [cited by applicant]
US 10515303B2 · Lie et al. · 2019 [cited by applicant]
US 10657438B2 · Lie et al. · 2020 [cited by applicant]
US 10664310B2 · Bokhari · 2020 [cited by applicant]
US 10740434B1 · Duong et al. · 2020 [cited by applicant]
US 10768856B1 · Diamant et al. · 2020 [cited by applicant]
US 10796198B2 · Franca-Neto · 2020 [cited by applicant]
US 10817042B2 · Desai et al. · 2020 [cited by applicant]
US 10853738B1 · Dockendorf et al. · 2020 [cited by applicant]
US 10970630B1 · Aimone · 2021 [cited by applicant]
US 11049013B1 · Duong et al. · 2021 [cited by applicant]
US 11138292B1 · Nair et al. · 2021 [cited by applicant]
US 11170289B1 · Duong et al. · 2021 [cited by applicant]
US 11250326B1 · Ko et al. · 2022 [cited by applicant]
US 11347297B1 · Ko et al. · 2022 [cited by applicant]
US 11423289B2 · Judd et al. · 2022 [cited by applicant]
US 11531868B1 · Duong et al. · 2022 [cited by applicant]
US 11537853B1 · Afzal · 2022 [cited by applicant]
US 11568227B1 · Ko et al. · 2023 [cited by applicant]
US 11586910B1 · Duong et al. · 2023 [cited by applicant]
US 11868867B1 · Afzal · 2024 [cited by applicant]
US 11868901B1 · Thomas et al. · 2024 [cited by applicant]
US 11977916B2 · Kim · 2024 [cited by applicant]
US 20040078403A1 · Scheuermann et al. · 2004 [cited by applicant]
US 20050021874A1 · Georgiou · 2005 [cited by examiner]
US 20110106741A1 · Denneau · 2011 [cited by applicant]
US 20110307685A1 · Song · 2011 [cited by applicant]
US 20150286525A1 · Singh · 2015 [cited by examiner]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160320994A1 · Chun · 2016 [cited by examiner]
US 20160342893A1 · Ross et al. · 2016 [cited by applicant]
US 20170011006A1 · Saber et al. · 2017 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170243110A1 · Esquivel et al. · 2017 [cited by applicant]
US 20170300828A1 · Feng et al. · 2017 [cited by applicant]
US 20170323196A1 · Gibson et al. · 2017 [cited by applicant]
US 20170344882A1 · Ambrose · 2017 [cited by applicant]
US 20180018559A1 · Yakopcic et al. · 2018 [cited by applicant]
US 20180025268A1 · Teig et al. · 2018 [cited by applicant]
US 20180046458A1 · Kuramoto · 2018 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180046905A1 · Li et al. · 2018 [cited by applicant]
US 20180046916A1 · Dally et al. · 2018 [cited by applicant]
US 20180101763A1 · Barnard et al. · 2018 [cited by applicant]
US 20180114569A1 · Strachan et al. · 2018 [cited by applicant]
US 20180121196A1 · Temam et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180164866A1 · Turakhia et al. · 2018 [cited by applicant]
US 20180181406A1 · Kuramoto · 2018 [cited by applicant]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180197068A1 · Narayanaswami et al. · 2018 [cited by applicant]
US 20180246855A1 · Redfern et al. · 2018 [cited by applicant]
US 20180285719A1 · Baum et al. · 2018 [cited by applicant]
US 20180285726A1 · Baum et al. · 2018 [cited by applicant]
US 20180285727A1 · Baum et al. · 2018 [cited by applicant]
US 20180285736A1 · Baum et al. · 2018 [cited by applicant]
US 20180293490A1 · Ma et al. · 2018 [cited by applicant]
US 20180293493A1 · Kalamkar et al. · 2018 [cited by applicant]
US 20180293691A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180300600A1 · Ma et al. · 2018 [cited by applicant]
US 20180307494A1 · Ould-Ahmed-Vall et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180307985A1 · Appu et al. · 2018 [cited by applicant]
US 20180308202A1 · Appu et al. · 2018 [cited by applicant]
US 20180314492A1 · Fais et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180322386A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180322387A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180329868A1 · Chen et al. · 2018 [cited by applicant]
US 20180365794A1 · Lee et al. · 2018 [cited by applicant]
US 20180373975A1 · Yu et al. · 2018 [cited by applicant]
US 20190012296A1 · Hsieh et al. · 2019 [cited by applicant]
US 20190026078A1 · Bannon et al. · 2019 [cited by applicant]
US 20190026237A1 · Talpes et al. · 2019 [cited by applicant]
US 20190026249A1 · Talpes et al. · 2019 [cited by applicant]
US 20190041961A1 · Desai et al. · 2019 [cited by applicant]
US 20190057036A1 · Mathuriya et al. · 2019 [cited by applicant]
US 20190073585A1 · Pu et al. · 2019 [cited by applicant]
US 20190087713A1 · Lamb et al. · 2019 [cited by applicant]
US 20190095776A1 · Kfir et al. · 2019 [cited by applicant]
US 20190114499A1 · Delaye et al. · 2019 [cited by applicant]
US 20190114534A1 · Teng · 2019 [cited by examiner]
US 20190138891A1 · Kim et al. · 2019 [cited by applicant]
US 20190147338A1 · Pau et al. · 2019 [cited by applicant]
US 20190156180A1 · Nomura et al. · 2019 [cited by applicant]
US 20190171927A1 · Diril et al. · 2019 [cited by applicant]
US 20190179635A1 · Jiao et al. · 2019 [cited by applicant]
US 20190180167A1 · Huang et al. · 2019 [cited by applicant]
US 20190187983A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20190196970A1 · Han et al. · 2019 [cited by applicant]
US 20190205094A1 · Diril et al. · 2019 [cited by applicant]
US 20190205358A1 · Diril et al. · 2019 [cited by applicant]
US 20190205736A1 · Bleiweiss et al. · 2019 [cited by applicant]
US 20190205739A1 · Liu et al. · 2019 [cited by applicant]
US 20190205740A1 · Judd et al. · 2019 [cited by applicant]
US 20190205780A1 · Sakaguchi · 2019 [cited by applicant]
US 20190236437A1 · Shin et al. · 2019 [cited by applicant]
US 20190236445A1 · Das et al. · 2019 [cited by applicant]
US 20190266217A1 · Arakawa et al. · 2019 [cited by applicant]
US 20190266479A1 · Singh et al. · 2019 [cited by applicant]
US 20190294413A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294959A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294968A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190303741A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303749A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303750A1 · Kumar et al. · 2019 [cited by applicant]
US 20190325296A1 · Fowers et al. · 2019 [cited by applicant]
US 20190332925A1 · Modha · 2019 [cited by applicant]
US 20190340493A1 · Coenen · 2019 [cited by examiner]
US 20190347559A1 · Kang et al. · 2019 [cited by applicant]
US 20190385046A1 · Cassidy et al. · 2019 [cited by applicant]
US 20200005131A1 · Nakahara et al. · 2020 [cited by applicant]
US 20200042856A1 · Datta et al. · 2020 [cited by applicant]
US 20200042859A1 · Mappouras et al. · 2020 [cited by applicant]
US 20200089506A1 · Power et al. · 2020 [cited by applicant]
US 20200117597A1 · Huang · 2020 [cited by examiner]
US 20200125926A1 · Choudhury · 2020 [cited by applicant]
US 20200134461A1 · Chai et al. · 2020 [cited by applicant]
US 20200234114A1 · Rakshit et al. · 2020 [cited by applicant]
US 20200257930A1 · Nahr et al. · 2020 [cited by applicant]
US 20200272907A1 · Jin · 2020 [cited by applicant]
US 20200301668A1 · Li · 2020 [cited by applicant]
US 20200301739A1 · Xu · 2020 [cited by applicant]
US 20200364545A1 · Shattil · 2020 [cited by applicant]
US 20200380344A1 · Lie et al. · 2020 [cited by applicant]
US 20210110236A1 · Shibata · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210182684A1 · Zappi · 2021 [cited by applicant]
US 20210232897A1 · Bichler · 2021 [cited by applicant]
US 20210241082A1 · Nagy et al. · 2021 [cited by applicant]
US 20210287074A1 · Coenen · 2021 [cited by applicant]
US 20220004854A1 · Lee · 2022 [cited by applicant]
US 20220121914A1 · Huang et al. · 2022 [cited by applicant]
US 20220335562A1 · Surti et al. · 2022 [cited by applicant]
US 20220414437A1 · Liu · 2022 [cited by applicant]
US 20230418666A1 · Puppala · 2023 [cited by examiner]
CN 108876698A · 2018 [cited by applicant]
CN 108280514B · 2020 [cited by applicant]
GB 2568086A · 2019 [cited by applicant]
WO 2020044527A1 · 2020 [cited by applicant]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Andri, Renzo, et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Mar. 14, 2017, 14 pages, IEEE, New York,… [cited by applicant]
Ardakani, Arash, et al., “Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks,” Proceedings of the 5th International Conference on Learning Representations (ICLR 2017), Apr.… [cited by applicant]
Bagherinezhad, Hessam, et al., “LCNN: Look-up Based Convolutional Neural Network,” Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 10 pages, IEEE, Honolulu, … [cited by applicant]
Bang, Suyoung, et al., “A 288 μW Programmable Deep-Learning Processor with 270KB On-Chip Weight Storage Using Non-Uniform Memory Hierarchy for Mobile Intelligence,” Proceedings of 2017 IEEE International Solid-State Cir… [cited by applicant]
Bong, Kyeongryeol, et al., “A 0.62mW Ultra-Low-Power Convolutional-Neural-Network Face-Recognition Processor and a CIS Integrated with Always-On Haar-Like Face Detector,” Proceedings of 2017 IEEE International Solid-Sta… [cited by applicant]
Boo, Yoonho, et al., “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations,” 2017 IEEE Workshop on Signal Processing Systems (SiPS), Oct. 3-5, 2017, 6 pages, IEEE, Lorie… [cited by applicant]
Chen, Yu-Hsin, et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA 2… [cited by applicant]
Chen, Yu-Hsin, et al., “Using Dataflow to Optimize Energy Efficiency of Deep Neural Network Accelerators,” IEEE Micro, Jun. 14, 2017, 10 pages, vol. 37, Issue 3, IEEE, New York, NY, USA. [cited by applicant]
Courbariaux, Matthieu, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or -1,” Mar. 17, 2016, 11 pages, arXiv:1602.02830v3, Computing Research Repository (CoRR… [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations,” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 15),… [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Fu, Yao, et al., “Embedded Vision with INT8 Optimization on Xilinx Devices,” WP490 (v1.0.1), Apr. 19, 2017, 15 pages, Xilinx, Inc., San Jose, CA, USA. [cited by applicant]
Guo, Yiwen, et al., “Network Sketching: Exploring Binary Structure in Deep CNNs,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 9 pages, IEEE, Honolulu, HI. [cited by applicant]
He, Zhezhi, et al., “Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy,” Jul. 20, 2018, 8 pages, arXiv: 1807.07948v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, N… [cited by applicant]
Hegde, Kartik, et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition,” Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA '18), Jun. 2-6, 2018, 14… [cited by applicant]
Huan, Yuxiang, et al., “A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity,” May 22, 2017, 5 pages, arXiv:1705.08009v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, … [cited by applicant]
Jouppi, Norman, P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17), Jun. 24-28, 2017, 17 pages, ACM, … [cited by applicant]
Judd, Patrick, et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing,” Apr. 29, 2017, 6 pages, arXiv:1705.00125v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, U… [cited by applicant]
Leng, Cong, et al., “Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM,” Proceedings of 32nd AAAI Conference on Artificial Intelligence (AAAI-18), Feb. 2-7, 2018, 16 pages, Association for the Advance… [cited by applicant]
Li, Fengfu, et al., “Ternary Weight Networks,” May 16, 2016, 9 pages, arXiv:1605.04711v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Merolla, Paul, et al., “Deep Neural Networks are Robust to Weight Binarization and Other Non-linear Distortions,” Jun. 7, 2016, 10 pages, arXiv:1606.01981v1, Computing Research Repository (CoRR)—Cornell University, Itha… [cited by applicant]
Moons, Bert, et al., “Envision: A 0.26-to-10TOPS/W Subword-Parallel Dynamic-Voltage-Accuracy-Frequency-Scalable Convolutional Neural Network Processor in 28nm FDSOI,” Proceedings of 2017 IEEE International Solid-State C… [cited by applicant]
Moshovos, Andreas, et al., “Exploiting Typical Values to Accelerate Deep Learning,” Computer, May 24, 2018, 13 pages, vol. 51-Issue 5, IEEE Computer Society, Washington, D.C. [cited by applicant]
Non-Published Commonly Owned Related U.S. Appl. No. 17/306,742 with similar specification, filed May 3, 2021, 78 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned Related U.S. Appl. No. 17/306,745 with similar specification, filed May 3, 2021, 78 pages, Perceive Corporation. [cited by applicant]
Park, Jongsoo, et al., “Faster CNNs with Direct Sparse Convolutions and Guided Pruning,” Jul. 28, 2017, 12 pages, arXiv: 1608.01409v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Proceedings of 2016 European Conference on Computer Vision (ECCV '16), Oct. 8-16, 2016, 17 pages, Lecture Note… [cited by applicant]
Ren, Mengye, et al., “SBNet: Sparse Blocks Network for Fast Inference,” Jan. 7, 2018, 10 pages, arXiv:1801.02108v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 12 pages, ICLR, Vanco… [cited by applicant]
Shin, Dongjoo, et al., “DNPU: An 8.1TOPS/W Reconfigurable CNN-RNN Processor for General-Purpose Deep Neural Networks,” Proceedings of 2017 IEEE International Solid-State Circuits Conference (ISSCC 2017), Feb. 5-7, 2017,… [cited by applicant]
Sim, Jaehyeong, et al., “A 1.42TOPS/W Deep Convolutional Neural Network Recognition Processor for Intelligent OE Systems,” Proceedings of 2016 IEEE International Solid-State Circuits Conference (ISSCC 2016), Jan. 31-Feb… [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Wang, Min, et al., “Factorized Convolutional Neural Networks,” 2017 IEEE International Conference on Computer Vision Workshops (ICCVW '17), Oct. 22-29, 2017, 9 pages, IEEE, Venice, Italy. [cited by applicant]
Wen, Wei, et al., “Learning Structured Sparsity in Deep Neural Networks,” Oct. 18, 2016, 10 pages, arXiv: 1608.03665v4, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Yang, Xuan, et al., “DNN Dataflow Choice Is Overrated,” Sep. 10, 2018, 13 pages, arXiv:1809.04070v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Zhang, Shijin, et al., “Cambricon-X: An Accelerator for Sparse Neural Networks,” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO '16), Oct. 15-19, 2016, 12 pages, IEEE, Taipei, Taiwan. [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv: 1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Abtahi, Tahmid, et al., “Accelerating Convolutional Neural Network With FFT on Embedded Hardware,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Sep. 2018, 14 pages, vol. 26, No. 9, IEEE. [cited by applicant]
Ardakani, Arash, et al., “An Architecture to Accelerate Convolution in Deep Neural Networks,” IEEE Transactions on Circuits and Systems I: Regular Papers, Oct. 17, 2017, 14 pages, vol. 65, No. 4, IEEE. [cited by applicant]
Iu, Shaoli, et al., “Cambricon: An Instruction Set Architecture for Neural Networks,” 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Jun. 18-22, 2016, 13 pages, IEEE, Seoul, South Korea. [cited by applicant]
Carbon, A., et al., “Pleura: A Scalable Energy-Efficient Programmable Hardware Accelerator for Neural Networks,” 2018 Design, Automation & Test in Europe Conference & Exhibition (Date 2018), Mar. 19-23, 2018, 6 pages, I… [cited by applicant]
Gokhale, Vinayak, et al., “Snowflake: A Model Agnostic Accelerator for Deep Convolutional Neural Networks,” Aug. 8, 2017, 11 pages, arXiv:1708.02579v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Jin, Canran,, et al., “Sparse Ternary Connect: Convolutional Neural Networks Using Ternarized Weights with Enhanced Sparsity,” 2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC), Jan. 22-25, 2018, 6… [cited by applicant]
Chen, Tianqi, et al., “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning,” Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI '18), Oct. 8-10, 2018, 17 pages, … [cited by applicant]
Han, Song, “Efficient Methods and Hardware for Deep Learning,” September 217, 125 pages, Stanford University, Palo Alto, CA, USA. [cited by applicant]