IP Library Granted Patent US 12,260,317
Granted Patent B1
US 12,260,317 · App. 16/525,460 · Granted Mar 25, 2025

Compiler for implementing gating functions for neural network configuration

Inventors: Brian Thomas (Vancouver, CA); Steven L. Teig (Menlo Park, CA)
Assignee: Amazon Technologies, Inc.
G06N3/063G06F17/16G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,317
App. No.
16/525,460
Granted
Mar 25, 2025
Kind
B1
Abstract

Some embodiments provide a compiler for optimizing the implementation of a machine-trained network (e.g., a neural network) on an integrated circuit (IC). The compiler of some embodiments receives a specification of a machine-trained network including multiple layers of computation nodes and generates a graph representing options for implementing the machine-trained network in the IC. In some embodiments, the compiler also generates instructions for gating operations. Gating operations, in some embodiments, include gating at multiple levels (e.g., gating of clusters, cores, or memory units). Gating operations conserve power in some embodiments by gating signals so that they do not reach the gated element or so that they are not propagated within the gated element. In some embodiments, a clock signal is gated such that a register that transmits data on a rising (or falling) edge of a clock signal is not triggered.

Claims (36)

1. A method for generating neural network program instructions for a neural network inference circuit (NNIC) to execute a neural network, the method comprising:

receiving a specification of the neural network comprising a plurality of layers of computation nodes; and

based on the received neural network specification, identifying elements of the NNIC that are not needed to execute each layer of a set of the layers, the NNIC elements comprising a plurality of clusters, each cluster comprising a set of dot product cores for (i) storing weight values and intermediate activation values during execution of the neural network and (ii) performing dot product computations between the weight values and the intermediate activation values; and

based on the received specification and the identified elements, generating program instructions for the NNIC to use to execute the neural network, the program instructions specifying gated elements for each layer of the set of layers in order to reduce power usage by the NNIC when executing the neural network, the gated elements for a first layer comprising at least one cluster of dot product cores that are gated during the execution of the first layer and not gated during the execution of a second layer.

2. The method of claim 1 , wherein an element gated during a particular layer does not propagate data during execution of the particular layer.

3. The method of claim 1 , wherein each dot product core comprises a set of memory units that store weight values and intermediate activation values.

4. The method of claim 1 , wherein the program instructions specify that a particular dot product core is gated for a third layer but that the cluster to which the particular dot product core belongs is not gated for the third layer.

5. The method of claim 3 , wherein the program instructions specify that a particular memory unit is gated for a third layer but that the dot product core to which the particular memory unit belongs is not gated for the third layer.

6. The method of claim 1 , wherein:

the received specification is used to determine an optimized implementation for the neural network using the NNIC;

the optimized implementation identifies a set of elements used to execute each layer of the neural network; and

identifying elements of the neural network inference circuit that are not needed to execute the layer comprises using the optimized implementation to identify the elements that are not needed to execute the layer.

7. The method of claim 3 , wherein:

when the program instructions specify to gate a particular cluster, the dot product cores belonging to the particular cluster and the memory units belonging to the dot product cores of the particular cluster are not separately gated; and

when the program instructions specify to gate a particular dot product core, the memory units of the particular dot product core are not separately gated.

8. The method of claim 1 , wherein gating an element comprises gating a clock signal.

9. The method of claim 8 , wherein gating the clock signal prevents the propagation of data by a set of registers triggered by the clock signal.

10. A non-transitory machine readable medium storing a program for execution by a set of processing units, the program for generating neural network program instructions for a neural network inference circuit (NNIC) to execute a neural network, the program comprising sets of instructions for:

receiving a specification of the neural network comprising a plurality of layers of computation nodes; and

based on the received neural network specification, identifying elements of the NNIC that are not needed to execute each layer of a set of the layers, the NNIC elements comprising a plurality of clusters, each cluster comprising a set of dot product cores for (i) storing weight values and intermediate activation values during execution of the neural network and (ii) performing dot product computations between the weight values and the intermediate activation values; and

based on the received specification and the identified elements, generating program instructions for the NNIC to use to execute the neural network, the program instructions specifying gated elements for each layer of the set of layers in order to reduce power usage by the NNIC when executing the neural network, the gated elements for a first layer comprising at least one cluster of dot product cores that are gated during the execution of the first layer and not gated during the execution of a second layer.

11. The non-transitory machine readable medium of claim 10 , wherein an element gated during a particular layer does not propagate data during execution of the particular layer.

12. The non-transitory machine readable medium of claim 10 , wherein each dot product core comprises a set of memory units that store weight values and intermediate activation values.

13. The non-transitory machine readable medium of claim 10 , wherein the program instructions specify that a particular dot product core is gated for a third layer but that the cluster to which the particular dot product core belongs is not gated for the third layer.

14. The non-transitory machine readable medium of claim 12 , wherein the program instructions specify that a particular memory unit is gated for a third layer but that the dot product core to which the particular memory unit belongs is not gated for the third layer.

15. The non-transitory machine readable medium of claim 10 , wherein:

the received specification is used to determine an optimized implementation for the neural network using the NNIC;

the optimized implementation identifies a set of elements used to execute each layer of the neural network; and

the set of instructions for identifying elements of the neural network inference circuit that are not needed to execute the layer comprises a set of instructions for using the optimized implementation to identify the elements that are not needed to execute the layer.

16. The non-transitory machine readable medium of claim 12 , wherein:

when the program instructions specify to gate a particular cluster, the dot product cores belonging to the particular cluster and the memory units belonging to the dot product cores of the particular cluster are not separately gated; and

when the program instructions specify to gate a particular dot product core, the memory units of the particular dot product core are not separately gated.

17. The non-transitory machine readable medium of claim 10 , wherein gating an element comprises gating a clock signal.

18. The non-transitory machine readable medium of claim 17 , wherein gating the clock signal prevents the propagation of data by a set of registers triggered by the clock signal.

19. The method of claim 1 , wherein the dot product cores of the cluster gated during the execution of the first layer do not perform any dot product operations during the first layer.

20. The method of claim 1 , wherein dot products computed during the first layer are small enough that not all of the clusters of the NNIC are required to be used during the first layer.

Assignments (3)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2019
From: THOMAS, BRIAN; TEIG, STEVEN L.
To: PERCEIVE CORPORATION
Reel/Frame 050057/0026 →
Continuity (2)
Provisional Application 62866599 · Jun 25, 2019
Provisional Application 62851082 · May 21, 2019
References Cited (174)
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 8468109B2 · Moussa et al. · 2013 [cited by applicant]
US 9710265B1 · Temam et al. · 2017 [cited by applicant]
US 9858636B1 · Lim et al. · 2018 [cited by applicant]
US 10445638B1 · Amirineni et al. · 2019 [cited by applicant]
US 10489478B2 · Lim et al. · 2019 [cited by applicant]
US 10515303B2 · Lie et al. · 2019 [cited by applicant]
US 10657438B2 · Lie et al. · 2020 [cited by applicant]
US 10732943B2 · Sun et al. · 2020 [cited by applicant]
US 10740434B1 · Duong et al. · 2020 [cited by applicant]
US 10768856B1 · Diamant et al. · 2020 [cited by applicant]
US 10796198B2 · Franca-Neto · 2020 [cited by applicant]
US 10817042B2 · Desai et al. · 2020 [cited by applicant]
US 11023360B2 · Gu et al. · 2021 [cited by applicant]
US 11132619B1 · Casas et al. · 2021 [cited by applicant]
US 11138292B1 · Nair et al. · 2021 [cited by applicant]
US 11423289B2 · Judd et al. · 2022 [cited by applicant]
US 11468145B1 · Duong et al. · 2022 [cited by applicant]
US 20040078403A1 · Scheuermann et al. · 2004 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342893A1 · Ross et al. · 2016 [cited by applicant]
US 20170011006A1 · Saber et al. · 2017 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170323196A1 · Gibson et al. · 2017 [cited by applicant]
US 20180032856A1 · Alvarez-Icaza et al. · 2018 [cited by applicant]
US 20180046458A1 · Kuramoto · 2018 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180046905A1 · Li et al. · 2018 [cited by applicant]
US 20180046916A1 · Dally et al. · 2018 [cited by applicant]
US 20180101763A1 · Barnard et al. · 2018 [cited by applicant]
US 20180107918A1 · Amir et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher · 2018 [cited by examiner]
US 20180164866A1 · Turakhia et al. · 2018 [cited by applicant]
US 20180181406A1 · Kuramoto · 2018 [cited by applicant]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi · 2018 [cited by examiner]
US 20180197068A1 · Narayanaswami et al. · 2018 [cited by applicant]
US 20180218518A1 · Yan et al. · 2018 [cited by applicant]
US 20180246855A1 · Redfern et al. · 2018 [cited by applicant]
US 20180285719A1 · Baum et al. · 2018 [cited by applicant]
US 20180285727A1 · Baum et al. · 2018 [cited by applicant]
US 20180285731A1 · Heifets et al. · 2018 [cited by applicant]
US 20180285736A1 · Baum et al. · 2018 [cited by applicant]
US 20180293057A1 · Sun et al. · 2018 [cited by applicant]
US 20180293490A1 · Ma et al. · 2018 [cited by applicant]
US 20180293493A1 · Kalamkar et al. · 2018 [cited by applicant]
US 20180293691A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180300600A1 · Ma et al. · 2018 [cited by applicant]
US 20180307494A1 · Ould-Ahmed-Vall et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180307985A1 · Appu et al. · 2018 [cited by applicant]
US 20180308202A1 · Appu et al. · 2018 [cited by applicant]
US 20180314492A1 · Fais · 2018 [cited by examiner]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180322386A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180322387A1 · Sridharan et al. · 2018 [cited by applicant]
US 20180365794A1 · Lee et al. · 2018 [cited by applicant]
US 20180373975A1 · Yu et al. · 2018 [cited by applicant]
US 20190026078A1 · Bannon et al. · 2019 [cited by applicant]
US 20190026237A1 · Talpes et al. · 2019 [cited by applicant]
US 20190026249A1 · Talpes et al. · 2019 [cited by applicant]
US 20190026625A1 · Vorenkamp et al. · 2019 [cited by applicant]
US 20190041961A1 · Desai · 2019 [cited by examiner]
US 20190042948A1 · Lee et al. · 2019 [cited by applicant]
US 20190057036A1 · Mathuriya et al. · 2019 [cited by applicant]
US 20190073585A1 · Pu et al. · 2019 [cited by applicant]
US 20190095776A1 · Kfir et al. · 2019 [cited by applicant]
US 20190101952A1 · Diamond · 2019 [cited by examiner]
US 20190114499A1 · Delaye et al. · 2019 [cited by applicant]
US 20190138891A1 · Kim et al. · 2019 [cited by applicant]
US 20190147338A1 · Pau et al. · 2019 [cited by applicant]
US 20190156180A1 · Nomura et al. · 2019 [cited by applicant]
US 20190171927A1 · Diril et al. · 2019 [cited by applicant]
US 20190180167A1 · Huang et al. · 2019 [cited by applicant]
US 20190187983A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20190196970A1 · Han et al. · 2019 [cited by applicant]
US 20190205358A1 · Diril et al. · 2019 [cited by applicant]
US 20190205736A1 · Bleiweiss et al. · 2019 [cited by applicant]
US 20190205739A1 · Liu et al. · 2019 [cited by applicant]
US 20190205740A1 · Judd et al. · 2019 [cited by applicant]
US 20190205780A1 · Sakaguchi · 2019 [cited by applicant]
US 20190228274A1 · Georgiadis et al. · 2019 [cited by applicant]
US 20190236437A1 · Shin et al. · 2019 [cited by applicant]
US 20190236445A1 · Das et al. · 2019 [cited by applicant]
US 20190286989A1 · Wang et al. · 2019 [cited by applicant]
US 20190294413A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294959A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190294968A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190303741A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303749A1 · Appuswamy et al. · 2019 [cited by applicant]
US 20190303750A1 · Kumar et al. · 2019 [cited by applicant]
US 20190325296A1 · Fowers et al. · 2019 [cited by applicant]
US 20190325309A1 · Flamant · 2019 [cited by applicant]
US 20190325314A1 · Bourges-Sevenier et al. · 2019 [cited by applicant]
US 20190332925A1 · Modha · 2019 [cited by applicant]
US 20190340500A1 · Olmschenk · 2019 [cited by applicant]
US 20190347559A1 · Kang et al. · 2019 [cited by applicant]
US 20190385046A1 · Cassidy et al. · 2019 [cited by applicant]
US 20200005131A1 · Nakahara · 2020 [cited by examiner]
US 20200042856A1 · Datta et al. · 2020 [cited by applicant]
US 20200042859A1 · Mappouras et al. · 2020 [cited by applicant]
US 20200089506A1 · Power et al. · 2020 [cited by applicant]
US 20200143226A1 · Georgiadis · 2020 [cited by applicant]
US 20200151088A1 · Gu et al. · 2020 [cited by applicant]
US 20200151572A1 · Gurumurthi · 2020 [cited by applicant]
US 20200210838A1 · Lo · 2020 [cited by examiner]
US 20200234114A1 · Rakshit et al. · 2020 [cited by applicant]
US 20200301668A1 · Li · 2020 [cited by applicant]
US 20200380344A1 · Lie et al. · 2020 [cited by applicant]
US 20210110236A1 · Shibata · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210241082A1 · Nagy et al. · 2021 [cited by applicant]
US 20220335562A1 · Surti et al. · 2022 [cited by applicant]
CN 108876698A · 2018 [cited by applicant]
CN 108280514B · 2020 [cited by applicant]
GB 2568086A · 2019 [cited by applicant]
WO 2020044527A1 · 2020 [cited by applicant]
Abtahi, T. et al., “Accelerating convolutional neural network with FFT on tiny cores” (Year: 2017). [cited by examiner]
Sze, V. et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey” (Year: 2017). [cited by examiner]
Chen, Tianqi, et al., “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning,” Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI '18), Oct. 8-10, 2018, 17 pages, … [cited by applicant]
Dua, Vivek, “A Mixed-integer Programming Approach for Optimal Configuration of Artificial Neural Networks,” Chemical Engineering Research and Design, Feb. 27, 2009, 6 pages, vol. 88, Elsevier B.V. [cited by applicant]
Sen, Sanchari, et al., “SPARCE: Sparsity aware General Purpose Core Extensions to Accelerate Deep Neural Networks,” Nov. 29, 2017, 13 pages, arXiv:1711.06315v2, Computer Research Repository (CoRR)—Cornell University, It… [cited by applicant]
Ardakani, Arash, et al., “An Architecture to Accelerate Convolution in Deep Neural Networks,” IEEE Transactions on Circuits and Systems I: Regular Papers, Oct. 17, 2017, 14 pages, vol. 65, No. 4, IEEE. [cited by applicant]
Wang, Peiqi, et al., “SNrram: An Efficient Sparse Neural Network Computation Architecture Based on Resistive Random-Access Memory,” DAC '18, Jun. 24-29, 2018, 7 pages, ACM, San Francisco, CA, USA. [cited by applicant]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Chen, Tianqi, et al., “TVM: End-to-End Optimization Stack for Deep Learning,” Feb. 12, 2018, 19 pages, arXiv:1802.04799v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Andri, Renzo, et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Mar. 14, 2017, 14 pages, IEEE, New York,… [cited by applicant]
Ardakani, Arash, et al., “Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks,” Proceedings of the 5th International Conference on Learning Representations (ICLR 2017), Apr.… [cited by applicant]
Bagherinezhad, Hessam, et al., “LCNN: Look-up Based Convolutional Neural Network,” Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 10 pages, IEEE, Honolulu, … [cited by applicant]
Bang, Suyoung, et al., “A 288 μW Programmable Deep-Learning Processor with 270KB On-Chip Weight Storage Using Non-Uniform Memory Hierarchy for Mobile Intelligence,” Proceedings of 2017 IEEE International Solid-State Cir… [cited by applicant]
Bong, Kyeongryeol, et al., “A 0.62mW Ultra-Low-Power Convolutional-Neural-Network Face-Recognition Processor and a CIS Integrated with Always-On Haar-Like Face Detector,” Proceedings of 2017 IEEE International Solid-Sta… [cited by applicant]
Chen, Yu-Hsin, et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Proceedings of 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA 2… [cited by applicant]
Chen, Yu-Hsin, et al., “Using Dataflow to Optimize Energy Efficiency of Deep Neural Network Accelerators,” IEEE Micro, Jun. 14, 2017, 10 pages, vol. 37, Issue 3, IEEE, New York, NY, USA. [cited by applicant]
Cho, Minsik, et al., “MEC: Memory-Efficient Convolution for Deep Neural Network,” Jun. 21, 2017, 10 pages, arXiv:1706.06873v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Courbariaux, Matthieu, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” Mar. 17, 2016, 11 pages, arXiv:1602.02830v3, Computing Research Repository (CoRR… [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations,” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 15),… [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Fu, Yao, et al., “Embedded Vision with INT8 Optimization on Xilinx Devices,” WP490 (v1.0.1), Apr. 19, 2017, 15 pages, Xilinx, Inc., San Jose, CA, USA. [cited by applicant]
Gao, Mingyu, et al., “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,” Proceedings of the 22nd International Conference on Architectural Support for Programming Languages and Operating Systems… [cited by applicant]
Guo, Yiwen, et al., “Network Sketching: Exploring Binary Structure in Deep CNNs,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 9 pages, IEEE, Honolulu, HI. [cited by applicant]
Hanlon, Jamie, “Why is So Much Memory Needed for Deep Neural Networks?,” Jan. 31, 2017, 6 pages, Graphcore, Bristol, United Kingdom, retrieved from https://www.graphcore.ai/posts/why-is-so-much-memory-needed-for-deep-ne… [cited by applicant]
He, Zhezhi, et al., “Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy,” Jul. 20, 2018, 8 pages, arXiv:1807.07948v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Hegde, Kartik, et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition,” Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA '18), Jun. 2-6, 2018, 14… [cited by applicant]
Horowitz, Mark, “Computing's Energy Problem (and what can we do about it),” 2014 IEEE International Solid-State Circuits Conference (ISSCC 2014), Feb. 9-13, 2014, 5 pages, IEEE, San Francisco, CA, USA. [cited by applicant]
Huan, Yuxiang, et al., “A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity,” May 22, 2017, 5 pages, arXiv:1705.08009v1, Computer Research Repository (CoRR)—Comell University, Ithaca, NY, U… [cited by applicant]
Jeon, Dongsuk, et al., “A 23-mW Face Recognition Processor with Mostly-Read 5T Memory in 40-nm CMOS,” IEEE Journal on Solid-State Circuits, Jun. 2017, 15 pages, vol. 52, No. 6, IEEE, New York, NY, USA. [cited by applicant]
Jouppi, Norman, P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17), Jun. 24-28, 2017, 17 pages, ACM, … [cited by applicant]
Judd, Patrick, et al., “Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing,” Apr. 29, 2017, 6 pages, arXiv:1705.00125v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, U… [cited by applicant]
Leng, Cong, et al., “Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM,” Proceedings of 32nd AAAI Conference on Artificial Intelligence (AAAI-18), Feb. 2-7, 2018, 16 pages, Association for the Advance… [cited by applicant]
Li, Fengfu, et al., “Ternary Weight Networks,” May 16, 2016, 9 pages, arXiv:1605.04711v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Merolla, Paul, et al., “Deep Neural Networks are Robust to Weight Binarization and Other Non-linear Distortions,” Jun. 7, 2016, 10 pages, arXiv:1606.01981v1, Computing Research Repository (CoRR)—Cornell University, Itha… [cited by applicant]
Moons, Bert, et al., “Envision: A 0.26-to-10TOPS/W Subword-Parallel Dynamic-Voltage-Accuracy-Frequency-Scalable Convolutional Neural Network Processor in 28nm FDSOI,” Proceedings of 2017 IEEE International Solid-State C… [cited by applicant]
Moshovos, Andreas, et al., “Exploiting Typical Values to Accelerate Deep Learning,” Computer, May 24, 2018, 13 pages, vol. 51—Issue 5, IEEE Computer Society, Washington, D.C. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 16/525,445, filed Jul. 29, 2019, 104 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 16/525,449, filed Jul. 29, 2019, 106 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 16/525,456, filed Jul. 29, 2019, 106 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 16/525,466, filed Jul. 29, 2019, 105 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 16/525,469, filed Jul. 29, 2019, 105 pages, Perceive Corporation. [cited by applicant]
Park, Jongsoo, et al., “Faster CNNs with Direct Sparse Convolutions and Guided Pruning,” Jul. 28, 2017, 12 pages, arXiv:1608.01409v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Pedram, Ardavan, et al., “Dark Memory and Accelerator-Rich System Optimization in the Dark Silicon Era,” Apr. 27, 2016, 8 pages, arXiv:1602.04183v3, Computer Research Repository (CoRR)—Comell University, Ithaca, NY, USA. [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Proceedings of 2016 European Conference on Computer Vision (ECCV '16), Oct. 8-16, 2016, 17 pages, Lecture Note… [cited by applicant]
Ren, Mengye, et al., “SBNet: Sparse Blocks Network for Fast Inference,” Jan. 7, 2018, 10 pages, arXiv:1801.02108v1, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Rutenbar, Rob A., et al., “Hardware Inference Accelerators for Machine Learning,” 2016 IEEE International Test Conference (ITC), Nov. 15-17, 2016, 39 pages, IEEE, Fort Worth, TX, USA. [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 12 pages, ICLR, Vanco… [cited by applicant]
Shin, Dongjoo, et al., “DNPU: An 8.1TOPS/W Reconfigurable CNN-RNN Processor for General-Purpose Deep Neural Networks,” Proceedings of 2017 IEEE International Solid-State Circuits Conference (ISSCC 2017), Feb. 5-7, 2017,… [cited by applicant]
Sim, Jaehyeong, et al., “A 1.42TOPS/W Deep Convolutional Neural Network Recognition Processor for Intelligent IoE Systems,” Proceedings of 2016 IEEE International Solid-State Circuits Conference (ISSCC 2016), Jan. 31 - … [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Wang, Min, et al., “Factorized Convolutional Neural Networks,” 2017 IEEE International Conference on Computer Vision Workshops (ICCVW '17), Oct. 22-29, 2017, 9 pages, IEEE, Venice, Italy. [cited by applicant]
Wen, Wei, et al., “Learning Structured Sparsity in Deep Neural Networks,” Oct. 18, 2016, 10 pages, arXiv: 1608.03665v4, Computer Research Repository (CoRR)ura onCornell University, Ithaca, NY, USA. [cited by applicant]
Yang, Tien-Ju, et al., “Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning,” Apr. 18, 2017, 9 pages, arXiv:1611.05128v4, Computer Research Repository (CoRR)—Cornell University, Ithaca, N… [cited by applicant]
Yang, Xuan, et al., “DNN Dataflow Choice Is Overrated,” Sep. 10, 2018, 13 pages, arXiv:1809.04070v1, Computer Research Repository (CoRR)ura onCornell University, Ithaca, NY, USA. [cited by applicant]
Zhang, Shijin, et al., “Cambricon-X: An Accelerator for Sparse Neural Networks,” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO '16), Oct. 15-19, 2016, 12 pages, IEEE, Taipei, Taiwan. [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv:1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Cited By (1)
US 12,693,990