IP Library Granted Patent US 12,198,042
Granted Patent B2
US 12,198,042 · App. 17/347,233 · Granted Jan 14, 2025

Binary neural network based central processing unit

Inventors: Jie Gu (Evanston, IL); Tianyu Jia (Evanston, IL)
Assignee: NORTHWESTERN UNIVERSITY
G06N3/065G06F15/7867G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,042
App. No.
17/347,233
Granted
Jan 14, 2025
Kind
B2
Abstract

Systems and methods for a unified reconfigurable neural central processing unit is provided. In one aspect, a neural central processing unit is in communication with a memory, wherein the neural central processing unit is configured to transition between a binary neural network accelerator mode and a central processing unit mode, wherein, in the binary neural network accelerator mode, the memory is configured as an image memory and weight memories, wherein, in the central processing unit mode, the memory is reconfigured, from the image memory and the weight memories, to a data cache.

Claims (35)

1. A system, comprising:

a memory;

a first layer in communication with the memory;

an instruction cache in communication with the first layer;

a second layer in communication with the first layer;

a register file in communication with the first layer and the second layer;

a third layer in communication with the first layer, the second layer, and the memory;

a fourth layer in communication with the third layer, the register file, and the memory; and

a result memory in communication with the fourth layer,

wherein, in a binary neural network accelerator mode, the memory is configured as an image memory and weight memories,

wherein, in a central processing unit mode, the memory is reconfigured, from the image memory and the weight memories, to a data cache.

2. The system of claim 1 , wherein the first layer comprises a plurality of XNOR neuron cells.

3. The system of claim 1 , wherein the first layer comprises a 32-bit adder.

4. The system of claim 3 , wherein the 32-bit adder is based on a RISC-V 32-bit base integer instruction set.

5. The system of claim 1 , wherein a portion of the first layer is configured as a program counter.

6. The system of claim 1 , wherein a portion of the first layer is configured to fetch instructions.

7. The system of claim 1 , wherein the second layer is configured to decode instructions into partial codes.

8. The system of claim 1 , wherein the third layer is configured to perform as an arithmetic logic unit.

9. The system of claim 1 , wherein the fourth layer is configured to read data from the data cache and to write data to the data cache.

10. The system of claim 1 , wherein transitioning between the binary neural network accelerator mode and central processing unit mode switches at zero-latency.

11. An edge device, comprising:

a memory; and

a neural central processing unit in communication with the memory,

wherein the neural central processing unit is configured to transition between a binary neural network accelerator mode and a central processing unit mode,

wherein, in the binary neural network accelerator mode, the memory is configured as an image memory and weight memories,

wherein, in the central processing unit mode, the memory is reconfigured, from the image memory and the weight memories, to a data cache.

12. The edge device of claim 11 , wherein the neural central processing unit comprises a first layer comprising a plurality of XNOR neuron cells.

13. The edge device of claim 12 , wherein the first layer comprises a 32-bit adder.

14. The edge device of claim 13 , wherein the 32-bit adder is based on a RISC-V 32-bit base integer instruction set.

15. The edge device of claim 12 , wherein a portion of the first layer is configured as a program counter.

16. The edge device of claim 12 , wherein a portion of the first layer is configured to fetch instructions.

17. The edge device of claim 11 , wherein the neural central processing unit comprises a second layer configured to decode instructions into partial codes.

18. The edge device of claim 11 , wherein the neural central processing unit comprises a third layer configured to perform as an arithmetic logic unit.

19. The edge device of claim 11 , wherein the neural central processing unit comprises a fourth layer configured to read data from a data cache of the neural central processing unit and to write data to the data cache of the neural central processing unit.

20. The edge device of claim 11 , wherein transitioning between the binary neural network accelerator mode and central processing unit mode switches at zero-latency.

Assignments (3)
CONFIRMATORY LICENSE Recorded Jan 30, 2025
From: NORTHWESTERN UNIVERSITY
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 070059/0680 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2024
From: GU, JIE; JIA, TIANYU
To: NORTHWESTERN UNIVERSITY
Reel/Frame 068599/0774 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2021
From: GU, JIE; JIA, TIANYU
To: NORTHWESTERN UNIVERSITY
Reel/Frame 056537/0207 →
Continuity (2)
Provisional Application 63039192 · Jun 15, 2020
Related Publication 20210390383A1 · Dec 16, 2021
References Cited (52)
US 20180121796A1 · Deisher · 2018 [cited by examiner]
Otseidu, Kofi et al., “Design and Optimization of Edge Computing Distributed Neural Processor for Biomedical Rehabilitation with Sensor Fusion,” International Conference on Computer-Aided Design (ICCAD), Nov. 2018, 8 Pa… [cited by applicant]
Blaauw, D. et al., “IoT Design Space Challenges: Circuits and Systems,” Symposium on VLSI Technology Digest of Technical Papers, 2014, 2 Pages. [cited by applicant]
Whatmough, Paul N. et al., “A 28nm SoC with a 1.2GHz 568nJ/Prediction Sparse Deep-Neural-Network Engine with >0.1 Timing Error Rate Tolerance for IoT Applications,” IEEE International Solid-State Circuits Conference, Fe… [cited by applicant]
Hempstead, Mark et al., “An Ultra Low Power System Architecture for Sensor Network Applications,” Proceedings of the 32nd International Symposium on Computer Architecture (ISCA'05), 2005, 12 Pages. [cited by applicant]
Sridhara, Srinivasa R. et al., “Microwatt Embedded Processor Platform for MedicalSystem-on-Chip Applications,” IEEE Journal of Solid-State Circuits, vol. 46, No. 4, Apr. 2011, 10 Pages. [cited by applicant]
Shi, Yao et al., “A 10mm3 Syringe-Implantable Near-Field Radio System on Glass Substrate,” IEEE International Solid-State Circuits Conference, Feb. 2016, 3 Pages. [cited by applicant]
Chen, Tianshi et al., “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning,” International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS),… [cited by applicant]
Chen, Yu-Hsin et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Jun. 2016, 13 Pages. [cited by applicant]
Jouppi, Norman P. et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” International Symposium on Computer Architecture (ISCA), Jun. 2017, 12 Pages. [cited by applicant]
Desoli, Giuseppe et al., “A 2.9TOPS/W Deep Convolutional Neural Network SoC in FD-SOI 28nm for Intelligent Embedded Systems,” IEEE International Solid-State Circuits Conference, Feb. 2017, 3 Pages. [cited by applicant]
Online Resource, Intel, “Neural Compute Engine: Hardware Based Acceleration for Deep Neural Networks”, https://www.movidius.com/MyriadX. [cited by applicant]
Song, Jinook et al., “An 11.5TOPS/W 1024-MAC Butterfly Structure Dual-Core Sparsity-Aware Neural Processing Unit in 8nm Flagship Mobile SoC,” IEEE International Solid-State Circuits Conference, Feb. 2019, 3 Pages. [cited by applicant]
Karnik, Tanay et al., “A cm-Scale Self-Powered Intelligent and Secure IoT Edge Mote Featuring an Ultra-Low-Power SoC in 14nm Tri-Gate CMOS,” IEEE International Solid-State Circuits Conference, Feb. 2018, 3 Pages. [cited by applicant]
Honkote, Vinayak et al., “A Distributed Autonomous and Collaborative Multi-Robot System Featuring a Low-Power Robot SoC in 22nm CMOS for Integrated Battery-Powered Minibots,” IEEE International Solid-State Circuits Conf… [cited by applicant]
Han, Song et al., “Deep Compression: Compressing Deep Neural Networks With Pruning, Trained Quantization and Huffman Coding,” International Conference on Learning Representations (ICLR), May 2016, 14 Pages. [cited by applicant]
Howard, Andrew G. et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” arXiv preprint arXiv:1704.04861v1, 2017, 9 Pages. [cited by applicant]
Han, Song et al., “EIE: Efficient Inference Engine on Compressed Deep Neural Network,” ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Jun. 2016, 12 Pages. [cited by applicant]
Parashar, Angshuman et al., “SCNN: An Accelerator for Compressed-sparseConvolutional Neural Networks,” International Symposium on Computer Architecture (ISCA), Jun. 2017, 14 Pages. [cited by applicant]
Moons, Bert et al., “ENVISION: A 0.26-to-10TOPS/W Subword-Parallel Dynamic-Voltage-Accuracy-Frequency-Scalable Convolutional Neural Network Processor in 28nm FDSOI,” IEEE International Solid-State Circuits Conference, F… [cited by applicant]
Ueyoshi, Kodai et al., “QUEST: A 7.49TOPS Multi-Purpose Log-Quantized DNN Inference Engine Stacked on 96MB 3D SRAM Using Inductive-Coupling Technology in 40nm CMOS,” IEEE International Solid-State Circuits Conference, F… [cited by applicant]
Abadi, Martin et al., “TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems,” arXiv preprint arXiv:1603.04467v2, Mar. 2016, 19 Pages. [cited by applicant]
Narayanan, Deepak et al., “Accelerating Deep Learning Workloads through Efficient Multi-Model Execution,” NeurJPS Workshop on Systems for Machine Learning, Dec. 2018, 8 Pages. [cited by applicant]
Wu, Carole-Jean et al., “Machine Learning at Facebook: Understanding Inference at the Edge,” IEEE International Symposium on High Performance Computer Architecture (HPCA), Feb. 2019, 14 Pages. [cited by applicant]
Sridhara, Srinivasa R., “Ultra-Low Power Microcontrollers for Portable, Wearable, and Implantable Medical Electronics,” Asia and South Pacific Design Automation Conference (ASP-DAC), Jan. 2011, 5 Pages. [cited by applicant]
Bol, David et al., “A 25MHz 7 μW/MHz Ultra-Low-Voltage Microcontroller SoC in 65nm LP/GP CMOS for Low-Carbon Wireless Sensor Nodes,” IEEE International Solid-State Circuits Conference, Feb. 2012, 3 Pages. [cited by applicant]
Online Resource, Nvidia, “Embedded Systems for Next-Generation Autonomous-Machines”, https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/. [cited by applicant]
Online Resource, Google, “Edge TPU”, https://cloud.google.com/edge-tpu/. [cited by applicant]
Hill, Mark D. et al., “Amdahl's Law in the Multicore Era,” Computer, vol. 41, No. 7, Jul. 2008, 6 Pages. [cited by applicant]
Esmaeilzadeh, Hadi et al., “Dark Silicon and the End of Multicore Scaling,” International Symposium on Computer Architecture (ISCA), Jun. 2011, 12 Pages. [cited by applicant]
Zhang, Jintao et al., “In-Memory Computation of a Machine-LearningClassifier in a Standard 6T SRAM Array,” IEEE Journal of Solid-State Circuits, vol. 52, No. 4, Apr. 2017, 10 Pages. [cited by applicant]
Jiang, Zhewei et al., “XNOR-SRAM: In-Memory Computing SRAM Macro for Binary/Ternary Deep Neural Networks,” IEEE Symposium on VLSI Technology Digest of Technical Papers, Jun. 2018, 2 Pages. [cited by applicant]
Eckert, Charles et al., “Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks,” ACM/IEEE 45th Annual International Symposium on Computer Architecture, Jun. 2018, 14 Pages. [cited by applicant]
Graves, Alex et al., “Neural Turing Machines,” arXiv preprint arXiv:1410.5401v2, 2014, 26 Pages. [cited by applicant]
Graves, Alex et al., “Symbolic Reasoning with Differentiable Neural Computers,” Nature, vol. 538, Oct. 2016, 64 Pages. [cited by applicant]
Jang, Hanhwi et al., “MnnFast: A Fast and Scalable System Architecture for Memory-Augmented Neural Networks,” International Symposium on Computer Architecture (ISCA), Jun. 2019, 14 Pages. [cited by applicant]
Stevens, Jacob R. et al., “Manna: An Accelerator for Memory-Augmented Neural Networks,” IEEE/ACM International Symposium on Microarchitecture (MICRO), Oct. 2019, 13 Pages. [cited by applicant]
Trask, Andrew et al., “Neural Arithmetic Logic Units,” arXiv preprint arXiv:1808.00508v1, Aug. 2018, 15 Pages. [cited by applicant]
Chen, Chixiao et al., “Exploring the Programmability for Deep Learning Processors: from Architecture to Tensorization,” Design Automation Conference (DAC), Jun. 2018, 6 Pages. [cited by applicant]
Putic, Mateja et al., “DyHard-DNN: Even More DNN Acceleration with Dynamic Hardware Reconfiguration,” Design Automation Conference (DAC), Jun. 2018, 6 Pages. [cited by applicant]
Hubara, Itay et al., “Binarized Neural Networks,” Advances in Neural Information Processing Systems (NIPS), Dec. 2016, 9 Pages. [cited by applicant]
Ando, Kota et al., “BRein Memory: A Single-Chip Binary/TernaryReconfigurable in-Memory Deep Neural Network Accelerator Achieving 1.4 TOPS at 0.6 W,” IEEE Journal of Solid-State Circuits, vol. 53, No. 4, Apr. 2018, 12 Pa… [cited by applicant]
Bankman, Daniel et al., “An Always-On 3.8 ÂμJ/86% CIFAR-10Mixed-Signal Binary CNN Processor With All Memory on Chip in 28-nm CMOS,” IEEE Journal of Solid-State Circuits, vol. 54, No. 1, Jan. 2019, 15 Pages. [cited by applicant]
Khwa, Win-San et al., “A 65nm 4Kb Algorithm-Dependent Computing-in-Memory SRAM Unit-Macro with 2.3ns and 55.8TOPS/W Fully Parallel Product-Sum Operation for Binary DNN Edge Processors,” IEEE International Solid-State Ci… [cited by applicant]
Park, Jeongwoo et al., “A 65nm 236.5nJ/Classification Neuromorphic Processor with 7.5% Energy Overhead On-Chip Learning Using Direct Spike-Only Feedback,” IEEE International Solid-State Circuits Conference, Feb. 2019, 3… [cited by applicant]
Waterman, Andrew et al., “The RISC-V Instruction Set Manual, vol. I: User-Level ISA, Document Version 2.2,” RISC-V Foundation, May 2017, 145 Pages. [cited by applicant]
Keller, Ben et al., “A RISC-V Processor SoC With Integrated Power Management at Submicrosecond Timescales in 28nm FD-SOI,” IEEE Journal of Solid-State Circuits, vol. 52, No. 7, Jul. 2017, 13 Pages. [cited by applicant]
Guthaus, Matthew R. et al., “MiBench: A Free, commercially representative embedded benchmark suite,” IEEE International Workshop on Workload Characterization, 2001, 12 Pages. [cited by applicant]
Mishiba, Kazu et al., “Image Resizing With Sift Feature Preservation,” IEEE International Conference on Image Processing (ICIP), 2013, 5 Pages. [cited by applicant]
Mitianoudis, Nikolaos et al., “Multi-Spectral Document Image Binarization Using Image Fusion and Background Subtraction Techniques,” IEEE International Conference on Image Processing (ICIP), 2014, 5 Pages. [cited by applicant]
Demirovic, Damir et al., “Performance of some image processing algorithms in TensorFlow,” International Conference on Systems, Signals and Image Processing (WSSIP), 2008, 4 Pages. [cited by applicant]
Atzori, Manfredo et al., “Building the NINAPRO Database: A Resource for the Biorobotics Community,” IEEE International Conference on Biomedical Robotics andBiomechatronics (BioRob), Jun. 2012, 8 Pages. [cited by applicant]