IP Library Granted Patent US 12,229,658
Granted Patent B2
US 12,229,658 · App. 16/933,859 · Granted Feb 18, 2025

Configurable processor for implementing convolution neural networks

Inventor: Pavel Sinha (Brossard, CA)
Assignee: AARISH TECHNOLOGIES
G06N3/063G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,229,658
App. No.
16/933,859
Granted
Feb 18, 2025
Kind
B2
Abstract

Configurable processors for implementing CNNs are provided. One such configurable CNN processor includes a plurality of core compute circuitry elements, each configured to perform a CNN function in accordance with a preselected dataflow graph, an active memory buffer, a plurality of connections between the active memory buffer and the plurality of core compute circuitry elements, each established in accordance with the preselected dataflow graph, a plurality of connections between the plurality of core compute circuitry elements, each established in accordance with the preselected dataflow graph, wherein the active memory buffer is configured to move data between the plurality of core compute circuitry elements via the active memory buffer in accordance with the preselected dataflow graph.

Claims (100)

1. A configurable processor dedicated to implementing convolution neural networks (CNNs) and implemented in a single integrated circuit die, comprising:

a plurality of core compute circuits on the die, each configured to perform a CNN function in accordance with a preselected dataflow graph;

an active memory buffer on the die;

a plurality of connections, on the die, between the active memory buffer and the plurality of core compute circuits, each established in accordance with the preselected dataflow graph; and

a plurality of connections, on the die, between the plurality of core compute circuits, each established in accordance with the preselected dataflow graph,

wherein the active memory buffer on the die is configured to store data from, and move data between, the plurality of core compute circuits in accordance with the preselected dataflow graph,

wherein each of the plurality of core compute circuits is configured to perform the CNN function in accordance with the preselected dataflow graph and without using an instruction set, and

wherein the active memory buffer is further configured to apply backpressure on a data generation source.

2. The configurable processor of claim 1 , wherein the preselected dataflow graph is based on a preselected CNN.

3. The configurable processor of claim 1 , wherein at least two of the plurality of core compute circuits are configured to operate asynchronously from one another.

4. The configurable processor of claim 1 , wherein the active memory buffer and each of the plurality of core compute circuits are configured to operate asynchronously from one another.

5. The configurable processor of claim 1 , wherein each of the plurality of core compute circuits is dedicated to performing the CNN function.

6. The configurable processor of claim 1 , wherein each of the plurality of core compute circuits is configured, prior to a runtime of the configurable processor, to perform the CNN function.

7. The configurable processor of claim 1 , wherein each of the plurality of core compute circuits is configured to compute a layer of the CNN function.

8. The configurable processor of claim 1 , wherein each of the plurality of core compute circuits is configured to compute an entire CNN.

9. The configurable processor of claim 1 , wherein each of the plurality of core compute circuits is configured to perform the CNN function for both inference and training.

10. The configurable processor of claim 1 , wherein each of the plurality of core compute circuits comprises a memory configured to store a weight used to perform the CNN function.

11. The configurable processor of claim 1 :

wherein the plurality of connections between the active memory buffer and the plurality of core compute circuits are established during a compile time and fixed during a runtime of the configurable processor; and

wherein the plurality of connections between the plurality of core compute circuits are established during the compile time and fixed during the runtime.

12. A processor array, comprising:

a plurality of the configurable processors of claim 1 ;

an interconnect circuitry; and

a plurality of connections between the plurality of configurable processors and/or the interconnect circuitry, each established in accordance with the preselected dataflow graph.

13. A system comprising:

a mobile industry processor interface camera serial interface (MIPI-CSI) source;

a MIPI-CSI sink;

a MIPI-CSI bus coupled between the MIPI-CSI source and the MIPI-CSI sink; and

the configurable processor of claim 1 disposed serially along the MIPI-CSI bus such that all data on the MIPI-CSI bus passes through the configurable processor.

14. The system of claim 13 , further comprising:

a non-MIPI-CSI output interface comprising at least one of a SPI, an I2C interface, or a UART interface; and

wherein the configurable processor is configured to send information to an external device using either the non-MIPI-CSI output interface or the MIPI-CSI bus.

15. A system comprising:

a sensor configured to generate sensor data;

the configurable processor of claim 1 directly coupled to the sensor and configured to generate processed data based on the sensor data; and

a wireless transmitter directly coupled to the configurable processor and configured to transmit at least a portion of the processed data.

16. The system of claim 15 :

wherein the sensor data comprises image data;

wherein the processed data comprises classification data generated based on the image data; and

wherein the wireless transmitter is configured to transmit the classification data.

17. A method for configuring a configurable processor dedicated to implementing convolution neural networks (CNNs), comprising:

receiving a preselected dataflow graph;

programming, prior to a runtime of the configurable processor, each of a plurality of core compute circuits of the configurable processor to perform a CNN function in accordance with the preselected dataflow graph;

programming, prior to the runtime, an active memory buffer of the configurable processor in accordance with the preselected dataflow graph;

programming a plurality of connections, of the configurable processor prior to the runtime, between the active memory buffer and the plurality of core compute circuits in accordance with the preselected dataflow graph;

programming a plurality of connections, of the configurable processor prior to the runtime, between the plurality of core compute circuits in accordance with the preselected dataflow graph;

programming, prior to the runtime, the active memory buffer to move data between the plurality of core compute circuits via the memory buffer in accordance with the preselected dataflow graph and to apply backpressure on a data generation source; and

operating the plurality of core compute circuits, at the runtime, to perform the CNN function without using an instruction set.

18. The method of claim 17 , further comprising operating the active memory buffer, at the runtime, without using an instruction set.

19. The method of claim 17 , wherein the preselected dataflow graph is based on a preselected CNN.

20. The method of claim 17 , further comprising operating at least two of the plurality of core compute circuits asynchronously from one another.

21. The method of claim 17 , further comprising operating the active memory buffer and each of the plurality of core compute circuits asynchronously from one another.

22. The method of claim 17 , wherein each of the plurality of core compute circuits is dedicated to performing the CNN function.

23. The method of claim 17 , further comprising:

performing, during the runtime, the CNN function at each of a respective one of the plurality of core compute circuits.

24. The method of claim 17 , further comprising:

computing, during the runtime, a layer of the CNN function at each of a respective one of the plurality of core compute circuits.

25. The method of claim 17 , further comprising:

computing, during the runtime, an entire CNN at least one of the plurality of core compute circuits.

26. The method of claim 17 :

wherein the plurality of connections between the active memory buffer and the plurality of core compute circuits are programmed during a compile time and fixed during the runtime; and

wherein the plurality of connections between the plurality of core compute circuits are programmed during the compile time and fixed during the runtime.

27. The method of claim 17 , wherein each of the plurality of core compute circuits is configured to perform the CNN function for both inference and training.

28. The method of claim 17 , wherein each of the plurality of core compute circuits comprises a memory configured to store a weight used to perform the CNN function.

29. A configurable processor dedicated to implementing convolution neural networks (CNNs) and implemented in a single integrated circuit die, comprising:

a plurality of means, on the die, for performing a CNN function in accordance with a preselected dataflow graph;

a means, on the die, for storing data;

a means, on the die, for establishing connections between the means for storing data and the plurality of means for performing the CNN function, in accordance with the preselected dataflow graph; and

a means, on the die, for establishing connections between the plurality of means for performing the CNN function, in accordance with the preselected dataflow graph,

wherein the means for storing data comprises a means for moving data between the plurality of means for performing the CNN function via the means for storing data in accordance with the preselected dataflow graph,

wherein each of the plurality of means for performing the CNN function is configured to perform the CNN function in accordance with the preselected dataflow graph and without using an instruction set, and

wherein the means for storing data is configured to apply backpressure on a data generation source.

30. A configurable processor dedicated to implementing convolution neural networks (CNNs), comprising:

a mobile industry processor interface camera serial interface (MIPI-CSI) input circuitry configured to be directly coupled to a MIPI-CSI source circuitry;

a MIPI-CSI output circuitry configured to be directly coupled to an application processor;

a MIPI-CSI bus coupled between the MIPI-CSI input circuitry and the MIPI-CSI output circuitry; and

a configurable CNN sub-processor implemented in a single integrated circuit die and disposed serially along the MIPI-CSI bus such that all data on the MIPI-CSI bus passes through the configurable processor, the configurable CNN sub-processor configured to:

receive image data from the MIPI-CSI source;

generate processed data based on the image data and without using an instruction set; and

provide the processed data to the application processor,

wherein the configurable CNN sub-processor comprises:

a plurality of core compute circuits on the die, each configured to perform a CNN function in accordance with a preselected dataflow graph; and

an active memory buffer on the die, wherein the active memory buffer on the die is configured to:

store data from, and move data on the die between the plurality of core compute circuits via the active memory buffer in accordance with the preselected dataflow graph; and

apply backpressure on a data generation source.

31. The configurable processor of claim 30 , wherein the configurable CNN sub-processor is further configured to generate the processed data based on the image data using a preselected CNN.

32. The configurable processor of claim 30 , wherein the configurable CNN sub-processor comprises a plurality of the configurable CNN sub-processors in a cascade configuration.

33. The configurable processor of claim 30 , wherein the configurable CNN sub-processor is configured to provide the processed data to the application processor via the MIPI-CSI bus.

34. The configurable processor of claim 30 , wherein the configurable CNN sub-processor further comprises:

a plurality of connections, on the die, between the active memory buffer and the plurality of core compute circuits, each established in accordance with the preselected dataflow graph; and

a plurality of connections, on the die, between the plurality of core compute circuits, each established in accordance with the preselected dataflow graph.

35. The configurable processor of claim 30 , further comprising:

a non-MIPI-CSI output interface comprising at least one of a SPI, an I2C interface, or a UART interface; and

wherein the configurable processor is configured to send information to the application processor using either the non-MIPI-CSI output interface or the MIPI-CSI bus.

36. The configurable processor of claim 1 :

wherein each of the plurality of core compute circuits configured to perform the CNN function in accordance with the preselected dataflow graph is further configured to generate intermediate results associated with performance of the CNN function; and

wherein the intermediate results are stored on the die and not in an external memory.

37. The configurable processor of claim 1 , wherein each of the plurality of core compute circuits are configured with a static configuration during a compile time of the configurable processor such that the static configuration does not change during a runtime of the configurable processor.

38. The configurable processor of claim 30 , wherein the active memory buffer comprises two or more ports configured to operate asynchronously from one another.

39. The configurable processor of claim 1 , wherein the active memory buffer comprises two or more ports configured to operate asynchronously from one another.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: SINHA, PAVEL
To: AARISH TECHNOLOGIES
Reel/Frame 065011/0500 →
Continuity (4)
Provisional Application 63025580 · May 15, 2020
Provisional Application 62941646 · Nov 27, 2019
Provisional Application 62876219 · Jul 19, 2019
Related Publication 20210034958A1 · Feb 4, 2021
References Cited (48)
US 10331983B1 · Yang · 2019 [cited by examiner]
US 20110206381A1 · Ji et al. · 2011 [cited by applicant]
US 20140180989A1 · Krizhevsky et al. · 2014 [cited by applicant]
US 20190205737A1 · Bleiweiss · 2019 [cited by examiner]
US 20200272779A1 · Boesch · 2020 [cited by examiner]
JP 2013008221A · 2013 [cited by applicant]
JP 2019003414A · 2019 [cited by applicant]
WO 2018193370A1 · 2018 [cited by applicant]
Pham, Phi-Hung, et al. “NeuFlow: Dataflow vision processing system-on-a-chip.” 2012 IEEE 55th International Midwest Symposium on Circuits and Systems (MWSCAS). IEEE, 2012. (Year: 2012). [cited by examiner]
Di Febbo, Paolo, et al. “Kcnn: Extremely-efficient hardware keypoint detection with a compact convolutional neural network.” Proceedings of the IEEE conference on computer vision and pattern recognition workshops. 2018.… [cited by examiner]
Frenger, Paul. “The Ultimate RISC: A zero-instruction computer.” ACM Sigplan Notices 35.2 (2000): 17-24. (Year: 2000). [cited by examiner]
Pinkevich, V. Yu, A. E. Platunov, and A. V. Penskoi. “The approach to design of problem-oriented reconfigurable hardware computational units.” 2020 Wave Electronics and its Application in Information and Telecommunicati… [cited by examiner]
Krizhevsky, Alex et al., “ImageNet Classification with Deep Convolutional Neural Networks”, Advances in Neural Information Processing Systems; 2012; https://proceedings.neurips.cc/paper/2012/file/c399862d3b9d6b76c8436e9… [cited by applicant]
Long, Jonathan et al., “Fully Convolutional Networks for Semantic Segmentation”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Nov. 14, 2014; https://arxiv.org/abs/1411.4038; 10 pages. [cited by applicant]
Vinyals, Oriol et al., “Show and Tell: A Neural Image Caption Generator”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Nov. 17, 2014; https://arxiv.org/abs/1411.4555; 9 pages. [cited by applicant]
Toshev, Alexander et al., “DeepPose: Human Pose Estimation via Deep Neural Networks”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Dec. 17, 2013; https://arxiv.org/abs/1312.4659; 9 page… [cited by applicant]
Lecun, Yann et al., “Gradient-Based Learning Applied to Document Recognition”, Proceedings of the IEEE; vol. 36, Issue 11; Nov. 1998; https://ieeexplore.ieee.org/document/726791; 46 pages. [cited by applicant]
Zeiler, Matthew D. et al., “Visualizing and Understanding Convolutional Networks”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Nov. 12, 2013; https://arxiv.org/abs/1311.2901; 11 pages. [cited by applicant]
Simonyan, Karen et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Sep. 4, 2014; https://arxiv.org/abs/1409.1556;… [cited by applicant]
Szegedy, Christian et al., “Going Deeper with Convolutions”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Sep. 17, 2014; https://arxiv.org/abs/1409.4842; 12 pages. [cited by applicant]
He, Kaiming et al., “Deep Residual Learning for Image Recognition”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Dec. 10, 2015; https://arxiv.org/abs/1512.03385; 12 pages. [cited by applicant]
Jaderberg, Max et al., “Spatial Transformer Networks”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Jun. 5, 2015; https://arxiv.org/abs/1506.02025; 15 pages. [cited by applicant]
Szegedy, Christian et al., “Going Deeper with Convolutions”, 2015 IEEE Conference on Computer Vision and Pattern Recognition; 2015; https://doi.ieeecomputersociety.org/10.1109/CVPR.2015.7298594; 9 pages. [cited by applicant]
He, Kaiming et al., “Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition”, IEEE Transactions on Pattern Analysis and Machine Intelligence; vol. 37, Issue 9; Sep. 1, 2015; https://ieeexplore.iee… [cited by applicant]
Iandola, Forrest N. et al., “SqueezeNet: AlexNet-Level Accuracy with 50x Fewer Parameters and <0.5MB Model Size”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Feb. 24, 2016; https://arx… [cited by applicant]
Wan, Lihong et al., “Face Recognition with Convolutional Neural Networks and Subspace Learning”, 2017 2nd International Conference on Image, Vision and Computing; Jun. 2-4, 2017; https://ieeexplore.ieee.org/document/798… [cited by applicant]
Canziani, Alfredo et al., “An Analysis of Deep Neural Network Models for Practical Applications”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; May 24, 2016; https://arxiv.org/abs/1605.0… [cited by applicant]
Strigl, Daniel et al., “Performance and Scalability of GPU-based Convolutional Neural Networks”, 2010 18th Euromicro Conference on Parallel, Distributed & Network-based Processing; Feb. 17-19, 2010; https://ieeexplore.i… [cited by applicant]
Ovtcharov, Kalin et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware”, Microsoft Research; Feb. 22, 2015; https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/CNN20Whitepap… [cited by applicant]
Andri, Renzo et al., “YodaNN: An Ultra-Low Power Convolutional Neural Network Accelerator Based on Binary Weights”, 2016 IEEE Computer Society Annual Symposium on VLSI; Jul. 11-13, 2016; https://ieeexplore.ieee.org/docu… [cited by applicant]
Jafri, Syed M. A. H. et al., “Can a Reconfigurable Architecture Beat ASIC as a CNN Accelerator?”, 2017 International Conference on Embedded Computer Systems: Architectures, Modeling, and Simulation; Jul. 17-20, 2017; ht… [cited by applicant]
Jouppi, Norman P. et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit”, Cornell University; Computer Science: Hardware Architecture; Apr. 16, 2017; https://arxiv.org/abs/1704.04760; 17 pages. [cited by applicant]
Courbariaux, Matthieu et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights During Propagations”, Cornell University; Computer Science: Machine Learning; Nov. 2, 2015; https://arxiv.org/abs/1511.0036… [cited by applicant]
Rastegari, Mohammad et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Mar. 16, 2016; https://arxiv.org… [cited by applicant]
Zhou, Shuchang et al., “DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients”, Cornell University; Computer Science: Neural and Evolutionary Computing; Jun. 20, 2016; https://arxiv… [cited by applicant]
Hubara, Itay et al., Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations; Cornell University; Computer Science: Neural and Evolutionary Computing; Sep. 22, 2016; https://arxiv.… [cited by applicant]
Lin, Darryl D. et al., “Fixed Point Quantization of Deep Convolutional Networks”, Cornell University; Computer Science: Machine Learning; Nov. 19, 2015; https://arxiv.org/abs/1511.06393?context=cs; 10 pages. [cited by applicant]
Mishra, Asit et al., “WRPN:Wide Reduced-Precision Networks”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Sep. 4, 2017; https://arxiv.org/abs/1709.01134; 11 pages. [cited by applicant]
Chen, Yu-Hsin et al., “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks”, IEEE Journal of Solid-State Circuits; vol. 52, Issue 1; Jan. 2017; https://ieeexplore.ieee.org/docu… [cited by applicant]
Moons, Bert et al., “A 0.3-2.6 TOPS/W Precision-Scalable Processor for Real-Time Large-Scale ConvNets”, Cornell University; Computer Science: Hardware Architecture; Jun. 16, 2016; https://arxiv.org/pdf/1606.05094.pdf; 2… [cited by applicant]
Moons, Bert et al., “14.5 Envision: A 0.26-to-10TOPS/W subword-parallel dynamic-voltage-accuracy-frequency-scalable Convolutional Neural Network processor in 28nm FDSOI”, 2017 IEEE International Solid-State Circuits Con… [cited by applicant]
Aimar, Alessandro et al., “NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Ju… [cited by applicant]
Groq, Inc., “Open Platform. Performance without Lock-In.”, Jan. 4, 2019; Last accessed Dec. 14, 2022 via Wayback Machine; https://web.archive.org/web/20190104174009/https://groq.com/; 2 pages (Website). [cited by applicant]
Sun, Baohua et al., “Ultra Power-Efficient CNN Domain Specific Accelerator with 9.3TOPS/Watt for Mobile and Embedded Applications”, Cornell University; Computer Science: Computer Vision and Pattern Recognition; Apr. 30,… [cited by applicant]
Dennis, Jack B. et al., “An Efficient Pipelined Dataflow Processor Architecture”, Supercomputing '88:Proceedings of the 1988 ACM/IEEE Conference on Supercomputing, vol. I; Nov. 14-18, 1988; https://ieeexplore.ieee.org/d… [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/IB2020/000609, dated Nov. 4, 2020, 13 pages. [cited by applicant]
Pham, Phi-Hung et al., “NeuFlow: Dataflow Vision Processing System-on-a-Chip”; 2012 IEEE 55th International Midwest Symposium on Circuits and Systems; 2012; https://ieeexplore.ieee.org/document/6292202; 4 pages. [cited by applicant]
Desoli, Giuseppe et al., “A 2.9TOPS/W Deep Convolutional Neural Network SoC in FD-SOI 28nm for Intelligent Embedded Systems”; 2017 IEEE International Solid-State Circuits Conference; 2017; https://ieeexplore.ieee.org/do… [cited by applicant]
Cited By (1)
US 12,493,574