IP Library › Granted Patent US 12,505,336
Granted Patent B2
US 12,505,336 · App. 17/243,136 · Granted Dec 23, 2025

Systolic-CNN: an OpenCL-defined scalable runtime-flexible programmable accelerator architecture for accelerating convolutional neural network inference in cloud/edge computing

Inventors: Akshay Dua (San Jose, CA); Fengbo Ren (Tempe, AZ)
Assignee: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
G06N3/063G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,336
App. No.
17/243,136
Granted
Dec 23, 2025
Kind
B2
Abstract

An OpenCL-defined scalable runtime-flexible programmable accelerator architecture for accelerating convolutional neural network (CNN) inference in cloud/edge computing is provided, referred to herein as Systolic-CNN. Existing OpenCL-defined programmable accelerators (e.g., field-programmable gate array (FPGA)-based accelerators) for CNN inference are insufficient due to limited flexibility for supporting multiple CNN models at runtime and poor scalability resulting in underutilized accelerator resources and limited computational parallelism. Systolic-CNN adopts a highly pipelined and paralleled one-dimensional (1-D) systolic array architecture, which efficiently explores both spatial and temporal parallelism for accelerating CNN inference on programmable accelerators (e.g., FPGAs). Systolic-CNN is highly scalable and parameterized, and can be easily adapted by users to efficiently utilize the coarse-grained computation resources for a given programmable accelerator. In addition, Systolic-CNN is runtime-flexible and can be time-shared to accelerate a variety of CNN models at runtime without the need to recompile the programmable accelerator kernel hardware or reprogram the programmable accelerator.

Claims (47)

1 . A method for accelerating a convolutional neural network (CNN) process on a programmable accelerator, the method comprising:

establishing on the programmable accelerator a convolution layer and additional layers which are runtime-flexible for a plurality of CNN models without recompiling the programmable accelerator using optimized values of at least three architectural parameters selected such that the programmable accelerator utilizes up to 100% of available computing resources and reduces runtime relative to using non-optimized values of the at least three architectural parameters, wherein each of the at least three architectural parameters has a different impact on a demand for the available computing resources;

receiving a first request to perform a first CNN inference process; and

at runtime, accelerating the first CNN inference process using the convolution layer and the additional layers with spatial and temporal parallel execution,

wherein the at least three architectural parameters comprise: pe_num, reuse_fac, and vec_fac, wherein:

the pe_num parameter defines a number of processing elements (PEs) in a one-dimensional (1-D) systolic array of PEs that perform temporally paralleled convolution in a deep pipeline and a parallelism of output feature maps (OFMs);

the reuse_fac parameter defines a parallelism of inner product (IP) units inside each of the PEs and a number of times an identical input feature map (IFM) is reused by each of the PEs for convolution computation within an identical OFM; and

the vec_fac parameter defines a single instruction multiple data (SIMD) width of a partial computation between a weight vector and IFM vector across different channels inside each IP unit in each of the PEs.

2 . The method of claim 1 , further comprising:

receiving a second request to perform a second CNN inference process; and

accelerating the second CNN inference process using the convolution layer and the additional layers without recompiling the programmable accelerator.

3 . The method of claim 2 , wherein the first CNN inference process uses at least one of the additional layers which the second CNN inference process does not use.

4 . The method of claim 2 , wherein the second CNN inference process performs a type of inference which is distinct from the first CNN inference process.

5 . The method of claim 1 , wherein accelerating the first CNN inference process comprises executing data-independent loops on the convolution layer in parallel spatially.

6 . The method of claim 5 , wherein executing the data-independent loops of the convolutional layer in parallel spatially comprises executing the data-independent loops over different processing elements (PEs) of the programmable accelerator.

7 . The method of claim 5 , wherein accelerating the first CNN inference process comprises executing data-dependent loops of the convolutional layer in parallel temporally.

8 . The method of claim 5 , wherein the additional layers comprise two or more of a batch normalization (BNORM) layer, a local response normalization (LRN) layer, a max pooling layer, an average pooling layer, an element-wise sum (ELTWISE) layer, and a rectified linear unit (ReLU) layer.

9 . The method of claim 1 , wherein the programmable accelerator comprises a field-programmable gate array (FPGA).

10 . A deep learning system, comprising:

a programmable accelerator; and

a memory storing instructions which, when executed, cause the programmable accelerator to:

establish processing resources on the programmable accelerator which are runtime-flexible for a plurality of convolutional neural network (CNN) models using optimized values of at least three architectural parameters selected such that the programmable accelerator utilizes up to 100% of available computing resources and reduces runtime relative to using non-optimized values of the at least three architectural parameters, wherein each of the at least three architectural parameters has a different impact on a demand for the available computing resources;

receive a request to perform a CNN inference process using one of the plurality of CNN models; and

perform the CNN inference process with the processing resources without recompiling the programmable accelerator,

wherein the at least three architectural parameters comprise: pe_num, reuse_fac, and vec_fac, wherein:

the pe_num parameter defines a number of processing elements (PEs) in a one-dimensional (1-D) systolic array of PEs that perform temporally paralleled convolution in a deep pipeline and a parallelism of output feature maps (OFMs);

the reuse_fac parameter defines a parallelism of inner product (IP) units inside each of the PEs and a number of times an identical input feature map (IFM) is reused by each of the PEs for convolution computation within an identical OFM; and

the vec_fac parameter defines a single instruction multiple data (SIMD) width of a partial computation between a weight vector and IFM vector across different channels inside each IP unit in each of the PEs.

11 . The deep learning system of claim 10 , wherein the processing resources comprise a convolution layer and additional layers.

12 . The deep learning system of claim 10 , wherein the programmable accelerator comprises a field-programmable gate array (FPGA).

13 . The deep learning system of claim 10 , wherein the deep learning system further comprises:

an off-chip memory coupled to the programmable accelerator; and

a host processor configured to provide the request to perform the CNN inference process to the programmable accelerator.

14 . The deep learning system of claim 13 , wherein the processing resources comprise a shift register-based input feature map (IFM) buffer for storing IFM data received from the off-chip memory.

15 . A convolutional neural network (CNN) accelerator architecture, comprising:

a one-dimensional (1-D) systolic array of processing elements (PEs) configured to execute a convolution layer of a CNN; and

an additional layer module configured to provide optional computations for the CNN;

wherein the CNN accelerator architecture is configured to accelerate a plurality of types of CNNs on a programmable accelerator at runtime without reconfiguring the programmable accelerator using optimized values of at least three architectural parameters selected such that the programmable accelerator utilizes up to 100% of available computing resources and reduces runtime relative to using non-optimized values of the at least three architectural parameters, wherein each of the at least three architectural parameters has a different impact on a demand for the available computing resources,

wherein the at least three architectural parameters comprise: pe_num, reuse_fac, and vec_fac, wherein:

the pe_num parameter defines a number of PEs in the 1-D systolic array of PEs that perform temporally paralleled convolution in a deep pipeline and a parallelism of output feature maps (OFMs);

the reuse_fac parameter defines a parallelism of inner product (IP) units inside each of the PEs and a number of times an identical input feature map (IFM) is reused by each of the PEs for convolution computation within an identical OFM; and

the vec_fac parameter defines a single instruction multiple data (SIMD) width of a partial computation between a weight vector and IFM vector across different channels inside each IP unit in each of the PEs.

16 . The CNN accelerator architecture of claim 15 , wherein the additional layer module comprises one or more of a batch normalization (BNORM) layer, a local response normalization (LRN) layer, a max pooling layer, an average pooling layer, an element-wise sum (ELTWISE) layer, and a rectified linear unit (ReLU) layer.

17 . The CNN accelerator architecture of claim 15 , wherein the additional layer module comprises multiple computation layers which are selected at runtime according to a type of CNN selected.

18 . The CNN accelerator architecture of claim 17 , wherein the multiple computation layers are cascaded such that a computation layer not selected provides data passthrough.

19 . The CNN accelerator architecture of claim 15 , further comprising a buffer configured to store an input feature map (IFM) for the CNN received from an external memory.

20 . The CNN accelerator architecture of claim 19 , wherein the buffer is a shift register-based buffer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2021
From: DUA, AKSHAY; REN, FENGBO
To: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
Reel/Frame 056787/0613 →
Continuity (2)
Provisional Application 63016434 · Apr 28, 2020
Related Publication 20210334636A1 · Oct 28, 2021
References Cited (20)
US 12039448B2 · Partovi Nia · 2024 [cited by examiner]
US 20200089506A1 · Power · 2020 [cited by examiner]
Venieris et al. (Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions. ACM Computing Surveys, vol. 51, No. 3, Article 56, published Jun. 2018, 39 pages). (Year: 2018). [cited by examiner]
Rao et al. (Runtime Network Routing for Efficient Image Classification, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, Issue: 10, published Oct. 26, 2018, pp. 2291-2304). (Year: 2018). [cited by examiner]
Aydonat, U. et al., “An OpenCL™ Deep Learning Accelerator on Arria 10,” Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA '17), Feb. 22-24, 2017, Monterey, CA, USA, ACM, 1… [cited by applicant]
Abdelouahab, K. et al., “Accelerating CNN inference on FPGAs: A Survey,” arXiv:1806.01683v1 [cs.DC], May 26, 2018, 30 pages. [cited by applicant]
Azizimazreah, A. et al., “Shortcut Mining: Exploiting Cross-layer Shortcut Reuse in DCNN Accelerators,” 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA), Feb. 16-20, 2019, Washington, D… [cited by applicant]
Colangelo, P. et al., “Exploration of Low Numeric Precision Deep Learning Inference Using Intel® FPGAs,” 2018 IEEE 26th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), Apr. 29-May … [cited by applicant]
Ma, Y. et al., “Optimizing Loop Operation and Dataflow in FPGA Acceleration of Deep Convolutional Neural Networks,” Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, Feb. 2017,… [cited by applicant]
Ma, Y. et al., “Optimizing the Convolution Operation to Accelerate Deep Neural Networks on FPGA,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 26, Issue 7, Jul. 2018, IEEE, 14 pages. [cited by applicant]
Nguyen, D. et al., “A High-Throughput and Power-Efficient FPGA Implementation of YOLO CNN for Object Detection,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 27, Issue 8, Aug. 2019, IEEE, 13 pa… [cited by applicant]
Redmon, J. et al., “YOLO9000: Better, Faster, Stronger,” 2017 IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21-26, 2017, IEEE, pp. 6517-6525. [cited by applicant]
Solovyev, R. et al., “FPGA Implementation of Convolutional Neural Networks with Fixed-Point Calculations,” arXiv:1808.09945v1 [cs.CV], Aug. 29, 2018, 9 pages. [cited by applicant]
Suda, N. et al., “Throughput-Optimized OpenCL-based FPGA Accelerator for Large-Scale Convolutional Neural Networks,” Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA'16),… [cited by applicant]
Wang, D. et al., “PipeCNN: An OpenCL-Based Open-Source FPGA Accelerator for Convolution Neural Networks,” 2017 International Conference on Field Programmable Technology (ICFPT), Dec. 11-13, 2017, Melbourne, VIC, Austral… [cited by applicant]
Wei, X. et al., “Automated Systolic Array Architecture Synthesis for High Throughput CNN Inference on FPGAs,” 2017 54th ACM/EDAC/IEEE Design Automation Conference (DAC '17), Jun. 18-22, 2017, Austin, TX, USA, ACM, 6 pag… [cited by applicant]
Zhang, C. et al., “Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks,” Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, Feb. 2015, ACM, pp. 161-1… [cited by applicant]
Zhang, J. et al., “Frequency Improvement ofSystolic Array-Based CNNs on FPGAs,” 2019 IEEE International Symposium on Circuits and Systems (ISCAS), May 2019, Sapporo, Japan, IEEE, 4 pages. [cited by applicant]
Zhao, R. et al., “Optimizing CNN-based Object Detection Algorithms on Embedded FPGA Platforms,” International Symposium on Applied Reconfigurable Computing, Mar. 2017, 12 pages. [cited by applicant]
Zohouri, H. et al., “Combined Spatial and Temporal Blocking for High-Performance Stencil Computation on FPGAs Using OpenCL,” 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA'18), Feb. 25-27… [cited by applicant]
Cited By (1)
US 12,633,102