IP Library › Granted Patent US 12,353,982
Granted Patent B2
US 12,353,982 · App. 17/291,309 · Granted Jul 8, 2025

Convolution block array for implementing neural network application and method using the same, and convolution block circuit

Inventor: Woon-Sik Suh (New Taipei, CN)
G06N3/063G06F1/03G06F7/50G06F7/523G06F7/53G06F7/5443G06F9/54G06N3/048G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,982
App. No.
17/291,309
Granted
Jul 8, 2025
Kind
B2
Abstract

A convolution block array for implementing neural network application, a method using the same, and a convolution block circuit are provided. The convolution block array includes a plurality of convolution block circuits configured to process a convolution operation of the neural network application, wherein each of the convolution block circuits comprises: a plurality of multiplier circuits configured to perform the convolution operation; and at least one adder circuit connected to the plurality of multiplier circuits and configured to perform an adding operation of results of the convolution operation and generate an output signal; where at least one of the convolution block circuits is configured to perform a biasing operation of the neural network application.

Claims (34)

1. A convolution block array for implementing a neural network application, comprising:

a plurality of convolution block circuits configured to process a convolution operation of the neural network application, wherein each of the convolution block circuits comprising:

a plurality of multiplier circuits configured to perform the convolution operation; and

at least one adder circuit connected to the plurality of multiplier circuits and configured to perform an adding operation of results of the convolution operation and generate an output signal;

wherein at least one of the convolution block circuits comprises at least an idle multiplier circuit and a convolution operation multiplier circuit and the one of the convolution block circuits is configured to concurrently perform a biasing operation and the convolution operation of the neural network application, such that the idle multiplier circuit performs the biasing operation and the convolution operation multiplier circuit performs the convolution operation, wherein a biasing coefficient passes to the at least one adder circuit through the idle multiplier circuit of the plurality of multiplier circuits, and the convolution operation is executed by multiplying feature values by weight coefficients and adding the biasing coefficient.

2. The convolution block array as claimed in claim 1 , wherein each of the convolution block circuits comprises four multiplier circuits, a first convolution adder circuit, a second convolution adder circuit, and a block adder circuit;

wherein two multiplier circuits of the four multiplier circuits are connected to the first convolution adder circuit, and another two multiplier circuits of the four multiplier circuits are connected to the second convolution adder circuit, and

the block adder circuit is connected to the first convolution adder circuit and the second convolution adder circuit.

3. The convolution block array as claimed in claim 1 , wherein the convolution block further comprises a latch connected to the at least one adder circuit and also connected to at least one downstream convolution block circuit;

wherein the latch is configured to transmit the output signal to the at least one downstream convolution block circuit or fed the output signal back to the at least one adder circuit.

4. The convolution block array as claimed in claim 3 , wherein the convolution block array further comprises:

a plurality of multiplexers connected to the plurality of multiplier circuits respectively, wherein the multiplier circuits are connected to the at least one adder circuit via respective multiplexer; and

a path controller connected to the plurality of multiplexers and also connected to at least one upstream convolution block circuit, wherein when a corresponding path of the path controller is enabled, an output signal from the at least one upstream convolution block circuit transmits to the at least one adder circuit via the path controller.

5. A method of implementing neural network application in a convolution block array, wherein the convolution block array comprising a plurality of convolution block circuits, and each of the convolution block circuits comprises a plurality of multiplier circuits and at least one adder circuit, and the method comprises:

S 10 , assigning a set of convolution block circuits in a first dimension to process a convolution operation of the neural network application;

S 20 , inputting a control signal to the set of convolution block circuits for controlling the set of convolution block circuits to perform M×M filter window involving N-bit based convolution operation, wherein M is an odd number and N is an integer greater than one;

S 30 , inputting feature values and filter values to the set of convolution block circuits;

S 40 , performing an N-bit multiplication function with the feature values and the filter values corresponding values of input images by the multiplier circuits of the convolution block circuit;

S 50 , adding results from the multiplier circuits by the at least one adder circuit; and

S 60 , generate a convolution output signal;

wherein a last one of the set of the convolution block circuits comprises at least an idle multiplier circuit and a convolution operation multiplier circuit and the last one of the convolution block circuits is configured to concurrently perform a biasing operation and the convolution operation of the neural network application, such that the idle multiplier circuit performs the biasing operation and the convolution operation multiplier circuit performs the convolution operation, wherein the filter values of one pixel comprise weight coefficients and a biasing coefficient, and in the last convolution block circuit for performing convolution operation of each pixel, the biasing coefficient passes to the at least one adder circuit through the idle multiplier circuit of the plurality of multiplier circuits.

6. The method as claimed in claim 5 , wherein the step S 10 comprises: according to a combination of the M×M filter window and the N-bit of the convolution operation, determine use a number of the convolution block circuits for performing the convolution operation of one pixel of the convolution block circuits, and the number of the convolution block circuits arranged in a line based.

7. The method as claimed in claim 5 , wherein after the step S 50 , the method comprises:

S 51 , adding the results by the at least one adder circuit of each convolution block circuit to generate a partial output signal;

S 52 , transmitting all of partial output signals to the last convolution block circuit; and

S 53 , adding the all of partial output signals by the at least one adder circuit of the last convolution block circuit, and generating the convolution output signal representing one pixel.

8. The method as claimed in claim 5 , wherein each of the convolution block circuits comprises a latch connected to the at least one adder circuit and also connected to the last convolution block circuit, and wherein after the step S 51 , the last convolution block circuit temporarily stores the partial output signal in the latch, and then in the step S 52 , the last convolution block circuit feeds the partial output signal back to its the at least one adder circuit.

9. A convolution block circuit, comprising:

four multiplier circuits configured to perform a M×M filter window involving N-bit based convolution operation;

a first convolution adder circuit connected to two multiplier circuits of the four multiplier circuits and configured to add results of the convolution operation from the two multiplier circuits;

a second convolution adder circuit connected to another two multiplier circuits of the four multiplier circuits and configured to add results of the convolution operation from the two multiplier circuits;

a block adder circuit connected to the first convolution adder circuit and the second convolution adder circuit and configured to perform a first adding operation and a second adding operation, wherein in the first adding operation, the block adder circuit adds results of partial convolution operations from the first convolution adder circuit and the second convolution adder circuit and a biasing coefficient, and generates a first convolution value, wherein the biasing coefficient transmits to the block adder circuit through an idle multiplier circuit of the four multiplier circuits; and

a latch connected to the block adder circuit configured to fed the first convolution value back to the block adder circuit;

wherein in response to the block adder circuit receive the first convolution value and other partial output signals from upstream convolution block circuits, the block adder circuit performs the second adding operation to add the first convolution value and the other partial output signals, and generates a convolution output signal, wherein the convolution block circuit defines a plurality of the four multiplier circuits as convolution operation multiplier circuits and is configured to concurrently perform a biasing operation and the convolution operation, such that the idle multiplier circuit performs the biasing operation and the convolution operation multiplier circuits perform the convolution operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2021
From: SUH, WOON-SIK
To: GENESYS LOGIC, INC.
Reel/Frame 056137/0941 →
Continuity (2)
Provisional Application 62756095 · Nov 6, 2018
Related Publication 20220027714A1 · Jan 27, 2022
References Cited (27)
US 10049323B1 · Kim et al. · 2018 [cited by applicant]
US 10394929B2 · Tsai et al. · 2019 [cited by applicant]
US 11803738B2 · Martin · 2023 [cited by applicant]
US 20180129935A1 · Kim et al. · 2018 [cited by applicant]
US 20180232621A1 · Du et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik · 2018 [cited by examiner]
US 20190266485A1 · Singh · 2019 [cited by examiner]
US 20190294413A1 · Vantrease · 2019 [cited by examiner]
CN 103985083A · 2014 [cited by applicant]
CN 106779060A · 2017 [cited by applicant]
CN 106845635A · 2017 [cited by applicant]
CN 107633297A · 2018 [cited by applicant]
CN 107844826A · 2018 [cited by applicant]
CN 107862374A · 2018 [cited by applicant]
CN 108171317A · 2018 [cited by applicant]
CN 108205701A · 2018 [cited by applicant]
CN 108229671A · 2018 [cited by applicant]
EP 3396534A2 · 2018 [cited by examiner]
GB 201718358A · 2017 [cited by applicant]
TW 201830232A · 2018 [cited by applicant]
Han et al., “Based on the neural network processing system and processing method for production line”, published on Mar. 30, 2018, Document ID: CN-10786374-A, pp. 11 (Year: 2018). [cited by examiner]
Gu et al., “For A Convolutional Neural Network-convolution Operation And Full Connection Arithmetic Circuit”, published on Nov. 6, 2018, Document ID: CN 108764467 A, pp. 24. (Year: 2018). [cited by examiner]
Chris Martin, “Development Of Sparsity In The Neural Network”, published on Jul. 26, 2019, but filed on Nov. 6, 2018, Document ID: CN 110059798 A, pp. 29. (Year: 2018). [cited by examiner]
Wang Xiaofeng:“Design of FPGA accelerator with high parallelism for convolution neural network”;Journal of Computer Applications;2021, 41(3):812-819;CN. [cited by applicant]
Wang Kun:“Convolutional neural network system design and hardware implementation in deep learning”;Technology Application Issue 5, 2018;School of Big Data and Information Engineering, Guizhou University, Guiyang, Guizho… [cited by applicant]
Yufei Ma : “Optimizing the Convolution Operation to Accelerate Deep Neural Networks on FPGA”,IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 26, No. 7, Jul. 2018;17875392;10.1109/TVLSI.2018.281560… [cited by applicant]
Haruyoshi Yonekawa : “On-chip Memory Based Binarized Convolutional Deep Neural Network Applying Batch Normalization Free Technique on an FPGA”,2017 IEEE International Parallel and Distributed Processing Symposium Worksh… [cited by applicant]