IP Library › Granted Patent US 12,614,058
Granted Patent B2
US 12,614,058 · App. 17/270,853 · Granted Apr 28, 2026

Architecture of a computer for calculating a convolution layer in a convolutional neural network

Inventors: Vincent Lorrain (Saulx-les-Chartreux, FR); Olivier Bichler (Meille-Eglise-en-Yvelines, FR); Mickael Guibert (Le Perreux-sur-Marne, FR)
Assignee: COMMISSARIAT A L'ENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES
G06N3/04G06F5/01G06F17/153G06F17/16G06N3/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,058
App. No.
17/270,853
Granted
Apr 28, 2026
Kind
B2
Abstract

A computer for computing a convolutional layer of an artificial neural network, includes at least one set of at least two partial sum computing modules connected in series, a storage member for storing the coefficients of at least one convolution filter, each partial sum computing module comprising at least one computing unit configured so as to carry out a multiplication of an item of input data of the computer and a coefficient of a convolution filter, followed by an addition of the output of the preceding partial sum computing module in the series, each set furthermore comprising, for each partial sum computing module except the first in the series, a shift register connected at input for storing the item of input data for the processing duration of the preceding partial sum computing modules in the series.

Claims (17)

1 . A computer for computing a convolutional layer of an artificial neural network, comprising:

at least one set of at least two partial sum computing modules connected in series, a storage member for storing the coefficients of at least one convolution filter, each partial sum computing module comprising several computing units each configured so as to carry out a multiplication of an item of input data of the computer and a coefficient of a convolution filter, followed by an addition of the output of the preceding partial sum computing module in the series or of a predefined value for the first partial sum computing module in the series,

each set furthermore comprising, for each partial sum computing module except the first in the series, a shift register connected at input for storing the item of input data for the processing duration of the preceding partial sum computing modules in the series, said shift register being connected at output to the next partial sum computing module in the series and being configured to introduce a latency equivalent to said processing duration to the data received by said next partial sum computing module in the series,

the computer furthermore comprising at least one accumulator connected at output of each set and a memory, the input data of the computer coming from at least two input matrices, each partial sum computing module being configured so as to receive, at input, the input data belonging to different input matrices and having the same coordinates in each input matrix,

in each clock cycle, a value of one of the at least two input matrices, read sequentially row by row, being received at input of one of the at least two partial sum computing modules, said value being received, in parallel, at input of each computing unit belonging to said one of the at least two partial sum computing modules, said computing units each receiving a coefficient of the convolution filter to implement a multiplication with said value in order to compute a different output neuron.

2 . The computer as claimed in claim 1 , configured so as to deliver, at output, for each input sub-matrix of dimension equal to that of the convolution filter, the value of a corresponding output neuron, the set of output neurons being arranged in at least one output matrix.

3 . The computer as claimed in claim 2 , wherein each partial sum computing module comprises at most a number of computing units equal to the dimension of the convolution filter.

4 . The computer as claimed in claim 2 , wherein each set comprises at most a number of partial sum computing modules equal to the number of input matrices.

5 . The computer as claimed in claim 2 , comprising at most a number of sets equal to the number of output matrices.

6 . The computer as claimed in claim 2 , wherein, for each received item of input data, each partial sum computing module is configured so as to compute a partial convolution result for all of the output neurons connected to the item of input data.

7 . The computer as claimed in claim 6 , wherein each partial sum computing module comprises multiple computing units, each one being configured so as to compute a partial convolution result for different output neurons of the other computing units.

8 . The computer as claimed in claim 6 , wherein each partial sum computing module is configured, for each received item of input data, so as to select, in the storage member, the coefficients of a convolution filter corresponding to the respective output neurons to be computed for each computing unit.

9 . The computer as claimed in claim 2 , wherein the input matrices are images.

10 . The computer as claimed in claim 1 , wherein the storage member has a two-dimensional toroidal topology.

11 . The computer as claimed in claim 1 , wherein the at least one accumulator connected at output of each set is configured so as to finalize a convolution computation in order to obtain the value of an output neuron from the partial sums delivered by the set the memory being used to save partial results of the convolution computation.

12 . The computer as claimed in claim 11 , wherein the addresses of the values stored in the memory are determined and configured to avoid two output neurons during the computation sharing the same memory block.

13 . The computer as claimed in claim 1 , furthermore comprising an activation module for activating an output neuron, connected at output of each accumulator.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2021
From: LORRAIN, VINCENT; BICHLER, OLIVIER; GUIBERT, MICKAEL
To: COMMISSARIAT A L'ENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES
Reel/Frame 057320/0323 →
Priority Claims (1)
FR 1857852 · Aug 31, 2018 · national
Continuity (1)
Related Publication 20210241071A1 · Aug 5, 2021
References Cited (38)
US 3748451A · Ingwersen · 1973 [cited by examiner]
US 8103606B2 · Moussa · 2012 [cited by examiner]
US 10635740B2 · Phelps · 2020 [cited by examiner]
US 20110208795A1 · Pajaniradja · 2011 [cited by examiner]
US 20140160135A1 · Krig · 2014 [cited by examiner]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180341495A1 · Culurciello · 2018 [cited by examiner]
CN 107025317A · 2017 [cited by examiner]
EP 1576494B1 · 2007 [cited by examiner]
EP 1576493B1 · 2008 [cited by examiner]
FR 3050846A1 · 2017 [cited by applicant]
GB 2554491A · 2018 [cited by examiner]
JP 2009245381A · 2009 [cited by examiner]
JP 5262248B2 · 2013 [cited by examiner]
JP 2018073103A · 2018 [cited by examiner]
KR 20180036587A · 2018 [cited by examiner]
WO WO2019227518A1 · 2019 [cited by examiner]
Chuan-Lin and Tse-Yun Feng, “Chapter 2—Parallel Architectures and Interconnection Networks”, published on Jan. 23, 2014 at https://www.cs.hunter.cuny.edu/˜sweiss/course_materials/csci493.65/lecture_notes/chapter02/pdf, … [cited by examiner]
James Garland, etc., “Low Complexity Multiply-Accumulate Units for Convolutional Neural Networks with Weight-Sharing”, published on May 1, 2018 via arXiv, id 1801.10219v3, retrieved Jun. 4, 2024. (Year: 2018). [cited by examiner]
Vinyak Gokhale, etc., “Snowflake: An Efficient Hardware Accelerator for Convolutional Neural Networks”, published in 2017 IEEE International Symposium on Circuits and Systems, pp. 1-4, 2017, retrieved Jun. 4, 2024. (Yea… [cited by examiner]
Sugil Lee, etc., “Double MAC on a DSP: Boosting the Performance of Convolutional Neural Networks on FPGAs”, published in print Apr. 5, 2018, and otherwise published via IEEE Transactions on Computer-Aided Design of Inte… [cited by examiner]
Chengbo Xue, etc., “A Reconfigurable Pipelined Architecture for Convolutional Neural Network Acceleration”, published via 2018 IEEE International Symposium on Circuits and Systems (ISCAS), May 27-30, 2018, Florence, Ita… [cited by examiner]
“Shift Registers: Serial-in, Serial-out”, published to https://www.allaboutcircuits.com/textbook/digital/chpt-12/serial-in-serial-out-shift-register on Jun. 21, 2005, retrieved Mar. 31, 2025. (Year: 2005). [cited by examiner]
“Shift Registers: Parallel-in, Serial-out”, published to https://www.allaboutcircuits.com/textbook/digital/chpt-12/parallel-in-serial-out-shift-register on Jun. 21, 2005, retrieved Mar. 31, 2025. (Year: 2005). [cited by examiner]
“LPM_SHIFTREG Megafunction”, published to https://people.ece.cornell.edu/land/courses/ece5760/DE1_SOC/lpm_shiftreg.pdf on Mar. 5, 2013, retrieved Mar. 31, 2025. (Year: 2013). [cited by examiner]
ECE410 at Michigan State University, “Binary Adder”, published to https://www.egr.msu.edu/classes/ece410/mason/files/Ch12.pdf on Aug. 10, 2011, retrieved Mar. 31, 2025. (Year: 2011). [cited by examiner]
Shannon Hilbert, “Verilog Shift Register”, published to https://www.bitweenie.com/listings/verilog-shift-register on Feb. 12, 2013, retrieved Mar. 31, 2025. (Year: 2013). [cited by examiner]
“Digital Electronics—Shift Registers”, published to https://www.tutorialspoint.com/digital-electronics/digital-electronics-shift-registers.htm on Sep. 13, 2016, retrieved Mar. 31, 2025. (Year: 2016). [cited by examiner]
Stylianos I. Venieris, etc., “Latency-Driven Design for FPGA-based Convolutional Neural Networks”, published via 2017 27th International Conference on Field Programmable Logic and Applications (FPL), Sep. 4-8, 2017, Ghe… [cited by examiner]
“Multi-Cycle Pipeline Operations”, published on Jun. 30, 2009 to https://ece-research.unm.edu/jimp/611/slides/chap3_6.html, retrieved Mar. 31, 2025. (Year: 2009). [cited by examiner]
Denis A. Gudovskiy, etc., “ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks”, published on Jun. 7, 2017 to arXiv, retrieved Mar. 31, 2025. (Year: 2017). [cited by examiner]
“Using Look-Up Tables as Shift Registers (SRL16) in Spartan-3 Generation FPGAs”, published to https://my.eng.utah.edu/˜cs3710/xilinx-docs/xapp465.pdf, dated May 20, 2005, retrieved Mar. 31, 2025. (Year: 2005). [cited by examiner]
Mark A. Horowitz, etc., “SPIM: A Pipelined 64×64-bit Iterative Multiplier”, published to IEEE Journal of Solid-State Circuits, vol. 24, No. 2, Apr. 1989, retrieved Oct. 3, 2025. (Year: 1989). [cited by examiner]
Fukushima, “Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position”, Biological Cybernetics, 36(4), pp. 193-202, 1980. [cited by applicant]
Zhang, et al, “Optimizing FPGA-based accelerator design for deep convolutional neural networks”, FPGA '15: Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, pp. 161-170, Feb. 2… [cited by applicant]
Rahman et al., “Efficient FPGA acceleration of convolutional neural networks using logical 3D compute array”, 2016 Design, Automation & Test in Europe Conference & Exhibition (Date), Mar. 2016. [cited by applicant]
Meloni, et al., “NEURAghe: Exploiting CPU-FPGA Synergies for Efficient and Flexible CNN Inference Acceleration on Zynq SoCs”, ACM Transactions on Reconfigurable Technology and Systems, vol. 1, No. 1, 2017. [cited by applicant]
Li, et al., “A high performance FPGA-based accelerator for large-scale convolutional neural networks”, 2016 26th International Conference on Field Programmable Logic and Applications (FPL), 2016. [cited by applicant]