IP Library › Granted Patent US 12,361,268
Granted Patent B2
US 12,361,268 · App. 17/461,626 · Granted Jul 15, 2025

Neural network hardware accelerator circuit with requantization circuits

Inventors: Giuseppe Desoli (San Fermo Della Battaglia, IT); Surinder Pal Singh (Noida, IN); Thomas Boesch (Rovio, CH)
Assignee: STMicroelectronics International N.V.
G06N3/063G06F9/5027G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,268
App. No.
17/461,626
Filed
Aug 30, 2021
Granted
Jul 15, 2025
Kind
B2
Art Unit
2129
USPC
706/15
Abstract

A convolutional neural network includes convolution circuitry. The convolution circuitry performs convolution operations on input tensor values. The convolutional neural network includes requantization circuitry that requantizes convolution values output from the convolution circuitry.

Claims (49)

1. A convolutional neural network (CNN), comprising: convolution circuitry, which, in

operation, receives an input tensor including a plurality of quantized input data values in an input quantization format and generates a plurality of intermediate data values by performing a convolution operation on the plurality of quantized input values of the input tensor;

first requantization circuitry coupled to an output of the convolution circuitry; and

second requantization circuitry different from the first requantization circuitry and coupled to the output of the convolution circuitry, wherein,

the first requantization circuitry, in operation, generates a first output tensor having a plurality of first quantized output values in a first output quantization format by performing a first requantization process on the generated plurality of intermediate data values;

and

the second requantization circuitry, in operation, generates a second output tensor having a plurality of second quantized output values in a second output quantization format by performing a second requantization process on the generated plurality of intermediate data values.

2. The CNN of claim 1 , wherein the first output quantization format is a scale/offset quantization format and the second output quantization format is a fixed point quantization format.

3. The CNN of claim 1 , wherein the input quantization format is a scale/offset format and the first output quantization format is the scale/offset quantization format.

4. The CNN of claim 1 , wherein the input quantization format is a scale/offset format and the first output quantization format is a fixed point quantization format.

5. The CNN of claim 1 , wherein the input quantization format is a fixed point quantization format and the first output quantization format is a scale/offset quantization format.

6. The CNN of claim 1 , wherein the input quantization format is a fixed point quantization format and the first output quantization format is the fixed point quantization format.

7. The CNN of claim 1 , comprising:

a stream link configured to receive the quantized input values; and

a subtractor positioned between the stream link and the convolution circuitry and configured to perform a subtraction operation on the quantized input values prior to the convolution operation.

8. The CNN of claim 7 , comprising a shifter coupled between the subtractor and the convolution circuitry and configured to adjust a number of bits of the quantized input values prior to the convolution operation.

9. The CNN of claim 1 , comprising pooling circuitry configured to generate a plurality of pooling values by performing a pooling operation on a plurality of second quantized input values; and

third requantization circuitry coupled to the pooling circuitry and configured to generate a plurality of third quantized output values by performing a third quantization process on the pooling values.

10. The CNN of claim 1 , further comprising activation circuitry configured to generate a plurality of activation values by performing an activation operation on a plurality of second quantized input values; and

third requantization circuitry coupled to the activation circuitry and configured to generate a plurality of third quantized output values by performing a third quantization process on the activation values.

11. A method, comprising:

receiving, at a first layer of a neural network, an input tensor including a plurality of quantized input data values;

generating intermediate data values from the input tensor values by performing a first operation on the quantized data values;

generating, at the first layer, a first output tensor including a plurality of first quantized output data values, the generating including by performing a first requantization process on the intermediate data values using first requantization circuitry; and

generating, at the first layer, a second output tensor including a plurality of second quantized output data values by performing a second requantization process on the generated intermediate data values using second requantization circuitry different from the first requantization circuitry.

12. The method of claim 11 , wherein the first operation is a convolution operation.

13. The method of claim 11 , wherein the first operation is a pooling operation.

14. The method of claim 11 , wherein the first operation is an activation operation.

15. The method of claim 11 , wherein the first quantized output data values are in a first quantization format and the second quantized output data values are in a second quantization format.

16. The method of claim 15 , wherein the first quantization format is a scale/offset format.

17. The method of claim 16 , wherein the second quantization format is a fixed point format.

18. An electronic device, comprising a neural network, the neural network including:

a stream link configured to provide an input tensor including a plurality of quantized input data values;

a hardware accelerator configured to receive the input tensor and to generate intermediate data values from the input tensor by performing an operation on the quantized input data values;

first requantization circuitry coupled to an output of the hardware accelerator and configured to generate a plurality of first quantized output data values by performing a first requantization operation on the generated intermediate data values; and

second requantization circuitry coupled to the output of the hardware accelerator and configured to generate a plurality of second quantized output data values by performing a second requantization operation on the generated intermediate data values.

19. The electronic device of claim 18 , wherein the hardware accelerator is a convolution accelerator.

20. The electronic device of claim 18 , wherein the first requantization circuitry includes arithmetic circuitry.

21. The electronic device of claim 18 , wherein the hardware accelerator is a pooling accelerator.

22. The electronic device of claim 18 , wherein the hardware accelerator is an activation accelerator.

23. A non-transitory computer-readable medium having contents which configure a hardware accelerator of convolutional neural network to perform a method, the method comprising:

receiving an input tensor including a plurality of quantized input data values;

generating intermediate data values from the input tensor values by performing a first operation on the quantized data values;

generating a first output tensor including a plurality of first quantized output data values, the generating including performing a first requantization process on the generated intermediate data values using first requantization circuitry; and

generating a second output tensor including a plurality of second quantized output data values by performing a second requantization process on the generated intermediate data values using second requantization circuitry different from the first requantization circuitry.

24. The non-transitory computer-readable medium of claim 23 , wherein the hardware accelerator is a convolution accelerator.

25. The method of claim 24 , wherein the first operation is a convolution operation.

26. The method of claim 23 , wherein the first operation is a pooling operation.

27. The method of claim 23 , wherein the first operation is an activation operation.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2024
From: BOESCH, THOMAS
To: STMICROELECTRONICS INTERNATIONAL N.V.
Reel/Frame 067061/0062 →
QUITCLAIM ASSIGNMENT Recorded Feb 15, 2024
From: STMICROELECTRONICS S.R.L.
To: STMICROELECTRONICS INTERNATIONAL N.V.
Reel/Frame 066601/0330 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2022
From: SINGH, SURINDER PAL
To: STMICROELECTRONICS INTERNATIONAL N.V.
Reel/Frame 059903/0985 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2022
From: DESOLI, GIUSEPPE; BOESCH, THOMAS
To: STMICROELECTRONICS S.R.L.
Reel/Frame 059904/0090 →
Continuity (1)
Related Publication 20230062910A1 · Mar 2, 2023
References Cited (136)
US 5768613A · Asghar · 1998 [cited by applicant]
US 9436637B2 · Kommanaboyina · 2016 [cited by applicant]
US 9779786B1 · Wu et al. · 2017 [cited by applicant]
US 9978014B2 · Lupon et al. · 2018 [cited by applicant]
US 10366050B2 · Henry et al. · 2019 [cited by applicant]
US 10372456B2 · Fowers et al. · 2019 [cited by applicant]
US 10402527B2 · Boesch et al. · 2019 [cited by applicant]
US 10657668B2 · Hassan et al. · 2020 [cited by applicant]
US 10949736B2 · Deisher et al. · 2021 [cited by applicant]
US 11227086B2 · Boesch et al. · 2022 [cited by applicant]
US 11270201B2 · Sridharan et al. · 2022 [cited by applicant]
US 11586907B2 · Singh et al. · 2023 [cited by applicant]
US 11610362B2 · Singh et al. · 2023 [cited by applicant]
US 20120303932A1 · Farabet et al. · 2012 [cited by applicant]
US 20140032465A1 · Modha · 2014 [cited by applicant]
US 20140281005A1 · Bhamidipati et al. · 2014 [cited by applicant]
US 20150170021A1 · Lupon et al. · 2015 [cited by applicant]
US 20150212955A1 · Easwaran · 2015 [cited by applicant]
US 20150278596A1 · Kilty et al. · 2015 [cited by applicant]
US 20160379109A1 · Chung et al. · 2016 [cited by applicant]
US 20170169315A1 · Vaca Castano et al. · 2017 [cited by applicant]
US 20180046458A1 · Kuramoto · 2018 [cited by applicant]
US 20180046485A1 · Maity et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180144214A1 · Hsieh et al. · 2018 [cited by applicant]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189641A1 · Boesch et al. · 2018 [cited by applicant]
US 20180189642A1 · Boesch et al. · 2018 [cited by applicant]
US 20180315155A1 · Park et al. · 2018 [cited by applicant]
US 20190042868A1 · Oesterreicher et al. · 2019 [cited by applicant]
US 20190155575A1 · Langhammer et al. · 2019 [cited by applicant]
US 20190205746A1 · Nurvitadhi et al. · 2019 [cited by applicant]
US 20190205758A1 · Zhu et al. · 2019 [cited by applicant]
US 20190251429A1 · Du et al. · 2019 [cited by applicant]
US 20190266479A1 · Singh et al. · 2019 [cited by applicant]
US 20190266485A1 · Singh et al. · 2019 [cited by applicant]
US 20190266784A1 · Singh et al. · 2019 [cited by applicant]
US 20190354846A1 · Mellempudi et al. · 2019 [cited by applicant]
US 20190370631A1 · Fais et al. · 2019 [cited by applicant]
US 20200133989A1 · Song et al. · 2020 [cited by applicant]
US 20200234137A1 · Chen et al. · 2020 [cited by applicant]
US 20200272779A1 · Boesch et al. · 2020 [cited by applicant]
US 20210073450A1 · Boesch et al. · 2021 [cited by applicant]
US 20210073569A1 · Gao et al. · 2021 [cited by applicant]
US 20210192833A1 · Singh et al. · 2021 [cited by applicant]
US 20210240440A1 · Langhammer et al. · 2021 [cited by applicant]
US 20210256346A1 · Desoli et al. · 2021 [cited by applicant]
US 20210264250A1 · Singh et al. · 2021 [cited by applicant]
US 20220121928A1 · Dong et al. · 2022 [cited by applicant]
US 20220188072A1 · Langhammer et al. · 2022 [cited by applicant]
US 20230153621A1 · Singh et al. · 2023 [cited by applicant]
US 20230186067A1 · Singh et al. · 2023 [cited by applicant]
CN 101739241A · 2010 [cited by applicant]
CN 101819679A · 2010 [cited by applicant]
CN 104484703A · 2015 [cited by applicant]
CN 105488565A · 2016 [cited by applicant]
CN 105518784A · 2016 [cited by applicant]
CN 106650655A · 2017 [cited by applicant]
CN 106779059A · 2017 [cited by applicant]
CN 107016521A · 2017 [cited by applicant]
CN 110214309A · 2019 [cited by applicant]
CN 209560950U · 2019 [cited by applicant]
CN 210428520U · 2020 [cited by applicant]
DE 10159331A1 · 2002 [cited by applicant]
EP 3346423A1 · 2018 [cited by applicant]
EP 3346424A1 · 2018 [cited by applicant]
EP 3346427A1 · 2018 [cited by applicant]
EP 3480740A1 · 2019 [cited by applicant]
JP 2002183111A · 2002 [cited by applicant]
KR 101947782B1 · 2019 [cited by applicant]
WO 2017017371A1 · 2017 [cited by applicant]
WO WO2019227322A1 · 2019 [cited by applicant]
WO WO2020249085A1 · 2020 [cited by applicant]
Wang, Jichen, Jun Lin, and Zhongfeng Wang. “Efficient hardware architectures for deep convolutional neural network.” IEEE Transactions on Circuits and Systems I: Regular Papers 65.6 (2017): 1941-1953. (Year: 2017). [cited by examiner]
Faraone, Julian, et al. “AddNet: Deep neural networks using FPGA-optimized multipliers.” IEEE Transactions on Very Large Scale Integration (VLSI) Systems 28.1 (2019): 115-128. (Year: 2019). [cited by examiner]
Wang et al., “3D Facial Reconstruction Based on 2.5D Carve System,” p. 165-167, 193, 2005. (with English Abstract). [cited by applicant]
Tian et al., “Automated localization of body part in CT images,” [cited by applicant]
Milletari et al., “Hough-CNN: Deep Learning for Segmentation of Deep Brain Regions in MRI and Ultrasound,” arXiv:1601.07014v3 [cs.CV], Jan. 31, 2016. (34 pages). [cited by applicant]
Yang et al., “Research on Deep Learning Acceleration Technique,” URL=http://www.c-s-a.org.cn, Special Issue, Sep. 25, 2016. (9 pages) (with English Abstract). [cited by applicant]
Hara et al., “Analysis of Function of Rectified Linear Unit Used in Deep learning,” European Union Conference Paper, Jul. 2015. (9 pages). [cited by applicant]
Lozito et al. “Microcontroller Based Maximum Power Point Tracking Through FCC and MLP Neural Networks,” [cited by applicant]
Lozito et al., “FPGA Implementations of Feed Forward Neural Network by Using Floating Point Hardware Accelerators,” [cited by applicant]
Sodre, “Fast-Track to Second Order Polynomials,” PowerPoint, UT-Austin, Aug. 2011. (12 pages). [cited by applicant]
Venieris et al., “fpgaConvNet: A Framework for Mapping Convolutional Neural Networks on FPGAs,” 2016 [cited by applicant]
Bhatele et al (Ed)., [cited by applicant]
Blanc-Talon et al (Ed)., [cited by applicant]
Brownlee, “A Gentle Introduction to Pooling Layers for Convolutional Neural Networks,” published online Apr. 22, 2019, downloaded on Dec. 11, 2019, from https://machinelearningmastery.com/pooling-layers-for convolutiona… [cited by applicant]
Chen et al., “DaDianNao: A Machine-Learning Supercomputer,” 47th Annual IEEE/ACM International Symposium on Microarchitecture, Cambridge, United Kingdom, Dec. 13-17, 2014, pp. 609-622. [cited by applicant]
Chen et al., “A High-Throughput Neural Network Accelerator,” IEEE Micro, 35:24-32, 2015. [cited by applicant]
Chen et al., “14.5: Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” IEEE International Solid-State Circuits Conference, San Francisco, California, Jan. 31-Feb. 4, 2016, p… [cited by applicant]
Choudhary et al., “NETRA: A Hierarchical and Partitionable Architecture for Computer Vision Systems,” [cited by applicant]
Cook, “Global Average Pooling Layers for Object Localization,” published online Apr. 9, 2019, downloaded on Dec. 11, 2019, from https://alexisbcook.github.io/2017/global-average-pooling-layers-for-object-localization/, … [cited by applicant]
Dai et al., “Deformable Convolutional Networks,” PowerPoint Presentation, International Conference on Computer Vision, Venice, Italy, Oct. 22-Oct. 29, 2017, 17 pages. [cited by applicant]
Dai et al., “Deformable Convolutional Networks,” Proceedings of the IEEE International Conference on Computer Vision :264-773, 2017. [cited by applicant]
Desoli et al., “14.1: A 2.9TOPS/W Deep Convolutional Neural Network SoC in FD-SOI 28nm for Intelligent Embedded Systems,” [cited by applicant]
Du et al., “ShiDianNao: Shifting Vision Processing Closer to the Sensor,” 2015 [cited by applicant]
Erdem et al., “Design Space Exploration for Orlando Ultra Low-Power Convolutional Neural Network SoC,” IEEE 29th International Conference on Application-specific Systems, Architectures and Processors, Milan, Italy, Jul.… [cited by applicant]
Github, “Building a quantization paradigm from first principles,” URL=https://github.com/google/gemmlowp/blob/master/doc/quantization.md, download date Jul. 29, 2021, 7 pages. [cited by applicant]
Github, “The low-precision paradigm in gemmlowp, and how it's implemented,” URL=https://github.com/google/gemmlowp/blob/master/doc/low-precision.md#efficient-handling-of-offsets, download date Jul. 29, 2021, 4 pages. [cited by applicant]
Gokhale et al., “A 240 G-ops/s Mobile Coprocessor for Deep Neural Networks (Invited Paper),” [cited by applicant]
Graf et al., “A Massively Parallel Digital Learning Processor,” [cited by applicant]
Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” International Conference on Learning Representations, San Juan, Puerto Rico, May 2-4, 2016, 14 page… [cited by applicant]
Hou et al., “An End-to-end 3D Convolutional Neural Network for Action Detection and Segmentation in Videos,” [cited by applicant]
Hou et al., “Tube Convolutional Neural Network (T-CNN) for Action Detection in Videos,” International Conference on Computer Vision, Venice Italy, Oct. 22-29, 2017, pp. 5822-5831. [cited by applicant]
Hu et al., “MaskRNN: Instance Level Video Object Segmentation,” 31st Conference on Neural Information Processing Systems, Long Beach, California, Dec. 4-9, 2017, 10 pages. [cited by applicant]
Jagannathan et al., “Optimizing Convolutional Neural Network on Dsp,” IEEE International Conference on Consumer Electronics, Jan. 7-11, 2016, Las Vegas, Nevada, pp. 371-372. [cited by applicant]
Jouppi et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” 44th International Symposium on Computer Architecture, Toronto, Canada, Jun. 26, 2017, 17 pages. [cited by applicant]
Kang et al., “T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos,” arXiv:1604.02532: Aug. 2017, 12 pages. [cited by applicant]
Kiningham, K. et al., “Design and Analysis of a Hardware CNN Accelerator,” Stanford University, 2017, 8 pages. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” Proceedings of the 25th International Conference on Neural Information Processing Systems 1:1097-1105, 2012 (9 pages). [cited by applicant]
Lascorz et al., “Tartan: Accelerating Fully-Connected and Convolutional Layers in Deep Learning Networks by Exploiting Numerical Precision Variability,” arXiv:1707.09068v1: Jul. 2017, 12 pages. [cited by applicant]
LeCun et al., “Gradient-Based Learning Applied to Document Recognition,” Proceedings of the IEEE 86(1):2278 2324, 1998. [cited by applicant]
Lin et al., “A Digital Circuit of Hyperbolic Tangent Sigmoid Function for Neural Networks,” [cited by applicant]
Lin et al., “Network In Network,” arXiv:1312.4400v3 [cs.NE], Mar. 4, 2014, 10 pages. [cited by applicant]
Meloni et al., “A High-Efficiency Runtime Reconfigurable IP for CNN Acceleration on a Mid-Range All-Programmable SoC,” International Conference on ReConFigurable Computing and FPGAs (ReConFig), Nov. 30-Dec. 2, 2016, Can… [cited by applicant]
Merritt, “AI Silicon Gets Mixed Report Card,” EE Times, published online Jan. 4, 2018, downloaded on Jan. 15, 2018, from https://www.eetimes.com/document.asp?doc_id=1332799&print-yes, 3 pages. [cited by applicant]
Moctar et al., “Routing Algorithms for FPGAS with Sparse Intra-Cluster Routing Crossbars,” 22nd International Conference on Field Programmable Logic and Applications (FPL), Aug. 29-31, 2012, Oslo, Norway, pp. 91-98. [cited by applicant]
NVIDIA Deep Learning Accelerator, “NVDLA,” downloaded on Dec. 12, from http://nvdla.org/, 2019, 5 pages. [cited by applicant]
Redmon, “YOLO: Real-Time Object Detection,” archived on Jan. 9, 2018, downloaded on Jul. 23, 2019, https://web.archive.org/web/20180109074144/https://pjreddie.com/darknet/yolo/, 11 pages. [cited by applicant]
Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” arXiv:1506.01497v3, Jan. 2016, 14 pages. [cited by applicant]
Salakhutdinov et al., “A Better Way to Pretrain Deep Boltzmann Machines,” Advances in Neural Processing Systems 25, Lake Tahoe, Nevada, Dec. 3-8, 2012, 9 pages. [cited by applicant]
Scardapane et al., “Kafnets: kernel-based non-parametric activation functions for neural networks,” arXiv:1707.04035v2, Nov. 2017, 35 pages. [cited by applicant]
Sim et al., “14.6: A 1.42TOPS/W Deep Convolutional Neural Network Recognition Processor for Intelligent IoE Systems,” International Solid-State Circuits Conference, San Francisco, Californai, Jan. 31-Feb. 4, 2016, pp. 2… [cited by applicant]
Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition,” International Conference on Learning Representations, San Diego, California, May 7-9, 2015, 14 pages. [cited by applicant]
Stenström, “Reducing Contention in Shared-Memory Multiprocessors,” [cited by applicant]
Stoutchinin et al., “Optimally Scheduling CNN Convolutions for Efficient Memory Access,” [cited by applicant]
TensorFlow “How to Quantize Neural Networks with TensorFlow,” archived on Sep. 25, 2017, downloaded on Jul. 23, 2019 from https://web.archive.org/web/20170925162122/https://www.tensorflow.org/performance/quantization, 1… [cited by applicant]
Tsang, “Review: DeconvNet—Unpooling Layer (Semantic Segmentation),” published online Oct. 8, 2018, downloaded on Dec. 12, 2019, from https://towardsdatascience.com/review-deconvnet-unpooling-layer semantic-segmentation-… [cited by applicant]
UFLDL Tutorial, “Pooling,” downloaded from http://deeplearning.stanford.edu/tutorial/supervised/Pooling/ on Dec. 12, 2019, 2 pages. [cited by applicant]
Vassiliadis et al., “Elementary Function Generators for Neural-Network Emulators,” [cited by applicant]
Vu et al., “Tube-CNN: Modeling temporal evolution of appearance for objest detection in video,” arXiv:1812.02619v1, Dec. 2018, 14 pages. [cited by applicant]
Wang et al.(Ed)., [cited by applicant]
Wang et al., “DLAU: A Scalable Deep Learning Accelerator Unit on FPGA,” [cited by applicant]
Wikipedia, “Convolutional neural network,” downloaded from https://en.wikipedia.org/wiki/Computer_vision on Dec. 12, 2019, 29 pages. [cited by applicant]
Xu et al., “R-C3D: Region Convolutional 3D Network for Temporal Activity Detection,” arXiv:1703.07814v2, Aug. 2017, 10 pages. [cited by applicant]
Zhong, K. et al., “Exploring the Potential of Low-bit Training of Convolutional Neural Networks,” [cited by applicant]