IP Library Granted Patent US 12,307,249
Granted Patent B2
US 12,307,249 · App. 18/004,802 · Granted May 20, 2025

Bit-parallel vector composability for neural acceleration

Inventors: Soroush Ghodrati (La Jolla, CA); Hadi Esmaeilzadeh (San Diego, CA)
Assignee: The Regents of the University of California
G06F9/30036G06F9/30018G06F9/30032G06F9/3822G06F9/3877
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,249
App. No.
18/004,802
Granted
May 20, 2025
Kind
B2
Abstract

Methods, apparatus and systems that relate to hardware accelerators of artificial neural network (ANN) performance that significantly reduce the energy and area costs associated with performing vector dot-product operations in the ANN training and inference tasks. Specifically, the methods, apparatus and systems reduce the cost of bit-level flexibility stemming from aggregation logic by amortizing related costs across vector elements and reducing complexity of the cooperating narrower bitwidth units.

Claims (43)

1. An apparatus for performing an energy-efficient dot-product operation between two input vectors, comprising:

a plurality of vector computation engines configurable into multiple groups of vector computation engines, wherein each vector computation engine from the plurality of vector computation engines comprises:

an array of multipliers connected through one or more add units and configured to

generate an output of the vector computation engine based on a dot-product

operation on a subset of bits of the two input vectors;

a plurality of shifters configured to shift the outputs of the vector computation engines based on a significance position, within the two input vectors, of a sub-vector associated with the outputs; and

an aggregator coupled to the plurality of shifters and configured to generate a scalar output for the energy-efficient dot-product operation based on aggregating the shifted outputs, wherein a number of the vector computation engines is based, in part, on bitwidths of datatypes of elements of the two input vectors,

wherein the multiple groups are operable according to levels including at least one level in which outputs generated by each of the multiple groups are aggregated to produce a final scalar result for the energy-efficient dot-product operation.

2. The apparatus of claim 1 , wherein the plurality of vector computation engines is configured to operate in parallel.

3. The apparatus of claim 1 , wherein one of the two vectors is a vector of weights of a neural network.

4. The apparatus of claim 1 , wherein the number of the vector computation engines is further based on a size of the subset of bits of the two input vectors.

5. The apparatus of claim 1 , wherein the number of the vector computation engines is based on a size of the subset of bits of the two input vectors and a maximum bitwidth of datatypes of elements of the two input vectors.

6. The apparatus of claim 1 , wherein each multiplier in the array of multipliers is configured to perform a 2-bit by 2-bit multiplication.

7. The apparatus of claim 6 , wherein the number of multipliers in the array of multipliers is 16.

8. The apparatus of claim 1 , wherein each multiplier in the array of multipliers is configured to perform a 4-bit by 4-bit multiplication.

9. The apparatus of claim 1 , wherein the plurality of vector computation engines is configurable into multiple groups of vector computation engines according to bitwidths of elements of the two input vectors.

10. The apparatus of claim 1 , wherein the plurality of shifters is configurable according to bitwidths of elements of the two input vectors.

11. The apparatus of claim 1 , wherein the multipliers of the array of multipliers and add units are configurable according to bitwidths of elements of the two input vectors.

12. The apparatus of claim 3 , comprising a memory configured to store the vector of weights of the neural network.

13. An apparatus for performing an energy-efficient dot-product operation between a first vector and a second vector, comprising:

a plurality of vector computation engines configurable into multiple groups of vector computation engines, the plurality of vector computation engines including,

a first vector computation engine comprising a first group of multipliers and add units, wherein the first vector computation engine, with the first group of multipliers and add units, is configured to produce a first dot-product of a third vector and a fourth vector, wherein each element of the third vector includes a subset of bits of a corresponding element of the first vector and each element of the fourth vector includes a subset of bits of a corresponding element of the second vector, and

a second vector computation engine comprising a second group of multipliers and add units, wherein the second vector computation engine, with the second group of multipliers and add units, is configured to produce a second dot-product of a fifth vector and a sixth vector, wherein each element of the fifth vector includes a subset of bits of a corresponding element of the first vector and each element of the sixth vector includes a subset of bits of a corresponding element of the second vector;

a first bit shifter configured to bit-shift the first dot-product by a first number of bits;

a second bit shifter configured to bit-shift the second dot-product by a second number of bits; and

an aggregator configured to aggregate the bit-shifted first dot-product and the bit-shifted second dot-product,

wherein the apparatus is configured to generate at least one of the first dot-product or the bit-shifted first dot-product in parallel with at least one of the second dot-product or the bit-shifted second dot-product,

wherein one of the first vector or the second vector is a vector of inputs of a layer of a neural network, wherein results of the first vector computation engine and the second vector computation engine are combined according to a bitwidth of the layer of the neural network,

wherein the multiple groups are operable according to levels including at least one level in which outputs generated by each of the multiple groups are aggregated to produce a final scalar result for the energy-efficient dot-product operation.

14. A method of performing an energy-efficient dot-product operation between two vectors, comprising:

generating partial dot-product outputs of a plurality of vector computation engines configurable into multiple groups of vector computation engines, wherein each partial dot-product output is based on performing a dot-product operation on a subset of bits of the two vectors using an array of multipliers connected through one or more add units;

shifting the partial dot-product outputs of the vector computation engines belonging to the multiple groups using bit shifters; and

aggregating the shifted partial dot-product outputs of each of the multiple groups using an aggregator to produce a scalar output for the energy-efficient dot-product operation,

wherein the generating partial dot-product outputs is performed in a parallel manner, wherein a number of the vector computation engines is based, in part, on bitwidths of datatypes of elements of the two vectors,

wherein the multiple groups are operable according to levels including at least one level in which outputs generated by each of the multiple groups are aggregated to produce a final scalar result for the energy-efficient dot-product operation.

15. The method of claim 14 , wherein one of the two vectors is a vector of weights of a neural network.

16. The method of claim 14 , wherein the number of the partial dot-product outputs is based on at least one of:

a size of the subset of bits of the two vectors and bitwidths of datatypes of components of the two vectors, or

a size of the subset of bits of the two vectors and a maximum bitwidth of datatypes of components of the two vectors.

17. The method of claim 14 , wherein each multiplier is configured to perform a 2-bit by 2-bit multiplication.

18. The method of claim 14 , wherein each multiplier is configured to perform a 4-bit by 4-bit multiplication.

19. The method of claim 14 , comprising configuring the shifter elements based on bitwidths of components of the two vectors.

20. The method of claim 14 , comprising configuring the multiplies and add units based on bitwidths of components of the two vectors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2024
From: GHODRATI, SOROUSH; ESMAEILZADEH, HADI
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 067837/0255 →
Continuity (2)
Provisional Application 63049982 · Jul 9, 2020
Related Publication 20230244484A1 · Aug 3, 2023
References Cited (35)
US 20010023425A1 · Oberman et al. · 2001 [cited by applicant]
US 20050219422A1 · Dorojevets · 2005 [cited by examiner]
US 20130124590A1 · Gunnam et al. · 2013 [cited by applicant]
US 20140207838A1 · Danne · 2014 [cited by examiner]
US 20190026249A1 · Talpes · 2019 [cited by examiner]
US 20190042252A1 · Kaul · 2019 [cited by examiner]
US 20200081711A1 · Bayat · 2020 [cited by examiner]
US 20200104131A1 · Liguori · 2020 [cited by examiner]
US 20210049463A1 · Ruff · 2021 [cited by examiner]
Ghodrati et al., “Bit-Parallel Vector Composability for Neural Acceleration,” https://arxiv.org/abs/2004.05333, submitted Apr. 11, 2020, 6 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US21/41167, mailed Oct. 20, 2021, 15 pages. [cited by applicant]
Albericio, J , “Cnvlutin: Ineffectual-Neuron-Free Deep Neural Network Computing”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, in ISCA, 2016, 6 pages. [cited by applicant]
Chen, Y , “DaDianNao: A Machine-Learning Supercomputer”, 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, in MICRO, 2014, 14 pages. [cited by applicant]
Chen, Y , “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, in ISCA 2016, 13 pages. [cited by applicant]
Chi, P , “PRIME: A Novel Processing-in-memory Architecture for Neural Network Computation in ReRAM-based Main Memory”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, in ISCA, 2016, 13 pages. [cited by applicant]
Choi, J , “PACT: Parameterized Clipping Activation for Quantized Neural Networks”, arXiv preprint arXiv:1805.06085, 2018, 15 pages. [cited by applicant]
Delmas Lascorz, A , “Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks”, Session: Machine Learning I ASPLOS'19, Apr. 13-17, 2019, Providence, RI, USA, in Proceedings of t… [cited by applicant]
Fowers, J , “A Configurable Cloud-Scale DNN Processor for Real-Time AI”, in Proceedings of the 45th Annual International Symposium on Computer Architecture, pp. 1-14, IEEE Press, 2018, 14 pages. [cited by applicant]
Gao, M , “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory”, in ASPLOS, 2017, 14 pages. [cited by applicant]
H.M. C. Consortium , “Hybridmemory cube specification 2.1”, Hybrid Memory Cube—HMC-30G-VSR PHY HMC Memory Features, Last Revision Jan. 2013., 132 pages. [cited by applicant]
Han, S , “EIE: Efficient Inference Engine on Compressed Deep Neural Network”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, in ISCA 2016, 12 pages. [cited by applicant]
Hubara, I , “Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations”, Journal of Machine Learning Research 18 (2018) 1-30 Submitted Sep. 2016; Revised Apr. 2017; Published Apr. 20… [cited by applicant]
Jouppi, N , “In-Datacenter Performance Analysis of a Tensor Processing Unit”, in ISCA, 2017, 12 pages. [cited by applicant]
Judd, P , “Stripes: Bit-Serial Deep Neural Network Computing”, Authorized licensed use limited to: Univ of Calif San Diego. Downloaded on Oct. 21, 2024 at 19:16:24 UTC from IEEE Xplore. Restrictions apply, in MICRO 2016… [cited by applicant]
Lee, J , “unpu: a 50.6TOPS/W Unified Deep Neural Network Accelerator with 1b-to-16b Fully-Variable Weight Bit-Precision”, ISSCC 2018 / SESSION 13 / Machine Learning and Signal Processing / 13.3, in ISSCC, 2018, 3 pages. [cited by applicant]
Li, S , “CACTI-P: Architecture-Level Modeling for SRAM-based Structures with Advanced Leakage Reduction Techniques”, in ICCAD, 2011, 8 pages. [cited by applicant]
Mishra, A , “WRPN: wide reduced-precision networks”, arXiv, 2017, 11 pages. [cited by applicant]
Parashar, A , “SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks”, in ISCA, 2017, 14 pages. [cited by applicant]
Reagen, B , “Minerva: Enabling Low-Power, Highly-Accurate Deep Neural Network Accelerators”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, in ISCA, 2016, 12 pages. [cited by applicant]
Shafiee, A , “ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, in ISCA, 2016, 13 pages. [cited by applicant]
Sharify, S , “Laconic Deep Learning Inference Acceleration”, in Proceedings of the 46th International Symposium on Computer Architecture, ACM, 2019, pp. 304-317. [cited by applicant]
Sharify, S , “Loom: Exploiting Weight and Activation Precisions to Accelerate Convolutional Neural Networks”, arXiv, 2017, 6 pages. [cited by applicant]
Sharma, H , “Bit Fusion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Networks”, 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture, 2018, 12 pages. [cited by applicant]
Zhang, S , “Cambricon-X: An Accelerator for Sparse Neural Networks”, Authorized licensed use limited to: Univ of Calif San Diego. Downloaded on Oct. 21, 2024 at 19:12:55 UTC from IEEE Xplore. Restrictions apply, in MICR… [cited by applicant]
Zhou, S , “DOREFA-NET: Training Low Bitwidth Convolutional Neural Networks With Low Bitwidth Gradients”, arXiv, 2016, 13 pages. [cited by applicant]