IP Library › Granted Patent US 12,271,439
Granted Patent B2
US 12,271,439 · App. 18/750,830 · Granted Apr 8, 2025

Flexible compute engine microarchitecture

Inventor: Mohammed Elneanaei Abdelmoneem Fouda (Irvine, CA)
G06F17/16G06F5/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,271,439
App. No.
18/750,830
Granted
Apr 8, 2025
Kind
B2
Abstract

A compute engine is described. The compute engine includes compute-in-memory (CIM) modules and may include an input buffer coupled with the CIM modules. The input buffer stores a vector. The CIM modules store weights corresponding to a matrix and perform a vector-matrix multiplication (VMM) for the matrix and the vector. The CIM modules further include storage cells and vector multiplication units (VMUs) coupled with the storage cell and, if present, the input buffer. The storage cells store the weights. The VMUs multiply, with the vector, at least a portion of each weight of a portion of the plurality of weights corresponding to a portion of the matrix. A set of VMUs performs multiplications for a first weight length and a second weight length different from the first weight length such that each VMU of the set performs multiplications for both the first weight length and the second weight length.

Claims (37)

1. A compute engine, comprising:

an input buffer configured to store a vector, the vector including a vector sign; and

a plurality of compute-in-memory (CIM) hardware modules coupled with the input buffer, configured to store a plurality of weights corresponding to a matrix, and configured to perform a vector-matrix multiplication (VMM) for the matrix and the vector, each of the plurality of weights including a weight sign, the plurality of CIM hardware modules further including:

a plurality of storage cells for storing the plurality of weights; and

a plurality of vector multiplication units (VMUs) coupled with the plurality of storage cells and the input buffer, each VMU of the plurality of VMUs configured to multiply, with the vector, at least a portion of a weight of a portion of the plurality of weights;

wherein a set of VMUs of the plurality of VMUs are configured to perform multiplications for a first weight length and a second weight length different from the first weight length such that each VMU of the set of VMUs performs the multiplications for both the first weight length and the second weight length, an indication of the vector sign being propagated to each VMU of the set of VMUs, the weight sign being propagated to a portion of the VMUs in the set of VMUs for the second weight length; and

wherein each VMU in the portion of the VMUs is configured to extend the weight sign for a remaining portion of the VMU for the second weight length.

2. The compute engine of claim 1 , wherein the set of VMUs includes a first VMU and a second VMU, the first VMU and the second VMU each being configured to perform the multiplications for the first weight length, the second weight length being twice the first weight length.

3. The compute engine of claim 1 , further comprising:

a combiner coupled with the set of VMUs, the combiner configured to combine a product of the at least the portion of each weight and an element of the vector for the second weight length.

4. The compute engine of claim 3 , wherein the combiner shifts the product for a first VMU of the set of VMUs by the first weight length relative to the product for a second VMU of the set of VMUs.

5. The compute engine of claim 1 , wherein each element of the vector in the input buffer is serialized for the multiplications.

6. The compute engine of claim 1 , wherein the first weight length corresponds to more significant bits of a weight and the second weight length corresponds to a combination of the more significant bits of the weight and less significant bits of the weight, the set of VMUs including a first VMU for the more significant bits of the weight and at least a second VMU for the less significant bits of the weight.

7. The compute engine of claim 1 , wherein the compute engine has a first output precision for the first weight length and a second output precision for the second weight length, the second output precision being a multiple of the first output precision.

8. The compute engine of claim 1 , wherein each VMU in the set of VMUs determines the multiplications using a lookup table.

9. The compute engine of claim 1 , wherein the input buffer is configured to store non-vector data; and wherein the plurality of CIM hardware modules are configured to multiply the non-vector data with the matrix.

10. The compute engine of claim 1 , wherein the plurality of VMUs are configured to perform multiplications for the first weight length only or the second weight length only for either a plurality of matrices or a plurality of vectors.

11. An accelerator, comprising:

a processor, and

a plurality of compute engines coupled with the processor, each of the plurality of compute engines including an input buffer configured to store a vector and a plurality of compute-in-memory (CIM) hardware modules, coupled with the input buffer, configured to store a plurality of weights corresponding to a matrix, and configured to perform a vector-matrix multiplication (VMM) for the matrix and the vector, the vector including a vector sign, each of the plurality of weights including a weight sign, the plurality of CIM hardware modules further including:

a plurality of storage cells for storing the plurality of weights;

a plurality of vector multiplication units (VMUs) coupled with the plurality of storage cells and the input buffer, each VMU of the plurality of VMUs configured to multiply, with the vector, at least a portion of a weight of a portion of the plurality of weights corresponding to a portion of the matrix;

wherein a set of VMUs of the plurality of VMUs are configured to perform multiplications for a first weight length and a second weight length different from the first weight length such that each VMU of the set of VMUs performs the multiplications for both the first weight length and the second weight length, an indication of the vector sign being propagated to each VMU of the set of VMUs, the weight sign being propagated to a portion of the VMUs in the set of VMUs for the second weight length; and

wherein each VMU in the portion of the VMUs is configured to extend the weight sign for a remaining portion of the VMU for the second weight length.

12. A method, comprising:

providing a vector to a compute engine including a plurality of compute-in-memory (CIM) hardware modules configured to store a plurality of weights corresponding to a matrix and configured to perform a vector-matrix multiplication (VMM) for the matrix, the plurality of CIM hardware modules further including a plurality of storage cells and a plurality of vector matrix multiplication units (VMUs), the plurality of storage cells storing the plurality of weights, the plurality of VMUs coupled with the plurality of storage cells, each VMU of the plurality of VMUs configured to multiply, with the vector, at least a portion of each weight of a portion of the plurality of weights corresponding to a portion of the matrix, the vector including a vector sign and each of the plurality of weights including a weight sign; and

performing a VMM of the vector and the matrix using the plurality of CIM hardware modules, a set of VMUs of the plurality of VMUs configured to perform multiplications for a first weight length and a second weight length different from the first weight length such that each VMU of the set of VMUs is capable of performing the multiplications for both the first weight length and the second weight length, the performing the VMM further including:

propagating an indication of the vector sign to each VMU of the set of VMUs;

propagating the weight sign to a portion of the VMUs in the set of VMUs for the second weight length; and

in each VMU of the portion of the VMUs, extending the weight sign for a remaining portion of the VMU for the second weight length.

13. The method of claim 12 , wherein the set of VMUs includes a first VMU and a second VMU, the first VMU and the second VMU each being configured to perform the multiplications for the first weight length, the second weight length being twice the first weight length.

14. The method of claim 12 , further comprising:

combining a product of the portion of each weight and an element of the vector for the second weight length.

15. The method of claim 14 , wherein the combining further includes:

shifting the product for a first VMU of the set of VMUs by the first weight length relative to the product for a second VMU of the set of VMUs.

16. The method of claim 12 , wherein the providing the vector further includes:

serializing each element of the vector.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2025
From: RAIN NEUROMORPHICS INC.
To: OPENAI OPCO, LLC
Reel/Frame 073238/0425 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2024
From: FOUDA, MOHAMMED ELNEANAEI ABDELMONEEM
To: RAIN NEUROMORPHICS INC.
Reel/Frame 068531/0573 →
Continuity (2)
Provisional Application 63523003 · Jun 23, 2023
Related Publication 20240427844A1 · Dec 26, 2024
References Cited (50)
US 11347477B2 · Sumbul · 2022 [cited by applicant]
US 20090177867A1 · Garde · 2009 [cited by applicant]
US 20100174883A1 · Lerner · 2010 [cited by applicant]
US 20150106311A1 · Birdwell · 2015 [cited by applicant]
US 20180189631A1 · Sumbul · 2018 [cited by applicant]
US 20190179795A1 · Huang · 2019 [cited by applicant]
US 20190205741A1 · Gupta · 2019 [cited by applicant]
US 20190340486A1 · Mills · 2019 [cited by applicant]
US 20190348110A1 · Sinangil · 2019 [cited by applicant]
US 20190362227A1 · Seshadri · 2019 [cited by applicant]
US 20200057938A1 · Lu · 2020 [cited by examiner]
US 20200207656A1 · Yoshioka · 2020 [cited by applicant]
US 20200293284A1 · Vantrease · 2020 [cited by applicant]
US 20200320403A1 · Daga · 2020 [cited by applicant]
US 20200410337A1 · Huang · 2020 [cited by applicant]
US 20210097431A1 · Olgiati · 2021 [cited by applicant]
US 20210158132A1 · Huynh · 2021 [cited by applicant]
US 20210224185A1 · Zhou · 2021 [cited by applicant]
US 20220004497A1 · Willcock · 2022 [cited by applicant]
US 20220019880A1 · Dasgupta · 2022 [cited by applicant]
US 20220114270A1 · Wang · 2022 [cited by applicant]
US 20220138286A1 · Zage · 2022 [cited by applicant]
US 20220164916A1 · Nurvitadhi · 2022 [cited by applicant]
US 20220207293A1 · Yao · 2022 [cited by applicant]
US 20220207656A1 · Yao · 2022 [cited by applicant]
US 20220244916A1 · Lee · 2022 [cited by examiner]
US 20220414432A1 · Banitalebi Dehkordi · 2022 [cited by applicant]
US 20230014565A1 · Ray · 2023 [cited by applicant]
US 20230047364A1 · Badaroglu · 2023 [cited by examiner]
US 20230074229A1 · Jia · 2023 [cited by applicant]
US 20230138695A1 · Kumar · 2023 [cited by applicant]
US 20230146647A1 · Byeon · 2023 [cited by applicant]
US 20230206044A1 · Ma · 2023 [cited by examiner]
US 20230259456A1 · Verma · 2023 [cited by applicant]
US 20230297580A1 · Sheng · 2023 [cited by applicant]
US 20230316060A1 · Jain · 2023 [cited by applicant]
US 20230359894A1 · Kim · 2023 [cited by applicant]
US 20240094986A1 · Lyubomirsky · 2024 [cited by examiner]
US 20240134606A1 · Yi · 2024 [cited by examiner]
US 20240169201A1 · Seok · 2024 [cited by examiner]
WO 2020190776 · 2020 [cited by applicant]
WO 2022029026 · 2022 [cited by applicant]
H. Mori et al., “A 4nm 6163-TOPS/W/b 4790-TOPS/mm2/b SRAM Based Digital-Computing-in-Memory Macro Supporting Bit-Width Flexibility and Simultaneous MAC and Weight Update,” 2023 ISSCC, San Francisco, CA, USA, Feb. 2023, … [cited by examiner]
K. Li et al., “A Precision-Scalable Energy-Efficient Bit-Split-and-Combination Vector Systolic Accelerator for NAS-Optimized DNNs on Edge,” 2022 Design, Automation & Test in Europe Conference & Exhibition (DATE), Antwer… [cited by examiner]
Y.-D. Chih et al., “16.4 An 89TOPS/W and 16.3TOPS/mm2 All-Digital SRAM-Based Full-Precision Compute-In Memory Macro in 22nm for Machine-Learning Edge Applications,” 2021 ISSCC, San Francisco, CA, USA, 2021, pp. 252-254,… [cited by examiner]
Kim et al., Moneta: A Processing-In-Memory-Based Hardware Platform for the Hybrid Convolutional Spiking Neural Network with Online Learning, Frontiers in Neuroscience, vol. 16, Apr. 11, 2022. [cited by applicant]
Korthikanti et al., Reducing Activation Recomputation in Large Transformer Models, May 10, 2022, pp. 1-17. [cited by applicant]
Lee et al., A 12nm 121-TOPS/W 41.6-TOPS/mm2 All Digital Full Precision SRAM-based Compute-in-Memory with Configurable Bit-width for AI Edge Applications, 2022 Symposium on VLSI Technology & Circuits Digest of Technical … [cited by applicant]
Song et al., PipeLayer: A Pipelined ReRAM-Based Accelerator for Deep Learning, 2017. [cited by applicant]
Lin et al., A Novel Voltage-Accumulation Vector-Matrix Multiplication Architecture using Resistor-Shunted Floating Gate Flash Memory Device for Low-Power and High-Density Neural Network Applications, 2018 IEEE Internati… [cited by applicant]