IP Library › Granted Patent US 12,536,118
Granted Patent B2
US 12,536,118 · App. 18/789,480 · Granted Jan 27, 2026

Tiled in-memory computing architecture

Inventors: Nawab Ali (Bellingham, WA); Muzaffer Kal (Redmond, WA); Alexander Almela Conklin (San Jose, CA); Burak Erbagci (San Jose, CA); Cagri Eryilmaz (San Francisco, CA); Mohammed Elneanaei Abdelmoneem Fouda (Irvine, CA)
Assignee: RAIN NEUROMORPHICS INC.
G06F13/28G06F2213/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,118
App. No.
18/789,480
Granted
Jan 27, 2026
Kind
B2
Abstract

A compute tile is described. The compute tile includes compute engines and a general-purpose (GP processor coupled with the compute engines. Each of the compute engines includes a compute-in-memory (CIM) hardware module. The CIM hardware module is configured to store weights corresponding to a matrix and to perform a vector-matrix multiplication (VMM) for the matrix. The GP processor is configured to control the compute engines, to receive output of the VMM for the matrix from the compute engines, and to perform a nonlinear operation on the output. The compute engines are addressable by data movement initiators. Data may be moved to and/or from the compute engines in data paths that bypass the GP processor.

Claims (48)

1 . A compute tile, comprising:

a plurality of compute engines, each of the plurality of compute engines including a compute-in-memory (CIM) hardware module, the CIM hardware module being configured to store a plurality of weights corresponding to a matrix and to perform a vector-matrix multiplication (VMM) for the matrix;

a general-purpose (GP) processor coupled with the plurality of compute engines and configured to control the plurality of compute engines, to receive output of the VMM for the matrix from each of the plurality of compute engines, and to perform a nonlinear operation on the output, the GP processor including a fixed function computing block usable in performing the nonlinear operation, the nonlinear operation corresponding to an activation function for the output of the VMM; and

a data conversion engine coupled with the plurality of compute engines and configured to convert data transferred to the plurality of compute engines from a first format to a second format, wherein the data was received from another compute tile in the first format;

wherein the plurality of compute engines is addressable by a plurality of data movement initiators and by the GP processor.

2 . The compute tile of claim 1 , wherein the plurality of compute engines and the plurality of data movement initiators are configured to at least one of move data to the plurality of compute engines in a data path bypassing the GP processor or move data from the plurality of compute engines in a data path bypassing the GP processor of the compute tile.

3 . The compute tile of claim 1 , further comprising:

a direct memory access (DMA) unit, the plurality of data movement initiators including at least one of the DMA unit or a component on an additional compute tile coupled with the compute tile.

4 . The compute tile of claim 3 , further comprising:

a local memory coupled with the plurality of compute engines and the GP processor, the DMA unit being configured to transfer data between the local memory and the plurality of compute engines in a data path bypassing the GP processor.

5 . The compute tile of claim 1 , wherein the data conversion engine includes at least one of a BFloat-Integer format converter or a reshape engine, the reshape engine configured to pad the data.

6 . The compute tile of claim 1 , further comprising:

a local memory coupled with the plurality of compute engines through a first bus, the GP processor being coupled with the plurality of compute engines through a second bus different from the first bus; and

a main bus coupled with the local memory and the GP processor.

7 . The compute tile of claim 1 , further comprising:

a local memory coupled with the plurality of compute engines by a first bus; and

a main bus coupled with the local memory, the plurality of compute engines, and the GP processor, the GP processor being coupled with the plurality of compute engines and the local memory through the main bus.

8 . The compute tile of claim 7 , wherein the plurality of compute engines is coupled to the main bus through an interconnect having an internal queueing.

9 . The compute tile of claim 1 , wherein the compute tile further includes:

a local compute engine memory having a first memory density; and

a cache controller coupled with the local compute engine memory;

wherein the CIM hardware module is coupled with the cache controller and has a second memory density less than the first memory density.

10 . A system, comprising:

a plurality of compute tiles, each of the plurality of compute tiles including a plurality of compute engines and a general-purpose (GP) processor, each of the plurality of compute engines including a compute-in-memory (CIM) hardware module, the CIM hardware module being configured to store a plurality of weights corresponding to a matrix and to perform a vector-matrix multiplication (VMM) for the matrix, the GP processor being coupled with the plurality of compute engines and being configured to control the plurality of compute engines, to receive output of the VMM for the matrix from each of the plurality of compute engines, and to perform a nonlinear operation on the output, the GP processor including a fixed function computing block usable in performing the nonlinear operation, the plurality of compute tiles including a first compute tile and a second compute tile, the nonlinear operation corresponding to an activation function for the VMM; and

a data conversion engine coupled with the plurality of compute engines of the first compute tile and configured to convert data transferred to the plurality of compute engines from a first format to a second format, wherein the data was received from the second compute tile in the first format;

wherein the plurality of compute engines is addressable by a plurality of data movement initiators on the plurality of compute tiles and by the GP processor, the plurality of data movement initiators being configured to move data to the plurality of compute engines on a the first compute tile of the plurality of compute tiles in a data path bypassing the GP processor on the first compute tile.

11 . A method, comprising:

providing an input vector to at least one compute engine, the at least one compute engine being part of a plurality of compute engines on a first compute tile of a plurality of compute tiles, the plurality of compute engines being coupled with a general-purpose (GP) processor on the first compute tile, the at least one compute engine storing a plurality of weights corresponding to a matrix, each of the plurality of compute engines including a compute-in-memory (CIM) hardware module, the CIM hardware module of the at least one compute engine being configured to store the plurality of weights in a plurality of storage cells and to perform a vector-matrix multiplication (VMM) of the matrix and the input vector, the at least one compute engine performing the VMM for the input vector and the matrix to provide an output, the CIM hardware module configured to perform the VMM by a bit-wise multiplication of the plurality of weights stored in the plurality of storage cells of the CIM hardware module and elements of the input vector;

applying, by the GP processor, a function to the output, the GP processor being configured to control the plurality of compute engines, the GP processor including a fixed function computing block usable in performing the function, the function being a nonlinear activation function for the VMM; and

converting data transferred to the plurality of compute engines of the first compute tile from a first format to a second format, wherein the data was received from a second compute tile of the plurality of compute tiles in the first format;

wherein the plurality of compute engines is addressable by a plurality of data movement initiators and by the GP processor such that the providing the input vector includes providing the input vector to the at least one compute engine by a data movement initiator of the plurality of data movement initiators in a data path that bypasses the GP processor or providing the input vector to the at least one compute engine by the GP processor.

12 . The method of claim 11 , further comprising:

moving data from the CIM hardware module to the at least one compute engine;

wherein the data is moved from the CIM hardware module to the at least one compute engine via at least one data path bypassing the GP processor.

13 . The method of claim 11 , wherein the providing the input vector further includes:

using at least one of a direct memory access (DMA) unit, the plurality of data movement initiators including the DMA unit, or a component on a third compute tile coupled with the first compute tile.

14 . The method of claim 11 , wherein a local memory is coupled with the plurality of compute engines through a first bus, the GP processor is coupled with the plurality of compute engines through a second bus different from the first bus, and a main bus is coupled with the local memory and the GP processor.

15 . The method of claim 11 , wherein a local memory is coupled with the plurality of compute engines by a first bus and wherein a main bus is coupled with the local memory, the plurality of compute engines, and the GP processor, the GP processor being coupled with the plurality of compute engines and the local memory through the main bus.

16 . The method of claim 15 , wherein the plurality of compute engines is coupled to the main bus through an interconnect having an internal queueing.

17 . The method of claim 11 , wherein each of the plurality of compute tiles further includes a local compute engine memory having a first memory density and a cache controller coupled with the local compute engine memory, and wherein the CIM hardware module is coupled with the cache controller and has a second memory density less than the first memory density.

18 . The compute tile of claim 1 , wherein the GP processor is reduced instruction set computer (RISC) processor configured as a separate data movement initiator providing data to the plurality of compute engines, the data including a vector for the VMM.

19 . The compute tile of claim 1 , wherein each of the plurality of compute engines includes a local update module configured to update at least one weight of the plurality of weights in the CIM hardware module.

20 . The compute tile of claim 1 , further comprising:

a local memory coupled with the GP processor;

wherein each compute engine of the plurality of compute engines is coupled with a local compute engine memory, the local compute engine memory having a first memory density;

wherein each compute engine of the plurality of compute engines is coupled to a cache controller, the cache controller being coupled with the local compute engine memory;

wherein the CIM hardware module of each compute engine is coupled with the cache controller and has a second memory density less than the first memory density.

21 . The compute tile of claim 1 , wherein the GP processor is a single GP processor for the compute tile, the single GP processor controlling the plurality of compute engines.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2025
From: RAIN NEUROMORPHICS INC.
To: OPENAI OPCO, LLC
Reel/Frame 073238/0425 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2024
From: ALI, NAWAB; KAL, MUZAFFER; CONKLIN, ALEXANDER ALMELA; ERBAGCI, BURAK; ERYILMAZ, CAGRI; FOUDA, MOHAMMED ELNEANAEI ABDELMONEEM
To: RAIN NEUROMORPHICS INC.
Reel/Frame 068903/0545 →
Continuity (4)
Provisional Application 63532254 · Aug 11, 2023
Provisional Application 63530229 · Aug 1, 2023
Provisional Application 63529921 · Jul 31, 2023
Related Publication 20250045224A1 · Feb 6, 2025
References Cited (58)
US 11347477B2 · Sumbul · 2022 [cited by applicant]
US 11475300B2 · Yao · 2022 [cited by applicant]
US 20090177867A1 · Garde · 2009 [cited by examiner]
US 20100174883A1 · Lerner · 2010 [cited by examiner]
US 20150106311A1 · Birdwell · 2015 [cited by applicant]
US 20180189631A1 · Sumbul · 2018 [cited by applicant]
US 20190102170A1 · Chen · 2019 [cited by applicant]
US 20190102359A1 · Knag · 2019 [cited by applicant]
US 20190179795A1 · Huang · 2019 [cited by applicant]
US 20190205741A1 · Gupta · 2019 [cited by examiner]
US 20190340486A1 · Mills · 2019 [cited by examiner]
US 20190348110A1 · Sinangil · 2019 [cited by applicant]
US 20190362227A1 · Seshadri · 2019 [cited by applicant]
US 20200057938A1 · Lu · 2020 [cited by applicant]
US 20200207656A1 · Yoshioka · 2020 [cited by applicant]
US 20200293284A1 · Vantrease · 2020 [cited by examiner]
US 20200320403A1 · Daga · 2020 [cited by applicant]
US 20200410337A1 · Huang · 2020 [cited by examiner]
US 20210097431A1 · Olgiati · 2021 [cited by applicant]
US 20210158132A1 · Huynh · 2021 [cited by examiner]
US 20210224185A1 · Zhou · 2021 [cited by applicant]
US 20210343343A1 · Teague · 2021 [cited by applicant]
US 20220004497A1 · Willcock · 2022 [cited by applicant]
US 20220019880A1 · Dasgupta · 2022 [cited by applicant]
US 20220114270A1 · Wang · 2022 [cited by examiner]
US 20220138286A1 · Zage · 2022 [cited by examiner]
US 20220164916A1 · Nurvitadhi · 2022 [cited by examiner]
US 20220207293A1 · Yao · 2022 [cited by applicant]
US 20220207656A1 · Yao · 2022 [cited by applicant]
US 20220244916A1 · Lee · 2022 [cited by applicant]
US 20220301605A1 · Mirhaj · 2022 [cited by applicant]
US 20220309328A1 · Saxena · 2022 [cited by applicant]
US 20220318610A1 · Seo · 2022 [cited by applicant]
US 20220414432A1 · Banitalebi Dehkordi · 2022 [cited by applicant]
US 20230014565A1 · Ray · 2023 [cited by examiner]
US 20230045840A1 · Chih · 2023 [cited by applicant]
US 20230047364A1 · Badaroglu · 2023 [cited by applicant]
US 20230074229A1 · Jia · 2023 [cited by examiner]
US 20230138695A1 · Kumar · 2023 [cited by applicant]
US 20230146647A1 · Byeon · 2023 [cited by applicant]
US 20230206044A1 · Ma · 2023 [cited by applicant]
US 20230259456A1 · Verma · 2023 [cited by examiner]
US 20230297580A1 · Sheng · 2023 [cited by applicant]
US 20230316060A1 · Jain · 2023 [cited by applicant]
US 20230359894A1 · Kim · 2023 [cited by applicant]
US 20240094986A1 · Lyubomirsky · 2024 [cited by applicant]
US 20240134606A1 · Yi · 2024 [cited by applicant]
US 20240169201A1 · Seok · 2024 [cited by applicant]
WO 2020190776 · 2020 [cited by applicant]
WO 2022029026 · 2022 [cited by applicant]
Kim et al., Moneta: A Processing-In-Memory-Based Hardware Platform for the Hybrid Convolutional Spiking Neural Network with Online Learning, Frontiers in Neuroscience, vol. 16, Apr. 11, 2022. [cited by applicant]
Korthikanti et al., Reducing Activation Recomputation in Large Transformer Models, May 10, 2022, pp. 1-17. [cited by applicant]
Lee et al., A 12nm 121-TOPS/W 41.6-TOPS/mm2 All Digital Full Precision SRAM-based Compute-in-Memory with Configurable Bit-width for AI Edge Applications, 2022 Symposium on VLSI Technology & Circuits Digest of Technical … [cited by applicant]
Song et al., PipeLayer: A Pipelined ReRAM-Based Accelerator for Deep Learning, 2017. [cited by applicant]
Lin et al., A Novel Voltage-Accumulation Vector-Matrix Multiplication Architecture using Resistor-Shunted Floating Gate Flash Memory Device for Low-Power and High-Density Neural Network Applications, 2018 IEEE Internati… [cited by applicant]
Chih et al., 16.4 An 89TOPS/W and 16.3TOPS/mm2 All-Digital SRAM-Based Full-Precision Compute-In Memory Macro in 22nm for Machine-Learning Edge Applications, In Proc. IEEE Int. Solid-State Circuits Conf.(ISSCC), vol. 64,… [cited by applicant]
Li et al., A Precision-Scalable Energy-Efficient Bit-Split-and-Combination Vector Systolic Accelerator for NAS-Optimized DNNs on Edge, 2022 Design, Automation & Test in Europe Conference & Exhibition, 2022, pp. 730-735. [cited by applicant]
Mori et al., A 4nm 6163-TOPS/W/b 4790-TOPS/mm2/b SRAM Based Digital-Computing-in-Memory Macro Supporting Bit-Width Flexibility and Simultaneous MAC and Weight Update, In 2023 IEEE International Solid-State Circuits Conf… [cited by applicant]