IP Library › Granted Patent US 12,748,722
Granted Patent B2
US 12,748,722 · App. 18/384,774 · Granted Sep 29, 2026

Compute in-memory architecture for continuous on-chip learning

Inventor: Mohammed Elneanaei Abdelmoneem Fouda (Irvine, CA)
Assignee: OpenAI OpCo, LLC
G06F15/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,722
App. No.
18/384,774
Granted
Sep 29, 2026
Kind
B2
Abstract

A system capable of providing on-chip learning comprising a processor and a plurality of compute engines coupled with the processor. Each of the compute engines including a compute-in-memory (CIM) hardware module and a local update module. The CIM hardware module stores a plurality of weights corresponding to a matrix and is configured to perform a vector-matrix multiplication for the matrix. The local update module is coupled with the CIM hardware module and configured to update at least a portion of the weights.

Claims (47)

1 . A system, comprising:

a processor; and

a plurality of compute engines coupled with the processor, the plurality of compute engines including a compute-in-memory (CIM) hardware module, a bit mixer, and a local update module, the CIM hardware module storing a plurality of weights corresponding to a matrix and configured to perform a vector-matrix multiplication (VMM) for the matrix, the local update module being coupled with the CIM hardware module and configured to update at least a portion of the plurality of weights,

wherein:

the bit mixer is coupled to the CIM and configured to combine results from the CIM into a weighted output by using exponential weights and bit slicing;

the CIM hardware module includes a plurality of cells for storing the plurality of weights; and

the local update module includes:

an adder configured to be selectively coupled with each of the plurality of cells, to receive a weight update, and to add the weight update with a weight of the plurality of weights for each of the plurality of cells; and

write circuitry coupled with the adder and the plurality of cells, the write circuitry configured to write a sum of the weight and the weight update to each of the plurality of cells, and

a compute engine of the plurality of compute engines includes:

a controller configured to provide a plurality of control signals to the CIM hardware module and the local update module,

a first portion of the plurality of control signals corresponding to an inference mode, a second portion of the plurality of control signals corresponding to a weight update mode,

wherein the weight update mode includes providing the weight update to the adder.

2 . The system of claim 1 , wherein the plurality of cells is selected from a plurality of analog static random access memory (SRAM) cells, a plurality of digital SRAM cells, and a plurality of resistive random access memory (RRAM) cells.

3 . The system of claim 2 , wherein the plurality of cells includes the plurality of analog SRAM cells, the CIM hardware module further including a capacitive voltage divider for each of the plurality of analog SRAM cells.

4 . The system of claim 1 , wherein the local update module further includes: a local batched weight update calculator coupled with the adder and configured to determine the weight update.

5 . The system of claim 1 , wherein each of the plurality of compute engines further includes: address circuitry configured to selectively couple the adder and the write circuitry with any of the plurality of cells.

6 . The system of claim 1 , wherein the plurality of weights includes at least one positive weight and at least one negative weight.

7 . The system of claim 1 , further comprising:

a scaled vector accumulation (SVA) unit coupled with the plurality of compute engines and the processor, the SVA unit configured to apply an activation function to an output of the plurality of compute engines.

8 . The system of claim 7 , wherein the SV A unit and the plurality of compute engines are in a plurality of tiles.

9 . The system of claim 1 , wherein the bit mixer is coupled to an analog-to-digital converter (ADC) that is coupled to an output cache, the output cache being configured to pass data by rows.

10 . The system of claim 1 , wherein the CIM module is configured to perform the VMM by mapping the plurality of weights to bipolar weights and multiplying the matrix with an input vector and scaling the multiplication result based on a symmetric range.

11 . The system of claim 1 , wherein the compute engine is further configured to comprise an input cache, a control unit, and an output cache.

12 . A machine learning system, comprising:

at least one processor; and

a plurality of tiles coupled with the at least one processor, the plurality of tiles including a plurality of compute engines and at least one scaled vector accumulation (SVA) unit, the SVA unit configured to apply an activation function to an output of the plurality of compute engines the plurality of compute engines being interconnected and coupled with the SV A unit, each of the plurality of compute engines including at least one compute-in-memory (CIM) hardware module, a controller, a bit mixer and at least one local update module, the at least one CIM hardware module including a plurality of static random access memory (SRAM) cells storing a plurality of weights corresponding to a matrix, the at least one CIM hardware module being configured to perform a vector-matrix multiplication (VMM) for the matrix, the at least one local update module being coupled with the at least one CIM hardware module and configured to update at least a portion of the plurality of weights, the controller being configured to provide a plurality of control signals to the at least one CIM hardware module and the at least one local update module, a first portion of the plurality of control signals corresponding to an inference mode, a second portion of the plurality of control signals corresponding to a weight update mode, wherein:

the bit mixer is coupled to the CIM and configured to combine results from the CIM into a weighted output by using exponential weights and bit slicing; and

a local update module of the at least one local update module includes:

an adder configured to be selectively coupled with each of the plurality of SRAM cells, to receive a weight update, and to add the weight update with a weight of the plurality of weights for each of the plurality of SRAM cells; and

write circuitry coupled with the adder and the plurality of SRAM cells, the write circuitry configured to write a sum of the weight and the weight update to each of the plurality of SRAM cells, wherein the weight update mode includes providing the weight update to the adder.

13 . The machine learning system of claim 12 , wherein each of the plurality of compute engines further includes address circuitry configured to selectively couple the adder and the write circuitry with each of the plurality of SRAM cells.

14 . A method, comprising:

providing an input vector to a plurality of compute engines coupled with a processor, the plurality of compute engines including a compute-in-memory (CIM) hardware module, a bit mixer, and a local update module, the CIM hardware module storing a plurality of weights corresponding to a matrix in a plurality of cells and configured to perform a vector-matrix multiplication (VMM), the local update module being coupled with the CIM hardware module and configured to update at least a portion of the plurality of weights, wherein the bit mixer is configured to combine results from the CIM into a weighted output by using exponential weights and bit slicing;

performing the VMM of the input vector and the matrix using the plurality of compute engines;

determining at least one weight update for the plurality of weights; and

locally updating the plurality of weights using the at least one weight update and the local update module, wherein:

a compute engine of the plurality of compute engines includes:

a controller configured to provide a plurality of control signals to the CIM hardware module and the local update module, a first portion of the plurality of control signals corresponding to an inference mode, a second portion of the plurality of control signals corresponding to a weight update mode, wherein the weight update mode includes providing the at least one weight update to an adder; and the locally updating of the plurality of weights includes:

adding, using the adder configured to be selectively coupled with each of the plurality of cells, the at least one weight update to a weight of at least a portion of the plurality of weights for each of the plurality of cells; and

writing, using write circuitry coupled with the adder and the plurality of cells, a sum of the weight and the weight update to each of the plurality of cells.

15 . The method of claim 14 , wherein the plurality of cells is selected from a plurality of analog static random access memory (SRAM) cells, a plurality of digital SRAM cells, and a plurality of resistive random access memory (RRAM) cells.

16 . The method of claim 14 , wherein the plurality of weights includes at least one positive weight and at least one negative weight.

17 . The method of claim 14 , further comprising: applying an activation function to an output of the plurality of compute engines.

18 . The method of claim 17 , wherein the applying further includes: using a scaled vector accumulation (SV A) unit coupled with the plurality of compute engines to apply the activation function to the output of the plurality of compute engines.

19 . The method of claim 14 , wherein the bit mixer is coupled with the CIM module and an analog-to-digital converter (ADC) that is coupled to an output cache, the output cache being configured to pass data by rows.

20 . The method of claim 14 , wherein the CIM module is configured to perform the VMM by mapping the plurality of weights to bipolar weights and multiplying the matrix with an input vector and scaling the multiplication result based on a symmetric range.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2025
From: RAIN NEUROMORPHICS INC.
To: OPENAI OPCO, LLC
Reel/Frame 073238/0425 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2024
From: FOUDA, MOHAMMED ELNEANAEI ABDELMONEEM
To: RAIN NEUROMORPHICS INC.
Reel/Frame 066137/0052 →
Continuity (2)
Provisional Application 63420437 · Oct 28, 2022
Related Publication 20240143541A1 · May 2, 2024
References Cited (58)
US 11347477B2 · Sumbul · 2022 [cited by applicant]
US 11475300B2 · Yao · 2022 [cited by examiner]
US 20090177867A1 · Garde · 2009 [cited by applicant]
US 20100174883A1 · Lerner · 2010 [cited by applicant]
US 20150106311A1 · Birdwell · 2015 [cited by applicant]
US 20180189631A1 · Sumbul · 2018 [cited by applicant]
US 20190102170A1 · Chen · 2019 [cited by examiner]
US 20190102359A1 · Knag · 2019 [cited by examiner]
US 20190179795A1 · Huang · 2019 [cited by applicant]
US 20190205741A1 · Gupta · 2019 [cited by applicant]
US 20190340486A1 · Mills · 2019 [cited by applicant]
US 20190348110A1 · Sinangil · 2019 [cited by applicant]
US 20190362227A1 · Seshadri · 2019 [cited by applicant]
US 20200057938A1 · Lu · 2020 [cited by applicant]
US 20200207656A1 · Yoshioka · 2020 [cited by applicant]
US 20200293284A1 · Vantrease · 2020 [cited by applicant]
US 20200320403A1 · Daga · 2020 [cited by applicant]
US 20200410337A1 · Huang · 2020 [cited by applicant]
US 20210097431A1 · Olgiati · 2021 [cited by applicant]
US 20210158132A1 · Huynh · 2021 [cited by applicant]
US 20210224185A1 · Zhou · 2021 [cited by examiner]
US 20210343343A1 · Teague · 2021 [cited by examiner]
US 20220004497A1 · Willcock · 2022 [cited by applicant]
US 20220019880A1 · Dasgupta · 2022 [cited by applicant]
US 20220114270A1 · Wang · 2022 [cited by applicant]
US 20220138286A1 · Zage · 2022 [cited by applicant]
US 20220164916A1 · Nurvitadhi · 2022 [cited by applicant]
US 20220207293A1 · Yao · 2022 [cited by applicant]
US 20220207656A1 · Yao · 2022 [cited by applicant]
US 20220244916A1 · Lee · 2022 [cited by examiner]
US 20220301605A1 · Mirhaj · 2022 [cited by examiner]
US 20220309328A1 · Saxena · 2022 [cited by examiner]
US 20220318610A1 · Seo · 2022 [cited by applicant]
US 20220414432A1 · Banitalebi Dehkordi · 2022 [cited by applicant]
US 20230014565A1 · Ray · 2023 [cited by applicant]
US 20230045840A1 · Chih · 2023 [cited by examiner]
US 20230047364A1 · Badaroglu · 2023 [cited by applicant]
US 20230074229A1 · Jia · 2023 [cited by applicant]
US 20230138695A1 · Kumar · 2023 [cited by applicant]
US 20230146647A1 · Byeon · 2023 [cited by applicant]
US 20230206044A1 · Ma · 2023 [cited by applicant]
US 20230259456A1 · Verma · 2023 [cited by applicant]
US 20230297580A1 · Sheng · 2023 [cited by applicant]
US 20230316060A1 · Jain · 2023 [cited by examiner]
US 20230359894A1 · Kim · 2023 [cited by applicant]
US 20240094986A1 · Lyubomirsky · 2024 [cited by applicant]
US 20240134606A1 · Yi · 2024 [cited by applicant]
US 20240169201A1 · Seok · 2024 [cited by applicant]
WO 2020190776 · 2020 [cited by applicant]
WO 2022029026 · 2022 [cited by applicant]
Chih et al., 16.4 An 89TOPS/W and 16.3TOPS/mm2 All-Digital SRAM-Based Full-Precision Compute—In Memory Macro in 22nm for Machine-Learning Edge Applications, In Proc. IEEE Int. Solid-State Circuits Conf.(ISSCC), vol. 64,… [cited by applicant]
Li et al., A Precision-Scalable Energy-Efficient Bit-Split-and-Combination Vector Systolic Accelerator for NAS-Optimized DNNs on Edge, 2022 Design, Automation & Test in Europe Conference & Exhibition, 2022, pp. 730-735. [cited by applicant]
Mori et al., A 4nm 6163-TOPS/W/b 4790-TOPS/mm2/b SRAM Based Digital-Computing-in-Memory Macro Supporting Bit-Width Flexibility and Simultaneous MAC and Weight Update, In 2023 IEEE International Solid-State Circuits Conf… [cited by applicant]
Song et al., PipeLayer: A Pipelined ReRAM-Based Accelerator for Deep Learning, 2017. [cited by applicant]
Korthikanti et al., Reducing Activation Recomputation in Large Transformer Models, May 10, 2022, pp. 1-17. [cited by applicant]
Lee et al., A 12nm 121-TOPS/W 41.6-TOPS/mm2 All Digital Full Precision SRAM-based Compute-in-Memory with Configurable Bit-width for AI Edge Applications, 2022 Symposium on VLSI Technology & Circuits Digest of Technical … [cited by applicant]
Kim et al., MONETA: A Processing-In-Memory-Based Hardware Platform for the Hybrid Convolutional Spiking Neural Network with Online Learning, Frontiers in Neuroscience, vol. 16, Apr. 11, 2022. [cited by applicant]
Lin et al., A Novel Voltage-Accumulation Vector-Matrix Multiplication Architecture using Resistor-Shunted Floating Gate Flash Memory Device for Low-Power and High-Density Neural Network Applications, 2018 IEEE Internati… [cited by applicant]