IP Library › Granted Patent US 12,482,520
Granted Patent B2
US 12,482,520 · App. 18/495,167 · Granted Nov 25, 2025

6T-SRAM-based digital computing-in-memory circuits supporting flexible input dimension

Inventors: Mingoo Seok (Tenafly, NJ); Jonghyun Oh (New York, NY)
Assignee: THE TRUSTEES OF COLUMBIA UNIVERSITY IN THE CITY OF NEW YORK
G11C11/4096G06F17/16G11C11/4094G06N3/063G11C7/1006G11C11/54
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,520
App. No.
18/495,167
Granted
Nov 25, 2025
Kind
B2
Abstract

Compute-in-memory (CIM) devices and methods for performing vector-matrix multiplication (VMM) are provided. The disclosed CIM device can include a static random access memory (SRAM) array. The SRAM array can include a plurality of column structures. Each column structure can include eight sub-column structures. Each sub-column structure can include at least one bitcell sharing a pair of a local bitline (LBL) and LBLb that can be connected to a pair of global bitlines (GBL) via switches. Each sub-column comprises at least one NOR gate. An even-numbered bitcell can include a wordline 1 (WL1) for a left access transistors, and an odd-numbered bitcell can include a wordline 2 (WL2) for a right access transistors. Every eight columns (8 columns) can be configured to share a hybrid compressor adder-tree (HCA), followed by a bit-first accumulation (BFA).

Claims (30)

1 . A compute-in-memory (CIM) device, comprising

a static random access memory (SRAM) array including

a plurality of column structures, wherein each column structure comprises eight sub-column structures;

wherein each sub-column structure comprises at least one bitcell sharing a pair of a local bitline (LBL) and local bitline bar (LBLb),

wherein the LBL and LBLb are connected to a pair of global bitlines (GBL) via switches,

wherein each sub-column comprises at least one NOR gate,

wherein an even-numbered bitcell comprises a wordline 1 (WL1) for a left access transistor, and an odd-numbered bitcell comprises a wordline 2 (WL2) for a right access transistor; and

wherein every eight columns (8 columns) are configured to share a hybrid compressor adder-tree (HCA) structure, followed by a bit-first accumulation (BFA) structure.

2 . The CIM device of claim 1 , wherein the CIM device is configured to perform a static dual wordline access without a pre-charging operation by accessing two consecutive bitcells in each sub-column using the LBL and the LBLb.

3 . The CIM device of claim 1 , wherein the HCA includes a plurality of 15:4 compressors followed by a 4b 8-input adder tree.

4 . The CIM device of claim 2 , wherein the adder tree comprises a carry-in port of ripple carry adders (RCA), wherein the RCA is 4b RCAs, 6b RCAs, 8b RCAs, or combinations thereof.

5 . The CIM device of claim 1 , wherein the BFA comprises a 23b RCA, a 30b register, and bi-directional shifters to perform the shift-and-accumulate on an output of HCA (partial sums).

6 . The CIM device of claim 5 , wherein the BFA is configured to accumulate the partial products across input bits first and then inputs.

7 . The CIM device of claim 1 , wherein the BFA is configured to a bi-directional bit-serial input operation.

8 . The CIM device of claim 1 , wherein the HCA and the BFA comprise area-efficient transmission-gate (TG)-based full adder cell (FA) and half adder (HA).

9 . The CIM device of claim 8 , wherein the FA is an input-inverted FA and/or wherein the HA is an input-inverted HA.

10 . The CIM device of claim 1 , wherein the BFA comprises a polarity-inverted multiplexer and a D-flip-flop (DFF).

11 . The CIM device of claim 1 , wherein the at least one NOR gate is configured to be a 1b multiplier.

12 . The CIM device of claim 1 , wherein the CIM device comprises 16 HCAs and 16 BFAs.

13 . The CIM device of claim 1 , wherein the SRAM array comprises a row peripheral configured to control a vector-matrix multiplication (VMM) operation and a SRAM Read/Write (R/W) operation.

14 . The CIM device of claim 1 , wherein the SRAM array comprises a column peripheral configured to control GBLs for the SRAM R/W operation.

15 . A method for performing vector-matrix multiplication comprising:

activating two consecutive wordline 1s (WL1s) in each sub-column of a compute-in-memory (CIM) device, wherein the WL1s are configured to transfer two weight bits, via LBL and LBLb, to the two NOR gates in the sub-column;

feeding corresponding two input activation bits via TLs to the NOR gates using a row peripheral of the CIM device;

generating a total of 16 8-b partial products using the column, where each column comprises eight sub-columns;

adding up the 16 partial products and producing partial sums using an HCA;

performing a shift-and-accumulate the partial sums using BFA;

producing a VMM result using the results from the shift-and-accumulate the partial sums.

16 . The method of claim 15 , wherein the VMM result is produced from an 8b 128×16d (dimension) VMM in 64 clock cycles.

17 . The method claim 15 , wherein the VMM result is a 23b 16d vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2026
From: SEOK, MINGOO; OH, JONGHYUN
To: THE TRUSTEES OF COLUMBIA UNIVERSITY IN THE CITY OF NEW YORK
Reel/Frame 073487/0951 →
Continuity (2)
Provisional Application 63427608 · Nov 23, 2022
Related Publication 20240170050A1 · May 23, 2024
References Cited (27)
US 5499215A · Hatta · 1996 [cited by examiner]
US 9846613B2 · Arslan et al. · 2017 [cited by applicant]
US 10579974B1 · Reed · 2020 [cited by applicant]
US 10616669B2 · Kumar et al. · 2020 [cited by applicant]
US 10762475B2 · Song et al. · 2020 [cited by applicant]
US 11061646B2 · Sumbul et al. · 2021 [cited by applicant]
US 20040223369A1 · Choi · 2004 [cited by examiner]
US 20050078542A1 · Oh · 2005 [cited by examiner]
US 20070002674A1 · Lee · 2007 [cited by examiner]
US 20080186795A1 · Lih · 2008 [cited by examiner]
US 20100014345A1 · Choi · 2010 [cited by examiner]
US 20110063898A1 · Ong · 2011 [cited by examiner]
US 20120243360A1 · Ferrant · 2012 [cited by examiner]
US 20210327474A1 · Seok · 2021 [cited by examiner]
Chih et al., “An 89TOPS/W and 16.3TOPS/mm2 all-digital SRAM-based full-precision compute-in memory macro in 22nm for machine-learning edge applications,” 2021 IEEE International Solid-State Circuits Conference (ISSCC) p… [cited by applicant]
Fujiwara et al., “A 5-nm 254-TOPS/W 221-TOPS/mm2 fully digital computing-in-memory macro supporting wide-range dynamic-voltage-frequency scaling and simultaneous MAC and write operations,” 2022 IEEE International Solid-… [cited by applicant]
Huang et al., “In-Memory computing architecture for a convolutional neural network based on spin orbit torque MRAM,” Electronics. 11(8):1245 (2022) 17 pgs. [cited by applicant]
Kiani et al., “A fully hardware-based memristive multilayer neural network,” Sci Adv 7:eabj4801 (2021). [cited by applicant]
Lee et al., “A 12nm 121-TOPS/W 41.6-TOPS/mm2 all digital full precision SRAM-based compute-in-memory with configurable bit-width for AI edge applications,” 2022 IEEE Symposium on VLSI Technology and Circuits (VLSI Techn… [cited by applicant]
Liu et al., “Two-dimensional materials for next-generation computing technologies,” Nat Nanotechnol 15(7):545-557 (2020). [cited by applicant]
Mackin et al., “Optimised weight programming for analogue memory-based deep neural networks,” Nat Commun 13:3765 (2022) 12 pgs. [cited by applicant]
Pereira et al., “Efficient hardware design and implementation of the voting scheme-based convolution,” Sensors (Basel) 22:2943 (2022) 18 pgs. [cited by applicant]
Tu et al., “A 28nm 29.2TFLOPS/W BF16 and 36.5TOPS/W INT8 Reconfigurable Digital CIM Processor with Unified FP/INT Pipeline and Bitwise In-Memory Booth Multiplication for Cloud Deep Learning Acceleration,” ISSCC, 2022, 3… [cited by applicant]
Wan et al., “A compute-in-memory chip based on resistive random-access memory,” Nature 608(7923):504-512 (2022). [cited by applicant]
Wang et al., “DIMC: 2219TOPS/W 2569F2/b Digital In-Memory Computing Macro in 28nm Based on Approximate Arithmetic Hardware,” ISSCC, 2022, 3 pgs. [cited by applicant]
Yan et al., “A 1.041-Mb/mm2 27.38-TOPS/W Signed-INT8 Dynamic-Logic-Based ADC-less SRAM Compute-In-Memory Macro in 28nm with Reconfigurable Bitwise Operation for AI and Embedded Applications,” ISSCC, 2022, 35 pgs. [cited by applicant]
Yu et al., “Evaluating architecture impact on system energy efficiency,” PLoS One 12(11):e0188428 (2017). [cited by applicant]