IP Library › Granted Patent US 12,608,439
Granted Patent B2
US 12,608,439 · App. 18/949,783 · Granted Apr 21, 2026

System and method of transposed matrix-vector multiplication

Inventors: Mohammed Elneanaei Abdelmoneem Fouda (Irvine, CA); Ramyad Hadidi (University Park, MD)
Assignee: OpenAl Opco, LLC
G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,439
App. No.
18/949,783
Granted
Apr 21, 2026
Kind
B2
Abstract

A system including a memory and hardware compute logic is described. The memory includes memory cells storing weights corresponding to a matrix. Hardware compute logic is coupled with the memory cells. The hardware compute logic is configured to perform a vector-matrix multiplication (VMM) for the matrix and for a matrix transpose for the weights being stationary for the memory cells.

Claims (43)

1 . A system, comprising:

a memory including a plurality of memory cells storing a plurality of weights corresponding to a matrix; and

hardware compute logic coupled with the plurality of memory cells and configured to perform a vector-matrix multiplication (VMM) for the matrix and for a matrix transpose for the plurality of weights being stationary for the plurality of memory cells, wherein the hardware compute logic includes:

multiplication circuitry for the plurality of memory cells, the multiplication circuitry multiplying at least a portion of a weight stored in a memory cell with a portion of an input vector to provide a product corresponding to the memory cell;

selection logic; and

a plurality of adder trees coupled with the selection logic and the memory, the plurality of adder trees for accumulating a sum of the product of each memory cell for a portion of the plurality of memory cells, the selection logic being configured to select a first portion of the plurality of adder trees for the VMM of the matrix and to select a second portion of the plurality of adder trees for the VMM of the matrix transpose, wherein the product is selectively provided to at least one of the first portion of the plurality of adder trees for the VMM of the matrix or to the second portion of the plurality of adder trees for the VMM of the matrix transpose,

wherein the selection logic further includes a plurality of input multiplexers, the plurality of input multiplexers being configured to route an element of the input vector to a corresponding multiplication circuitry based on whether the VMM is for the matrix or the matrix transpose.

2 . The system of claim 1 , wherein the selection logic further includes:

input logic configured to provide the portion of the input vector with the memory cell for the matrix or the matrix transpose.

3 . The system of claim 1 , wherein each adder tree of the plurality of adder trees shares at least one adder with another adder tree of the plurality of adder trees.

4 . The system of claim 3 , wherein the selection logic further includes a plurality of multiplexers, each multiplexer of the plurality of multiplexers is coupled with the at least one adder for selecting a product corresponding to the matrix or the matrix transpose.

5 . The system of claim 1 , wherein the selection logic is configured to provide, to the plurality of adder trees, the product for the matrix or for the matrix transpose for each memory cell of the plurality of memory cells.

6 . The system of claim 1 , wherein each of the plurality of memory cells is selected from a digital static random access memory (SRAM) memory cell and an analog SRAM memory cell.

7 . The system of claim 1 , wherein the selection logic is further configured to:

receive a control signal indicating whether the VMM is for the matrix or the matrix transpose; and

provide the control signal to the plurality of input multiplexers to control routing of elements of the input vector to corresponding multiplication circuitry based on whether the VMM is for the matrix or the matrix transpose.

8 . The system of claim 4 , wherein each multiplexer of the plurality of multiplexers is configured to select between a first product and a second product for providing to the at least one shared adder based on whether the VMM is for the matrix or the matrix transpose.

9 . The system of claim 1 , wherein the plurality of input multiplexers are arranged such that each input multiplexer corresponds to a respective multiplication circuitry, and each input multiplexer selects which element of the input vector is provided to its corresponding multiplication circuitry based on whether the VMM is for the matrix or the matrix transpose.

10 . A hardware accelerator, comprising:

at least one processor; and

at least one compute tile including and coupled with the at least one processor, the at least one compute tile including a memory and hardware compute logic, the memory including a plurality of memory cells storing a plurality of weights corresponding to a matrix, the hardware compute logic being coupled with the plurality of memory cells and configured to perform a vector-matrix multiplication (VMM) for the matrix and for a matrix transpose for the plurality of weights being stationary for the plurality of memory cells, wherein the hardware compute logic includes:

multiplication circuitry for the plurality of memory cells, the multiplication circuitry multiplying at least a portion of a weight stored in a memory cell with a portion of an input vector to provide a product corresponding to the memory cell;

selection logic; and

a plurality of adder trees coupled with the selection logic and the memory, the plurality of adder trees for accumulating a sum of the product of each memory cell for a portion of the plurality of memory cells, the selection logic being configured to select a first portion of the plurality of adder trees for the VMM of the matrix and to select a second portion of the plurality of adder trees for the VMM of the matrix transpose, wherein the product is selectively provided to at least one of the first portion of the plurality of adder trees for the VMM of the matrix or to the second portion of the plurality of adder trees for the VMM of the matrix transpose,

wherein the selection logic further includes a plurality of input multiplexers, the plurality of input multiplexers being configured to route an element of the input vector to a corresponding multiplication circuitry based on whether the VMM is for the matrix or the matrix transpose.

11 . The hardware accelerator of claim 10 , wherein the selection logic further includes:

input logic configured to provide the portion of the input vector with the memory cell for the matrix or the matrix transpose.

12 . The hardware accelerator of claim 10 , wherein each adder tree of the plurality of adder trees shares at least one adder with another adder tree of the plurality of adder trees.

13 . The hardware accelerator of claim 12 , wherein the selection logic further includes a plurality of multiplexers, each multiplexer of the plurality of multiplexers is coupled with the at least one adder for selecting a product corresponding to the matrix or the matrix transpose.

14 . The hardware accelerator of claim 10 , wherein the selection logic is configured to provide, to the plurality of adder trees, the product for the matrix or for the matrix transpose for each memory cell of the plurality of memory cells.

15 . The hardware accelerator of claim 10 , wherein each of the plurality of memory cells is selected from a digital static random access memory (SRAM) memory cell and an analog SRAM memory cell.

16 . The hardware accelerator of claim 10 , wherein the memory and the hardware compute logic of each of the at least one compute tile are configured as a compute-in-memory (CIM) hardware module.

17 . A method, comprising:

providing an input vector to a plurality of compute engines coupled with a processor, the plurality of compute engines including a compute-in-memory (CIM) hardware module, the CIM hardware module including a memory and hardware compute logic, the memory including a plurality of memory cells storing a plurality of weights corresponding to a matrix, the hardware compute logic coupled with the plurality of memory cells and configured to perform a vector-matrix multiplication (VMM) for the matrix and for a matrix transpose for the plurality of weights being stationary for the plurality of memory cells, the hardware compute logic including:

multiplication circuitry for the plurality of memory cells, the multiplication circuitry multiplying at least a portion of a weight stored in a memory cell with a portion of the input vector to provide a product corresponding to the memory cell;

selection logic; and

a plurality of adder trees coupled with the selection logic and the memory, the plurality of adder trees for accumulating a sum of the product of each memory cell for a portion of the plurality of memory cells, the selection logic being configured to select a first portion of the plurality of adder trees for the VMM of the matrix and to select a second portion of the plurality of adder trees for the VMM of the matrix transpose, wherein the product is selectively provided to at least one of the first portion of the plurality of adder trees for the VMM of the matrix or to the second portion of the plurality of adder trees for the VMM of the matrix transpose,

selecting the matrix or the matrix transpose; and

performing the VMM of the input vector and the matrix or the matrix transpose based on the selecting and using at least one of the plurality of compute engines,

wherein the selection logic further includes a plurality of input multiplexers, the plurality of input multiplexers being configured to route an element of the input vector to a corresponding multiplication circuitry based on whether the VMM is for the matrix or the matrix transpose.

18 . The method of claim 17 , wherein each adder tree of the plurality of adder trees shares at least one adder with another adder tree of the plurality of adder trees.

19 . The method of claim 17 , wherein the selection logic is configured to provide, to the plurality of adder trees, the product for the matrix or for the matrix transpose for each memory cell of the plurality of memory cells.

20 . The method of claim 17 , further comprising providing the input vector to an input buffer prior to providing the input vector to the plurality of compute engines, the input buffer temporarily storing elements of the input vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: FOUDA, MOHAMMED ELNEANAEI ABDELMONEEM; HADIDI, RAMYAD
To: RAIN NEUROMORPHICS INC.
Reel/Frame 069772/0796 →
Continuity (3)
Continuation 18891921 · Sep 20, 2024
Provisional Application 63539753 · Sep 21, 2023
Related Publication 20250103680A1 · Mar 27, 2025
References Cited (5)
US 20190004997A1 · Cohen · 2019 [cited by examiner]
US 20220019407A1 · Chih · 2022 [cited by examiner]
US 20220188628A1 · Rasch · 2022 [cited by examiner]
H. Jiang et al., “CIMAT: a transpose SRAM-based compute-in-memory architecture for deep neural network on-chip training” in Proceedings of the International Symposium on Memory Systems (MEMSYS '19). New York, NY, USA, 4… [cited by examiner]
Y. Luo and S. Yu, “AILC: Accelerate On-Chip Incremental Learning With Compute-in-Memory Technology,” in IEEE Transactions on Computers, vol. 70, No. 8, pp. 1225-1238, Aug. 1, 2021, doi: 10.1109/TC.2021.3053199. (Year: 2… [cited by examiner]