IP Library › Patent Application 19575128
Patent Application
App. No. 19/575,128

SYSTEM AND METHOD OF TRANSPOSED MATRIX-VECTOR MULTIPLICATION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/575,128
Abstract

A system including a memory and hardware compute logic is described. The memory includes memory cells storing weights corresponding to a matrix. Hardware compute logic is coupled with the memory cells. The hardware compute logic is configured to perform a vector-matrix multiplication (VMM) for the matrix and for a matrix transpose for the weights being stationary for the memory cells.

Claims (40)

1 . A system, comprising:

a memory including a plurality of memory cells storing a plurality of weights corresponding to a matrix; and

hardware compute logic coupled to the plurality of memory cells and configured to selectively perform a vector-matrix multiplication (VMM) using the matrix and the matrix transpose with the plurality of weights remaining stored in the plurality of memory cells, wherein the hardware compute logic includes:

multiplication circuitry associated with the plurality of memory cells and coupled to the memory, the multiplication circuitry configured to multiply at least a portion of a weight stored in a memory cell with a portion of an input vector to output a result corresponding to the memory cell;

a plurality of adder trees coupled to the memory; and

a multiplexer module coupled to the multiplication circuitry and the plurality of adder trees, the multiplexer module configured to route the result to a corresponding adder tree of the plurality of adder trees, wherein the multiplexer module is configured to implement a first routing configuration corresponding to the VMM using the matrix and a second routing configuration corresponding to the VMM using the matrix transpose.

2 . The system of claim 1 , wherein each of the plurality of memory cells is selected from a digital static random access memory (SRAM) memory cell and an analog SRAM memory cell.

3 . The system of claim 1 , wherein the hardware compute logic further includes a plurality of input multiplexers, each input multiplexer having a first input configuration corresponding to the VMM using the matrix and a second input configuration corresponding to the VMM using the matrix transpose.

4 . The system of claim 1 , wherein the hardware compute logic further includes a shift and accumulate module coupled to the plurality of adder trees.

5 . The system of claim 1 , wherein the hardware compute logic further includes an input buffer configured to receive the input vector and an output buffer configured to provide a result of the VMM.

6 . The system of claim 1 , wherein the multiplexer module is configured to receive a control signal, the control signal selecting between the first routing configuration and the second routing configuration, wherein

in the first routing configuration, results from memory cells are routed to adder trees for accumulation corresponding to the VMM using the matrix, and

in the second routing configuration, results from memory cells are routed to adder trees for accumulation corresponding to the VMM using the matrix transpose.

7 . The system of claim 1 , wherein the system is included in a compute tile, the compute tile further comprising at least one processor coupled to the system, the at least one processor configured to provide the input vector and to select between the first routing configuration and the second routing configuration.

8 . The system of claim 7 , wherein the compute tile further includes a local update module configured to update the plurality of weights stored in the plurality of memory cells.

9 . The system of claim 1 , wherein the plurality of adder trees are configured to accumulate a sum of the results corresponding to a portion of the plurality of memory cells.

10 . The system of claim 1 , wherein the memory is configured to store a rectangular matrix of weights.

11 . A hardware accelerator, comprising:

at least one processor; and

at least one compute tile coupled to the at least one processor, each of the at least one compute tile including a memory and hardware compute logic, the memory including a plurality of memory cells storing a plurality of weights corresponding to a matrix, the hardware compute logic being coupled to the plurality of memory cells and configured to selectively perform a vector-matrix multiplication (VMM) using the matrix and the matrix transpose with the plurality of weights remaining stored in the plurality of memory cells, wherein the hardware compute logic includes:

multiplication circuitry associated with the plurality of memory cells and coupled to the memory, the multiplication circuitry configured to multiply at least a portion of a weight stored in a memory cell with a portion of an input vector to output a result corresponding to the memory cell;

a plurality of adder trees coupled to the memory; and

a multiplexer module coupled to the multiplication circuitry and the plurality of adder trees, the multiplexer module configured to route the result to a corresponding adder tree of the plurality of adder trees, wherein the multiplexer module is configured to implement a first routing configuration corresponding to the VMM using the matrix and a second routing configuration corresponding to the VMM using the matrix transpose.

12 . The hardware accelerator of claim 11 , wherein each of the plurality of memory cells is selected from a digital static random access memory (SRAM) memory cell and an analog SRAM memory cell.

13 . The hardware accelerator of claim 11 , wherein the hardware compute logic further includes a plurality of input multiplexers, each input multiplexer having a first input configuration corresponding to the VMM using the matrix and a second input configuration corresponding to the VMM using the matrix transpose.

14 . The hardware accelerator of claim 11 , wherein the multiplexer module is configured to receive a control signal, the control signal selecting between the first routing configuration and the second routing configuration, wherein

in the first routing configuration, results from memory cells are routed to adder trees for accumulation corresponding to the VMM using the matrix, and

in the second routing configuration, results from memory cells are routed to adder trees for accumulation corresponding to the VMM using the matrix transpose.

15 . The hardware accelerator of claim 11 , wherein each of the at least one compute tile further includes a local update module configured to update the plurality of weights stored in the plurality of memory cells.

16 . A method, comprising:

providing an input vector to hardware compute logic of a compute-in-memory (CIM) hardware module, the CIM hardware module including a memory and the hardware compute logic, the memory including a plurality of memory cells storing a plurality of weights corresponding to a matrix, the hardware compute logic coupled to the plurality of memory cells and configured to selectively perform a vector-matrix multiplication (VMM) using the matrix and the matrix transpose with the plurality of weights remaining stored in the plurality of memory cells, the hardware compute logic including:

multiplication circuitry associated with the plurality of memory cells and coupled to the memory, the multiplication circuitry configured to multiply at least a portion of a weight stored in a memory cell with a portion of the input vector to output a result corresponding to the memory cell;

a plurality of adder trees coupled to the memory; and

a multiplexer module coupled to the multiplication circuitry and the plurality of adder trees;

selecting a first routing configuration of the multiplexer module corresponding to the VMM using the matrix or a second routing configuration of the multiplexer module corresponding to the VMM using the matrix transpose; and

performing the VMM using the selected routing configuration, wherein the multiplexer module routes the result to a corresponding adder tree of the plurality of adder trees based on the selected routing configuration.

17 . The method of claim 16 , wherein selecting the first routing configuration or the second routing configuration comprises receiving a control signal indicating the selected routing configuration.

18 . The method of claim 16 , further comprising providing the input vector to an input buffer prior to providing the input vector to the hardware compute logic, the input buffer temporarily storing elements of the input vector.

19 . The method of claim 16 , further comprising updating the plurality of weights stored in the plurality of memory cells using a local update module.

20 . The method of claim 16 , wherein the memory is configured to store a rectangular matrix of weights.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2026
From: FOUDA, MOHAMMED ELNEANAEI ABDELMONEEM; HADIDI, RAMYAD
To: RAIN NEUROMORPHICS INC.
Reel/Frame 074194/0776 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2026
From: RAIN NEUROMORPHICS INC.
To: OPENAI OPCO, LLC
Reel/Frame 074194/0862 →