SYSTEM AND METHOD OF TRANSPOSED MATRIX-VECTOR MULTIPLICATION
A system including a memory and hardware compute logic is described. The memory includes memory cells storing weights corresponding to a matrix. Hardware compute logic is coupled with the memory cells. The hardware compute logic is configured to perform a vector-matrix multiplication (VMM) for the matrix and for a matrix transpose for the weights being stationary for the memory cells.
1 . A system, comprising:
a memory storing a plurality of weights corresponding to a matrix; and
compute logic coupled to the memory and configured to selectively perform a first vector-matrix multiplication (VMM) using the matrix and a second VMM using a transpose of the matrix,
wherein the compute logic comprises:
multiplication circuitry coupled to the memory, the multiplication circuitry configured to output results corresponding to multiplications of at least a portion of the plurality of weights and at least a portion of input vectors;
a plurality of adder trees configured to accumulate the results; and
selection logic comprising a plurality of input multiplexers configured to route elements of the input vector to the multiplication circuitry,
wherein the selection logic is configured to:
select a first portion of the plurality of adder trees and provide the results to the first portion in response to the compute logic performing the first VMM; and
select a second portion of the plurality of adder trees and provide the results to the second portion in response to the compute logic performing the second VMM.
2 . The system of claim 1 , wherein the selection logic further comprises input logic configured to provide the portion of the input vector to the multiplication circuitry in response to the compute logic performing the first VMM or the second VMM.
3 . The system of claim 1 , wherein each adder tree of the plurality of adder trees shares at least one adder with another adder tree of the plurality of adder trees.
4 . The system of claim 3 , wherein the selection logic further comprises a plurality of output multiplexers, the plurality of output multiplexers being coupled to the at least one adder.
5 . The system of claim 4 , wherein the plurality of output multiplexers are configured to select between a first result and a second result based on whether the compute logic performs the first VMM or the second VMM.
6 . The system of claim 1 , wherein the memory and the compute logic are configured as a compute-in-memory (CIM) hardware module.
7 . The system of claim 1 , wherein
the memory comprises a plurality of memory cells; and
the plurality of memory cells comprise digital Static Random Access Memory (SRAM) memory cells and analog SRAM memory cells.
8 . The system of claim 1 , wherein the selection logic is further configured to:
receive a control signal indicating whether the compute logic is performing the first VMM or the second VMM; and
provide the control signal to the plurality of input multiplexers to control routing of elements of the input vector.
9 . The system of claim 1 , wherein
the plurality of input multiplexers are arranged to correspond to respective multiplication circuitry, and
each input multiplexer routes a respective element of the input vector to its corresponding multiplication circuitry based on whether the compute logic is performing the first VMM or the second VMM.
10 . A hardware accelerator, comprising:
at least one processor; and
at least one compute tile coupled to the at least one processor, the at least one compute tile including a memory storing a plurality of weights corresponding to a matrix, and compute logic coupled to the memory and configured to selectively perform a first vector-matrix multiplication (VMM) using the matrix and a second VMM using a transpose of the matrix, wherein the compute logic comprises:
multiplication circuitry coupled to the memory, the multiplication circuitry configured to output results corresponding to multiplications of at least a portion of the plurality of weights and at least a portion of input vectors;
a plurality of adder trees configured to accumulate the results; and
selection logic comprising a plurality of input multiplexers configured to route elements of the input vector to the multiplication circuitry,
wherein the selection logic is configured to:
select a first portion of the plurality of adder trees and provide the results to the first portion in response to the compute logic performing the first VMM; and
select a second portion of the plurality of adder trees and provide the results to the second portion in response to the compute logic performing the second VMM.
11 . The hardware accelerator of claim 10 , wherein the selection logic further comprises input logic configured to provide the portion of the input vector to the multiplication circuitry in response to the compute logic performing the first VMM or the second VMM.
12 . The hardware accelerator of claim 10 , wherein each adder tree of the plurality of adder trees shares at least one adder with another adder tree of the plurality of adder trees.
13 . The hardware accelerator of claim 12 , wherein the selection logic further comprises a plurality of output multiplexers, the plurality of output multiplexers being coupled to the at least one adder.
14 . The hardware accelerator of claim 10 , wherein the selection logic is further configured to:
receive a control signal indicating whether the compute logic is performing the first VMM or the second VMM; and
provide the control signal to the plurality of input multiplexers to control routing of elements of the input vector.
15 . The hardware accelerator of claim 10 , wherein
the memory comprises a plurality of memory cells; and
the plurality of memory cells comprise digital Static Random Access Memory (SRAM) memory cells and analog SRAM memory cells.
16 . The hardware accelerator of claim 10 , wherein the memory and the compute logic of each of the at least one compute tile are configured as a compute-in-memory (CIM) hardware module.
17 . A method, comprising:
providing an input vector to compute logic coupled to a processor, the compute logic being configured as a compute-in-memory (CIM) hardware module, the CIM hardware module comprising a memory storing a plurality of weights corresponding to a matrix, and the compute logic coupled to the memory and configured to selectively perform a first vector-matrix multiplication (VMM) using the matrix and a second VMM using a transpose of the matrix, wherein the compute logic comprises:
multiplication circuitry coupled to the memory, the multiplication circuitry configured to output results corresponding to multiplications of at least a portion of the plurality of weights and at least a portion of input vectors;
a plurality of adder trees configured to accumulate the results; and
selection logic comprising a plurality of input multiplexers configured to route elements of the input vector to the multiplication circuitry, wherein the selection logic is configured to:
select a first portion of the plurality of adder trees and provide the results to the first portion in response to the compute logic performing the first VMM; and
select a second portion of the plurality of adder trees and provide the results to the second portion in response to the compute logic performing the second VMM;
selecting the first VMM or the second VMM; and
performing the first VMM or the second VMM based on the selecting.
18 . The method of claim 17 , wherein each adder tree of the plurality of adder trees shares at least one adder with another adder tree of the plurality of adder trees.
19 . The method of claim 17 , further comprising:
receiving, by the selection logic, a control signal indicating whether the compute logic is performing the first VMM or the second VMM; and
providing, by the selection logic, the control signal to the plurality of input multiplexers to control routing of elements of the input vector.
20 . The method of claim 17 , further comprising providing the input vector to an input buffer prior to providing the input vector to the compute logic, the input buffer temporarily storing elements of the input vector.