TILED IN-MEMORY COMPUTING ARCHITECTURE
A compute tile is described. The compute tile includes compute engines and a general-purpose (GP processor coupled with the compute engines. Each of the compute engines includes a compute-in-memory (CIM) hardware module. The CIM hardware module is configured to store weights corresponding to a matrix and to perform a vector-matrix multiplication (VMM) for the matrix. The GP processor is configured to control the compute engines, to receive output of the VMM for the matrix from the compute engines, and to perform a nonlinear operation on the output. The compute engines are addressable by data movement initiators. Data may be moved to and/or from the compute engines in data paths that bypass the GP processor.
1 . A computing device comprising:
a first compute tile comprising at least one compute engine configured to generate data in a first format, wherein the at least one compute engine corresponds to a compute-in-memory (CIM) module;
a second compute tile comprising at least one additional compute engine configured to generate additional data in a second format, wherein the at least one additional compute engine corresponds to an additional CIM module; and
a data conversion engine configured to convert data transferred from the at least one compute engine of the first compute tile to the at least one additional compute engine of the second compute tile from the first format to the second format.
2 . The computing device of claim 1 , wherein the data conversion engine comprises a reshape engine configured to perform padding on the data during transfer from the first compute tile to the second compute tile.
3 . The computing device of claim 1 , wherein:
the data conversion engine is configured to perform an im2col transformation on the data in the first format during the transfer to generate an im2col-formatted activation;
and the data conversion engine comprises a gather engine configured to form im2col data on-the-fly during the transfer.
4 . The computing device of claim 1 , wherein the data conversion engine is configured to perform a transpose operation on the data during the transfer from the first compute tile to the second compute tile.
5 . The computing device of claim 1 , wherein the data conversion engine comprises a buffer configured to augment the data while the data in the first format is in-flight between tiles.
6 . The computing device of claim 1 , further comprising a direct memory access (DMA) unit, wherein: the DMA unit is configured to orchestrate movement of the data between the first compute tile and the second compute tile, and the data conversion engine is configured to perform at least one of padding or data re shaping while the data is moved by the DMA unit.
7 . The computing device of claim 1 , wherein:
the data conversion engine is disposed on at least one of a mesh_out connection or a mesh_in connection;
the data conversion engine comprises a BFloat-to-integer format converter; and
the first format and the second format comprise different numeric formats.
8 . The computing device of claim 1 , wherein the data conversion engine is configured to:
convert data retrieved in an integer format to a BFloat format for processing by the at least one additional compute engine; and
convert output data in the BFloat format to the integer format for storage in memory or transfer off-tile.
9 . The computing device of claim 1 , wherein the data conversion engine is configured to selectively perform one or more operations comprising padding, reshaping, transposing, gathering, and im2col based on control instructions specifying a type of data-augmentation.
10 . The computing device of claim 1 , wherein the data conversion engine comprises at least one of:
a float to integer converter;
or a reshape engine configured to perform padding operations.
11 . A method comprising:
generating, by at least one compute engine of a first compute tile, data in a first format, the at least one compute engine including a compute-in-memory (CIM) module;
transferring the data from the first compute tile to a second compute tile;
converting, by a data conversion engine, the data from the first format to a second format during the transfer;
providing the converted data to at least one additional compute engine of the second compute tile, the at least one additional compute engine including an additional CIM module; and
generating, by the at least one additional compute engine, additional data in the second format based on the converted data.
12 . The method of claim 11 , wherein converting comprises reshaping the data to perform padding while the data is transferred between the first compute tile and the second compute tile.
13 . The method of claim 11 , wherein:
converting comprises performing an im2col transformation on the data during the transfer to generate an im2col-formatted activation; and
performing the im2col transformation comprises gathering elements of the data on-the-fly during the transfer.
14 . The method of claim 11 , wherein converting comprises:
at least one of: performing a transpose operation on the data during the transfer; or
buffering the data while the data is in-flight between tiles.
15 . The method of claim 11 , further comprising orchestrating the transfer using a direct memory access (DMA) unit, wherein converting comprises performing at least one of padding or data reshaping while the DMA unit moves the data.
16 . The method of claim 11 , further comprising converting output data from the at least one additional compute engine from the BFloat format to the integer format for storage in memory or transfer off-tile.
17 . The method of claim 11 , wherein converting comprises selectively performing one or more operations including padding, reshaping, transposing, gathering, and im2col based on control instructions specifying a type of data augmentation to apply during the transfer.
18 . A device, comprising:
a plurality of compute tiles comprising a plurality of compute engines comprising compute-in-memory (CIM) modules;
a general-purpose (GP) processor coupled to the plurality of compute engines; and
a data conversion engine coupled to the plurality of compute engines, the data conversion engine being configured to convert data transferred from one of the plurality of compute tiles, the data conversion engine being configured to convert the data from a first format to a second format.
19 . The device of claim 18 , further comprising:
a data movement initiator configured to transfer the converted data to the plurality of compute engines via a data path that bypasses the GP processor, wherein: the plurality of compute engines is addressable by both the data movement initiator and the GP processor; and
the data movement initiator is configured to transfer data while bypassing the GP processor by transferring the data via buses coupled directly to interconnects coupled to the plurality of compute engines
20 . The device of claim 18 , wherein the plurality of compute tiles further comprise:
a local memory coupled with the plurality of compute engines and the GP processor; and
direct memory access (DMA) unit configured to transfer data between the local memory and the plurality of compute engines in a data path that bypasses the GP processor.