Nibble block format
A matrix multiplication system and method are provided. The system includes a memory that stores one or more weight tensors, a processor and a matrix multiply accelerator (MMA). The processor converts each weight tensor into an encoded block set that is stored in the memory. Each encoded block set includes a number of encoded blocks, and each encoded block includes a data field and an index field. The MMA converts each encoded block set into a reconstructed weight tensor, and convolves each reconstructed weight tensor and an input data tensor to generate an output data matrix.
1 . A system, comprising:
a first memory configured to store one or more encoded block sets corresponding to one or more weight tensors, each weight tensor including a number of n-bit integer weights;
a matrix multiply accelerator including a second memory;
a processor, operatively coupled to the first memory and the second memory, the processor configured to:
store an encoded block set from the first memory to the second memory, the encoded block including a data field and an index field, the data field including a number of encoded weights, the index field including an index associated with each of the plurality of n-bit integer weights, where:
the index field includes a plurality of indices, each index indicating an encoding type of an associated weight of the plurality of n-bit integer weights, the encoding type selected from:
a zero magnitude type, indicating a weight with value zero;
a small magnitude type, indicating a weight with zero-valued most significant n/2 bits; and
a large magnitude type, indicating a weight with zero-valued least significant n/2 bits; and
the data field includes:
the most significant n/2 bits of each weight of the large magnitude encoding type; and
the least significant n/2 bits of each weight of the small magnitude encoding type;
where the matrix multiply accelerator (MMA) is configured to:
read, from the second memory, an encoding type of a weight from the index field of an encoded block;
when the encoding type is large magnitude encoding type or the small magnitude encoding type:
read n/2 bits from the data field of the encoded block; and
reconstruct an n-bit integer weight from the n/2 bits; and
when the encoding type is the zero magnitude encoding type:
reconstruct an n-bit integer weight with value zero.
2 . The system according to claim 1 , where the processor is further configured to:
for each weight tensor:
generate, based on the weight tensor, a basic block matrix set including a number of basic block matrices, each basic block matrix including a number of weights,
generate, based on the basic block matrix set, an encoded block set, the encoded block set including a number of encoded blocks, each encoded block including a data field and an index field, the data field including a number of encoded weights, the index field including an index associated with each weight in the basic block matrix, the number of encoded weights being less than the number of weights in the basic block matrix, each encoded block having a same size, and
store the encoded block set in the memory;
and where the matrix multiply accelerator (MMA) is further configured to:
convert each encoded block set into a reconstructed weight tensor having a number of weights equal to the number of weights of the respective weight tensor, and
convolve each reconstructed weight tensor and an input data tensor to generate an output data matrix,
the MMA further including a controller and an array of processing elements (PEs) coupled to the controller, where the controller of the MMA is configured to:
convert the reconstructed weight tensor to a converted weight matrix, and
convert the input data tensor to a converted input data matrix; and
the array of PEs of the MMA configured to multiply the converted weight matrix and the converted input data matrix, each PE including a multiply-and-accumulate (MAC) circuit configured to generate a dot product between one row of the converted weight matrix and one column of the converted input data matrix.
3 . The system according to claim 2 , where:
each weight tensor has a height, a width and a depth equal to a number of input channels;
each basic block matrix has a width of 1 and a height equal to the number of input channels; and
each basic block matrix includes one weight from each input channel.
4 . The system according to claim 2 , where said generate the encoded block set includes:
for each weight in the basic block matrix:
determine an encoding type for the weight;
generate an index for the weight based on the encoding type;
generate an encoded weight based on the encoding type and the weight;
add the index to the index field of the encoded block; and
when the encoding type is not a zero magnitude weight type, add the encoded weight to the data field of the encoded block.
5 . The system according to claim 4 , where:
said determine an encoding type for the weight is based on a lower threshold value and an upper threshold value.
6 . The system according to claim 5 , where said determine an encoding type for the weight includes:
select the zero magnitude weight type when the weight has a zero value or the weight has a non-zero value that is less than or equal to the lower threshold value;
select the small magnitude weight type when the weight has a non-zero value that is greater than the lower threshold value and less than or equal to the upper threshold value; and
select the large magnitude weight type when the weight has a non-zero value that is greater than the upper threshold value.
7 . The system according to claim 2 , where:
each weight in the reconstructed weight tensor has a corresponding weight in the respective weight tensor that has a same value;
said convert each encoded block set into a reconstructed weight tensor includes:
generate, based on the encoded block set, the basic block matrix set, and
generate, based on the basic block matrix set, the reconstructed weight tensor;
said convolve each reconstructed weight tensor and an input data tensor includes:
convert the reconstructed weight tensor, based on a convolution operation, to the converted weight matrix,
convert the input data tensor, based on the convolution operation, to the converted input data matrix, and
multiply the converted weight matrix and the converted input data matrix to generate the output data matrix.
8 . The system according to claim 1 , where:
the encoding type includes a full magnitude weight type that is an n-bit integer element.
9 . A system comprising:
a memory configured to store one or more weight tensors, each weight tensor including a number of weights;
a processor, coupled to the memory, configured to:
for each weight tensor:
generate, based on the weight tensor, a basic block matrix set including a number of basic block matrices, each basic block matrix including a number of weights,
generate, based on the basic block matrix set, an encoded block set, the encoded block set including a number of encoded blocks, each encoded block including a data field and an index field, the data field including a number of encoded weights, the index field including an index associated with each weight in the basic block matrix, the number of encoded weights being less than the number of weights in the basic block matrix, each encoded block having a same size, and
store the encoded block set in the memory; and
a matrix multiply accelerator (MMA), coupled to the processor and the memory, configured to:
convert each encoded block set into a reconstructed weight tensor having a number of weights equal to the number of weights of the respective weight tensor, and
convolve each reconstructed weight tensor and an input data tensor to generate an output data matrix,
where each weight in the reconstructed weight tensor has a corresponding weight in the respective weight tensor that has a same value;
said convert each encoded block set into a reconstructed weight tensor includes:
generate, based on the encoded block set, the basic block matrix set, and
generate, based on the basic block matrix set, the reconstructed weight tensor;
said convolve each reconstructed weight tensor and an input data tensor includes:
convert the reconstructed weight tensor, based on a convolution operation, to a converted weight matrix,
convert the input data tensor, based on the convolution operation, to a converted input data matrix, and
multiply the converted weight matrix and the converted input data matrix to generate the output data matrix,
where the MMA includes:
a second memory;
a controller configured to:
convert the reconstructed weight tensor to the converted weight matrix, and
convert the input data tensor to the converted input data matrix;
a first register configured to store at least a portion of the converted input data matrix;
a second register configured to store at least a portion of the converted weight matrix;
a third register configured to store at least a portion of the output data matrix; and
an array of processing elements (PEs), coupled to the controller and the first, second and third registers, configured to multiply the converted weight matrix and the converted input data matrix, each PE including a multiply-and-accumulate (MAC) circuit configured to generate a dot product between one row of the converted weight matrix and one column of the converted input data matrix.
10 . A computer-based method, comprising:
at a processor coupled to a first memory storing one or more encoded block sets generated from one or more weight tensors:
store an encoded block set from the first memory to a second memory of a matrix multiply accelerator (MMA), the encoded block set generated from a weight tensor,
the encoded block set including a number of encoded blocks, each encoded block including a data field and an index field, the data field including a number of encoded weights, the index field including an index associated with each of a plurality of n-bit weights of the weight tensor, where
each weight tensor includes n-bit integer elements;
the index indicates an encoding type of the associated weight, the encoding type including a zero magnitude type, a small magnitude type and a large magnitude type,
the zero magnitude type is an n/2-bit integer element;
the small magnitude type is an n/2-bit integer element; and
the large magnitude type is an n/2-bit integer element;
at the matrix multiply accelerator (MMA):
reading, from the second memory, an encoding type of a weight from the index field of an encoded block;
when the encoding type is large magnitude encoding type or the small magnitude encoding type:
reading n/2 bits from the data field of the encoded block; and
reconstructing an n-bit integer weight from the n/2 bits; and
when the encoding type is the zero magnitude encoding type:
reconstructing an n-bit integer weight with value zero.
11 . The computer-based method according to claim 10 , further comprising:
for each weight tensor of one or more weight tensors:
generating, based on the weight tensor, a basic block matrix set including a number of basic block matrices, each basic block matrix including a number of weights; and
generating the encoded block set based on the basic block matrix set,
where:
each weight tensor has a height, a width and a depth equal to a number of input channels;
each basic block matrix has a width of 1 and a height equal to the number of input channels; and
each basic block matrix includes one weight from each input channel.
12 . The computer-based method according to claim 10 , further comprising at the MMA:
converting each encoded block set in the second memory into a reconstructed weight tensor having a number of weights equal to the number of weights of the respective weight tensor,
convolving each reconstructed weight tensor and an input data tensor to generate an output data matrix, at a controller of the MMA:
converting the reconstructed weight tensor to the converted weight matrix, and
converting the input data tensor to the converted input data matrix, and at an array of processing elements (PEs) of the MMA:
multiplying the converted weight matrix and the converted input data matrix, with each PE of the array of PEs including a multiply-and-accumulate (MAC) circuit, the MAC circuit of each PE generating a dot product between one row of the converted weight matrix and one column of the converted input data matrix.
13 . The computer-based method according to claim 10 , where:
the encoding type includes a full magnitude weight type that is an n-bit integer element.
14 . The computer-based method according to claim 10 , further comprising:
for each weight tensor of one or more weight tensors:
generating, based on the weight tensor, a basic block matrix set including a number of basic block matrices, each basic block matrix including a number of weights; and
generating the encoded block set based on the basic block matrix set,
where generating an encoded block of the encoded block set includes:
for each weight in the basic block matrix of the basic block matrix set:
determining an encoding type for the weight;
generating an index for the weight based on the encoding type;
generating an encoded weight based on the encoding type and the weight;
adding the index to the index field of the encoded block; and
when the encoding type is not a zero magnitude weight type, adding the encoded weight to the data field of the encoded block.
15 . The computer-based method according to claim 14 , where:
said determining an encoding type for the weight is based on a lower threshold value and an upper threshold value.
16 . The computer-based method according to claim 15 , where said determining an encoding type for the weight includes:
selecting the zero magnitude weight type when the weight has a zero value or the weight has a non-zero value that is less than or equal to the lower threshold value;
selecting the small magnitude weight type when the weight has a non-zero value that is greater than the lower threshold value and less than or equal to the upper threshold value; and
selecting the large magnitude weight type when the weight has a non-zero value that is greater than the upper threshold value.
17 . The computer-based method according to claim 16 , where:
each weight in the reconstructed weight tensor has a corresponding weight in the respective weight tensor that has a same value.
18 . The computer-based method according to claim 16 , where:
said converting each encoded block set into a reconstructed weight tensor includes:
generating, based on the encoded block set, the basic block matrix set, and
generating, based on the basic block matrix set, the reconstructed weight tensor;
said convolving each reconstructed weight tensor and the input data tensor includes:
converting the reconstructed weight tensor, based on a convolution operation, to a converted weight matrix,
converting the input data tensor, based on the convolution operation, to a converted input data matrix, and
multiplying the converted weight matrix and the converted input data matrix to generate the output data matrix.