IP Library › Granted Patent US 12,632,509
Granted Patent B2
US 12,632,509 · App. 17/470,470 · Granted May 19, 2026

Nibble block format

Inventors: Paul Nicholas Whatmough (Cambridge, MA); Zhi-Gang Liu (Westford, MA); Matthew Mattina (Boylston, MA)
Assignee: Arm Limited
G06F17/16G06F7/50G06F7/523G06F7/5443G06F9/5027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,509
App. No.
17/470,470
Granted
May 19, 2026
Kind
B2
Abstract

A matrix multiplication system and method are provided. The system includes a memory that stores one or more weight tensors, a processor and a matrix multiply accelerator (MMA). The processor converts each weight tensor into an encoded block set that is stored in the memory. Each encoded block set includes a number of encoded blocks, and each encoded block includes a data field and an index field. The MMA converts each encoded block set into a reconstructed weight tensor, and convolves each reconstructed weight tensor and an input data tensor to generate an output data matrix.

Claims (145)

1 . A system, comprising:

a first memory configured to store one or more encoded block sets corresponding to one or more weight tensors, each weight tensor including a number of n-bit integer weights;

a matrix multiply accelerator including a second memory;

a processor, operatively coupled to the first memory and the second memory, the processor configured to:

store an encoded block set from the first memory to the second memory, the encoded block including a data field and an index field, the data field including a number of encoded weights, the index field including an index associated with each of the plurality of n-bit integer weights, where:

the index field includes a plurality of indices, each index indicating an encoding type of an associated weight of the plurality of n-bit integer weights, the encoding type selected from:

a zero magnitude type, indicating a weight with value zero;

a small magnitude type, indicating a weight with zero-valued most significant n/2 bits; and

a large magnitude type, indicating a weight with zero-valued least significant n/2 bits; and

the data field includes:

the most significant n/2 bits of each weight of the large magnitude encoding type; and

the least significant n/2 bits of each weight of the small magnitude encoding type;

where the matrix multiply accelerator (MMA) is configured to:

read, from the second memory, an encoding type of a weight from the index field of an encoded block;

when the encoding type is large magnitude encoding type or the small magnitude encoding type:

read n/2 bits from the data field of the encoded block; and

reconstruct an n-bit integer weight from the n/2 bits; and

when the encoding type is the zero magnitude encoding type:

reconstruct an n-bit integer weight with value zero.

2 . The system according to claim 1 , where the processor is further configured to:

for each weight tensor:

generate, based on the weight tensor, a basic block matrix set including a number of basic block matrices, each basic block matrix including a number of weights,

generate, based on the basic block matrix set, an encoded block set, the encoded block set including a number of encoded blocks, each encoded block including a data field and an index field, the data field including a number of encoded weights, the index field including an index associated with each weight in the basic block matrix, the number of encoded weights being less than the number of weights in the basic block matrix, each encoded block having a same size, and

store the encoded block set in the memory;

and where the matrix multiply accelerator (MMA) is further configured to:

convert each encoded block set into a reconstructed weight tensor having a number of weights equal to the number of weights of the respective weight tensor, and

convolve each reconstructed weight tensor and an input data tensor to generate an output data matrix,

the MMA further including a controller and an array of processing elements (PEs) coupled to the controller, where the controller of the MMA is configured to:

convert the reconstructed weight tensor to a converted weight matrix, and

convert the input data tensor to a converted input data matrix; and

the array of PEs of the MMA configured to multiply the converted weight matrix and the converted input data matrix, each PE including a multiply-and-accumulate (MAC) circuit configured to generate a dot product between one row of the converted weight matrix and one column of the converted input data matrix.

3 . The system according to claim 2 , where:

each weight tensor has a height, a width and a depth equal to a number of input channels;

each basic block matrix has a width of 1 and a height equal to the number of input channels; and

each basic block matrix includes one weight from each input channel.

4 . The system according to claim 2 , where said generate the encoded block set includes:

for each weight in the basic block matrix:

determine an encoding type for the weight;

generate an index for the weight based on the encoding type;

generate an encoded weight based on the encoding type and the weight;

add the index to the index field of the encoded block; and

when the encoding type is not a zero magnitude weight type, add the encoded weight to the data field of the encoded block.

5 . The system according to claim 4 , where:

said determine an encoding type for the weight is based on a lower threshold value and an upper threshold value.

6 . The system according to claim 5 , where said determine an encoding type for the weight includes:

select the zero magnitude weight type when the weight has a zero value or the weight has a non-zero value that is less than or equal to the lower threshold value;

select the small magnitude weight type when the weight has a non-zero value that is greater than the lower threshold value and less than or equal to the upper threshold value; and

select the large magnitude weight type when the weight has a non-zero value that is greater than the upper threshold value.

7 . The system according to claim 2 , where:

each weight in the reconstructed weight tensor has a corresponding weight in the respective weight tensor that has a same value;

said convert each encoded block set into a reconstructed weight tensor includes:

generate, based on the encoded block set, the basic block matrix set, and

generate, based on the basic block matrix set, the reconstructed weight tensor;

said convolve each reconstructed weight tensor and an input data tensor includes:

convert the reconstructed weight tensor, based on a convolution operation, to the converted weight matrix,

convert the input data tensor, based on the convolution operation, to the converted input data matrix, and

multiply the converted weight matrix and the converted input data matrix to generate the output data matrix.

8 . The system according to claim 1 , where:

the encoding type includes a full magnitude weight type that is an n-bit integer element.

9 . A system comprising:

a memory configured to store one or more weight tensors, each weight tensor including a number of weights;

a processor, coupled to the memory, configured to:

for each weight tensor:

generate, based on the weight tensor, a basic block matrix set including a number of basic block matrices, each basic block matrix including a number of weights,

generate, based on the basic block matrix set, an encoded block set, the encoded block set including a number of encoded blocks, each encoded block including a data field and an index field, the data field including a number of encoded weights, the index field including an index associated with each weight in the basic block matrix, the number of encoded weights being less than the number of weights in the basic block matrix, each encoded block having a same size, and

store the encoded block set in the memory; and

a matrix multiply accelerator (MMA), coupled to the processor and the memory, configured to:

convert each encoded block set into a reconstructed weight tensor having a number of weights equal to the number of weights of the respective weight tensor, and

convolve each reconstructed weight tensor and an input data tensor to generate an output data matrix,

where each weight in the reconstructed weight tensor has a corresponding weight in the respective weight tensor that has a same value;

said convert each encoded block set into a reconstructed weight tensor includes:

generate, based on the encoded block set, the basic block matrix set, and

generate, based on the basic block matrix set, the reconstructed weight tensor;

said convolve each reconstructed weight tensor and an input data tensor includes:

convert the reconstructed weight tensor, based on a convolution operation, to a converted weight matrix,

convert the input data tensor, based on the convolution operation, to a converted input data matrix, and

multiply the converted weight matrix and the converted input data matrix to generate the output data matrix,

where the MMA includes:

a second memory;

a controller configured to:

convert the reconstructed weight tensor to the converted weight matrix, and

convert the input data tensor to the converted input data matrix;

a first register configured to store at least a portion of the converted input data matrix;

a second register configured to store at least a portion of the converted weight matrix;

a third register configured to store at least a portion of the output data matrix; and

an array of processing elements (PEs), coupled to the controller and the first, second and third registers, configured to multiply the converted weight matrix and the converted input data matrix, each PE including a multiply-and-accumulate (MAC) circuit configured to generate a dot product between one row of the converted weight matrix and one column of the converted input data matrix.

10 . A computer-based method, comprising:

at a processor coupled to a first memory storing one or more encoded block sets generated from one or more weight tensors:

store an encoded block set from the first memory to a second memory of a matrix multiply accelerator (MMA), the encoded block set generated from a weight tensor,

the encoded block set including a number of encoded blocks, each encoded block including a data field and an index field, the data field including a number of encoded weights, the index field including an index associated with each of a plurality of n-bit weights of the weight tensor, where

each weight tensor includes n-bit integer elements;

the index indicates an encoding type of the associated weight, the encoding type including a zero magnitude type, a small magnitude type and a large magnitude type,

the zero magnitude type is an n/2-bit integer element;

the small magnitude type is an n/2-bit integer element; and

the large magnitude type is an n/2-bit integer element;

at the matrix multiply accelerator (MMA):

reading, from the second memory, an encoding type of a weight from the index field of an encoded block;

when the encoding type is large magnitude encoding type or the small magnitude encoding type:

reading n/2 bits from the data field of the encoded block; and

reconstructing an n-bit integer weight from the n/2 bits; and

when the encoding type is the zero magnitude encoding type:

reconstructing an n-bit integer weight with value zero.

11 . The computer-based method according to claim 10 , further comprising:

for each weight tensor of one or more weight tensors:

generating, based on the weight tensor, a basic block matrix set including a number of basic block matrices, each basic block matrix including a number of weights; and

generating the encoded block set based on the basic block matrix set,

where:

each weight tensor has a height, a width and a depth equal to a number of input channels;

each basic block matrix has a width of 1 and a height equal to the number of input channels; and

each basic block matrix includes one weight from each input channel.

12 . The computer-based method according to claim 10 , further comprising at the MMA:

converting each encoded block set in the second memory into a reconstructed weight tensor having a number of weights equal to the number of weights of the respective weight tensor,

convolving each reconstructed weight tensor and an input data tensor to generate an output data matrix, at a controller of the MMA:

converting the reconstructed weight tensor to the converted weight matrix, and

converting the input data tensor to the converted input data matrix, and at an array of processing elements (PEs) of the MMA:

multiplying the converted weight matrix and the converted input data matrix, with each PE of the array of PEs including a multiply-and-accumulate (MAC) circuit, the MAC circuit of each PE generating a dot product between one row of the converted weight matrix and one column of the converted input data matrix.

13 . The computer-based method according to claim 10 , where:

the encoding type includes a full magnitude weight type that is an n-bit integer element.

14 . The computer-based method according to claim 10 , further comprising:

for each weight tensor of one or more weight tensors:

generating, based on the weight tensor, a basic block matrix set including a number of basic block matrices, each basic block matrix including a number of weights; and

generating the encoded block set based on the basic block matrix set,

where generating an encoded block of the encoded block set includes:

for each weight in the basic block matrix of the basic block matrix set:

determining an encoding type for the weight;

generating an index for the weight based on the encoding type;

generating an encoded weight based on the encoding type and the weight;

adding the index to the index field of the encoded block; and

when the encoding type is not a zero magnitude weight type, adding the encoded weight to the data field of the encoded block.

15 . The computer-based method according to claim 14 , where:

said determining an encoding type for the weight is based on a lower threshold value and an upper threshold value.

16 . The computer-based method according to claim 15 , where said determining an encoding type for the weight includes:

selecting the zero magnitude weight type when the weight has a zero value or the weight has a non-zero value that is less than or equal to the lower threshold value;

selecting the small magnitude weight type when the weight has a non-zero value that is greater than the lower threshold value and less than or equal to the upper threshold value; and

selecting the large magnitude weight type when the weight has a non-zero value that is greater than the upper threshold value.

17 . The computer-based method according to claim 16 , where:

each weight in the reconstructed weight tensor has a corresponding weight in the respective weight tensor that has a same value.

18 . The computer-based method according to claim 16 , where:

said converting each encoded block set into a reconstructed weight tensor includes:

generating, based on the encoded block set, the basic block matrix set, and

generating, based on the basic block matrix set, the reconstructed weight tensor;

said convolving each reconstructed weight tensor and the input data tensor includes:

converting the reconstructed weight tensor, based on a convolution operation, to a converted weight matrix,

converting the input data tensor, based on the convolution operation, to a converted input data matrix, and

multiplying the converted weight matrix and the converted input data matrix to generate the output data matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2021
From: WHATMOUGH, PAUL NICHOLAS; LIU, ZHI-GANG; MATTINA, MATTHEW
To: ARM LIMITED
Reel/Frame 057720/0510 →
Continuity (1)
Related Publication 20230076138A1 · Mar 9, 2023
References Cited (12)
US 5200828A · Jang · 1993 [cited by examiner]
US 7599975B1 · Donovan · 2009 [cited by examiner]
US 11977885B2 · Maiyuran · 2024 [cited by examiner]
US 20060038879A1 · Kremen · 2006 [cited by examiner]
US 20190182484A1 · Kalevo · 2019 [cited by examiner]
US 20200342632A1 · Frumkin · 2020 [cited by examiner]
US 20210150770A1 · Appu et al. · 2021 [cited by applicant]
A.J. Hussain, Ali Al-Fayadh, Naeem Radi, “Image compression techniques: a survey in lossless and lossy algorithms”, Mar. 9, 2018, Elsevier B.V., Neurocomputing 300 (2018) 44-69 (Year: 2018). [cited by examiner]
Pothos et al., “Deep Learning Inference with Dynamic Graphs on Heterogeneous Platforms”, Int J Parallel Prog 49, 158â176. https://doi.org/10.1007/s10766-020-00654-2 (Year: 2020). [cited by examiner]
Bratt, Ian, “Arm's First-Generation Machine Learning Processor”, Hot Chips conference, Cupertino, CA, Aug. 19-20, 2018. [cited by applicant]
Liu et al., “Systolic Tensor Array: an Efficient Structured-Sparse GEMM Accelerator for Mobile CNN Inference”, IEEE Computer Architecture Letters, https://arxiv.org/pdf/2005.08098.pdf, Mar. 12, 2020. [cited by applicant]
Liu et al., “Layerwise Sparse Coding for Pruned Deep Neural Networks with Extreme Compression Ratio,” Proceedings of the AAAI Conference on Artificial Intelligence, 34(04), 2020, 4900-4907, https://doi.org/10.1609/aaai.… [cited by applicant]