IP Library › Granted Patent US 12,438,553
Granted Patent B2
US 12,438,553 · App. 18/465,495 · Granted Oct 7, 2025

Methods, systems, articles of manufacture, and apparatus to decode zero-value-compression data vectors

Inventors: Gautham Chinya (Sunnyvale, CA); Debabrata Mohapatra (San Jose, CA); Arnab Raha (San Jose, CA); Huichu Liu (Santa Clara, CA); Cormac Brick (San Francisco, CA)
Assignee: Intel Corporation
H03M7/3082G06F16/2237G06N3/063G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,438,553
App. No.
18/465,495
Granted
Oct 7, 2025
Kind
B2
Abstract

Methods, systems, articles of manufacture, and apparatus are disclosed to decode zero-value-compression data vectors. An example apparatus includes: a buffer monitor to monitor a buffer for a header including a value indicative of compressed data; a data controller to, when the buffer includes compressed data, determine a first value of a sparse select signal based on (1) a select signal and (2) a first position in a sparsity bitmap, the first value of the sparse select signal corresponding to a processing element that is to process a portion of the compressed data; and a write controller to, when the buffer includes compressed data, determine a second value of a write enable signal based on (1) the select signal and (2) a second position in the sparsity bitmap, the second value of the write enable signal corresponding to the processing element that is to process the portion of the compressed data.

Claims (73)

1. An apparatus for performing a multiply-accumulate (MAC) operation, the apparatus comprising:

a memory to:

store a compressed tensor, the compressed tensor comprising one or more nonzero-valued elements in a tensor associated with the MAC operation, the tensor associated with the MAC operation further comprises one or more zero-valued elements, and

store a sparsity bitmap, the sparsity bitmap encoding one or more positions of the one or more nonzero-valued elements in the tensor associated with the MAC operation;

a processing element to perform the MAC operation using a nonzero-valued element stored in the memory; and

a multiplexer associated with the processing element, the multiplexer to:

receive a signal generated based on the sparsity bitmap,

select, based on the signal, the nonzero-valued element from the one or more non-zero elements, and

transmit the nonzero-valued element from the memory to the processing element.

2. The apparatus of claim 1 , further comprising:

an additional processing element to perform the MAC operation using an additional nonzero-valued element stored in the memory; and

an additional multiplexer associated with the additional processing element, the additional multiplexer to:

receive an additional signal generated based on the sparsity bitmap,

select, based on the additional signal, the additional nonzero-valued element from the one or more nonzero-valued elements, and

transmit the additional nonzero-valued element from the memory to the additional processing element.

3. The apparatus of claim 1 , wherein the tensor associated with the MAC operation is an activation tensor.

4. The apparatus of claim 3 , further comprising:

another memory to:

store a compressed weight tensor, the compressed weight tensor comprising one or more nonzero-valued weights in a weight tensor, the weight tensor further comprises one or more zero-valued weights, and

store a weight sparsity bitmap, the weight sparsity bitmap encoding one or more positions of the one or more nonzero-valued weights in the weight tensor,

wherein:

a nonzero-valued weight is selected from the one or more nonzero-valued weights based on the weight sparsity bitmap and transmitted to the processing element, and

the processing element is to perform the MAC operation using the nonzero-valued weight.

5. The apparatus of claim 1 , wherein the tensor associated with the MAC operation is a weight tensor.

6. The apparatus of claim 5 , further comprising:

another memory to:

store a compressed activation tensor, the compressed activation tensor comprising one or more nonzero-valued activations in an activation tensor, the activation tensor further comprises one or more zero-valued activations, and

store an activation sparsity bitmap, the activation sparsity bitmap encoding one or more positions of the one or more nonzero-valued activation in the activation tensor;

wherein:

a nonzero-valued activation is selected from the one or more nonzero-valued activations based on the activation sparsity bitmap and transmitted to the processing element, and

the processing element is to perform the MAC operation using the nonzero-valued activation.

7. The apparatus of claim 1 , wherein the one or more nonzero-valued elements are stored at consecutive memory addresses of the memory.

8. The apparatus of claim 1 , wherein the sparsity bitmap and the compressed tensor are stored at consecutive memory addresses.

9. The apparatus of claim 1 , wherein the MAC operation is further associated with another tensor, and the another tensor comprises a plurality of other elements.

10. The apparatus of claim 1 , wherein the one or more nonzero-valued elements and the one or more zero-valued elements are in different channels.

11. A method for performing a multiply-accumulate (MAC) operation, the method comprising:

storing a compressed tensor in a memory, the compressed tensor comprising one or more nonzero-valued elements in a tensor associated with the MAC operation, the tensor associated with the MAC operation further comprises one or more zero-valued elements;

storing a sparsity bitmap in the memory, the sparsity bitmap encoding one or more positions of the one or more nonzero-valued elements in the tensor associated with the MAC operation;

providing a signal to a multiplexer, wherein the signal corresponds to a nonzero-valued element in the memory and is generated using the sparsity bitmap, and the multiplexer is associated with a processing element; and

transmitting, by the multiplexer based on the signal, the nonzero-valued element from the memory to the processing element, wherein the processing element is to perform the MAC operation using the nonzero-valued element.

12. The method of claim 11 , further comprising:

providing an additional signal to an additional multiplexer, wherein the additional signal corresponds to an additional nonzero-valued element in the memory and is generated using the sparsity bitmap, and the additional multiplexer is associated with an additional processing element; and

transmitting, by the additional multiplexer based on the additional signal, the additional nonzero-valued element from the memory to the additional processing element, wherein the additional processing element is to perform the MAC operation using the additional nonzero-valued element.

13. The method of claim 11 , further comprising:

storing a compressed weight tensor in another memory, the compressed weight tensor comprising one or more nonzero-valued weights in a weight tensor, the weight tensor further comprises one or more zero-valued weights; and

storing a weight sparsity bitmap in the another memory, the weight sparsity bitmap encoding one or more positions of the one or more nonzero-valued weights in the weight tensor,

wherein:

a nonzero-valued weight is selected from the one or more nonzero-valued weights based on the weight sparsity bitmap and transmitted to the processing element,

the processing element is to perform the MAC operation using the nonzero-valued weight, and

the tensor associated with the MAC operation is an activation tensor.

14. The method of claim 11 , further comprising:

storing a compressed activation tensor in another memory, the compressed activation tensor comprising one or more nonzero-valued activations in an activation tensor, the activation tensor further comprises one or more zero-valued activations; and

storing an activation sparsity bitmap in the another memory, the activation sparsity bitmap encoding one or more positions of the one or more nonzero-valued activation in the activation tensor,

wherein:

a nonzero-valued activation is selected from the one or more nonzero-valued activations based on the activation sparsity bitmap and transmitted to the processing element,

the processing element is to perform the MAC operation using the nonzero-valued activation, and

the tensor associated with the MAC operation is a weight tensor.

15. The method of claim 11 , wherein the one or more nonzero-valued elements are stored at consecutive memory addresses of the memory.

16. The method of claim 11 , wherein the sparsity bitmap and the compressed tensor are stored at consecutive memory addresses.

17. The method of claim 11 , wherein the MAC operation is further associated with another tensor, and the another tensor comprises a plurality of other elements.

18. The method of claim 11 , wherein the one or more nonzero-valued elements and the one or more zero-valued elements are in different channels.

19. One or more non-transitory computer-readable media storing instructions executable to perform operations for performing a multiply-accumulate (MAC) operation, the operations comprising:

storing a compressed tensor in a memory, the compressed tensor comprising one or more nonzero-valued elements in a tensor associated with the MAC operation, the tensor associated with the MAC operation further comprises one or more zero-valued elements;

storing a sparsity bitmap in the memory, the sparsity bitmap encoding one or more positions of the one or more nonzero-valued elements in the tensor associated with the MAC operation;

providing a signal to a multiplexer, wherein the signal corresponds to a nonzero-valued element in the memory and is generated using the sparsity bitmap, and the multiplexer is associated with a processing element; and

transmitting, by the multiplexer based on the signal, the nonzero-valued element from the memory to the processing element, wherein the processing element is to perform the MAC operation using the nonzero-valued element.

20. The one or more non-transitory computer-readable media of claim 19 , wherein the operations further comprise:

storing a compressed weight tensor in another memory, the compressed weight tensor comprising one or more nonzero-valued weights in a weight tensor, the weight tensor further comprises one or more zero-valued weights; and

storing a weight sparsity bitmap in the another memory, the weight sparsity bitmap encoding one or more positions of the one or more nonzero-valued weights in the weight tensor,

wherein:

a nonzero-valued weight is selected from the one or more nonzero-valued weights based on the weight sparsity bitmap and transmitted to the processing element,

the processing element is to perform the MAC operation using the nonzero-valued weight, and

the tensor associated with the MAC operation is an activation tensor.

Continuity (2)
Continuation 16832804 · Mar 27, 2020
Related Publication 20240022259A1 · Jan 18, 2024
References Cited (13)
US 10594338B1 · Lew · 2020 [cited by examiner]
US 10637500B2 · Chen et al. · 2020 [cited by applicant]
US 20120254592A1 · San Adrian et al. · 2012 [cited by applicant]
US 20150370667A1 · Bonas · 2015 [cited by examiner]
CN 104008533A · 2014 [cited by applicant]
CN 107944555A · 2018 [cited by applicant]
TW I550512B · 2016 [cited by applicant]
TW 201915835A · 2019 [cited by applicant]
Aimar et al., “NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps”, arXiv:1706.01406v2 [cs.CV] Mar. 6, 2018, Doi: 10.1109/TNNLS.2018.2852335, 13 pages. [cited by applicant]
Minsoo et al., “Compressing DMA Engine: Leveraging Activation Sparsity for Training Deep Neural Networks”, 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA), arXiv:1705.01626v1 [cs.LG] M… [cited by applicant]
Parashar et al., SCNN: An accelerator for compressed-sparse convolutional neural networks, Proceedings of the 44th Annual International Symposium on Computer Architecture, arXiv:1708.04485v1 [cs.NE] May 23, 2017, ISCA, … [cited by applicant]
Udit et al., “MSR: A Modular Accelerator for Sparse RNNs”, 2019 28th International Conference on Parallel Architectures and Compilation Techniques (PACT), arXiv: 1908.08976v1 [eess.SP] Aug. 23, 2019, IEEE, Sep. 23, 2019… [cited by applicant]
Yu et al., Spring: A Sparsity-Aware Reduced-Precision Monolithic 3D CNN Accelerator Architecture for Training and Inference, arXiv:1909.00557v2 [cs.ar], Feb. 3, 2020, 14 pages. [cited by applicant]