IP Library › Granted Patent US 12,321,857
Granted Patent B2
US 12,321,857 · App. 17/357,924 · Granted Jun 3, 2025

Methods and apparatus to perform machine-learning model operations on sparse accelerators

Inventors: Martin Power (Dublin, IE); Kevin Brady (Newry, GB); Niall Hanrahan (Galway, IE); Martin-Thomas Grymel (Leixlip, IE); David Bernard (Kilcullen, IE); Gary Baugh (Bray, IE)
Assignee: Intel Corporation
G06N3/08G06F17/15G06F17/153G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,857
App. No.
17/357,924
Granted
Jun 3, 2025
Kind
B2
Abstract

Methods, apparatus, systems and articles of manufacture are disclosed to perform machine-learning model operations on sparse accelerators. An example apparatus includes first circuitry, second circuitry to generate sparsity data based on an acceleration operation, and third circuitry to instruct one or more data buffers to provide at least one of activation data or weight data based on the sparsity data to the first circuitry, the first circuitry to execute the acceleration operation based on the at least one of the activation data or the weight data.

Claims (75)

1. An apparatus comprising:

first circuitry comprising one or more multiply accumulators;

a buffer to store first sparsity data, the first sparsity data indicating sparsity in input data of a first convolution type;

second circuitry to generate second sparsity data for a second convolution type;

one or more other buffers to store input data of a convolution; and

third circuitry to:

select sparsity data between the first sparsity data and the second sparsity data based on a type of a convolution,

identify one or more data elements in the input data of the convolution based on the selected sparsity data, and

transfer the one or more data elements from the one or more other data buffers to the first circuitry,

wherein the first circuitry is to execute at least part of the convolution by performing one or more multiply-accumulate operations on the one or more data elements.

2. The apparatus of claim 1 , wherein the second sparsity data includes a sparse weight vector, and the apparatus further comprises:

fourth circuitry to identify the type of the convolution based on configuration information, the sparse weight vector generated based on the configuration information,

wherein the third circuitry is further to generate a sparsity bit mask based on a sparse activation vector and the sparse weight vector and is to identify the one or more data elements by using the sparsity bit mask.

3. The apparatus of claim 1 , wherein the selected sparsity data includes at least one of activation sparsity data and weight sparsity data, the one or more data buffers include a weight data buffer and an activation data buffer, and the third circuitry is to:

generate a combined sparsity bit mask based on the activation sparsity data and the weight sparsity data;

identify one or more weights stored in the weight data buffer and one or more activations stored in the activation data buffer based on the combined sparsity bit mask;

transfer the one or more activations from the activation data buffer to the first circuitry; and

transfer the one or more weights from the weight data buffer to the first circuitry.

4. The apparatus of claim 1 , wherein the second sparsity data includes a sparse weight vector, the type of the convolution is depthwise convolution, and

the sparse weight vector includes a sequence of elements, in which an element has a different value from other elements.

5. The apparatus of claim 1 , wherein the second sparsity data includes a sparse weight vector, the type of the convolution is grouped convolution, and

the sparse weight vector includes a sequence of elements, in which a group of consecutive elements in the sequence has a different value from other elements in the sequence.

6. The apparatus of claim 1 , wherein the type of the convolution is elementwise addition, and wherein:

the second circuitry is to

generate the second sparsity data by generating a plurality of sparse weight vectors, each sparse weight vector including a non-zero element and one or more other zero elements.

7. The apparatus of claim 1 , wherein the second sparsity data includes a sparse weight matrix, the type of the convolution is dilated convolution, and

the second circuitry is to generate the sparse weight matrix based on a dimension of a kernel of the convolution and a dilation parameter of the convolution.

8. A non-transitory computer readable medium storing instructions executable to perform operations, the operations comprising:

storing first sparsity data, the first sparsity data indicating sparsity in input data of a first convolution type;

generating second sparsity data for a second convolution type;

selecting sparsity data between the first sparsity data and the second sparsity data based on a type of a convolution;

identifying one or more data elements in the input data of the convolution based on the selected sparsity data;

transferring the one or more data elements from one or more data buffers to one or more multiply accumulators; and

executing, by the one or more multiply accumulators, at least part of the convolution by performing one or more multiply-accumulate operations on the one or more data elements.

9. The non-transitory computer readable medium of claim 8 , wherein the second sparsity data includes a sparse weight vector, and the operations further comprise:

identifying the type of the convolution based on configuration information, the sparse weight vector generated based on the configuration information

generating a sparsity bit mask based on a sparse activation vector and the sparse weight vector,

wherein the one or more data elements are identified by using the sparsity bit mask.

10. The non-transitory computer readable medium of claim 8 , wherein the selected sparsity data includes at least one of activation sparsity data and weight sparsity data, the one or more data buffers include a weight data buffer and an activation data buffer, and the operations further comprise:

generating a combined sparsity bit mask based on the activation sparsity data and the weight sparsity data;

identifying one or more weights stored in the weight data buffer and one or more activations stored in the activation data buffer based on the combined sparsity bit mask;

transferring the one or more activations from the activation data buffer to the one or more multiply accumulators; and

transferring the one or more weights from the weight data buffer to the one or more multiply accumulators.

11. The non-transitory computer readable medium of claim 8 , wherein the second sparsity data includes a sparse weight vector, the type of the convolution is depthwise convolution, and

the sparse weight vector includes a sequence of elements, in which an element has a different value from other elements.

12. The non-transitory computer readable medium of claim 8 , wherein the second sparsity data includes a sparse weight vector, the type of the convolution a grouped convolution, and

the sparse weight vector includes a sequence of elements, in which a group of consecutive elements in the sequence has a different value from other elements in the sequence.

13. The non-transitory computer readable medium of claim 8 , wherein the type of the convolution is elementwise addition, and generating the second sparsity data comprises:

generating a plurality of sparse weight vectors, each sparse weight vector including a non-zero element and one or more other zero elements.

14. The non-transitory computer readable medium of claim 8 , wherein the second sparsity data includes a sparse weight matrix, the type of the convolution is a dilated convolution, and generating the second sparsity data comprises:

generating the sparse weight matrix based on a dimension of a kernel of the convolution and a dilation parameter of the convolution.

15. A method comprising:

storing first sparsity data, the first sparsity data indicating sparsity in input data of a first convolution type;

generating second sparsity data for a second convolution type;

selecting sparsity data between the first sparsity data and the second sparsity data based on a type of a convolution;

identifying one or more data elements in the input data of the convolution based on the selected sparsity data;

transferring the one or more data elements from one or more data buffers to one or more multiply accumulators; and

executing, by the one or more multiply accumulators, at least part of the convolution by performing one or more multiply-accumulate operations on the one or more data element.

16. The method of claim 15 , wherein the second sparsity data includes a sparse weight vector, and the method further comprises:

identifying the type of the convolution based on configuration information, the sparse weight vector generated based on the configuration information

generating a sparsity bit mask based on a sparse activation vector and the sparse weight vector,

wherein the one or more data elements are identified by using the sparsity bit mask.

17. The method of claim 15 , wherein the selected sparsity data includes at least one of activation sparsity data and weight sparsity data, the one or more data buffers include a weight data buffer and an activation data buffer, and the method further comprises:

generating a combined sparsity bit mask based on the activation sparsity data and the weight sparsity data;

identifying one or more weights stored in the weight data buffer and one or more activations stored in the activation data buffer based on the combined sparsity bit mask;

transferring the one or more activations from the activation data buffer to the one or more multiply accumulators; and

transferring the one or more weights from the weight data buffer to the one or more multiply accumulators.

18. The method of claim 15 , wherein the second sparsity data includes a sparse weight vector, the type of the convolution is depthwise convolution, and

the sparse weight vector includes a sequence of elements, in which an element has a different value from other elements.

19. The method of claim 15 , wherein the second sparsity data includes a sparse weight vector, the type of the convolution is grouped convolution, and

the sparse weight vector includes a sequence of elements, in which a group of consecutive elements in the sequence has a different value from other elements in the sequence.

20. The method of claim 15 , wherein the type of the convolution is elementwise addition, and the method further comprises:

generating a plurality of sparse weight vectors, each sparse weight vector including a non-zero element and one or more other zero elements.

21. The method of claim 15 , wherein the second sparsity data includes a sparse weight matrix, the type of the convolution is dilated convolution, and the method further comprises:

generating the sparse weight matrix based on a dimension of a kernel of the convolution and a dilation parameter of the convolution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2021
From: POWER, MARTIN; BRADY, KEVIN; HANRAHAN, NIALL; GRYMEL, MARTIN-THOMAS; BERNARD, DAVID; BAUGH, GARY
To: INTEL CORPORATION
Reel/Frame 056738/0106 →
Continuity (1)
Related Publication 20210319317A1 · Oct 14, 2021
References Cited (40)
US 10768895B2 · Connor et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 11803736B1 · Meyer · 2023 [cited by examiner]
US 12182695B1 · Meyer · 2024 [cited by examiner]
US 20200125953A1 · Yoo · 2020 [cited by examiner]
US 20210004668A1 · Moshovos et al. · 2021 [cited by applicant]
US 20210224640A1 · Nakahara · 2021 [cited by examiner]
US 20210319317A1 · Power et al. · 2021 [cited by applicant]
US 20220188600A1 · Li · 2022 [cited by examiner]
US 20220405553A1 · Li · 2022 [cited by examiner]
CN 111626410 · 2020 [cited by applicant]
GB 2568102A · 2019 [cited by applicant]
L. Bai et al., DepthNet: Real-Time LiDAR Point Cloud Depth Completion for Autonomous Vehicles, IEEE Access, DOI 10.1109/ACCESS.2020.3045681, 2020 (Year: 2020). [cited by examiner]
B.C Lai et al., Enhancing Utilization of SIMD-Like Accelerator for Sparse Convolutional Neural Networks, IEEE Transactions of Very Large Scale Integration (VLSI) Systems, vol. 27, No. 5, 2019 (Year: 2019). [cited by examiner]
R. Zhao et.al., Learning Grouped Convolution for Efficient Domain Adaptation, arXiv:1811.09341v1 [cs.CV], 2018 (Year: 2018). [cited by examiner]
International Searching Authority, “Invitation to Correct Defects in the International Application,” issued in connection with International Patent Application No. PCT/US2022/021590, mailed on Apr. 6, 2022, 4 pages. [cited by applicant]
International Searching Authority, “Written Opinion of the International Searching Authority,” issued in connection with International Patent Application No. PCT/US2022/021590, mailed on Jun. 30, 2022, 4 pages. [cited by applicant]
International Searching Authority, “International Search Report,” issued in connection with International Patent Application No. PCT/US2022/021590, mailed on Jun. 30, 2022, 4 pages. [cited by applicant]
International Searching Authority, “International Preliminary Report on Patentability,” issued in connection with International Patent Application No. PCT/US2022/021590, mailed on Jan. 4, 2024, 6 pages. [cited by applicant]
“Keem Bay—Overview.” [Online]. Available: https://www.intel.com/content/www/us/en/secure/design/confidential/products-andsolutions/processors-and-chipsets/keem-bay/overview.html. [Accessed: Apr. 12, 2021] (9 pages). [cited by applicant]
V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient Processing of Deep Neural Networks,” Synth. Lect. Comput. Archit., vol. 15, No. 2, pp. 1-341, Jun. 2020 (343 pages). [cited by applicant]
V. Nair and G. E. Hinton, “Rectified Linear Units Improve Restricted Boltzmann Machines,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10), Jun. 21-24, 2010, Haifa, Israel, 2010, pp. 807… [cited by applicant]
Y.-H. Chen, T.-J. Yang, J. S. Emer, and V. Sze, “Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices,” IEEE J. Emerg. Sel. Top. Circuits Syst., vol. 9, No. 2, pp. 292-308, Jun. 2019 (1… [cited by applicant]
Y.-H. Chen, J. Emer, V. Sze, Y.-H. Chen, J. Emer, and V. Sze, “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” in 2016 ACM/IEEE 43rd Annual International Symposium on Co… [cited by applicant]
S. Han et al., “EIE: Efficient Inference Engine on Compressed Deep Neural Network,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016, pp. 243-254 (12 pages). [cited by applicant]
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-Neuron-Free Deep Neural Network Computing,” in Proceedings—2016 ACM/IEEE 43rd International Symposium on Computer … [cited by applicant]
“What Is Sparsity in AI Inference and Machine Learning? | NVIDIA Blog.” [Online]. Available: https://blogs.nvidia.com/blog/2020/05/14/sparsity-ai-inference/. [Accessed: Apr. 13, 2021] (4 pages). [cited by applicant]
J.-S. Park et al., “9.5 A 6K-MAC Feature-Map-Sparsity-Aware Neural Processing Unit in 5nm Flagship Mobile SoC,” in 2021 IEEE International Solid- State Circuits Conference (ISSCC), 2021, pp. 152-154 (3 pages). [cited by applicant]
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proceedings—30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR, Apr. 4, 2017, pp. 1800-1807 (8 pages). [cited by applicant]
A. G. Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.” Apr. 17, 2017, arXiv:1704.04861v1 (9 pages). [cited by applicant]
L. Kaiser, A. N. Gomez, and F. Chollet, “Depthwise Separable Convolutions for Neural Machine Translation,” arXiv, Jun. 16, 2017 (10 pages). [cited by applicant]
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” arXiv:1512.03385v1, Dec. 10, 2015 (12 pages). [cited by applicant]
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in Proceedings—30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Apr. 11, 2017,… [cited by applicant]
F. Yu and V. Koltun, “Multi-Scale Context Aggregation by Dilated Convolutions,” 4th Int. Conf. Learn. Represent. ICLR Apr. 30, 2016—Conf. Track Proc., arXiv:1511.07122v3 (13 pages). [cited by applicant]
X. Cui, K. Zheng, L. Gao, B. Zhang, D. Yang, and J. Ren, “Multiscale Spatial-Spectral Convolutional Network with Image-Based Framework for Hyperspectral Imagery Classification,” Remote Sens., vol. 11, No. 19, p. 2220, S… [cited by applicant]
Grand Keynote: What DL Hardware Will We Need ?—Overview. [Online]. Available: https://slideslive.com/38923058/grand-keynote-what-dl-hardware-will-we-need. [Accessed: Jun. 24, 2021] (2 pages). [cited by applicant]
Weijie You et al., “RSNN: A Software/Hardware Co-Optimized Framework for Sparse Convolutional Neural Networks on FPGAs”, IEEE Access, vol. 9, Dec. 24, 2020, 12 pages. [cited by applicant]
Xiao Liu et al., ‘Layerwise Sparse Coding for Pruned Deep Neural Networks with Extreme Compression Ratio’, Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, No. 04, Apr. 3, 2020, 8 pages. [cited by applicant]
International Searching Authority, “International Search Report and Written Opinion,” mailed in connection with International Patent Application No. PCT/US2022/021590, on Jun. 30, 2022, 8 pages. [cited by applicant]
Liu et al., “S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN Acceleration,” arXiv:2017.07983v1 [cs.AR] Jul. 16, 2021, 13 pages. [cited by applicant]