IP Library Granted Patent US 12,367,255
Granted Patent B2
US 12,367,255 · App. 18/626,599 · Granted Jul 22, 2025

Machine learning architecture support for block sparsity

Inventor: Omid Azizi (Redwood City, CA)
Assignee: Intel Corporation
G06F17/16G06F9/3001G06F9/30036G06F9/3016G06F9/3802G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,255
App. No.
18/626,599
Granted
Jul 22, 2025
Kind
B2
Abstract

This disclosure relates matrix operation acceleration for different matrix sparsity patterns. A matrix operation accelerator may be designed to perform matrix operations more efficiently for a first matrix sparsity pattern rather than for a second matrix sparsity pattern. A matrix with the second sparsity pattern may be converted to a matrix with the first sparsity pattern and provided to the matrix operation accelerator. By rearranging the rows and/or columns of the matrix, the sparsity pattern of the matrix may be converted to a sparsity pattern that is suitable for computation with the matrix operation accelerator.

Claims (46)

1. A method of making an executable neural network, the method comprising operations for:

generating a first matrix having a first sparsity from a third matrix having a second sparsity by performing a transformation of the third matrix from the second sparsity to the first sparsity;

generating a second matrix by performing a matrix operation in the executable neural network on the first matrix;

performing a transformation of the second matrix to generate a result matrix; and

outputting the result matrix.

2. The method of claim 1 , wherein performing the transformation of the third matrix comprises:

rearranging a plurality of columns of the third matrix or a plurality of rows of the third matrix.

3. The method of claim 2 , wherein rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix comprises:

rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix based on a read order of the third matrix from a memory that stores the third matrix.

4. The method of claim 3 , wherein the read order comprises a coprime stride.

5. The method of claim 2 , wherein rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix comprises:

rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix based on a size of the third matrix or the second sparsity.

6. The method of claim 1 , wherein performing the transformation of the second matrix comprises:

performing an opposite of the transformation of the third matrix by rearranging columns of the second matrix in an order that is inverse to a rearrangement of columns of the third matrix during the transformation of the third matrix.

7. The method of claim 1 , wherein the first sparsity is finer-grained than the second sparsity.

8. One or more non-transitory computer-readable media storing instructions executable to perform a method of making an executable neural network, the method comprising operations for:

generating a first matrix having a first sparsity from a third matrix having a second sparsity by performing a transformation of the third matrix from the second sparsity to the first sparsity;

generating a second matrix by performing a matrix operation in the executable neural network on the first matrix;

performing a transformation of the second matrix to generate a result matrix; and

outputting the result matrix.

9. The one or more non-transitory computer-readable media of claim 8 , wherein performing the transformation of the third matrix comprises:

rearranging a plurality of columns of the third matrix or a plurality of rows of the third matrix.

10. The one or more non-transitory computer-readable media of claim 9 , wherein rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix comprises:

rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix based on a read order of the third matrix from a memory that stores the third matrix.

11. The one or more non-transitory computer-readable media of claim 10 , wherein the read order comprises a coprime stride.

12. The one or more non-transitory computer-readable media of claim 9 , wherein rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix comprises:

rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix based on a size of the third matrix or the second sparsity.

13. The one or more non-transitory computer-readable media of claim 8 , wherein performing the transformation of the second matrix comprises:

performing an opposite of the transformation of the third matrix by rearranging columns of the second matrix in an order that is inverse to a rearrangement of columns of the third matrix during the transformation of the third matrix.

14. The one or more non-transitory computer-readable media of claim 8 , wherein the first sparsity is finer-grained than the second sparsity.

15. An apparatus, comprising:

a computer processor for executing computer program instructions; and

a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform a method of making an executable neural network, the method comprising operations for:

generating a first matrix having a first sparsity from a third matrix having a second sparsity by performing a transformation of the third matrix from the second sparsity to the first sparsity,

generating a second matrix by performing a matrix operation in the executable neural network on the first matrix,

performing a transformation of the second matrix to generate a result matrix, and

outputting the result matrix.

16. The apparatus of claim 15 , wherein performing the transformation of the third matrix comprises:

rearranging a plurality of columns of the third matrix or a plurality of rows of the third matrix.

17. The apparatus of claim 16 , wherein rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix comprises:

rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix based on a read order of the third matrix from a memory that stores the third matrix.

18. The apparatus of claim 17 , wherein the read order comprises a coprime stride.

19. The method of apparatus of claim 16 , wherein rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix comprises:

rearranging the plurality of columns of the third matrix or the plurality of rows of the third matrix based on a size of the third matrix or the second sparsity.

20. The apparatus of claim 15 , wherein performing the transformation of the second matrix comprises:

performing an opposite of the transformation of the third matrix by rearranging columns of the second matrix in an order that is inverse to a rearrangement of columns of the third matrix during the transformation of the third matrix.

Continuity (3)
Continuation 17481064 · Sep 21, 2021
Continuation 16370094 · Mar 29, 2019
Related Publication 20240330402A1 · Oct 3, 2024
References Cited (3)
US 20140211039A1 · Herman · 2014 [cited by examiner]
US 20170147301A1 · Rong · 2017 [cited by examiner]
US 20180121388A1 · Rennich · 2018 [cited by applicant]