IP Library Granted Patent US 11,392,829
Granted Patent B1
US 11,392,829 · App. 16/373,301 · Granted Jul 19, 2022

Managing data sparsity for neural networks

Inventors: Jeff Pool (Durham, NC); Ganesh Venkatesh (San Jose, CA); Jorge Albericio Latorre (San Jose, CA); Jack Choquette (Palo Alto, CA); Ronny Krashinsky (San Francisco, CA); John Tran (Denver, CO); Feng Xie (Shanghai, CN); Ming Y. Siu (Santa Clara, CA); Manan Patel (San Jose, CA)
Assignee: NVIDIA Corporation
G06N3/08G06F17/16G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,392,829
App. No.
16/373,301
Granted
Jul 19, 2022
Kind
B1
Abstract

Approaches in accordance with various embodiments provide for the processing of sparse matrices for mathematical and programmatic operations. In particular, various embodiments enforce sparsity constraints for performing sparse matrix multiply-add instruction (MMA) operations. Deep neural networks can exhibit significant sparsity in the data used in operations, both in the activations and weights. The computational load can be reduced by excluding zero-valued data elements. A sparsity constraint is applied across all submatrices of a sparse matrix, providing fine-grained structured sparsity that is evenly distributed across the matrix. The matrix may then be compressed since a minimum number of elements of the matrix are known to have zero value. Matrix operations are then performed using these matrices.

Claims (24)

1. A processor comprising:

one or more arithmetic logic units (ALUs) to perform one or more matrix multiply operations on one or more sparse matrices corresponding to one or more neural networks, wherein the sparsity of the one or more sparse matrices is constrained by a minimum percentage of zero values, which is more than zero percent of the one or more sparse matrices, or a maximum percentage of non-zero values, which is less than one hundred percent of the one or more sparse matrices.

2. The processor of claim 1 , wherein the one or more matrix multiply operations are matrix multiply-add (MMA) operations, and wherein the one or more sparse matrices correspond to at least one of weights or activations of the one or more neural networks.

3. The processor of claim 1 , wherein the one or more ALUs are further to compress the one or more sparse matrices by excluding one or more of the zero values from the one or more sparse matrices and reducing a number of rows or a number of columns of the one or more sparse matrices.

4. The processor of claim 3 , wherein the one or more ALUs are further to store metadata corresponding to the one or more zero values excluded from the one or more compressed matrices.

5. The processor of claim 1 , wherein the one or more ALUs are further to determine a plurality of submatrices of the one or more sparse matrices, and further constrain the sparsity across the plurality of submatrices.

6. The processor of claim 5 , wherein the plurality of submatrices are generated in a compressed sparse row (CSR) format, compressed sparse column (CSC), or a compressed sparse block (CSB) format.

7. The processor of claim 1 , wherein the one or more ALUs are further to ensure that a first operand of the one or more matrix multiply operations corresponds to the one or more sparse matrices and a second operand of the one or more matrix multiply operations corresponds to one or more dense matrices.

8. A system comprising:

one or more computers comprising one or more processors to train one or more neural networks, in which one or more sparse matrices of weight values are to be constrained based, at least in part, on a minimum percentage of zero values, which is more than zero percent of the one or more sparse matrices, or a maximum percentage of non-zero values, which is less than one hundred percent of the one or more sparse matrices.

9. The system of claim 8 , wherein the one or more matrix multiply operations are matrix multiply-add (MMA) operations, and wherein the one or more sparse matrices correspond to at least one of weights or activations of the one or more neural networks.

10. The system of claim 8 , wherein the one or more processors are further to compress the one or more sparse matrices by excluding one or more of the zero values from the one or more sparse matrices and reducing a number of rows or a number of columns of the one or more sparse matrices.

11. The system of claim 10 , wherein the one or more processors are further to store metadata corresponding to the one or more zero values excluded from the one or more compressed matrices.

12. The system of claim 8 , wherein the one or more processors are further to determine a plurality of submatrices of the one or more sparse matrices, and further constrain the sparsity across the plurality of submatrices.

13. The system of claim 12 , wherein the plurality of submatrices are generated in a compressed sparse row (CSR) format, or a compressed sparse column (CSC) format, or a compressed sparse block (CSB).

14. The system of claim 8 , wherein the one or more processors are further to ensure that a first operand of the one or more matrix multiply operations corresponds to the one or more sparse matrices and a second operand of the one or more matrix multiply operations corresponds to one or more dense matrices.

15. A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

train one or more neural networks, in which one or more sparse matrices of weight values are to be constrained based, at least in part, on a minimum percentage of zero values, which is more than zero percent of the one or more sparse matrices, or a maximum percentage of non-zero values, which is less than one hundred percent of the one or more sparse matrices.

16. The machine-readable medium of claim 15 , wherein the one or more matrix multiply operations are matrix multiply-add (MMA) operations, and wherein the one or more sparse matrices correspond to at least one of weights or activations of the one or more neural networks.

17. The machine-readable medium of claim 15 , wherein the instructions when executed further cause the one or more processors to compress the one or more sparse matrices by excluding one or more of the zero values from the one or more sparse matrices and reducing a number of rows or a number of columns of the one or more sparse matrices.

18. The machine-readable medium of claim 17 ,

wherein the instructions when executed further cause the one or more processors to store metadata corresponding to the one or more zero values excluded from the one or more compressed matrices.

19. The machine-readable medium of claim 15 , wherein the instructions when executed further cause the one or more processors to determine a plurality of submatrices of the one or more sparse matrices, and further constrain the sparsity across the plurality of submatrices.

20. The machine-readable medium of claim 15 , wherein the instructions when executed further cause the one or more processors to ensure that a first operand of the one or more matrix multiply operations corresponds to the one or more sparse matrices and a second operand of the one or more matrix multiply operations corresponds to one or more dense matrices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2019
From: POOL, JEFF; VENKATESH, GANESH; LATORRE, JORGE ALBERICIO; CHOQUETTE, JACK; KRASHINSKY, RONNY; TRAN, JOHN; XIE, FUNG; SIU, MICHAEL; PATEL, MANAN
To: NVIDIA CORPORATION
Reel/Frame 048773/0403 →
Continuity (1)
Provisional Application 62665665 · May 2, 2018
Cited By (11)
US 12,231,152 US 12,248,367 US 12,430,543 US 12,443,835 US 12,499,357 US 12,530,624 US 12,585,928 US 12,645,458 US 12,670,233 US 12,711,196 US 12,711,393