IP Library › Granted Patent US 12,602,252
Granted Patent B2
US 12,602,252 · App. 18/240,281 · Granted Apr 14, 2026

Sparse matrix multiplication in a neural network

Inventors: Jorge Albericio Latorre (Brooklyn, NY); Chong Yu (Shanghai, CN)
Assignee: NVIDIA Corporation
G06F9/5027G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,252
App. No.
18/240,281
Granted
Apr 14, 2026
Kind
B2
Abstract

Apparatuses, systems, and methods to enable matrix multiplication acceleration by modifying an input to apply sparsity through sparse activation filtering. In at least one embodiment, a neural network modifies pixels within an image through sparse activation filtering to enable use of one or more matrix multiplication acceleration units to perform a sparse patch embedding operation.

Claims (28)

1 . A processor, comprising:

one or more circuits to:

modify one or more values corresponding to one or more pixels within one or more images to zero according to a sparsity constraint; and

cause one or more operations to be performed, using one or more matrix multiplication acceleration units, on one or more operands representing pixels including the one or more modified pixels, wherein the one or more matrix multiplication acceleration units are configured to accelerate operation based on the sparsity constraint.

2 . The processor of claim 1 , wherein the one or more matrix multiplication acceleration units comprise a sparse tensor core, the sparse tensor core comprising circuitry to perform accelerated general matrix multiplication (GEMM).

3 . The processor of claim 1 , wherein the one or more matrix multiplication acceleration units are to prune and compress at least one matrix.

4 . The processor of claim 1 , wherein the one or more circuits are to determine an active portion and an inactive portion of an input image, and are to perform a patch embedding operation using the one or more matrix multiplication acceleration units based, at least in part, on the active portion of the input image.

5 . The processor of claim 4 , wherein the one or more circuits are to divide the input image into one or more square patches and convert the one or more square patches to a patch matrix comprising flattened row vectors representing the active portion and all-zero row vectors representing the inactive portion.

6 . The processor of claim 4 , wherein the one or more circuits are to determine the active portion and the inactive portion of the input image according to an activation threshold.

7 . The processor of claim 4 , wherein the patch embedding operation comprises using at least one index matrix to identify locations of non-zero values in a matrix used to generate vector tokens.

8 . The processor of claim 4 , wherein the patch embedding operation uses a row-sparse patch matrix and a sparse weight matrix to generate final patch embedding.

9 . A system, comprising:

one or more processors to:

modify one or more values corresponding to one or more pixels within one or more images to zero according to a sparsity constraint; and

cause one or more operations to be performed, using one or more matrix multiplication acceleration units, on one or more operands representing pixels including the one or more modified pixels, wherein the one or more matrix multiplication acceleration units are configured to accelerate operation based on the sparsity constraint.

10 . The system of claim 9 , wherein the one or more matrix multiplication acceleration units comprises a sparse tensor core, the sparse tensor core comprising circuitry to perform accelerated general matrix multiplication (GEMM).

11 . The system of claim 9 , wherein the one or more matrix multiplication acceleration units prunes and compresses at least one matrix.

12 . The system of claim 9 , wherein the one or more processors are to determine an active portion and an inactive portion of an input image, and compute a patch embedding operation using the one or more matrix multiplication acceleration units based, at least in part, on the active portion of the input image.

13 . The system of claim 12 , wherein the input image is divisible into one or more square patches, wherein one or more portions corresponding to the one or more square patches are converted to a patch matrix comprising flattened row vectors representing the active portion and all-zero row vectors representing the inactive portion.

14 . The system of claim 12 , wherein the one or more processors are to use at least one index matrix to identify locations of non-zero values in a matrix used to generate vector tokens.

15 . A method, comprising:

modifying one or more values corresponding to one or more pixels within one or more images to zero according to a sparsity constraint; and

causing one or more operations to be performed, using one or more matrix multiplication acceleration units, on one or more operands representing pixels including the one or more modified pixels, wherein the one or more matrix multiplication acceleration units are configured to accelerate operation based on the sparsity constraint.

16 . The method of claim 15 , further comprising using the one or more matrix multiplication acceleration units to perform accelerated GEMM.

17 . The method of claim 15 , wherein the matrix multiplication acceleration unit prunes and compresses at least one matrix.

18 . The method of claim 15 , further comprising determining an active portion and an inactive portion of an input image, and compute a patch embedding operation using the one or more matrix multiplication acceleration units based, at least in part, on the active portion of the input image.

19 . The method of claim 18 , further comprising generating flattened row vectors representing the active portion and all-zero row vectors representing the inactive portion.

20 . The method of claim 18 , wherein the patch embedding operation includes using at least one index matrix to identify locations of non-zero values in a matrix used to generate vector tokens.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2023
From: ALBERICIO LATORRE, JORGE; YU, CHONG
To: NVIDIA CORPORATION
Reel/Frame 064894/0943 →
Continuity (2)
Continuation PCTCN2023110915 · Aug 3, 2023
Related Publication 20250045107A1 · Feb 6, 2025
References Cited (10)
US 11995884B1 · Sammoura · 2024 [cited by examiner]
US 20130028516A1 · Warfield · 2013 [cited by examiner]
US 20150030232A1 · Parkhomenko et al. · 2015 [cited by applicant]
US 20190295228A1 · Liu · 2019 [cited by examiner]
US 20220358748A1 · Beer Mohideen et al. · 2022 [cited by applicant]
US 20230074229A1 · Jia et al. · 2023 [cited by applicant]
CN 109993293A · 2019 [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetic”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/CN2023/110915, mailed Apr. 10, 2024, filed Aug. 3, 2023, 7 pages. [cited by applicant]
Mishra et al., “Accelerating Sparse Deep Neural Networks,” Apr. 16, 2021, 18 pages. [cited by applicant]