IP Library › Granted Patent US 12,645,925
Granted Patent B2
US 12,645,925 · App. 17/521,840 · Granted Jun 2, 2026

Dual-sparse neural processing unit with multi-dimensional routing of non-zero values

Inventors: Jong Hoon Shin (San Jose, CA); Ali Shafiee Ardestani (San Jose, CA); Joseph H. Hassoun (Los Gatos, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06N3/063G06F7/5443G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,925
App. No.
17/521,840
Granted
Jun 2, 2026
Kind
B2
Abstract

A general matrix-matrix (GEMM) accelerator core includes first and second buffers, a control logic circuit, and a first processing element (PE). The first buffer receives a elements of a first matrix A of activation values. The second buffer receives b elements of a second matrix B of weight values. The control logic circuit replaces a zero-valued a element in a first column of the first buffer with a nonzero-valued a element that is within a maximum borrowing distance of a location of the zero-valued a element in the first column of the first buffer. The PE receives a elements from the first column of the first buffer including the nonzero-valued element a selected to replace the zero-valued a element and receives b elements from locations in the second buffer that correspond to locations in the first buffer from where the a elements have been received by the PE.

Claims (43)

1 . A general matrix-matrix (GEMM) accelerator core, comprising:

a first buffer comprising K 0 rows and K 1 columns of locations, the first buffer being configured to receive a elements of a first matrix A of activation values, and K 0 and K 1 being integers greater than 1;

a second buffer comprising K 1 rows and K 0 columns of locations, the second buffer being configured to receive b elements of a second matrix B of weight values;

a control logic circuit coupled to the first buffer, the control logic circuit being configured to select a first nonzero-valued a element based a first zero-valued a element being in a first column of the first buffer and to replace the first zero-valued a element with the first nonzero-value a element, the first nonzero-valued a element being selected to replace the first zero-valued a element in the first column of the first buffer being within a maximum borrowing distance of a first location of the first zero-valued a element in the first column of the first buffer; and

a first processing element (PE) comprising an array of K 0 multipliers, the first PE being associated with the first buffer and the second buffer, and being configured to receive a elements from the first column of the first buffer including the first nonzero-valued element a selected to replace the first zero-valued element a and to receive b elements from locations in the second buffer that correspond to locations in the first buffer from where the a elements have been received by the first PE.

2 . The GEMM accelerator core of claim 1 , wherein the first PE is further configured to multiply the a elements received from the first column of the first buffer and the b elements received from the second buffer.

3 . The GEMM accelerator core of claim 1 , wherein the control logic circuit is further configured to indicate a b element at a location in the second buffer that corresponds to a location of the first nonzero-valued a element selected to replace the first zero-valued a element.

4 . The GEMM accelerator core of claim 1 , wherein the maximum borrowing distance of the first location comprises a predetermined distance from the first location that is in at least one direction of at least one of three dimensions.

5 . The GEMM accelerator core of claim 1 , wherein the control logic circuit is further configured to select the first nonzero-valued a element to replace the first zero-valued a element based on the first nonzero-valued a element having a fewest number of possibilities of replacing a zero-valued a element as compared to a number of possibilities of other nonzero-valued elements that are within the maximum borrowing distance of the first location of the first zero-valued a element.

6 . The GEMM accelerator core of claim 1 , wherein the second matrix B is preprocessed to replace zero-valued elements of matrix B with nonzero-valued elements of second matrix B.

7 . The GEMM accelerator core of claim 6 , wherein the control logic circuit is further configured to pair nonzero-valued elements of the first matrix A with corresponding nonzero-valued elements of the second matrix B.

8 . The GEMM accelerator core of claim 1 , further comprising:

a third buffer comprising K 0 rows and K 1 columns of locations, the third buffer configured to receive a elements of the first matrix A of activation values; and

a second PE comprising an array of K 0 multipliers, the second PE being associated with the third buffer and the second buffer,

wherein the control logic circuit is coupled to the third buffer, the control logic circuit being configured to select a second nonzero-valued a element in the first buffer based a second zero-valued a element being in a first column of the third buffer and to replace the second zero-valued a element in the first column of the third buffer with the second nonzero-value a element in the first buffer, the second nonzero-valued a element being selected to replace the second zero-valued a element in the first column of the third buffer being within a maximum borrowing distance of a second location of the second zero-valued a element in the first column of the third buffer, and

wherein the second PE is configured to receive a elements from the first column of the third buffer including the second nonzero-valued a element selected to replace the second zero-value a element and to receive b elements from locations in the second buffer that correspond to locations in the third and the first buffers from where the a elements have been received by the second PE.

9 . The GEMM accelerator core of claim 8 , wherein the second PE is further configured to multiply the a elements received from the first column of the third buffer and the b elements received from the second buffer.

10 . The GEMM accelerator core of claim 8 , wherein the maximum borrowing distance of the second location in the first column of the third buffer comprises a predetermined distance from the second location in the first column of the third buffer that is in at least one direction of at least one of three dimensions.

11 . A general matrix-matrix (GEMM) accelerator core, comprising:

a first buffer comprising K 0 rows and K 1 columns of locations, the first buffer being configured to receive elements a of a first matrix A of activation values, and K 0 and K 1 being integers greater than 1;

a second buffer comprising K 1 rows and K 0 columns of locations, the second buffer being configured to receive elements b of a second matrix B of weight values;

a control logic circuit coupled to the first buffer, the control logic circuit being configured to select a first nonzero-valued a element based a first zero-valued a element being in a first column of the first buffer, the first nonzero-valued a element being selected based on the first nonzero-valued a element having a fewest number of possibilities of replacing a zero-valued a element as compared to a number of possibilities of other nonzero-valued elements that are within a maximum borrowing distance of a first location of the first zero-valued a element in the first column of the first buffer; and

a first processing element (PE) comprising an array of K 0 multipliers, the first PE being associated with the first buffer and the second buffer, and the first PE being configured to receive a elements from the first column of the first buffer including the first nonzero-valued a element selected to replace the first zero-valued a element and to receive b elements from locations in the second buffer that correspond to locations in the first buffer from where the a elements have been received by the first PE.

12 . The GEMM accelerator core of claim 11 , wherein the maximum borrowing distance of the first location comprises a predetermined distance from the first location that is in at least one direction of at least one of three dimensions.

13 . The GEMM accelerator core of claim 12 , wherein the second matrix B is preprocessed to replace zero-valued elements of matrix B with nonzero-valued elements of second matrix B.

14 . The GEMM accelerator core of claim 13 , further comprising:

a third buffer comprising K 0 rows and K 1 columns of locations, the third buffer being configured to receive elements a of a first matrix A of activation values; and

a second PE comprising an array of K 0 multipliers, the second PE being associated with the third buffer and the second buffer,

wherein the control logic circuit is coupled to the third buffer, and the control logic circuit being further configured to select a second nonzero-valued a element in the first buffer based a second zero-valued a element being in a first column of the third buffer and to replace the second zero-valued a element in the first column of the third buffer with the second nonzero-value a element in the first buffer, the second nonzero-valued a element being selected from within the maximum borrowing distance of a second location of the second zero-valued a element, and

wherein the second PE is configured to receive a elements from the first column of the third buffer including the second nonzero-valued a element selected to replace the second zero-value a element in the first column of the third buffer and to receive b elements from locations in the second buffer that correspond to locations in the third and the first buffers from where the a elements have been received by the second PE.

15 . A general matrix-matrix (GEMM) accelerator core, comprising:

a first buffer comprising K 0 rows and K 1 columns of locations, the first buffer being configured to receive a elements of a first matrix A of activation values, and K 0 and K 1 being integers greater than 1;

a second buffer comprising K 1 rows and K 0 columns of locations, the second buffer being configured to receive b elements of a second matrix B of weight values, the second matrix B being preprocessed to replace zero-valued elements of matrix B with nonzero-valued elements of second matrix B;

a third buffer comprising K 0 rows and K 1 columns of locations, the third buffer being configured to receive a elements of the first matrix A of activation values;

a control logic circuit coupled to the first buffer and the third buffer, the control logic circuit being configured to select a first nonzero-valued a element based a first zero-valued a element being in a first column of the first buffer and to replace the first zero-valued a element with the first nonzero-value a element, the first nonzero-valued a element being selected from within a maximum borrowing distance of a first location of the first zero-valued a element in the first column of the first buffer, and the control logic circuit being further configured to select a second nonzero-valued a element in the first buffer based a second zero-valued a element being in a first column of the third buffer and to replace the second zero-valued a element with the second nonzero-value a element, the second nonzero-valued a element being selected from within a maximum borrowing distance of a second location of the second zero-valued a element;

a first processing element (PE) comprising an array of K 0 multipliers, the first PE being associated with the first buffer and the second buffer, and being configured to receive a elements from the first column of the first buffer including the first nonzero-valued element a selected to replace the first zero-valued element a and to receive b elements from locations in the second buffer that correspond to locations in the first buffer from where the a elements have been received by the first PE; and

a second PE comprising an array of K 0 multipliers, the first PE being associated with the third buffer and the second buffer, the second PE being configured to receive a elements from the first column of the third buffer including the second nonzero-valued a element selected to replace the second zero-value a element and to receive b elements from locations in the second buffer that correspond to locations in the third and the first buffer from where the a elements have been received by the second PE.

16 . The GEMM accelerator core of claim 15 , wherein the first PE is further configured to multiply the a elements received from the first column of the first buffer and the b elements received from the second buffer, and

wherein the second PE is further configured to multiply the a elements received from the first column of the third buffer and the b elements received from the second buffer.

17 . The GEMM accelerator core of claim 15 , wherein the maximum borrowing distance of the first location comprises a predetermined distance from the first location that is in at least one direction of at least one of three dimensions, and comprises the predetermined distance from the second location that is in at least one direction of at least one of three dimensions.

18 . The GEMM accelerator core of claim 15 , wherein the second matrix B is preprocessed to replace zero-valued elements of matrix B with nonzero-valued elements of second matrix B.

19 . The GEMM accelerator core of claim 15 , wherein the control logic circuit is further configured to select the first nonzero-valued element a to replace the first zero-valued a element based on the first nonzero-valued a element having a fewest number of possibilities of replacing a zero-valued a element as compared to a number of possibilities of other nonzero-valued elements that are within the maximum borrowing distance of the first location of the first zero-valued a element.

20 . The GEMM accelerator core of claim 15 , wherein the control logic circuit is further configured to select the second nonzero-valued element a to replace the second zero-valued a element based on the second nonzero-valued a element having a fewest number of possibilities of replacing a zero-valued a element as compared to a number of possibilities of other nonzero-valued elements that are within the maximum borrowing distance of the second location of the second zero-valued a element.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2023
From: SHIN, JONG HOON; SHAFIEE ARDESTANI, ALI; HASSOUN, JOSEPH H.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064569/0449 →
Continuity (2)
Provisional Application 63113820 · Nov 13, 2020
Related Publication 20220156568A1 · May 19, 2022
References Cited (10)
US 11100193B2 · Gu et al. · 2021 [cited by applicant]
US 20180315158A1 · Nurvitadhi · 2018 [cited by examiner]
US 20190041961A1 · Desai et al. · 2019 [cited by applicant]
US 20210110508A1 · Mellempudi · 2021 [cited by examiner]
JP 2020091853A · 2020 [cited by applicant]
KR 20200070088A · 2020 [cited by applicant]
Delmas Lascorz, Alberto et al., “Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks,” Proceedings of the Twenty-Fourth International Conference on Architectural Support fo… [cited by applicant]
Gondimalla, Ashish et al., “SparTen: A Sparse Tensor Accelerator for Convolutional Neural Networks,” Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, 2019, pp. 151-165. [cited by applicant]
Office Action for U.S. Appl. No. 17/521,846, mailed May 1, 2025. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/521,846, mailed Dec. 23, 2025. [cited by applicant]