IP Library › Granted Patent US 12,639,398
Granted Patent B2
US 12,639,398 · App. 18/757,003 · Granted May 26, 2026

Scalable sparse matrix multiply acceleration using systolic arrays with feedback inputs

Inventors: Subramaniam Maiyuran (Gold River, CA); Jorge Parra (El Dorado Hills, CA); Supratim Pal (Bangalore, IN); Ashutosh Garg (Folsom, CA); Shubra Marwaha (Folsom, CA); Chandra Gurram (Folsom, CA); Darin Starkey (Roseville, CA); Durgesh Borkar (Folsom, CA); Varghese George (Folsom, CA)
Assignee: Intel Corporation
G06F17/16G06F9/3001G06F9/30036G06F9/30038G06F9/30145G06F15/8046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,398
App. No.
18/757,003
Granted
May 26, 2026
Kind
B2
Abstract

Described herein is a graphics processor including a plurality of processing clusters coupled with a host interface, each processing cluster comprising a plurality of multiprocessors, the plurality of multiprocessors interconnected via a data interconnect, and each multiprocessor comprising sparse matrix multiply acceleration hardware including a systolic processing array with feedback inputs.

Claims (48)

1 . A graphics processor comprising:

a host interface; and

a plurality of processing clusters coupled with the host interface, each processing cluster of the plurality of processing clusters comprising a plurality of multiprocessors, the plurality of multiprocessors interconnected via a data interconnect, and each multiprocessor of the plurality of multiprocessors comprising sparse matrix multiply acceleration hardware including a multi-stage processing pipeline configured to:

read, via a first source operand, multiple data elements of a first matrix into memory of the sparse matrix multiply acceleration hardware;

read, via a second source operand, multiple data elements of a second matrix into the memory of the sparse matrix multiply acceleration hardware;

detect non-zero values within the multiple data elements of the second matrix;

group the non-zero values within the multiple data elements of the second matrix into a group that includes one or more data elements;

provide a data element of the group to a corresponding stage of the multi-stage processing pipeline;

multiply a provided data element of the group with multiple data elements of the first matrix to generate a set of products;

sum the set of products and accumulate a sum of the set of products with an accumulator value; and

write the accumulator value to a next stage of the multi-stage processing pipeline.

2 . The graphics processor as in claim 1 , wherein a number of data elements of the group that includes the one or more data elements is to correspond with a number of stages in the multi-stage processing pipeline.

3 . The graphics processor as in claim 1 , wherein to write the accumulator value to the next stage of the multi-stage processing pipeline, the multi-stage processing pipeline is to write a pipeline feedback value to a first stage of the multi-stage processing pipeline.

4 . The graphics processor as in claim 1 , wherein to provide a data element of the group to a corresponding stage of the multi-stage processing pipeline, the multi-stage processing pipeline is configured to broadcast the data element to multiple channels of a processing element of the corresponding stage.

5 . The graphics processor as in claim 1 , wherein the multi-stage processing pipeline is configured to detect the non-zero values within the memory of the sparse matrix multiply acceleration hardware.

6 . The graphics processor as in claim 1 , wherein the multi-stage processing pipeline includes multiple pipeline paths and the sparse matrix multiply acceleration hardware is configured to provide data elements of the first matrix and the second matrix to the multiple pipeline paths via shared hardware circuitry associated with the first source operand and separate hardware circuitry associated with the second source operand.

7 . A method of performing a dot product operation on a set of input matrices via a hardware matrix multiply accelerator having a multi-stage processing pipeline, the method comprising:

reading, via a first source operand, multiple data elements of a first matrix into memory of the hardware matrix multiply accelerator;

reading, via a second source operand, multiple data elements of a second matrix into the memory of the hardware matrix multiply accelerator;

detecting non-zero values within the multiple data elements of the second matrix;

grouping the non-zero values within the multiple data elements of the second matrix into a group including one or more data elements;

providing a data element of the group to a corresponding stage of the multi-stage processing pipeline;

multiplying, a provided data element of the group with multiple data elements of the first matrix to generate a set of products;

summing the set of products and accumulating a sum of the set of products with an accumulator value; and

writing the accumulator value to a next stage of the multi-stage processing pipeline.

8 . The method as in claim 7 , wherein a number of data elements of the group including one or more data elements corresponds with a number of stages in the multi-stage processing pipeline of the hardware matrix multiply accelerator.

9 . The method as in claim 7 , wherein writing the accumulator value to the next stage of the multi-stage processing pipeline includes writing a pipeline feedback value to a first stage of the multi-stage processing pipeline.

10 . The method as in claim 7 , wherein providing a data element of the group to a corresponding stage of the multi-stage processing pipeline includes broadcasting the data element to multiple channels of a processing element of the corresponding stage.

11 . The method as in claim 7 , wherein detecting the non-zero values within the multiple data elements of the second matrix includes detecting the non-zero values within the memory of the hardware matrix multiply accelerator.

12 . The method as in claim 7 , wherein the multi-stage processing pipeline of the hardware matrix multiply accelerator includes multiple pipeline paths.

13 . The method as in claim 12 , further comprising providing data elements of the first matrix and the second matrix to the multiple pipeline paths via shared hardware circuitry associated with the first source operand and separate hardware circuitry associated with the second source operand.

14 . A graphics processing system comprising:

a memory device; and

a graphics processor coupled with the memory device via a host interface, the graphics processor comprising a plurality of processing clusters, each processing cluster of the plurality of processing clusters comprising a plurality of multiprocessors, the plurality of multiprocessors interconnected via a data interconnect, and each multiprocessor of the plurality of multiprocessors comprising sparse matrix multiply acceleration hardware including a multi-stage processing pipeline configured to:

read, via a first source operand, multiple data elements of a first matrix into memory of the sparse matrix multiply acceleration hardware;

read, via a second source operand, multiple data elements of a second matrix into the memory of the sparse matrix multiply acceleration hardware;

detect non-zero values within the multiple data elements of the second matrix;

group the non-zero values within the multiple data elements of the second matrix into a group that includes one or more data elements;

provide a data element of the group to a corresponding stage of the multi-stage processing pipeline;

multiply a provided data element of the group with multiple data elements of the first matrix to generate a set of products;

sum the set of products and accumulate a sum of the set of products with an accumulator value; and

write the accumulator value to a next stage of the multi-stage processing pipeline.

15 . The graphics processing system as in claim 14 , wherein a number of data elements of the group that includes the one or more data elements is to correspond with a number of stages in the multi-stage processing pipeline.

16 . The graphics processing system as in claim 14 , wherein to write the accumulator value to the next stage of the multi-stage processing pipeline, the multi-stage processing pipeline is to write a pipeline feedback value to a first stage of the multi-stage processing pipeline.

17 . The graphics processing system as in claim 14 , wherein to provide a data element of the group to a corresponding stage of the multi-stage processing pipeline, the multi-stage processing pipeline is configured to broadcast the data element to multiple channels of a processing element of the corresponding stage.

18 . The graphics processing system as in claim 14 , wherein the multi-stage processing pipeline is configured to detect the non-zero values within the memory of the sparse matrix multiply acceleration hardware.

19 . The graphics processing system as in claim 14 , wherein the multi-stage processing pipeline includes multiple pipeline paths.

20 . The graphics processing system as in claim 19 , wherein the sparse matrix multiply acceleration hardware is configured to provide data elements of the first matrix and the second matrix to the multiple pipeline paths via shared hardware circuitry associated with the first source operand and separate hardware circuitry associated with the second source operand.

Priority Claims (1)
IN 202041019059 · May 5, 2020 · national
Continuity (4)
Continuation 18301386 · Apr 17, 2023
Continuation 17527882 · Nov 16, 2021
Continuation 16913800 · Jun 26, 2020
Related Publication 20240427847A1 · Dec 26, 2024
References Cited (25)
US 11204977B2 · Maiyuran et al. · 2021 [cited by applicant]
US 11347477B2 · Sumbul · 2022 [cited by examiner]
US 20140258689A1 · Song · 2014 [cited by examiner]
US 20160140084A1 · Daga · 2016 [cited by examiner]
US 20180067899A1 · Rub · 2018 [cited by examiner]
US 20180121388A1 · Rennich · 2018 [cited by examiner]
US 20180189234A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180189675A1 · Nurvitadhi · 2018 [cited by examiner]
US 20180330192A1 · Atasu · 2018 [cited by examiner]
US 20190042542A1 · Narayanamoorthy · 2019 [cited by examiner]
US 20190294413A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190347125A1 · Sankaran · 2019 [cited by examiner]
US 20200104692A1 · Hill · 2020 [cited by examiner]
US 20200160181A1 · Zlateski · 2020 [cited by examiner]
US 20210349966A1 · Maiyuran et al. · 2021 [cited by applicant]
US 20220188600A1 · Li · 2022 [cited by examiner]
CN 110851779A · 2020 [cited by applicant]
CN 113610697 · 2021 [cited by applicant]
DE 102020131666 · 2021 [cited by applicant]
JP 7728639 · 2025 [cited by applicant]
TW 201826122A · 2018 [cited by applicant]
TW 202143031 · 2021 [cited by applicant]
Office Action for TW Application No. 109145287, mailed May 10, 2024, 9 pages (no translation). [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/527,882, mailed Dec. 20, 2022, 12 pages. [cited by applicant]
Notice of Allowance for TW Application No. 109145287, mailed Sep. 12, 2024, 3 pages. [cited by applicant]