IP Library › Granted Patent US 12,039,001
Granted Patent B2
US 12,039,001 · App. 18/301,386 · Granted Jul 16, 2024

Scalable sparse matrix multiply acceleration using systolic arrays with feedback inputs

Inventors: Subramaniam Maiyuran (Gold River, CA); Jorge Parra (El Dorado Hills, CA); Supratim Pal (Bangalore, IN); Ashutosh Garg (Folsom, CA); Shubra Marwaha (Folsom, CA); Chandra Gurram (Folsom, CA); Darin Starkey (Roseville, CA); Durgesh Borkar (Folsom, CA); Varghese George (Folsom, CA)
Assignee: Intel Corporation
G06F17/16G06F9/3001G06F9/30145G06F15/8046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,039,001
App. No.
18/301,386
Granted
Jul 16, 2024
Kind
B2
Abstract

Described herein is a graphics processor including a plurality of processing clusters coupled with a host interface, each processing cluster comprising a plurality of multiprocessors, the plurality of multiprocessors interconnected via a data interconnect, and each multiprocessor comprising sparse matrix multiply acceleration hardware including a systolic processing array with feedback inputs.

Claims (36)

1. A graphics processor comprising:

a host interface;

a plurality of processing clusters coupled with the host interface, each processing cluster of the plurality of processing clusters comprising a plurality of multiprocessors, the plurality of multiprocessors interconnected via a data interconnect, and each multiprocessor of the plurality of multiprocessors comprising sparse matrix multiply acceleration hardware including:

a systolic processing array with feedback inputs, the systolic processing array comprising multiple processing array modules, the multiple processing array modules including a processing array module having a first number of pipeline paths, the first number of pipeline paths having a second number of pipeline stages,

wherein a first pipeline stage is configurable to receive feedback output from a final pipeline stage and the processing array module includes pipeline paths configured with shared circuitry to read data elements associated with a first source input and separate circuitry to read data elements associated with a second source input.

2. The graphics processor as in claim 1 , wherein the processing array module includes pipeline paths configured with first circuitry to read data elements associated with a first source input and second circuitry to read data elements associated with a second source input.

3. The graphics processor as in claim 2 , wherein the processing array module include third circuitry configured to detect non-zero data elements in the second source input and selectively perform dot product operations based on the non-zero data elements of the second source input and data elements of the first source input that correspond with the non-zero data elements of the second input.

4. The graphics processor as in claim 3 , wherein the processing array module include a pipeline path including separate output hardware for each pipeline stage.

5. The graphics processor as in claim 4 , wherein the processing array module include a first pipeline path configurable to execute a first dot product instruction having a first set of inputs and a second pipeline path configurable to execute a second dot product instruction having a second set of inputs.

6. The graphics processor as in claim 4 , the pipeline path including first output hardware configured to selectably write output of the first pipeline stage to a location selected from one of output memory and a second pipeline stage.

7. The graphics processor as in claim 6 , the pipeline path including second output hardware configured to selectably write output of the final pipeline stage a location selected from one of output memory and the first pipeline stage.

8. The graphics processor as in claim 7 , wherein the second pipeline stage is the final pipeline stage.

9. A method comprising:

receiving non-zero source values and a calculation depth for an instruction to be executed by a matrix accelerator of a general-purpose graphics processor, the calculation depth to specify a number of pipeline stages to use to calculate a dot product in response to the instruction;

evaluating a write enable mask to determine a set of enabled parallel processing channels;

generating a set of products based on an elementwise multiply of source input elements;

calculating a sum of the set of products and adding the sum to a value in an accumulator;

at a processing element at a non-final calculation layer of the matrix accelerator, outputting the accumulator value to a next layer of the matrix accelerator; and

at a processing element at a final calculation layer of the matrix accelerator, outputting the calculated sum to a specified destination register.

10. The method of claim 9 , further comprising:

receiving an initial accumulator value along with the non-zero source values and the calculation depth; and

storing the initial value to the accumulator before adding the sum to the value in the accumulator.

11. The method of claim 9 , wherein outputting the accumulator value to a next layer of the matrix accelerator includes providing a feedback value to a processing element at an initial stage of the processing pipeline.

12. The method of claim 9 , further comprising configuring the write enable mask based on a predicate mask received with the instruction.

13. A data processing system comprising:

a host processor; and

a graphics processor coupled with the host processor via a host interface, the graphics processor comprising a plurality of processing clusters, each processing cluster of the plurality of processing clusters comprising a plurality of multiprocessors, the plurality of multiprocessors interconnected via a data interconnect, and each multiprocessor of the plurality of multiprocessors comprising sparse matrix multiply acceleration hardware including:

a systolic processing array with feedback inputs, the systolic processing array comprising multiple processing array modules, the multiple processing array modules including a processing array module having a first number of pipeline paths, the first number of pipeline paths having a second number of pipeline stages,

wherein a first pipeline stage is configurable to receive feedback output from a final pipeline stage and the processing array module includes pipeline paths configured with shared circuitry to read data elements associated with a first source input and separate circuitry to read data elements associated with a second source input.

14. The data processing system as in claim 13 , wherein the processing array module includes pipeline paths configured with first circuitry to read data elements associated with a first source input and second circuitry to read data elements associated with a second source input.

15. The data processing system as in claim 14 , wherein the processing array module include third circuitry configured to detect non-zero data elements in the second source input and selectively perform dot product operations based on the non-zero data elements of the second source input and data elements of the first source input that correspond with the non-zero data elements of the second input.

16. The data processing system as in claim 15 , wherein the processing array module include a pipeline path including separate output hardware for each pipeline stage.

17. The data processing system as in claim 16 , wherein the processing array module include a first pipeline path configurable to execute a first dot product instruction having a first set of inputs and a second pipeline path configurable to execute a second dot product instruction having a second set of inputs.

18. The data processing system as in claim 16 , the pipeline path including first output hardware configured to selectably write output of the first pipeline stage to a location selected from one of output memory and a second pipeline stage.

19. The data processing system as in claim 18 , the pipeline path including second output hardware configured to selectably write output of the final pipeline stage a location selected from one of output memory and the first pipeline stage.

20. The data processing system as in claim 19 , wherein the second pipeline stage is the final pipeline stage.

Priority Claims (1)
IN 202041019059 · May 5, 2020 · national
Continuity (3)
Continuation 17527882 · Nov 16, 2021
Continuation 16913800 · Jun 26, 2020
Related Publication 20230281272A1 · Sep 7, 2023