IP Library › Granted Patent US 11,620,723
Granted Patent B2
US 11,620,723 · App. 17/870,169 · Granted Apr 4, 2023

Handling pipeline submissions across many compute units

Inventors: Balaji Vembu (Folsom, CA); Altug Koker (El Dorado Hills, CA); Joydeep Ray (Folsom, CA)
Assignee: Intel Corporation
G06T1/20G06T15/005G06T2200/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,620,723
App. No.
17/870,169
Granted
Apr 4, 2023
Kind
B2
Abstract

One embodiment provides a graphics processor including a plurality of processing clusters, each processing cluster including a plurality of multiprocessors and a data interconnect coupled to the plurality of multiprocessors. At least one multiprocessor of the plurality of multiprocessors is configured to share data with another multiprocessor over the data interconnect.

Claims (57)

1. A graphics processor comprising:

a plurality of processing clusters, each processing cluster comprising a plurality of multiprocessors;

a data interconnect coupled to the plurality of multiprocessors, at least one multiprocessor of the plurality of multiprocessors to share data with another multiprocessor over the data interconnect;

scheduling hardware logic to schedule a collection of threads for execution, the collection of threads comprising a plurality of thread groups to be executed on a corresponding plurality of multiprocessors of the processing cluster; and

a first multiprocessor of the plurality of multiprocessors to execute a first thread group to produce a result, and to share the result with a second thread group executed on a second multiprocessor via the data interconnect coupled to the plurality of multiprocessors;

the first multiprocessor comprising:

single-instruction multiple thread (SIMT) execution circuitry to simultaneously execute instructions of the first thread group, the first thread group comprising a plurality of cooperating threads, the cooperating threads comprising one or more producer threads to produce data for consumption by one or more consumer threads of the cooperating threads.

2. The graphics processor of claim 1 , further comprising:

a counter to store a count value associated with execution of threads in the first thread group, the counter to be updated in response to a change in activity of a thread in the first thread group.

3. The graphics processor of claim 2 , wherein the change in activity includes the thread completing execution.

4. The graphics processor of claim 2 , wherein, in accordance with a barrier operation, the thread in the first thread group is to wait for one or more other threads in the first thread group to complete execution.

5. The graphics processor of claim 4 , wherein the SIMT execution circuitry is to cause the thread to sleep while waiting.

6. The graphics processor of claim 5 , wherein the thread is to include dot product instructions.

7. The graphics processor of claim 1 , wherein the second multiprocessor comprises:

single-instruction multiple thread (SIMT) execution circuitry to simultaneously execute instructions of the second thread group using the result.

8. The graphics processor of claim 1 , further comprising:

a local memory to share data between threads in the first thread group.

9. The graphics processor of claim 1 , wherein the SIMT execution circuitry comprises a plurality of functional units or cores.

10. The graphics processor of claim 9 , wherein the plurality of functional units or cores comprises a number of floating point units and/or integer units, the graphics processor further comprising:

a memory interface to couple the plurality of multiprocessors to a system memory.

11. The graphics processor of claim 10 , wherein the system memory comprises a 3D stacked memory.

12. The graphics processor of claim 11 , further comprising:

coherency hardware logic to maintain coherency of the shared result.

13. A method comprising:

scheduling a collection of threads for execution on a graphics processor including a plurality of processing clusters, each processing cluster comprising a plurality of multiprocessors and a data interconnect coupled to the plurality of multiprocessors;

executing, via a first multiprocessor of the plurality of multiprocessors, a first thread group to produce a result, wherein the first multiprocessor comprises single-instruction multiple thread (SIMT) execution circuitry configured to simultaneously execute instructions of the first thread group, the first thread group comprises a plurality of cooperating threads, the cooperating threads comprising one or more producer threads to produce data for consumption by one or more consumer threads of the cooperating threads; and

sharing the result with a second thread group executed on a second multiprocessor of the plurality of multiprocessors via the data interconnect coupled to the plurality of multiprocessors, the second multiprocessor comprising single-instruction multiple thread (SIMT) execution circuitry to simultaneously execute instructions of the second thread group using the result.

14. The method of claim 13 , further comprising:

storing a count value to a counter, the count value associated with execution of threads in the first thread group; and

updating the counter in response to a change in activity of a thread in the first thread group.

15. The method of claim 14 , wherein the change in activity includes the thread completing execution.

16. The method of claim 14 , wherein, in accordance with a barrier operation, the thread in the first thread group is to wait for one or more other threads in the first thread group to complete execution.

17. The method of claim 16 , wherein the SIMT execution circuitry is to cause the thread to sleep while waiting.

18. The method of claim 17 , wherein the thread is to include dot product instructions.

19. A data processing system comprising:

a system interconnect; and

a graphics processor coupled with the system interconnect, the graphics processor including a plurality of processing clusters, each processing cluster comprising a plurality of multiprocessors, the graphics processor additionally including:

a data interconnect coupled to the plurality of multiprocessors, at least one multiprocessor of the plurality of multiprocessors to share data with another multiprocessor over the data interconnect;

scheduling hardware logic to schedule a collection of threads for execution, the collection of threads comprising a plurality of thread groups to be executed on a corresponding plurality of multiprocessors of the processing cluster; and

a first multiprocessor of the plurality of multiprocessors to execute a first thread group to produce a result, and to share the result with a second thread group executed on a second multiprocessor via the data interconnect coupled to the plurality of multiprocessors, the first multiprocessor comprising:

single-instruction multiple thread (SIMT) execution circuitry to simultaneously execute instructions of the first thread group, the first thread group comprising a plurality of cooperating threads, the cooperating threads comprising one or more producer threads to produce data for consumption by one or more consumer threads of the cooperating threads.

20. The data processing system of claim 19 , further comprising:

a counter to store a count value associated with execution of threads in the first thread group, the counter to be updated in response to a change in activity of a thread in the first thread group.

21. The data processing system of claim 20 , wherein the change in activity includes the thread completing execution.

22. The data processing system of claim 20 , wherein, in accordance with a barrier operation, the thread in the first thread group is to wait for one or more other threads in the first thread group to complete execution.

23. The data processing system of claim 22 , wherein the SIMT execution circuitry is to cause the thread to sleep while waiting.

24. The data processing system of claim 23 , wherein the thread is to include dot product instructions.

25. The data processing system of claim 19 , wherein the second multiprocessor comprises:

single-instruction multiple thread (SIMT) execution circuitry to simultaneously execute instructions of the second thread group using the result.

26. The data processing system of claim 19 , further comprising:

a local memory to share data between threads in the first thread group.

27. The data processing system of claim 19 , wherein the SIMT execution circuitry comprises a plurality of functional units or cores.

28. The data processing system of claim 27 , wherein the plurality of functional units or cores comprises a number of floating point units and/or integer units, the graphics processor further comprising:

a memory interface to couple the plurality of multiprocessors to a system memory.

29. The data processing system of claim 28 , wherein the system memory comprises a 3D stacked memory.

30. The data processing system of claim 29 , further comprising:

coherency hardware logic to maintain coherency of the shared result.

Continuity (7)
Continuation 17591152 · Feb 2, 2022
Continuation 17197126 · Mar 10, 2021
Continuation 16834902 · Mar 30, 2020
Continuation 16446946 · Jun 20, 2019
Continuation 16150012 · Oct 2, 2018
Continuation 15493233 · Apr 21, 2017
Related Publication 20220358618A1 · Nov 10, 2022