IP Library › Granted Patent US 11,803,934
Granted Patent B2
US 11,803,934 · App. 17/591,152 · Granted Oct 31, 2023

Handling pipeline submissions across many compute units

Inventors: Balaji Vembu (Folsom, CA); Altug Koker (El Dorado Hills, CA); Joydeep Ray (Folsom, CA)
Assignee: Intel Corporation
G06T1/20G06T15/005G06T2200/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,803,934
App. No.
17/591,152
Granted
Oct 31, 2023
Kind
B2
Abstract

One embodiment provides an apparatus comprising an interconnect fabric comprising one or more fabric switches, a plurality of memory interfaces coupled to the interconnect fabric to provide access to a plurality of memory devices, an input/output (IO) interface coupled to the interconnect fabric to provide access to IO devices, an array of multiprocessors coupled to the interconnect fabric, scheduling circuitry to distribute a plurality of thread groups across the array of multiprocessors, each thread group comprising a plurality of threads and each thread comprising a plurality of instructions to be executed by at least one of the multiprocessors, and a first multiprocessor of the array of multiprocessors to be assigned to process a first thread group comprising a first plurality of threads, the first multiprocessor comprising a plurality of parallel execution circuits.

Claims (32)

1. A graphics processor comprising:

first circuitry to distribute a plurality of thread groups associated with program code for execution by the graphics processor; and

a processing cluster including:

a plurality of streaming multiprocessors;

a data crossbar configured to interconnect the plurality of streaming multiprocessors; and

second circuitry to receive a thread group from the first circuitry, divide the thread group into a plurality of thread sub-groups, and distribute the plurality of thread sub-groups to the plurality of streaming multiprocessors, wherein the second circuitry is to retire a first thread sub-group of a first thread group and launch a second thread sub-group of the plurality of thread sub-groups.

2. The graphics processor as in claim 1 , wherein the second circuitry includes a thread group manager.

3. The graphics processor as in claim 2 , wherein the thread group manager includes a group identifier manager to allocate a group identifier for the first thread group and a sub-group manager to allocate a sub-group identifier for the first thread sub-group and the second thread sub-group.

4. The graphics processor as in claim 3 , wherein the thread group manager is configured to track a number of active sub-groups of a thread group.

5. The graphics processor as in claim 4 , wherein the thread group manager is additionally configured to track a number of active threads of each sub-group of the thread group.

6. The graphics processor as in claim 5 , wherein the thread group manager includes a launch unit to launch each thread of a thread sub-group on the plurality of streaming multiprocessors.

7. The graphics processor as in claim 6 , the second circuitry additionally including a sub-group retirement queue and wherein to retire the first thread sub-group of the first thread group includes to store the sub-group identifier for the first thread sub-group to the sub-group retirement queue.

8. The graphics processor as in claim 7 , the second circuitry additionally including a run queue and wherein to launch the second thread sub-group of the first thread group includes to store the sub-group identifier for the second thread sub-group to the run queue.

9. The graphics processor as in claim 8 , wherein the thread group manager is to additionally launch a second thread sub-group of second thread group after the first thread sub-group of the first thread group is retired.

10. The graphics processor as in claim 9 , wherein the second circuitry is to configure the plurality of streaming multiprocessors to execute multiple thread sub-groups associated with multiple different thread groups.

11. A data processing system comprising:

a memory device; and

a graphics processor coupled with the memory device, the graphics processor comprising:

first circuitry to distribute a plurality of thread groups associated with program code for execution by the graphics processor; and

a processing cluster including:

a plurality of streaming multiprocessors;

a data crossbar configured to interconnect the plurality of streaming multiprocessors; and

second circuitry to receive a thread group from the first circuitry, divide the thread group into a plurality of thread sub-groups, and distribute the plurality of thread sub-groups to the plurality of streaming multiprocessors, wherein the second circuitry is to retire a first thread sub-group of a first thread group and launch a second thread sub-group of the plurality of thread sub-groups.

12. The data processing system as in claim 11 , wherein the second circuitry includes a thread group manager.

13. The data processing system as in claim 12 , wherein the thread group manager includes a group identifier manager to allocate a group identifier for the first thread group and a sub-group manager to allocate a sub-group identifier for the first thread sub-group and the second thread sub-group.

14. The data processing system as in claim 13 , wherein the thread group manager is configured to track a number of active sub-groups of a thread group.

15. The data processing system as in claim 14 , wherein the thread group manager is additionally configured to track a number of active threads of each sub-group of the thread group.

16. The data processing system as in claim 15 , wherein the thread group manager includes a launch unit to launch each thread of a thread sub-group on the plurality of streaming multiprocessors.

17. The data processing system as in claim 16 , the second circuitry additionally including a sub-group retirement queue and wherein to retire the first thread sub-group of the first thread group includes to store the sub-group identifier for the first thread sub-group to the sub-group retirement queue.

18. The data processing system as in claim 17 , the second circuitry additionally including a run queue and wherein to launch the second thread sub-group of the first thread group includes to store the sub-group identifier for the second thread sub-group to the run queue.

19. The data processing system as in claim 18 , wherein the thread group manager is to additionally launch a second thread sub-group of second thread group after the first thread sub-group of the first thread group is retired.

20. The data processing system as in claim 19 , wherein the second circuitry is to configure the plurality of streaming multiprocessors to execute multiple thread sub-groups associated with multiple different thread groups.

Continuity (6)
Continuation 17197126 · Mar 10, 2021
Continuation 16834902 · Mar 30, 2020
Continuation 16446946 · Jun 20, 2019
Continuation 16150012 · Oct 2, 2018
Continuation 15493233 · Apr 21, 2017
Related Publication 20220230269A1 · Jul 21, 2022