IP Library Granted Patent US 11,397,585
Granted Patent B2
US 11,397,585 · App. 17/173,923 · Granted Jul 26, 2022

Scheduling of threads for execution utilizing load balancing of thread groups

Inventors: Balaji Vembu (Folsom, CA); Abhishek R. Appu (El Dorado Hills, CA); Joydeep Ray (Folsom, CA); Altug Koker (El Dorado Hills, CA)
Assignee: Intel Corporation
G06F9/3851G06F9/46G06F9/4843G06F9/4881G06F9/5027G06F9/522G06F9/545G06F12/0866G06F12/0897G06F15/16G06F15/76G06T1/20G06T1/60G06F2209/5018G06T2200/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,397,585
App. No.
17/173,923
Granted
Jul 26, 2022
Kind
B2
Abstract

An apparatus to facilitate thread scheduling is disclosed. The apparatus includes logic to store barrier usage data based on a magnitude of barrier messages in an application kernel and a scheduler to schedule execution of threads across a plurality of multiprocessors based on the barrier usage data.

Claims (38)

1. An apparatus comprising:

one or more processors including a graphics processor, the one or more processors to analyze an application kernel to determine a magnitude of barrier messages in the application kernel and generate barrier usage data having a value corresponding to the magnitude of barrier messages in the application kernel; and

a memory to store the barrier usage data;

wherein the graphics processor includes:

a plurality of workgroup processors, each workgroup processor including a plurality of compute units for execution of threads in a plurality of wavefronts, each wavefront including a plurality of threads, and

a scheduler to schedule a plurality of wavefronts for execution by the plurality of workgroup processors according to a scheduling policy, the scheduling policy being based at least in part on the barrier usage data, wherein the scheduler is to prioritize scheduling of a set of wavefronts of the plurality of wavefronts to a same workgroup processor of the plurality of workgroup processors upon a determination that the barrier usage data indicates a high magnitude of barrier messages in the set of wavefronts.

2. The apparatus of claim 1 , wherein the memory is to store the barrier usage data as thread metadata for the application kernel.

3. The apparatus of claim 1 , wherein the scheduling policy is further based on load balancing of wavefronts across the plurality of workgroup processors.

4. The apparatus of claim 3 , wherein the scheduler is to perform load balancing by performing a first cost function based on wavefronts scheduled for each of the plurality of workgroup processors.

5. The apparatus of claim 4 , wherein the scheduler is to perform load balancing by further performing a second cost function based at least in part on the barrier usage data.

6. The apparatus of claim 5 , wherein scheduling the wavefronts to the plurality of workgroup processors includes minimizing a sum of values including the first cost function and the second cost function.

7. The apparatus of claim 1 , wherein the one or more processors are to schedule the set of wavefronts to two or more workgroup processors of the plurality of workgroup processors upon a determination that the barrier usage data indicates a low magnitude of barrier messages in the set of wavefronts.

8. A method comprising:

analyzing an application kernel to determine a magnitude of barrier messages in the application kernel;

generating barrier usage data having a value corresponding to the magnitude of barrier messages in the application kernel;

receiving threads in a plurality of wavefronts for scheduling to a plurality of workgroup processors of a graphics processor, each wavefront including a plurality of threads; and

scheduling execution of a plurality of wavefronts to workgroup processors of the plurality of workgroup processors according to a scheduling policy, the scheduling policy being based at least in part on the barrier usage data, wherein scheduling includes prioritizing of a set of wavefronts of the plurality of wavefronts to a same workgroup processor of the plurality of workgroup processors upon a determination that the barrier usage data indicates a high magnitude of barrier messages in the set of wavefronts.

9. The method of claim 8 , further comprising storing the barrier usage data in a computer memory.

10. The method of claim 9 , wherein storing the barrier usage data includes storing the barrier usage data as thread metadata for the application kernel.

11. The method of claim 8 , wherein the scheduling policy is based both on the barrier usage data and on load balancing of thread groups across the plurality of workgroup processors.

12. The method of claim 11 , further comprising:

performing a first cost function based on wavefronts scheduled for each of the plurality of workgroup processors.

13. The method of claim 12 , further comprising:

performing a second cost function based at least in part on the barrier usage data.

14. The method of claim 13 , wherein scheduling the plurality of wavefronts to the plurality of workgroup processors includes minimizing a sum of values including the first cost function and the second cost function.

15. A non-transitory computer readable medium having instructions, which when executed by one or more processors, cause the processors to perform operations comprising:

analyzing an application kernel to determine a magnitude of barrier messages in the application kernel;

generating barrier usage data having a value corresponding to the magnitude of barrier messages in the application kernel;

receiving threads in a plurality of wavefronts for scheduling to a plurality of workgroup processors of a graphics processor, each wavefront including a plurality of threads; and

scheduling execution of the plurality of wavefronts to workgroup processors of the plurality of workgroup processors according to a scheduling policy, the scheduling policy being based at least in part on the barrier usage data, wherein scheduling includes prioritizing of a set of wavefronts of the plurality of wavefronts to a same workgroup processor of the plurality of workgroup processors upon a determination that the barrier usage data indicates a high magnitude of barrier messages in the set of wavefronts.

16. The computer readable medium of claim 15 , having instructions, which when executed by the one or more processors, further cause the processors to perform operations comprising:

storing the barrier usage data in a computer memory.

17. The computer readable medium of claim 15 , wherein the scheduling policy is based both on the barrier usage data and on load balancing of wavefronts across the plurality of workgroup processors.

18. The computer readable medium of claim 15 , having instructions, which when executed by the one or more processors, further cause the processors to perform operations comprising:

performing a first cost function based on wavefronts scheduled for each of the plurality of workgroup processors.

19. The computer readable medium of claim 18 , having instructions, which when executed by the one or more processors, further cause the processors to perform operations comprising:

performing a second cost function based at least in part on the barrier usage data.

20. The computer readable medium of claim 19 , wherein scheduling the plurality of wavefronts to the plurality of workgroup processors includes minimizing a sum of values including the first cost function and the second cost function.

Continuity (4)
Continuation 16825129 · Mar 20, 2020
Continuation 16388444 · Apr 18, 2019
Continuation 15477017 · Apr 1, 2017
Related Publication 20210373899A1 · Dec 2, 2021