IP Library Granted Patent US 12,333,339
Granted Patent B2
US 12,333,339 · App. 17/399,784 · Granted Jun 17, 2025

Affinity-based graphics scheduling

Inventors: Andrew M. Havlir (Orlando, FL); Ajay Simha Modugala (Orlando, FL); Benjamin Bowman (London, GB); Yunjun Zhang (Orlando, FL)
Assignee: Apple Inc.
G06F9/5027G06F9/4881
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,339
App. No.
17/399,784
Granted
Jun 17, 2025
Kind
B2
Abstract

Techniques are disclosed relating to affinity-based scheduling of graphics work. In disclosed embodiments, first and second groups of graphics processor sub-units may share respective first and second caches. Distribution circuitry may receive a software-specified set of graphics work and a software-indicated mapping of portions of the set of graphics work to groups of graphics processor sub-units. The distribution circuitry may assign subsets of the set of graphics work based on the mapping. This may improve cache efficiency, in some embodiments, by allowing graphics work that accesses the same memory areas to be assigned to the same group of sub-units that share a cache.

Claims (68)

1. An apparatus, comprising:

first and second groups of graphics processor sub-units, wherein the first group of sub-units shares a first cache and the second group of sub-units shares a second cache;

distribution circuitry configured to:

receive a software-specified set of graphics work and a software-indicated mapping of portions of the set of graphics work to groups of graphics processor sub-units; and

assign, based on the mapping, a first subset of the set of graphics work to the first group of graphics sub-units and a second subset of the set of graphics work to the second group of graphics sub-units; and

work sharing control circuitry configured to:

determine that another group of sub-units that share a cache have dispatched all of their assigned portion for the set of graphics work;

update a hardware-implemented finite state machine based on the determination, wherein the hardware finite state machine is implemented using sequential circuitry and combinational circuitry; and

override, based on the update to the hardware-implemented finite state machine, the software-indicated mapping of portions of the set of graphics work to groups of graphics processor sub-units, wherein the override re-assigns part of the first subset of the set of graphics work from the first group of graphics units to the other group of sub-units and the part of the first subset would not have been assigned to the other group of sub-units under the software-indicated mapping.

2. The apparatus of claim 1 , wherein the apparatus includes control circuitry configured to store, in configuration registers, multiple mappings of portions of sets of graphics work to groups of graphics processor sub-units.

3. The apparatus of claim 1 , wherein the set of graphics work is a compute kernel.

4. The apparatus of claim 3 , wherein the distribution circuitry includes:

primary kernel walker circuitry configured to determine the portions of the compute kernel;

first group walker circuitry configured to iterate portions of the compute kernel assigned to the first group of graphics sub-units to determine batches of workgroups; and

second group walker circuitry configured to iterate portions of the compute kernel assigned to the second group of graphics sub-units to determine batches of workgroups.

5. The apparatus of claim 4 , wherein the distribution circuitry further includes:

group walker arbitration circuitry configured to select from among batches of workgroups determined by the first and second group walker circuitry; and

sub-unit assign circuitry configured to assign batches selected by the group walker arbitration circuitry to one or more graphics sub-units in the group of sub-units corresponding to the selected group walker circuitry.

6. The apparatus of claim 1 , wherein the apparatus supports mappings of portions of sets of graphics work to groups of graphics processor sub-units for multiple dimensionalities including single-dimension, two-dimensional, and three-dimensional.

7. The apparatus of claim 1 , further comprising:

circuitry that implements a plurality of logical slots, wherein a set of sub-units in the first and second groups each implement multiple distributed hardware slots; and

control circuitry configured to:

assign the set of graphics work to a first logical slot;

determine a distribution rule for the set of graphics work that indicates to whether distribute to all of the graphics processor sub-units in the set or to distribute to only a portion of the graphics processor sub-units; and

determine mappings between the first logical slot and respective sets of one or more distributed hardware slots based on the distribution rule and based on the mappings of portions of the set of graphics work to groups of graphics processor sub-units.

8. A non-transitory computer-readable storage medium having stored thereon design information that specifies a design of at least a portion of a hardware integrated circuit in a format recognized by a semiconductor fabrication system that is configured to use the design information to produce the circuit according to the design, wherein the design information specifies that the circuit includes:

first and second groups of graphics processor sub-units, wherein the first group of sub-units shares a first cache and the second group of sub-units shares a second cache;

distribution circuitry configured to:

receive a software-specified set of graphics work and a software-indicated mapping of portions of the set of graphics work to groups of graphics processor sub-units; and

assign, based on the mapping, a first subset of the set of graphics work to the first group of graphics sub-units and a second subset of the set of graphics work to the second group of graphics sub-units; and

work sharing control circuitry configured to:

determine that another group of sub-units that share a cache have dispatched all of their assigned portion for the set of graphics work;

update a hardware-implemented finite state machine based on the determination wherein the hardware finite state machine is implemented using sequential circuitry and combinational circuitry; and

override, based on the update to the hardware-implemented finite state machine, the software-indicated mapping of portions of the set of graphics work to groups of graphics processor sub-units, wherein the override re-assigns part of the first subset of the set of graphics work from the first group of graphics units to the other group of sub-units and the part of the first subset would not have been assigned to the other group of sub-units under the software-indicated mapping.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the circuit includes control circuitry configured to store, in configuration registers, multiple mappings of portions of sets of graphics work to groups of graphics processor sub-units.

10. The non-transitory computer-readable storage medium of claim 8 , wherein the distribution circuitry includes:

primary walker circuitry configured to determine the portions of the set of graphics work;

first group walker circuitry configured to iterate portions of the set of graphics work assigned to the first group of graphics sub-units to determine subsets of the set of graphics work; and

second group walker circuitry configured to iterate portions of the set of graphics work assigned to the second group of graphics sub-units to determine subsets of the set of graphics work.

11. The non-transitory computer-readable storage medium of claim 10 , wherein the distribution circuitry further includes:

group walker arbitration circuitry configured to select from among subsets determined by the first and second group walker circuitry; and

sub-unit assign circuitry configured to assign batches selected by the group walker arbitration circuitry to one or more graphics sub-units in the group of sub-units corresponding to the selected group walker circuitry.

12. The non-transitory computer-readable storage medium of claim 8 , wherein the circuit supports mappings of portions of sets of graphics work to groups of graphics processor sub-units for multiple dimensionalities of sets of graphics work including single-dimension, two-dimensional, and three-dimensional.

13. The non-transitory computer-readable storage medium of claim 8 , wherein the circuit further comprises:

circuitry that implements a plurality of logical slots, wherein a set of sub-units in the first and second groups each implement multiple distributed hardware slots; and

control circuitry configured to:

assign the set of graphics work to a first logical slot;

determine a distribution rule for the set of graphics work that indicates to whether distribute to all of the graphics processor sub-units in the set or to distribute to only a portion of the graphics processor sub-units; and

determine mappings between the first logical slot and respective sets of one or more distributed hardware slots based on the distribution rule and based on the mappings of portions of the set of graphics work to groups of graphics processor sub-units.

14. A method, comprising:

receiving, by distribution circuitry of a graphics processor that includes first and second groups of graphics processor sub-units, a software-specified set of graphics work and a software-indicated mapping of portions of the set of graphics work to groups of graphics processor sub-units, wherein the first group of sub-units shares a first cache and the second group of sub-units shares a second cache;

assigning, by the distribution circuitry based on the mapping, a first subset of the set of graphics work to the first group of graphics sub-units and a second subset of the set of graphics work to the second group of graphics sub-units;

determining, by work sharing control circuitry of the graphics processor, that another group of sub-units that share a cache have dispatched all of their assigned portion for the set of graphics work;

updating, by the work sharing control circuitry, a hardware-implemented finite state machine based on the determination, wherein the hardware finite state machine is implemented using sequential circuitry and combinational circuitry; and

overriding, by the work sharing control circuitry based on the update to the hardware-implemented finite state machine, the software-indicated mapping of portions of the set of graphics work to groups of graphics processor sub-units, wherein the override re-assigns part of the first subset of the set of graphics work from the first group of graphics units to the other group of sub-units and the part of the first subset would not have been assigned to the other group of sub-units under the software-indicated mapping.

15. The method of claim 14 , further comprising:

disabling affinity-based work distribution in one or more modes of operation.

16. The method of claim 14 , further comprising:

assigning subsets of sets of graphics work to groups of graphics sub-units, based on multiple different mappings, wherein the multiple different mappings include mappings for at least two dimensionalities of sets of graphics work.

17. The method of claim 14 , wherein:

the graphics processor implements a plurality of logical slots;

the sub-units in the first and second groups each implement multiple distributed hardware slots; and

the method further includes specifying one or more software overrides to at least partially control a mapping between logical slots and respective sets of one or more distributed hardware slots for the set of graphics work.

18. The method of claim 17 , wherein the one or more software overrides include at least one of the following overrides:

mask information that indicates which sub-units are available to the set of graphics work;

a specified distribution rule that indicates distribution breadth;

group information that indicates a group of sub-units on which the set of graphics work should be deployed; and

policy information that indicates a scheduling policy.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2021
From: HAVLIR, ANDREW M.; MODUGALA, AJAY SIMHA; BOWMAN, BENJAMIN; ZHANG, YUNJUN
To: APPLE INC.
Reel/Frame 057151/0334 →
Continuity (1)
Related Publication 20230047481A1 · Feb 16, 2023
References Cited (88)
US 6088779A · Bharadhwaj · 2000 [cited by applicant]
US 6332191B1 · Witt · 2001 [cited by applicant]
US 6363471B1 · Meier · 2002 [cited by applicant]
US 6481251B1 · Meier · 2002 [cited by applicant]
US 6519678B1 · Basham et al. · 2003 [cited by applicant]
US 7406587B1 · Zhang · 2008 [cited by applicant]
US 8115773B2 · Swift et al. · 2012 [cited by applicant]
US 8120608B2 · Jiao et al. · 2012 [cited by applicant]
US 8174534B2 · Jiao · 2012 [cited by applicant]
US 8626141B2 · Davies-Moore · 2014 [cited by examiner]
US 9952987B2 · Guddeti · 2018 [cited by examiner]
US 10026145B2 · Bourd et al. · 2018 [cited by applicant]
US 10325341B2 · Vembu et al. · 2019 [cited by applicant]
US 10635439B2 · Alsup · 2020 [cited by examiner]
US 11074109B2 · Valerio · 2021 [cited by examiner]
US 11714995B2 · Wang · 2023 [cited by applicant]
US 12086644B2 · Havlir · 2024 [cited by examiner]
US 20070030280A1 · Paltashev et al. · 2007 [cited by applicant]
US 20080126751A1 · Mizrachi · 2008 [cited by examiner]
US 20080266296A1 · Ramey · 2008 [cited by examiner]
US 20080303846A1 · Brichter · 2008 [cited by applicant]
US 20100123717A1 · Jiao · 2010 [cited by applicant]
US 20100251016A1 · Abernathy · 2010 [cited by applicant]
US 20110302585A1 · Dice · 2011 [cited by applicant]
US 20120320070A1 · Arvo · 2012 [cited by applicant]
US 20130021353A1 · Drebin et al. · 2013 [cited by applicant]
US 20130064275A1 · Mourad et al. · 2013 [cited by applicant]
US 20130290971A1 · Chen et al. · 2013 [cited by applicant]
US 20140123145A1 · Barrow-Williams et al. · 2014 [cited by applicant]
US 20140373028A1 · Lyashevsky et al. · 2014 [cited by applicant]
US 20140380003A1 · Hsu · 2014 [cited by examiner]
US 20150235341A1 · Mei · 2015 [cited by examiner]
US 20160062894A1 · Schwetman, Jr. et al. · 2016 [cited by applicant]
US 20160266901A1 · Winkel et al. · 2016 [cited by applicant]
US 20160349832A1 · Jiao · 2016 [cited by examiner]
US 20170069054A1 · Ramadoss · 2017 [cited by examiner]
US 20170178386A1 · Redshaw et al. · 2017 [cited by applicant]
US 20180046512A1 · Kang · 2018 [cited by examiner]
US 20180060995A1 · Doyle · 2018 [cited by applicant]
US 20180089881A1 · Johnson · 2018 [cited by applicant]
US 20180349204A1 · Liu · 2018 [cited by applicant]
US 20190019267A1 · Suresh · 2019 [cited by applicant]
US 20190155660A1 · McQuighan et al. · 2019 [cited by applicant]
US 20190188148A1 · Dong · 2019 [cited by applicant]
US 20190206018A1 · Khodakovsky et al. · 2019 [cited by applicant]
US 20190318446A1 · Ray et al. · 2019 [cited by applicant]
US 20190369707A1 · Iwamoto · 2019 [cited by applicant]
US 20190370059A1 · Puthoor et al. · 2019 [cited by applicant]
US 20200043123A1 · Dash et al. · 2020 [cited by applicant]
US 20200104180A1 · Banerjee · 2020 [cited by examiner]
US 20200175645A1 · King et al. · 2020 [cited by applicant]
US 20200183700A1 · Battle et al. · 2020 [cited by applicant]
US 20200183779A1 · Cariello · 2020 [cited by applicant]
US 20200201690A1 · Sankaralingam · 2020 [cited by examiner]
US 20200293380A1 · Ashbaugh · 2020 [cited by examiner]
US 20200310883A1 · Valerio · 2020 [cited by examiner]
US 20200312006A1 · Du et al. · 2020 [cited by applicant]
US 20200356366A1 · Battle · 2020 [cited by applicant]
US 20200401529A1 · Zhang · 2020 [cited by examiner]
US 20200410628A1 · Shah et al. · 2020 [cited by applicant]
US 20210104010A1 · Bakalash · 2021 [cited by applicant]
US 20210110506A1 · Prakash et al. · 2021 [cited by applicant]
US 20210158450A1 · Davis · 2021 [cited by examiner]
US 20210294646A1 · Hassaan · 2021 [cited by examiner]
US 20210294648A1 · Champigny · 2021 [cited by examiner]
US 20220207813A1 · Godey · 2022 [cited by applicant]
US 20220308938A1 · Iyer et al. · 2022 [cited by applicant]
CN 102982505A · 2013 [cited by applicant]
CN 103124957A · 2013 [cited by applicant]
CN 110322390A · 2019 [cited by applicant]
CN 110764901A · 2020 [cited by applicant]
CN 112527513A · 2021 [cited by applicant]
JP 2008538620A · 2008 [cited by applicant]
JP 2008276740A · 2008 [cited by applicant]
JP 2014523021A · 2014 [cited by applicant]
JP 2020109625A · 2020 [cited by applicant]
JP 2020113252A · 2020 [cited by applicant]
KR 1020120031759A · 2012 [cited by applicant]
KR 1020200086613A · 2020 [cited by applicant]
KR 20200086613A · 2020 [cited by applicant]
WO 2019165428A1 · 2019 [cited by applicant]
Zahaf et al.; “Contention-Aware GPU Partitioning and Task-to-Partition Allocation for Real-TimeWorkloads”; Apr. 9, 2021, Association for Computing Machinery; https://doi.org/10.1145/3453417.3453439; (Zahaf_2021.pdf; pp.… [cited by examiner]
Jain et al.; “Fractional GPUs: Software-based Compute and Memory Bandwidth Reservation for GPUs”; Carnegie Mellon University; 2019 IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS); (Jain_2019.pdf… [cited by examiner]
Machi Xue et al., “Scalable GPU Virtualization with Dynamic Sharing of Graphics Memory Space,” IEE Transactions on Parallel and Distributed Systems, vol. 29, No. 8, Aug. 2018, pp. 1823-1836. [cited by applicant]
Owen Harrison, Practical Symmetric Key Cryptography on Modern Graphics Hardware. (Year: 2008). [cited by applicant]
International Search Report and Written Opinion in PCT Appl. No. PCT/US2022/038061 mailed Nov. 4, 2022, 10 pages. [cited by applicant]
Office Action in Japanese Appl. No. 2024-508431 mailed Feb. 25, 2025, 3 pages. [cited by applicant]
Kuo, Hsien-Kai et al., Thread Affinity Mapping for Irregular Data Access on Shared CacheGPGPU, Proceedings of the 17th Asia and South Pacific Design Automation Conference, IEEE, Jan. 30, 2012, pp. 659-664, https://ieeex… [cited by applicant]