IP Library Granted Patent US 12,518,340
Granted Patent B2
US 12,518,340 · App. 18/450,964 · Granted Jan 6, 2026

Geometry kick distribution in graphics processor

Inventors: Arjun Thottappilly (Oviedo, FL); Steven Fishwick (St Albans, GB); Jason D. Carroll (Oviedo, FL)
Assignee: Apple Inc.
G06T1/20G06F9/5061G06F2209/503
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,340
App. No.
18/450,964
Granted
Jan 6, 2026
Kind
B2
Abstract

Disclosed techniques relate to parsing and assigning sets of geometry work to distributed hardware slots. In some embodiments, graphics control circuitry implements a plurality of logical slots. Control circuitry may assign a parse version of a set of geometry work to distributed hardware slots of one or more of the graphics processor sub-units that each implement multiple distributed hardware slots. Control circuitry may determine a number of segments for the set of geometry work based on execution of the parse version and assign determined segments to distributed hardware slots of respective graphics processor sub-units for execution. Stitch circuitry may stitch results of the segments processed by the assigned distributed hardware slots.

Claims (60)

1 . An apparatus, comprising:

a set of graphics processor sub-units that each implement multiple distributed hardware slots; and

control circuitry configured to:

assign a parse version of a set of geometry work to one or more distributed hardware slots of one or more of the graphics processor sub-units;

determine a number of segments for the set of geometry work based on execution of the parse version; and

assign an execution version of determined segments of the set of geometry work to distributed hardware slots of respective graphics processor sub-units for execution; and

stitch circuitry configured to stitch results of the execution version of the segments of the set of geometry work executed by the assigned one or more distributed hardware slots.

2 . The apparatus of claim 1 , wherein the stitch circuitry includes:

stitch control circuitry configured to:

assign stitch work for one or more first data structure categories to hardware stitch slots in primary control circuitry; and

assign stitch work for one or more second data structure categories to distributed hardware stitch slots in respective graphics processor sub-units, wherein the distributed hardware stitch slots include respective memory interfaces to access a memory that stores data structures to be stitched.

3 . The apparatus of claim 2 , wherein the one or more first data structure categories utilize less memory space than the one or more second data structure categories.

4 . The apparatus of claim 2 , wherein:

one or more first data structure categories include a layer identifier cache and a list of closed pages; and

one or more second data structure categories include tile region array headers.

5 . The apparatus of claim 1 , wherein the control circuitry is configured to:

assign parse work for the set of geometry work to at most one distributed hardware slot of a given graphics processor sub-unit; and

assign segment execution work to at most one distributed hardware slot of a given graphics processor sub-unit.

6 . The apparatus of claim 1 , wherein the control circuitry is configured to:

assign the parse version to all graphics processor sub-units in a set of graphics processor sub-units; and

serially assign segments of the determined number of segments to available graphics processor sub-units in the set.

7 . The apparatus of claim 1 , wherein the stitch circuitry is configured to stitch results on multiple graphics processor sub-units that were assigned a segment for the set of geometry work.

8 . The apparatus of claim 1 , wherein the control circuitry is configured to allow different sets of geometry work to execute in parallel on different distributed hardware slots only if the different sets of geometry work share a parameter buffer.

9 . The apparatus of claim 1 , wherein the control circuitry is configured to dynamically change the number of distributed hardware slots assigned to the set of geometry work during execution of the set of geometry work.

10 . The apparatus of claim 1 , wherein one or more graphics processor sub-units of the set are configured to execute fragment work that operates on the stitched results.

11 . A non-transitory computer readable storage medium having stored thereon design information that specifies a design of at least a portion of a hardware integrated circuit in a format recognized by a semiconductor fabrication system that is configured to use the design information to produce the circuit according to the design, wherein the design information specifies that the circuit includes:

a set of graphics processor sub-units that each implement multiple distributed hardware slots; and

control circuitry configured to:

assign a parse version of a set of geometry work to one or more distributed hardware slots of one or more of the graphics processor sub-units;

determine a number of segments for the set of geometry work based on execution of the parse version; and

assign an execution version of determined segments of the set of geometry work to distributed hardware slots of respective graphics processor sub-units for execution; and

stitch circuitry configured to stitch results of the execution version of the segments of the set of geometry work executed by the assigned one or more distributed hardware slots.

12 . The non-transitory computer readable storage of claim 11 , wherein the stitch circuitry includes:

stitch control circuitry configured to:

assign stitch work for one or more first data structure categories to hardware stitch slots in primary control circuitry; and

assign stitch work for one or more second data structure categories to distributed hardware stitch slots in respective graphics processor sub-units, wherein the distributed hardware stitch slots include respective memory interfaces to access a memory that stores data structures to be stitched.

13 . The non-transitory computer readable storage of claim 12 , wherein:

one or more first data structure categories include a layer identifier cache and a list of closed pages; and

one or more second data structure categories include tile region array headers.

14 . The non-transitory computer readable storage of claim 11 , wherein the control circuitry is configured to:

assign parse work for the set of geometry work to at most one distributed hardware slot of a given graphics processor sub-unit; and

assign segment execution work to at most one distributed hardware slot of a given graphics processor sub-unit.

15 . The non-transitory computer readable storage of claim 11 , wherein the control circuitry is configured to:

assign the parse version to all graphics processor sub-units in a set of graphics processor sub-units; and

serially assign segments of the determined number of segments to available graphics processor sub-units in the set.

16 . The non-transitory computer readable storage of claim 11 , wherein the control circuitry is configured to dynamically change the number of distributed hardware slots assigned to the set of geometry work during execution of the set of geometry work.

17 . A method, comprising:

assigning, by a computing device, a parse version of a set of geometry work to one or more distributed hardware slots of one or more graphics processor sub-units;

determining, by the computing device, a number of segments for the set of geometry work based on execution of the parse version;

assigning, by the computing device, an execution version of determined segments of the set of geometry work to distributed hardware slots of respective graphics processor sub-units for execution; and

stitching, by the computing device, results of the execution version of the segments of the set of geometry work executed by the assigned one or more distributed hardware slots.

18 . The method of claim 17 , further comprising:

assigning, by the computing device, work for one or more first data structure categories to hardware stitch slots in primary control circuitry; and

assigning, by the computing device, work for one or more second data structure categories to distributed hardware stitch slots in respective graphics processor sub-units, wherein the distributed hardware stitch slots include respective memory interfaces to access a memory that stores data structures to be stitched.

19 . The method of claim 18 , wherein:

one or more first data structure categories include a layer identifier cache and a list of closed pages; and

one or more second data structure categories include tile region array headers.

20 . The method of claim 18 , further comprising:

assigning, by the computing device, parse work for the set of geometry work to at most one distributed hardware slot of a given graphics processor sub-unit; and

assigning, by the computing device, segment execution work to at most one distributed hardware slot of a given graphics processor sub-unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2023
From: THOTTAPPILLY, ARJUN; FISHWICK, STEVEN; CARROLL, JASON D.
To: APPLE INC.
Reel/Frame 064613/0538 →
Continuity (2)
Provisional Application 63484893 · Feb 14, 2023
Related Publication 20240273667A1 · Aug 15, 2024
References Cited (40)
US 5664200A · Barlow et al. · 1997 [cited by applicant]
US 6247064B1 · Alferness et al. · 2001 [cited by applicant]
US 7664942B1 · Tremblay · 2010 [cited by applicant]
US 9552206B2 · Johnson et al. · 2017 [cited by applicant]
US 9582320B2 · Holt et al. · 2017 [cited by applicant]
US 10733693B2 · Schluessler et al. · 2020 [cited by applicant]
US 10761822B1 · Borkovic et al. · 2020 [cited by applicant]
US 10956359B2 · Targowski et al. · 2021 [cited by applicant]
US 11021944B2 · Zheng et al. · 2021 [cited by applicant]
US 20040252711A1 · Romano et al. · 2004 [cited by applicant]
US 20130021353A1 · Drebin et al. · 2013 [cited by applicant]
US 20170329646A1 · Drebin et al. · 2017 [cited by applicant]
US 20180315157A1 · Ould-Ahmed-Vall · 2018 [cited by examiner]
US 20180349146A1 · Iwamoto · 2018 [cited by examiner]
US 20190019267A1 · Suresh · 2019 [cited by applicant]
US 20190108671A1 · Rollingson · 2019 [cited by examiner]
US 20200042321A1 · Genden · 2020 [cited by applicant]
US 20200098160A1 · Havlir et al. · 2020 [cited by applicant]
US 20200112708A1 · Morgan · 2020 [cited by examiner]
US 20200151847A1 · Schluessler et al. · 2020 [cited by applicant]
US 20200310883A1 · Valerio · 2020 [cited by examiner]
US 20200334889A1 · Rollingson et al. · 2020 [cited by applicant]
US 20200401529A1 · Zhang · 2020 [cited by examiner]
US 20210241418A1 · Vembu et al. · 2021 [cited by applicant]
US 20210294660A1 · Uralsky et al. · 2021 [cited by applicant]
US 20220005148A1 · Cerny et al. · 2022 [cited by applicant]
US 20220020108A1 · Uhrenholt et al. · 2022 [cited by applicant]
US 20220051476A1 · Woop · 2022 [cited by examiner]
US 20220261950A1 · Howson · 2022 [cited by examiner]
US 20220397809A1 · Talpade · 2022 [cited by examiner]
US 20230048951A1 · Fishwick et al. · 2023 [cited by applicant]
US 20230050061A1 · Havlir et al. · 2023 [cited by applicant]
US 20230051906A1 · Havlir et al. · 2023 [cited by applicant]
US 20230305978A1 · Davis · 2023 [cited by examiner]
CN 111767080A · 2020 [cited by applicant]
CN 112527513A · 2021 [cited by applicant]
JP 2021099786A · 2021 [cited by applicant]
WO 2018235124A1 · 2018 [cited by applicant]
WO 2023018529A1 · 2023 [cited by applicant]
Markus Steinberger, “On Dynamic Scheduling for the GPU and its Applications in Computer Graphics and Beyond,” IEEE Computer Graphics and Applications, vol. 38, Issue: 3, May 31, 2018, pp. 119-130. [cited by applicant]