IP Library › Granted Patent US 12,724,615
Granted Patent B2
US 12,724,615 · App. 18/911,522 · Granted Sep 1, 2026

Granular source read scheduling for instruction execution

Inventors: David K. Li (Austin, TX); Ana Lucia Rescala Loper (Austin, TX); Arun Bansal (Austin, TX); Chance C. Coats (Austin, TX); Zeran Zhu (Austin, TX)
Assignee: Apple Inc.
G06F9/3888G06F9/3836G06F9/5022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,615
App. No.
18/911,522
Granted
Sep 1, 2026
Kind
B2
Abstract

Techniques are disclosed relating to accessing source data in single-instruction multiple-thread (SIMT) pipelines. In some embodiments, multiple categories of operand resource circuits are configured to provide operands for instructions executed by processor pipeline circuitry. Per-resource arbitration circuitry may arbitrate between the SIMT execution slots for access to different operand resources. Source access circuitry may access operand data from operand resources based on source capture commands and source control circuitry may prior to a first SIMT group winning arbitration for all its operands, send a source capture command to the source access circuitry in response to the first SIMT group winning arbitration at the per-resource arbitration circuitry for a first operand resource. Instruction control circuitry may send an instruction release command down the processor pipeline circuitry for the first SIMT group, in response to the first SIMT group winning arbitration for all its operands.

Claims (74)

1 . An apparatus, comprising:

processor pipeline circuitry;

first and second categories of operand resource circuits configured to provide operand data for instructions to be executed by the processor pipeline circuitry;

wherein the processor pipeline circuitry includes:

first resource arbitration circuitry configured to arbitrate between instructions of single-instruction multiple-thread (SIMT) groups for access to the first operand resource category;

second resource arbitration circuitry configured to arbitrate between SIMT groups for access to the second operand resource category;

source access circuitry configured to access the operand data from the operand resource circuits based on source capture commands; and

source control circuitry configured to, prior to an instruction of a first SIMT group winning arbitration for all its operands and in response the instruction winning arbitration at the first resource arbitration circuitry for a first operand resource of the operand resource circuits, send a source capture command to the source access circuitry to access operand data from the first operand resource for the instruction; and

instruction control circuitry configured to generate an instruction release command for the instruction, in response to the instruction winning arbitration for all its operands, including at the second resource arbitration circuitry.

2 . The apparatus of claim 1 , wherein the processor pipeline circuitry further includes:

a source landing stage configured to buffer accessed operand data from the source access circuitry, for the instruction of the first SIMT group, prior to the instruction release command.

3 . The apparatus of claim 1 , wherein the operand resource circuits include:

multiple operand cache read ports;

multiple data cache read ports; and

uniform storage circuitry.

4 . The apparatus of claim 1 , wherein the first resource arbitration circuitry implements a different arbitration scheme than the second resource arbitration circuitry.

5 . The apparatus of claim 1 , wherein:

the processor pipeline circuitry includes multiple pipelines that each include multiple SIMT slots; and

the first resource arbitration circuitry is configured to arbitrate among instructions of SIMT groups from multiple pipelines.

6 . The apparatus of claim 1 , further comprising:

first-stage scheduler circuitry configured to arbitrate among SIMT groups to assign SIMT groups to SIMT execution slots; and

second-stage scheduling circuitry configured to arbitrate among instructions in SIMT execution slots for assignment to execution resources, wherein the second-stage scheduler circuitry includes the first resource arbitration circuitry.

7 . The apparatus of claim 1 , further comprising:

buffer circuitry configured to:

buffer multiple instructions per SIMT group; and

provide a next instruction for arbitration for an operand resource in a next cycle subsequent to another instruction from the same SIMT group winning arbitration for the operand resource; and

landing buffer circuitry configured to buffer operand data retrieved for the multiple instructions per SIMT group.

8 . The apparatus of claim 1 , wherein:

the operand resource circuits include one or more operand caches that are dynamically managed by hardware; and

control circuitry is configured to:

lock a given entry in at least one of the one or more operand caches in response to a hit for the given entry; and

determine to unlock the given entry in response to an unlock event.

9 . The apparatus of claim 1 , further comprising:

fixed-function circuitry configured to control the processor pipeline circuitry to perform operations for at least one of the following types of programs:

graphics shader programs; and

machine learning programs.

10 . The apparatus of claim 1 , wherein the apparatus is a computing device that further includes:

a display; and

network interface circuitry.

11 . A method, comprising:

arbitrating, by a computing system, between instructions of single-instruction multiple-thread (SIMT) execution groups for access to a first operand resource category;

arbitrating, by the computing system, between instructions of SIMT groups for access to a second operand resource category;

sending, by the computing system down pipeline circuitry of the computing system to source access circuitry of the computing system, prior to an instruction of a first SIMT slot winning arbitration for all its operands, a source capture command in response to the instruction winning arbitration for a first operand resource; and

accessing, by the source access circuitry, operand data from one or more operand resources based on the source capture command; and

generating, by the computing system subsequent to the accessing, an instruction release command for the instruction, in response to the instruction winning arbitration for all its operands.

12 . The method of claim 11 , further comprising:

buffering, by the computing system, accessed operand data from the source access circuitry, for the instruction of the first SIMT slot, prior to the instruction release command.

13 . The method of claim 11 , wherein the operand resources include:

multiple operand cache read ports; and

multiple data cache read ports.

14 . The method of claim 11 , wherein the arbitrating applies different arbitration schemes for different categories of operand resource circuits.

15 . The method of claim 11 , wherein:

the pipeline circuitry includes multiple pipelines that each include multiple SIMT slots; and

the arbitrating includes arbitrating among SIMT slots from multiple pipelines.

16 . The method of claim 15 , further comprising:

arbitrating among SIMT groups to assign SIMT groups to SIMT execution slots.

17 . The method of claim 11 , further comprising:

buffering, by the computing system, multiple instructions per SIMT slot; and

providing, by the computing system, a next instruction for arbitration for an operand resource in a next cycle subsequent to another instruction from the same SIMT slot winning arbitration for the operand resource; and

buffering, by the computing system, operand data retrieved for the multiple instructions per SIMT slot.

18 . The method of claim 11 , wherein the operand resources include one or more operand caches that are dynamically managed by hardware, the method further comprising:

locking a given entry in at least one of the one or more operand caches in response to a hit for the given entry; and

determining to unlock the given entry in response to an unlock event.

19 . A non-transitory computer-readable medium having instructions of a hardware description programming language stored thereon that, when processed by a computing system, program the computing system to generate a computer simulation model, wherein the model represents a hardware circuit that includes:

processor pipeline circuitry;

multiple first and second categories of operand resource circuits configured to provide operand data for instructions to be executed by the processor pipeline circuitry;

wherein the processor pipeline circuitry includes:

first resource arbitration circuitry configured to arbitrate between instructions of single-instruction multiple-thread (SIMT) groups for access to the first operand resource category;

second resource arbitration circuitry configured to arbitrate between SIMT groups for access to the second operand resource category;

source access circuitry configured to access the operand data from the operand resources based on source capture commands; and

source control circuitry configured to, prior to an instruction of a first SIMT group winning arbitration for all its operands and in response the instruction winning arbitration at the first resource arbitration circuitry for a first operand resource of the operand resource circuits, send a source capture command to the source access circuitry to access operand data from the first operand resource for the instruction; and

instruction control circuitry configured to generate an instruction release command for the instruction first SIMT group, in response to the instruction winning arbitration for all its operands, including at the second resource arbitration circuitry.

20 . The non-transitory computer-readable medium of claim 19 , wherein the processor pipeline circuitry further includes:

a source landing stage configured to buffer accessed operand data from the source access circuitry, for the instruction of the first SIMT group, prior to the instruction release command.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2024
From: LI, DAVID K.; LOPER, ANA LUCIA RESCALA; BANSAL, ARUN; COATS, CHANCE C.; ZHU, ZERAN
To: APPLE INC.
Reel/Frame 068861/0977 →
Continuity (2)
Provisional Application 63696562 · Sep 19, 2024
Related Publication 20260079715A1 · Mar 19, 2026
References Cited (52)
US 5974438A · Neufeld · 1999 [cited by applicant]
US 6105051A · Borkenhagen et al. · 2000 [cited by applicant]
US 6145054A · Mehrotra et al. · 2000 [cited by applicant]
US 7237093B1 · Musoll et al. · 2007 [cited by applicant]
US 7418576B1 · Lindholm et al. · 2008 [cited by applicant]
US 7478388B1 · Chen et al. · 2009 [cited by applicant]
US 7895415B2 · Gonzalez · 2011 [cited by applicant]
US 8095778B1 · Golla · 2012 [cited by applicant]
US 8533719B2 · Fedorova · 2013 [cited by applicant]
US 10027587B1 · O'Brien et al. · 2018 [cited by applicant]
US 10089114B2 · Kulkarni et al. · 2018 [cited by applicant]
US 10585670B2 · Abdallah · 2020 [cited by applicant]
US 10642618B1 · Hakewill · 2020 [cited by applicant]
US 11301298B2 · Varma · 2022 [cited by applicant]
US 11314562B2 · Dice · 2022 [cited by applicant]
US 20040059896A1 · Kossman et al. · 2004 [cited by applicant]
US 20040060052A1 · Brown et al. · 2004 [cited by applicant]
US 20050114856A1 · Eickemeyer et al. · 2005 [cited by applicant]
US 20050210158A1 · Cowperthwaite et al. · 2005 [cited by applicant]
US 20060179280A1 · Jensen et al. · 2006 [cited by applicant]
US 20070103476A1 · Huang et al. · 2007 [cited by applicant]
US 20070143582A1 · Coon et al. · 2007 [cited by applicant]
US 20090210677A1 · Luick · 2009 [cited by applicant]
US 20100083267A1 · Adachi et al. · 2010 [cited by applicant]
US 20110276784A1 · Gewirtz et al. · 2011 [cited by applicant]
US 20120079503A1 · Dally et al. · 2012 [cited by applicant]
US 20130166881A1 · Choquette et al. · 2013 [cited by applicant]
US 20140160126A1 · Legakis et al. · 2014 [cited by applicant]
US 20140282566A1 · Lindholm · 2014 [cited by applicant]
US 20150046684A1 · Mehrara et al. · 2015 [cited by applicant]
US 20160246728A1 · Ron et al. · 2016 [cited by applicant]
US 20160292814A1 · Holland · 2016 [cited by applicant]
US 20180046577A1 · Chen et al. · 2018 [cited by applicant]
US 20180181491A1 · DeLaurier et al. · 2018 [cited by applicant]
US 20190018676A1 · Alexander · 2019 [cited by applicant]
US 20190340019A1 · Brewer · 2019 [cited by applicant]
US 20190370059A1 · Puthoor et al. · 2019 [cited by applicant]
US 20200043123A1 · Dash · 2020 [cited by applicant]
US 20200293450A1 · Vemulapalli et al. · 2020 [cited by applicant]
US 20210073450A1 · Boesch · 2021 [cited by applicant]
US 20210326139A1 · Gupta · 2021 [cited by examiner]
US 20210342158A1 · Brewer · 2021 [cited by applicant]
US 20210382717A1 · Jiang · 2021 [cited by examiner]
US 20220043413A1 · Heydari · 2022 [cited by applicant]
US 20220091657A1 · Tsien · 2022 [cited by applicant]
US 20220100484A1 · Abu-Ghazaleh et al. · 2022 [cited by applicant]
US 20220206876A1 · Beckmann et al. · 2022 [cited by applicant]
US 20230367676A1 · Jeyapaul et al. · 2023 [cited by applicant]
US 20230401132A1 · Patle et al. · 2023 [cited by applicant]
US 20240419447A1 · Chaudhari · 2024 [cited by examiner]
WO 2008061154A3 · 2008 [cited by applicant]
WO 2013098643A2 · 2013 [cited by applicant]