Granular source read scheduling for instruction execution
Techniques are disclosed relating to accessing source data in single-instruction multiple-thread (SIMT) pipelines. In some embodiments, multiple categories of operand resource circuits are configured to provide operands for instructions executed by processor pipeline circuitry. Per-resource arbitration circuitry may arbitrate between the SIMT execution slots for access to different operand resources. Source access circuitry may access operand data from operand resources based on source capture commands and source control circuitry may prior to a first SIMT group winning arbitration for all its operands, send a source capture command to the source access circuitry in response to the first SIMT group winning arbitration at the per-resource arbitration circuitry for a first operand resource. Instruction control circuitry may send an instruction release command down the processor pipeline circuitry for the first SIMT group, in response to the first SIMT group winning arbitration for all its operands.
1 . An apparatus, comprising:
processor pipeline circuitry;
first and second categories of operand resource circuits configured to provide operand data for instructions to be executed by the processor pipeline circuitry;
wherein the processor pipeline circuitry includes:
first resource arbitration circuitry configured to arbitrate between instructions of single-instruction multiple-thread (SIMT) groups for access to the first operand resource category;
second resource arbitration circuitry configured to arbitrate between SIMT groups for access to the second operand resource category;
source access circuitry configured to access the operand data from the operand resource circuits based on source capture commands; and
source control circuitry configured to, prior to an instruction of a first SIMT group winning arbitration for all its operands and in response the instruction winning arbitration at the first resource arbitration circuitry for a first operand resource of the operand resource circuits, send a source capture command to the source access circuitry to access operand data from the first operand resource for the instruction; and
instruction control circuitry configured to generate an instruction release command for the instruction, in response to the instruction winning arbitration for all its operands, including at the second resource arbitration circuitry.
2 . The apparatus of claim 1 , wherein the processor pipeline circuitry further includes:
a source landing stage configured to buffer accessed operand data from the source access circuitry, for the instruction of the first SIMT group, prior to the instruction release command.
3 . The apparatus of claim 1 , wherein the operand resource circuits include:
multiple operand cache read ports;
multiple data cache read ports; and
uniform storage circuitry.
4 . The apparatus of claim 1 , wherein the first resource arbitration circuitry implements a different arbitration scheme than the second resource arbitration circuitry.
5 . The apparatus of claim 1 , wherein:
the processor pipeline circuitry includes multiple pipelines that each include multiple SIMT slots; and
the first resource arbitration circuitry is configured to arbitrate among instructions of SIMT groups from multiple pipelines.
6 . The apparatus of claim 1 , further comprising:
first-stage scheduler circuitry configured to arbitrate among SIMT groups to assign SIMT groups to SIMT execution slots; and
second-stage scheduling circuitry configured to arbitrate among instructions in SIMT execution slots for assignment to execution resources, wherein the second-stage scheduler circuitry includes the first resource arbitration circuitry.
7 . The apparatus of claim 1 , further comprising:
buffer circuitry configured to:
buffer multiple instructions per SIMT group; and
provide a next instruction for arbitration for an operand resource in a next cycle subsequent to another instruction from the same SIMT group winning arbitration for the operand resource; and
landing buffer circuitry configured to buffer operand data retrieved for the multiple instructions per SIMT group.
8 . The apparatus of claim 1 , wherein:
the operand resource circuits include one or more operand caches that are dynamically managed by hardware; and
control circuitry is configured to:
lock a given entry in at least one of the one or more operand caches in response to a hit for the given entry; and
determine to unlock the given entry in response to an unlock event.
9 . The apparatus of claim 1 , further comprising:
fixed-function circuitry configured to control the processor pipeline circuitry to perform operations for at least one of the following types of programs:
graphics shader programs; and
machine learning programs.
10 . The apparatus of claim 1 , wherein the apparatus is a computing device that further includes:
a display; and
network interface circuitry.
11 . A method, comprising:
arbitrating, by a computing system, between instructions of single-instruction multiple-thread (SIMT) execution groups for access to a first operand resource category;
arbitrating, by the computing system, between instructions of SIMT groups for access to a second operand resource category;
sending, by the computing system down pipeline circuitry of the computing system to source access circuitry of the computing system, prior to an instruction of a first SIMT slot winning arbitration for all its operands, a source capture command in response to the instruction winning arbitration for a first operand resource; and
accessing, by the source access circuitry, operand data from one or more operand resources based on the source capture command; and
generating, by the computing system subsequent to the accessing, an instruction release command for the instruction, in response to the instruction winning arbitration for all its operands.
12 . The method of claim 11 , further comprising:
buffering, by the computing system, accessed operand data from the source access circuitry, for the instruction of the first SIMT slot, prior to the instruction release command.
13 . The method of claim 11 , wherein the operand resources include:
multiple operand cache read ports; and
multiple data cache read ports.
14 . The method of claim 11 , wherein the arbitrating applies different arbitration schemes for different categories of operand resource circuits.
15 . The method of claim 11 , wherein:
the pipeline circuitry includes multiple pipelines that each include multiple SIMT slots; and
the arbitrating includes arbitrating among SIMT slots from multiple pipelines.
16 . The method of claim 15 , further comprising:
arbitrating among SIMT groups to assign SIMT groups to SIMT execution slots.
17 . The method of claim 11 , further comprising:
buffering, by the computing system, multiple instructions per SIMT slot; and
providing, by the computing system, a next instruction for arbitration for an operand resource in a next cycle subsequent to another instruction from the same SIMT slot winning arbitration for the operand resource; and
buffering, by the computing system, operand data retrieved for the multiple instructions per SIMT slot.
18 . The method of claim 11 , wherein the operand resources include one or more operand caches that are dynamically managed by hardware, the method further comprising:
locking a given entry in at least one of the one or more operand caches in response to a hit for the given entry; and
determining to unlock the given entry in response to an unlock event.
19 . A non-transitory computer-readable medium having instructions of a hardware description programming language stored thereon that, when processed by a computing system, program the computing system to generate a computer simulation model, wherein the model represents a hardware circuit that includes:
processor pipeline circuitry;
multiple first and second categories of operand resource circuits configured to provide operand data for instructions to be executed by the processor pipeline circuitry;
wherein the processor pipeline circuitry includes:
first resource arbitration circuitry configured to arbitrate between instructions of single-instruction multiple-thread (SIMT) groups for access to the first operand resource category;
second resource arbitration circuitry configured to arbitrate between SIMT groups for access to the second operand resource category;
source access circuitry configured to access the operand data from the operand resources based on source capture commands; and
source control circuitry configured to, prior to an instruction of a first SIMT group winning arbitration for all its operands and in response the instruction winning arbitration at the first resource arbitration circuitry for a first operand resource of the operand resource circuits, send a source capture command to the source access circuitry to access operand data from the first operand resource for the instruction; and
instruction control circuitry configured to generate an instruction release command for the instruction first SIMT group, in response to the instruction winning arbitration for all its operands, including at the second resource arbitration circuitry.
20 . The non-transitory computer-readable medium of claim 19 , wherein the processor pipeline circuitry further includes:
a source landing stage configured to buffer accessed operand data from the source access circuitry, for the instruction of the first SIMT group, prior to the instruction release command.