IP Library › Granted Patent US 12,386,617
Granted Patent B2
US 12,386,617 · App. 17/481,448 · Granted Aug 12, 2025

Gathering payload from arbitrary registers for send messages in a graphics environment

Inventors: Supratim Pal (Folsom, CA); Chandra Gurram (Folsom, CA); Fan-Yin Tzeng (Fremont, CA); Subramaniam Maiyuran (Gold River, CA); Guei-Yuan Lueh (San Jose, CA); Timothy R. Bauer (Hillsboro, CA); Vikranth Vemulapalli (Folsom, CA); Wei-Yu Chen (San Jose, CA)
Assignee: INTEL CORPORATION
G06F9/3012G06F9/30036G06F9/30105G06F9/3826G06F9/3851G06F9/3854G06F9/3858G06F9/3887G06F9/3888G06F9/38885G06F12/0223G06T1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,386,617
App. No.
17/481,448
Filed
Sep 22, 2021
Granted
Aug 12, 2025
Kind
B2
Art Unit
2183
USPC
712/225
Abstract

An apparatus to facilitate gathering payload from arbitrary registers for send messages in a graphics environment is disclosed. The apparatus includes processing resources comprising execution circuitry to receive a send gather message instruction identifying a number of registers to access for a send message and identifying IDs of a plurality of individual registers corresponding to the number of registers; decode a first phase of the send gather message instruction; based on decoding the first phase, cause a second phase of the send gather message instruction to bypass an instruction decode stage; and dispatch the first phase subsequently followed by dispatch of the second phase to a send pipeline. The apparatus can also perform an immediate move of the IDs of the plurality of individual registers to an architectural register of the execution circuitry and include a pointer to the architectural register in the send gather message instruction.

Claims (39)

1. A processor comprising:

processing resources comprising execution circuitry to:

receive a send gather message instruction identifying a number of registers to access for a send message and identifying IDs of a plurality of individual registers corresponding to the number of registers;

decode a first phase of the send gather message instruction;

based on decoding the first phase of the send gather message instruction, cause a second phase of the send gather message instruction to bypass an instruction decode stage of the execution circuitry; and

dispatch the first phase of the send gather message instruction subsequently followed by dispatch of the second phase of the send gather message instruction to a send pipeline of the execution circuitry.

2. The processor of claim 1 , wherein the dispatch of the first phase and the second phase of the send gather message instruction are dispatched to the send pipeline without any intervening dispatched messages between the first phase and the second phase in the send pipeline.

3. The processor of claim 1 , wherein the plurality of individual registers comprise general register file (GRF) registers.

4. The processor of claim 1 , wherein the second phase of the send gather message instruction comprises the IDs of the plurality of individual registers.

5. The processor of claim 1 , wherein the first phase of the send gather message instruction identifies the number of the registers to access for the send message.

6. The processor of claim 1 , wherein the plurality of individual registers are non-contiguous.

7. The processor of claim 1 , wherein the first phase of the send gather message instruction identifies a destination shared function ID, a function-specific encoding of an operation of the send gather message instruction, and a destination register for a writeback response to the send gather message instruction.

8. The processor of claim 1 , wherein the processor comprises a graphics processing unit (GPU).

9. The processor of claim 1 , wherein the processor is at least one of a single instruction multiple data (SIMD) machine or a single instruction multiple thread (SIMT) machine.

10. A method comprising:

receiving, by execution circuitry of a graphics processor, a send gather message instruction identifying a number of registers to access for a send message and identifying IDs of a plurality of individual registers corresponding to the number of registers;

decoding a first phase of the send gather message instruction;

based on decoding the first phase of the send gather message instruction, causing a second phase of the send gather message instruction to bypass an instruction decode stage of the execution circuitry; and

dispatching the first phase of the send gather message instruction subsequently followed by dispatch of the second phase of the send gather message instruction to a send pipeline of the execution circuitry.

11. The method of claim 10 , wherein the dispatch of the first phase and the second phase of the send gather message instruction are dispatched to the send pipeline without any intervening dispatched messages between the first phase and the second phase in the send pipeline.

12. The method of claim 10 , wherein the plurality of individual registers comprise general register file (GRF) registers.

13. The method of claim 10 , wherein the first phase of the send gather message instruction identifies the number of the registers to access for the send message, and wherein the second phase of the send gather message instruction comprises the IDs of the plurality of individual registers.

14. The method of claim 10 , wherein the plurality of individual registers are non-contiguous.

15. The method of claim 10 , wherein the first phase of the send gather message instruction identifies a destination shared function ID, a function-specific encoding of an operation of the send gather message instruction, and a destination register for a writeback response to the send gather message instruction.

16. A system comprising:

a memory to store a block of data; and

a processor coupled to the memory, the processor comprising processing resources, the processing resources comprising execution circuitry to:

receive a send gather message instruction identifying a number of registers to access for a send message and identifying IDs of a plurality of individual registers corresponding to the number of registers;

decode a first phase of the send gather message instruction;

based on decoding the first phase of the send gather message instruction, cause a second phase of the send gather message instruction to bypass an instruction decode stage of the execution circuitry; and

dispatch the first phase of the send gather message instruction subsequently followed by dispatch of the second phase of the send gather message instruction to a send pipeline of the execution circuitry.

17. The system of claim 16 , wherein the execution circuitry is further to:

perform an immediate move of the IDs of the plurality of individual registers to an architectural register of the execution circuitry;

include a pointer to the architectural register in the send gather message instruction; and

dispatch the send gather message instruction to a send pipeline of the execution circuitry;

wherein the architectural register is a scalar register.

18. The system of claim 17 , wherein the send gather message instruction identifies the number of the registers to access for the send gather message instruction and the architectural register.

19. The system of claim 16 , wherein the plurality of individual registers are non-contiguous, and wherein the IDs comprise pointers.

20. The system of claim 16 , wherein the send gather message instruction identifies a destination shared function ID, a function-specific encoding of an operation of the send gather message instruction, and a destination register for a writeback response to the send gather message instruction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2023
From: PAL, SUPRATIM; GURRAM, CHANDRA; TZENG, FAN-YIN; MAIYURAN, SUBRAMANIAM; LUEH, GUEI-YUAN; BAUER, TIMOTHY R., DR.; VEMULAPALLI, VIKRANTH; CHEN, WEI-YU
To: INTEL CORPORATION
Reel/Frame 062960/0434 →
Continuity (1)
Related Publication 20230088743A1 · Mar 23, 2023
References Cited (9)
US 20050053012A1 · Moyer · 2005 [cited by examiner]
US 20080059760A1 · Sachs · 2008 [cited by examiner]
US 20080222392A1 · Fuchs · 2008 [cited by examiner]
US 20120254588A1 · Adrian · 2012 [cited by examiner]
US 20130212353A1 · Mimar · 2013 [cited by examiner]
US 20140317382A1 · Segelken · 2014 [cited by examiner]
US 20180165210A1 · Sethuraman · 2018 [cited by examiner]
US 20230088743A1 · Pal et al. · 2023 [cited by applicant]
DE 102022119733A1 · 2023 [cited by applicant]