IP Library › Granted Patent US 11,868,777
Granted Patent B2
US 11,868,777 · App. 17/123,270 · Granted Jan 9, 2024

Processor-guided execution of offloaded instructions using fixed function operations

Inventors: John Kalamatianos (Boxborough, MA); Michael T. Clark (Austin, TX); Marius Evers (Santa Clara, CA); William L. Walker (Fort Collins, CO); Paul Moyer (Fort Collins, CO); Jay Fleischman (Fort Collins, CO); Jagadish B. Kotra (Austin, TX)
Assignee: ADVANCED MICRO DEVICES, INC.
G06F9/30181G06F9/30043G06F9/30098G06F9/30138G06F9/3834G06F9/3877G06F9/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,868,777
App. No.
17/123,270
Granted
Jan 9, 2024
Kind
B2
Abstract

Processor-guided execution of offloaded instructions using fixed function operations is disclosed. Instructions designated for remote execution by a target device are received by a processor. Each instruction includes, as an operand, a target register in the target device. The target register may be an architected virtual register. For each of the plurality of instructions, the processor transmits an offload request in the order that the instructions are received. The offload request includes the instruction designated for remote execution. The target device may be, for example, a processing-in-memory device or an accelerator coupled to a memory.

Claims (42)

1. A method of processor-guided execution of offloaded instructions using fixed function operations, the method comprising:

receiving one or more instructions designated for remote execution by a target device; and

transmitting, for each of the one or more instructions, an offload request, the offload request including a pointer to an entry, within a command buffer at the target device, identifying an opcode corresponding to the instruction designated for remote execution.

2. The method of claim 1 , wherein each instruction of the one or more instructions includes, as an operand, a target register in the target device; and wherein a processor implements an instruction set architecture extension that identifies the target register as a virtual register.

3. The method of claim 1 , wherein each of the one or more instructions includes an opcode from a group of opcodes in an instruction set architecture extension implemented by a processor; and wherein the group of opcodes in the instruction set architecture extension consists of a remote load opcode, a remote computation opcode, and a remote store opcode.

4. The method of claim 1 , wherein transmitting, for each of the one or more instructions, an offload request, the offload request including the instruction designated for remote execution includes:

generating a memory address for an instruction designated for remote execution; and

coupling the memory address with the offload request.

5. The method of claim 1 , wherein transmitting, for each of the one or more instructions, an offload request, the offload request including the instruction designated for remote execution includes:

obtaining local data for the instruction designated for remote execution; and

coupling the local data with the offload request.

6. The method of claim 1 , wherein transmitting, for each of the one or more instructions, an offload request, the offload request including the instruction designated for remote execution includes:

buffering the offload requests until after an oldest instruction in the one or more instructions has retired; and

transmitting, for each of the one or more instructions in an order received, an offload request.

7. The method of claim 1 further comprising performing a cache operation on one or more caches that contain an entry corresponding to a memory address included in the offload request, wherein the cache operation includes at least one of invalidating a cache entry containing clean data and flushing a cache entry containing dirty data.

8. The method of claim 7 , wherein the cache operation is performed on a plurality of caches that contain an entry corresponding to a memory address included in the offload request, and wherein the plurality of caches are distributed across a plurality of core clusters each including a plurality of processor cores.

9. The method of claim 1 , wherein the target device is a processing-in-memory device.

10. The method of claim 1 , wherein the target device is an accelerator coupled to a memory device, and wherein the entry is one of a plurality of entries included within the command buffer that identifies a plurality of opcodes.

11. A multicore processor comprising:

two or more processor cores;

at least one cache shared by the two or more processor cores; and

at least one memory controller configured for communication with a target device

wherein the two or more processor cores are configured to:

receive one or more instructions designated for remote execution by the target device; and

transmit, for each of the one or more instructions, an offload request, the offload request including a pointer to an entry, within a command buffer at the target device, identifying an opcode corresponding to the instruction designated for remote execution.

12. The processor of claim 11 , wherein each instruction of the one or more instructions includes, as an operand, a target register in the target device; and wherein the processor implements an instruction set architecture extension that identifies the target register as a virtual register.

13. The processor of claim 11 , wherein each of the one or more instructions includes an opcode from a group of opcodes in an instruction set architecture extension implemented by the processor; and wherein the group of opcodes in the instruction set architecture extension consists of a remote load opcode, a remote computation opcode, and a remote store opcode.

14. The processor of claim 11 , wherein transmitting, for each of the one or more instructions, an offload request, the offload request including the instruction designated for remote execution includes:

buffering the offload requests until after an oldest instruction in the one or more instructions has retired; and

transmitting, for each of the one or more instructions in an order received, an offload request.

15. The processor of claim 11 , wherein the two or more processor cores are further configured to perform a cache operation on one or more caches that contain an entry corresponding to a memory address included in the offload request, wherein the cache operation includes at least one of invalidating a cache entry containing clean data and flushing a cache entry containing dirty data.

16. A system comprising:

a processing-in-memory (PIM) device; and

a multicore processor coupled to the PIM device, the processor configured to:

receive one or more instructions designated for remote execution by the PIM device; and

transmit, for each of the one or more instructions, an offload request, the offload request including a pointer to an entry, within a command buffer of the PIM device, identifying an opcode corresponding to the instruction designated for remote execution.

17. The system of claim 16 , wherein each instruction of the one or more instructions includes, as an operand, a target register in the PIM device; and wherein the processor implements an instruction set architecture extension that identifies the target register as a virtual register.

18. The system of claim 16 , wherein each of the one or more instructions includes an opcode from a group of opcodes in an instruction set architecture extension implemented by the processor; and wherein the group of opcodes in the instruction set architecture extension consists of a remote load opcode, a remote computation opcode, and a remote store opcode.

19. The system of claim 16 , wherein transmitting, for each of the one or more instructions, an offload request, the offload request including the instruction designated for remote execution includes:

buffering the offload requests until after an oldest instruction in the one or more instructions has retired; and

transmitting, for each of the one or more instructions in an order received, an offload request.

20. The system of claim 16 , wherein the processor is further configured to perform a cache operation on one or more caches that contain an entry corresponding to a memory address included in the offload request, wherein the cache operation includes at least one of invalidating a cache entry containing clean data and flushing a cache entry containing dirty data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2021
From: KALAMATIANOS, JOHN; CLARK, MICHAEL T.; EVERS, MARIUS; WALKER, WILLIAM L.; MOYER, PAUL; FLEISCHMAN, JAY; KOTRA, JAGADISH B.
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 056963/0551 →
Continuity (1)
Related Publication 20220188117A1 · Jun 16, 2022
Cited By (7)
US 12,197,378 US 12,265,470 US 12,455,826 US 12,474,897 US 12,498,931 US 12,596,650 US 12,681,866