IP Library Granted Patent US 10,649,956
Granted Patent B2
US 10,649,956 · App. 15/477,027 · Granted May 12, 2020

Engine to enable high speed context switching via on-die storage

Inventors: Altug Koker (El Dorado Hills, CA); Prasoonkumar Surti (Folsom, CA); David Puffer (Tempe, AZ); Subramaniam Maiyuran (Gold River, CA); Guei-Yuan Lueh (San Jose, CA); Abhishek R. Appu (El Dorado Hills, CA); Joydeep Ray (Folsom, CA); Balaji Vembu (Folsom, CA); Tomer Bar-On (Petah Tikva, IL); Andrew T. Lauritzen (Victoria, CA); Hugues Labbe (Folsom, CA); John G. Gierach (Hillsboro, OR); Gabor Liktor (San Francisco, CA)
Assignee: INTEL CORPORATION
G06F16/13G06F9/30G06F9/38G06F9/3836G06F9/461G06F16/113G06F16/172G06F12/0831G06F12/1036G06F12/1045G06F2201/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,649,956
App. No.
15/477,027
Granted
May 12, 2020
Kind
B2
Abstract

In an example, an apparatus comprises a plurality of execution units, and a first memory communicatively couple to the plurality of execution units, wherein the first shared memory is shared by the plurality of execution units and a copy engine to copy context state data from at least a first of the plurality of execution units to the first shared memory. Other embodiments are also disclosed and claimed.

Claims (58)

1. A graphics multiprocessor comprising:

an instruction cache to receive a stream of instructions;

an instruction unit to dispatch instructions in the stream of instructions for execution;

a general-purpose graphics processing compute block comprising a plurality of graphics processing cores, the graphics processing cores comprising a plurality of execution units to execute the instructions;

a first shared memory communicatively coupled to the plurality of execution units, wherein the first shared memory is shared by the plurality of execution units; and

a copy engine to:

receive a signal from a scheduler indicating an initiation of a preemption process;

stop an execution of an existing context on at least a first of the plurality of execution units;

copy context state data from the existing context on the at least a first of the plurality of execution units to the first shared memory;

generate a signal to indicate that a context preemption process is complete; and

upload the context data from the first shared memory to a second memory, separate from the first shared memory.

2. The graphics multiprocessor of claim 1 , further comprising:

a scheduler to initiate a preemption process.

3. The graphics multiprocessor of claim 1 , further comprising:

a first communication fabric to communicatively couple the plurality of execution units to the first shared memory.

4. The graphics multiprocessor of claim 3 , wherein the copy engine is coupled to the first shared memory by the first communication fabric.

5. The graphics multiprocessor of claim 4 , wherein the first shared memory comprises a L3 cache.

6. The graphics multiprocessor of claim 1 , further comprising:

a second communication fabric to communicatively couple the copy engine to the second memory.

7. The graphics multiprocessor of claim 1 , the copy engine to:

copy context state data from the existing context on at least a first of the plurality of execution units to the first shared memory in parallel with executing a new context on the plurality of execution units.

8. The graphics multiprocessor of claim 1 , further comprising logic, at least partially including hardware logic, to:

restore the first shared memory.

9. The graphics multiprocessor of claim 1 , wherein the plurality of execution units and a first general register file are on a single integrated circuit.

10. An electronic device, comprising:

a central processing unit;

a graphics multiprocessor communicatively coupled to the central processing unit, comprising:

an instruction cache to receive a stream of instructions;

an instruction unit to dispatch instructions in the stream of instructions for execution;

a general-purpose graphics processing compute block comprising a plurality of graphics processing cores, the graphics processing cores comprising a plurality of execution units to execute the instructions;

a first shared memory communicatively coupled to the plurality of execution units, wherein the first shared memory is shared by the plurality of execution units; and

a copy engine to:

receive a signal from a scheduler indicating an initiation of a preemption process;

stop an execution of an existing context on the at least a first of the plurality of execution units;

copy context state data from the existing context on at least a first of the plurality of execution units to the first shared memory;

generate a signal to indicate that a context preemption process is complete; and

upload the context data from the first shared memory to a second memory, separate from the first shared memory.

11. The electronic device of claim 10 , further comprising:

a scheduler to initiate a preemption process.

12. The electronic device of claim 10 , further comprising:

a first communication fabric to communicatively couple the plurality of execution units to the first shared memory.

13. The electronic device of claim 12 , wherein the copy engine is coupled to the first shared memory by the first communication fabric.

14. The electronic device of claim 13 , wherein the first shared memory comprises a L3 cache.

15. The electronic device of claim 10 , further comprising:

a second communication fabric to communicatively couple the copy engine to the second memory.

16. The electronic device of claim 15 , the copy engine to:

copy context state data from the existing context on at least a first of the plurality of execution units to the first shared memory in parallel with executing a new context on the plurality of execution units.

17. The electronic device of claim 10 , further comprising logic, at least partially including hardware logic, to:

restore the first shared memory.

18. The electronic device of claim 10 , wherein the plurality of execution units and a first general register file are on a single integrated circuit.

19. A method comprising:

receiving, in a copy engine of a general-purpose graphics processing compute block comprising a plurality of graphics processing cores, the graphics processing cores comprising a processor having a plurality of execution units to execute the instructions a signal from a scheduler indicating an initiation of a preemption process;

stopping an execution of an existing context on the at least a first of the plurality of execution units;

copying context state data from the existing context on at least a first of the plurality of execution units to a first shared memory;

generating a signal to indicate that a context preemption process is complete; and

uploading the context data from the first shared memory to a second memory, separate from the first shared memory.

20. The method of claim 19 , further comprising:

copying context state data from an existing context on the at least a first of the plurality of execution units to the first shared memory in parallel with executing a new context on the plurality of execution units.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2018
From: KOKER, ALTUG; SURTI, PRASOONKUMAR; PUFFER, DAVID; MAIYURAN, SUBRAMANIAM; LUEH, GUEI-YUAN; APPU, ABHISHEK R.; RAY, JOYDEEP; VEMBU, BALAJI; BAR-ON, TOMER; LAURITZEN, ANDREW T.; LABBE, HUGUES; GIERACH, JOHN G.; LIKTOR, GABOR
To: INTEL CORPORATION
Reel/Frame 046936/0829 →
Continuity (1)
Related Publication 20180285374A1 · Oct 4, 2018
Cited By (2)
US 12,293,090 US 12,399,734