IP Library Granted Patent US 11,861,759
Granted Patent B2
US 11,861,759 · App. 17/580,352 · Granted Jan 2, 2024

Memory prefetching in multiple GPU environment

Inventors: Joydeep Ray (Folsom, CA); Aravindh Anantaraman (Folsom, CA); Valentin Andrei (San Jose, CA); Abhishek R. Appu (El Dorado Hills, CA); Nicolas Galoppo von Borries (Portland, OR); Varghese George (Folsom, CA); Altug Koker (El Dorado Hills, CA); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Mike Macpherson (Portland, OR); Subramaniam Maiyuran (Gold River, CA)
Assignee: INTEL CORPORATION
G06T1/20G06F9/3802G06F9/3877G06T1/60G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,861,759
App. No.
17/580,352
Granted
Jan 2, 2024
Kind
B2
Abstract

Embodiments are generally directed to memory prefetching in multiple GPU environment. An embodiment of an apparatus includes multiple processors including a host processor and multiple graphics processing units (GPUs) to process data, each of the GPUs including a prefetcher and a cache; and a memory for storage of data, the memory including a plurality of memory elements, wherein the prefetcher of each of the GPUs is to prefetch data from the memory to the cache of the GPU; and wherein the prefetcher of a GPU is prohibited from prefetching from a page that is not owned by the GPU or by the host processor.

Claims (31)

1. An apparatus comprising:

a plurality of processors including a host processor and a plurality of graphics processing units (GPUs) to process data including at least a first graphics processing unit (GPU), each of the plurality of GPUs including a prefetcher and one or more caches; and

a memory for storage of data;

wherein the prefetcher of each of the plurality of GPUs is to prefetch data from the memory to a cache of the respective GPU;

wherein a prefetch operation by the first GPU includes the prefetcher of the first GPU issuing a gather/scatter prefetch message including a plurality of prefetch addresses; and

wherein the plurality of processors are to parse the gather/scatter prefetch message and issue a prefetch message for each of the plurality of prefetch addresses.

2. The apparatus of claim 1 , wherein the gather/scatter prefetch message includes an entry for each of the plurality of prefetch addresses, the entry to indicate a cache level for prefetching.

3. The apparatus of claim 2 , wherein the gather/scatter prefetch message includes a plurality of different cache levels within the gather/scatter prefetch message.

4. The apparatus of claim 1 , wherein the plurality of prefetch addresses includes noncontiguous addresses.

5. The apparatus of claim 1 , wherein the prefetcher of the first GPU is to send a notification to a thread in a core of the first GPU when a prefetch for the thread is complete.

6. The apparatus of claim 1 , wherein the prefetchers of the plurality of GPUs are to prefetch data from the memory to a cache of each respective GPU in execution of a multi-GPU workload.

7. One or more non-transitory computer-readable storage mediums having stored thereon executable computer program instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

processing a workload in a computing system including a plurality of processors, the plurality of processors including a host processor and a plurality of graphics processing units (GPUs), each of the plurality of GPUs including a prefetcher and one or more caches;

prefetching data by a first graphics processing unit (GPU) of the plurality of GPUs from a memory of the computing system to a cache of the first GPU wherein prefetching data by the first GPU includes the prefetcher of the first GPU issuing a gather/scatter prefetch message including a plurality of prefetch addresses; and

parsing the gather/scatter prefetch message and issuing a prefetch message for each of the plurality of prefetch addresses.

8. The one or more computer-readable storage mediums of claim 7 , wherein the gather/scatter prefetch message includes an entry for each of the plurality of prefetch addresses, the entry to indicate a cache level for prefetching.

9. The one or more computer-readable storage mediums of claim 8 , wherein the gather/scatter prefetch message includes a plurality of different cache levels within the gather/scatter prefetch message.

10. The one or more computer-readable storage mediums of claim 7 , wherein the plurality of prefetch addresses includes noncontiguous addresses.

11. The one or more computer-readable storage mediums of claim 7 , wherein the instructions further include instructions for:

sending a notification to a thread in a core of the first GPU when a prefetch for the thread is complete.

12. The one or more computer-readable storage mediums of claim 7 , wherein the workload is a multi-GPU workload, and wherein the prefetchers of the plurality of GPUs are to prefetch data from the memory to a cache of each respective GPU in execution of the multi-GPU workload.

13. A method comprising:

processing a workload in a computing system including a plurality of processors, the plurality of processors including a host processor and a plurality of graphics processing units (GPUs), each of the plurality of GPUs including a prefetcher and one or more caches;

prefetching data by a first graphics processing unit (GPU) of the plurality of GPUs from a memory of the computing system to a cache of the first GPU wherein prefetching data by the first GPU includes the prefetcher of the first GPU issuing a gather/scatter prefetch message including a plurality of prefetch addresses; and

parsing the gather/scatter prefetch message and issuing a prefetch message for each of the plurality of prefetch addresses.

14. The method of claim 13 , wherein the gather/scatter prefetch message includes an entry for each of the plurality of prefetch addresses, the entry to indicate a cache level for prefetching.

15. The method of claim 14 , wherein the gather/scatter prefetch message includes a plurality of different cache levels within the gather/scatter prefetch message.

16. The method of claim 13 , wherein the plurality of prefetch addresses includes noncontiguous addresses.

17. The method of claim 13 , further comprising:

sending a notification to a thread in a core of the first GPU when a prefetch for the thread is complete.

18. The method of claim 13 , wherein the workload is a multi-GPU workload, and wherein the prefetchers of the plurality of GPUs are to prefetch data from the memory to a cache of each respective GPU in execution of the multi-GPU workload.

Continuity (2)
Continuation 16355274 · Mar 15, 2019
Related Publication 20220222767A1 · Jul 14, 2022