IP Library Granted Patent US 12,373,912
Granted Patent B2
US 12,373,912 · App. 18/511,074 · Granted Jul 29, 2025

Prefetch status notification for memory prefetching

Inventors: Joydeep Ray (Folsom, CA); Aravindh Anantaraman (Folsom, CA); Valentin Andrei (San Jose, CA); Abhishek R. Appu (El Dorado Hills, CA); Nicolas Galoppo von Borries (Portland, OR); Varghese George (Folsom, CA); Altug Koker (El Dorado Hills, CA); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Mike Macpherson (Portland, OR); Subramaniam Maiyuran (Gold River, CA)
Assignee: INTEL CORPORATION
G06T1/20G06F9/3802G06F9/3877G06T1/60G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,912
App. No.
18/511,074
Granted
Jul 29, 2025
Kind
B2
Abstract

Embodiments are generally directed to memory prefetching in multiple GPU environment. An embodiment of an apparatus includes multiple processors including a host processor and multiple graphics processing units (GPUs) to process data, each of the GPUs including a prefetcher and a cache; and a memory for storage of data, the memory including a plurality of memory elements, wherein the prefetcher of each of the GPUs is to prefetch data from the memory to the cache of the GPU; and wherein the prefetcher of a GPU is prohibited from prefetching from a page that is not owned by the GPU or by the host processor.

Claims (32)

1. An apparatus comprising:

one or more processors including a graphics processing unit (GPU) to process data, the GPU including:

one or more processing cores, including a first core,

a prefetcher, and

one or more caches; and

a memory for storage of data;

wherein the prefetcher is to prefetch data from the memory to a cache of the one or more caches;

wherein, upon completion of a prefetch operation by the GPU to prefetch data for a first thread running on the first core, the prefetcher is to issue a prefetch status notification to the first thread.

2. The apparatus of claim 1 , wherein issuance of the prefetch status notification indicates that the data prefetched for the first thread has been loaded into the cache.

3. The apparatus of claim 1 , wherein the GPU is to synchronize execution of the first thread with one or more other threads based at least in part on the prefetch status notification.

4. The apparatus of claim 1 , wherein the GPU is to throttle one or more prefetches for the first thread based at least in part on the prefetch status notification.

5. The apparatus of claim 1 , wherein the prefetch status notification is a one-bit flag.

6. The apparatus of claim 1 , wherein the first core is a shader core.

7. One or more non-transitory computer-readable storage mediums having stored thereon executable computer program instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

performing, by a prefetcher of a graphics processing unit (GPU) in a computing system, a prefetch operation to prefetch data for a first thread running on a first core of the GPU, the data being prefetched from a computer memory to a cache of the GPU; and

upon completion of the prefetch operation, issuing, by the prefetcher, a prefetch status notification to the first thread.

8. The one or more computer-readable storage mediums of claim 7 , wherein issuance of the prefetch status notification indicates that the data prefetched for the first thread has been loaded into the cache.

9. The one or more computer-readable storage mediums of claim 7 , wherein the instructions further include instructions for:

synchronizing execution of the first thread with one or more other threads based at least in part on the prefetch status notification.

10. The one or more computer-readable storage mediums of claim 7 , wherein the instructions further include instructions for:

throttling one or more prefetches for the first thread based at least in part on the prefetch status notification.

11. The one or more computer-readable storage mediums of claim 7 , wherein the prefetch status notification is a one-bit flag.

12. The one or more computer-readable storage mediums of claim 7 , wherein the first core is a shader core.

13. A method comprising:

performing, by a prefetcher of a graphics processing unit (GPU) in a computing system, a prefetch operation to prefetch data for a first thread running on a first core of the GPU, the data being prefetched from a computer memory to a cache of the GPU; and

upon completion of the prefetch operation, issuing, by the prefetcher, a prefetch status notification to the first thread.

14. The method of claim 13 , wherein issuance of the prefetch status notification indicates that the data prefetched for the first thread has been loaded into the cache.

15. The method of claim 13 , further comprising:

synchronizing execution of the first thread with one or more other threads based at least in part on the prefetch status notification.

16. The method of claim 13 , further comprising:

throttling one or more prefetches for the first thread based at least in part on the prefetch status notification.

17. The method of claim 13 , wherein the prefetch status notification is a one-bit flag.

Continuity (3)
Continuation 17580352 · Jan 20, 2022
Continuation 16355274 · Mar 15, 2019
Related Publication 20240161226A1 · May 16, 2024
References Cited (44)
US 6507894B1 · Hoshi · 2003 [cited by examiner]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 8429351B1 · Yu · 2013 [cited by examiner]
US 9912957B1 · Surti et al. · 2018 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891156B1 · Zhao et al. · 2021 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 11232533B2 · Ray · 2022 [cited by examiner]
US 20070288697A1 · Keltcher · 2007 [cited by applicant]
US 20130151787A1 · Riguer · 2013 [cited by applicant]
US 20130318306A1 · Gonion · 2013 [cited by applicant]
US 20130346697A1 · Alexander et al. · 2013 [cited by applicant]
US 20140149632A1 · Kannan et al. · 2014 [cited by applicant]
US 20140207871A1 · Miloushev et al. · 2014 [cited by applicant]
US 20150074373A1 · Sperber et al. · 2015 [cited by applicant]
US 20150221063A1 · Kim et al. · 2015 [cited by applicant]
US 20150278099A1 · Jain et al. · 2015 [cited by applicant]
US 20150310580A1 · Kumar · 2015 [cited by applicant]
US 20150339233A1 · Kapil et al. · 2015 [cited by applicant]
US 20150378920A1 · Gierach et al. · 2015 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20170091103A1 · Smelyanskiy · 2017 [cited by applicant]
US 20170177349A1 · Yount · 2017 [cited by applicant]
US 20170177360A1 · Gokhale · 2017 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180074963A1 · Mehta · 2018 [cited by applicant]
US 20190197760A1 · Cho · 2019 [cited by applicant]
US 20190266695A1 · Rao et al. · 2019 [cited by applicant]
US 20200057717A1 · Jayasena et al. · 2020 [cited by applicant]
CN 106687927A · 2017 [cited by examiner]
CN 106776047A · 2017 [cited by examiner]
CN 113424163A · 2021 [cited by applicant]
EP 3938916A1 · 2022 [cited by applicant]
WO 2020190424A1 · 2020 [cited by applicant]
Dreslinski, “Analysis of Hardware Prefetching Across Virtual Page Boundaries”, Dept. of Electrical Engineering and Computer Science 2260 Hayward Ave, Ann Arbor, MI 48109-2121, Copyright 2007 ACM, CF '07, May 7-9, 2007, … [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Notification of Transmittal of the International Preliminary Report on Patentability of the International Searching Authority for PCT Application No. PCT/US2020/017739, mailed Sep. 30, 2021 , 7 pages. [cited by applicant]
Notification of Transmittal of the International Search Report and the Written Opinion of the international Searching Authority for PCT Application No. PCT/US2020/017739, mailed Jun. 16, 2020, 10 pages. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Stephen Junking, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Xiao, “Unified Virtual Memory Support for Deep CNN Accelerator on Soc FPGA”, College of Computer, National University of Defense Technology, Changsha 410073, China, Springer International Publishing Switzerland, 2015, p… [cited by applicant]