IP Library Granted Patent US 12,657,128
Granted Patent B2
US 12,657,128 · App. 18/391,346 · Granted Jun 16, 2026

Data prefetching for graphics data processing

Inventors: Vikranth Vemulapalli (Folsom, CA); Lakshminarayanan Striramassarma (Folsom, CA); Mike MacPherson (Portland, OR); Aravindh Anantaraman (Folsom, CA); Ben Ashbaugh (Folsom, CA); Murali Ramadoss (Folsom, CA); William B. Sadler (Folsom, CA); Jonathan Pearce (Portland, OR); Scott Janus (Loomis, CA); Brent Insko (Portland, OR); Vasanth Ranganathan (El Dorado Hills, CA); Kamal Sinha (Folsom, CA); Arthur Hunter, Jr. (Cameron Park, CA); Prasoonkumar Surti (Folsom, CA); Nicolas Galoppo von Borries (Portland, OR); Joydeep Ray (Folsom, CA); Abhishek R. Appu (El Dorado Hills, CA); ElMoustapha Ould-Ahmed-Vall (Chandler, AZ); Altug Koker (El Dorado Hills, CA); Sungye Kim (Folsom, CA); Subramaniam Maiyuran (Gold River, CA); Valentin Andrei (San Jose, CA)
Assignee: Intel Corporation
G06F12/0862G06T1/20G06T1/60G06F2212/602G06F2212/608
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,128
App. No.
18/391,346
Granted
Jun 16, 2026
Kind
B2
Abstract

Embodiments are generally directed to data prefetching for graphics data processing. An embodiment of an apparatus includes one or more processors including one or more graphics processing units (GPUs); and a plurality of caches to provide storage for the one or more GPUs, the plurality of caches including at least an L1 cache and an L3 cache, wherein the apparatus to provide intelligent prefetching of data by a prefetcher of a first GPU of the one or more GPUs including measuring a hit rate for the Li cache; upon determining that the hit rate for the L1 cache is equal to or greater than a threshold value, limiting a prefetch of data to storage in the L3 cache, and upon determining that the hit rate for the L1 cache is less than a threshold value, allowing the prefetch of data to the L1 cache.

Claims (28)

1 . An apparatus comprising:

one or more processors including one or more graphics processing units (GPUs), the one or more GPUs including a hardware preprocessor; and

a memory for storage of data for the one or more processors;

wherein the hardware preprocessor is to:

access a table of IP (Internet Protocol) addresses for utilization by a kernel, the kernel to be processed by the one or more GPUs, and

speculatively prefetch one or more IP addresses from the table of IP addresses to one or more caches, the hardware preprocessor to commence the speculative prefetching prior to thread execution performed by the one or more GPUs for the kernel.

2 . The apparatus of claim 1 , wherein the hardware preprocessor is sharable among multiple processing resources of the one or more GPUs.

3 . The apparatus of claim 2 , wherein the multiple processing resources include multiple execution units.

4 . The apparatus of claim 1 , wherein data for the table of IP addresses is loaded based on software generated sequences for utilization by the kernel.

5 . The apparatus of claim 1 , wherein data for the table of IP addresses is loaded based on a stride generated to identify IP addresses associated with the kernel.

6 . One or more non-transitory computer-readable storage mediums having stored thereon executable computer program instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

accessing, utilizing a hardware preprocessor of one or more processors in a computing system, a table of IP (Internet Protocol) addresses for utilization by a kernel, the one or more processors including one or more graphics processing units (GPUs), the kernel to be processed by the one or more GPUs; and

speculatively prefetching, by the hardware preprocessor, one or more IP addresses from the table of IP addresses to one or more caches, the hardware preprocessor to commence the speculative prefetching prior to thread execution for the kernel performed by the one or more GPUs.

7 . The one or more computer-readable storage mediums of claim 6 , wherein the hardware preprocessor is sharable among multiple processing resources of the one or more GPUs.

8 . The one or more computer-readable storage mediums of claim 7 , wherein the multiple processing resources include multiple execution units.

9 . The one or more computer-readable storage mediums of claim 6 , further comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

loading data for the table of IP addresses based on software generated sequences for utilization by the kernel.

10 . The one or more computer-readable storage mediums of claim 6 , further comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

loading data for the table of IP addresses based on a stride generated to identify IP addresses associated with the kernel.

11 . A method comprising:

accessing, utilizing a hardware preprocessor of one or more processors in a computing system, a table of IP (Internet Protocol) addresses for utilization by a kernel, the one or more processors including one or more graphics processing units (GPUs), the kernel to be processed by the one or more GPUs; and

speculatively prefetching, by the hardware preprocessor, one or more IP addresses from the table of IP addresses to one or more caches, the hardware preprocessor to commence the speculative prefetching prior to thread execution for the kernel performed by the one or more GPUs.

12 . The method of claim 11 , wherein the hardware preprocessor is sharable among multiple processing resources of the one or more GPUs.

13 . The method of claim 12 , wherein the multiple processing resources include multiple execution units.

14 . The method of claim 11 , further comprising:

loading data for the table of IP addresses based on software generated sequences for utilization by the kernel.

15 . The method of claim 11 , further comprising:

loading data for the table of IP addresses based on a stride generated to identify IP addresses associated with the kernel.

Continuity (4)
Continuation 17865666 · Jul 15, 2022
Continuation 17161465 · Jan 28, 2021
Continuation 16355015 · Mar 15, 2019
Related Publication 20240256456A1 · Aug 1, 2024
References Cited (38)
US 6898674B2 · Maiyuran et al. · 2005 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 20060041723A1 · Hakura et al. · 2006 [cited by applicant]
US 20100115250A1 · Kriegel et al. · 2010 [cited by applicant]
US 20120066455A1 · Punyamurtula et al. · 2012 [cited by applicant]
US 20130246708A1 · Ono et al. · 2013 [cited by applicant]
US 20140149677A1 · Jayasena et al. · 2014 [cited by applicant]
US 20150149714A1 · Pugsley et al. · 2015 [cited by applicant]
US 20150221063A1 · Kim et al. · 2015 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20170371795A1 · Wang et al. · 2017 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20190294546A1 · Agarwal et al. · 2019 [cited by applicant]
US 20200117462A1 · Jin et al. · 2020 [cited by applicant]
CN 113454609A · 2021 [cited by applicant]
DE 112020000902 · 2021 [cited by applicant]
EP 3972187A1 · 2022 [cited by applicant]
WO 2020190429A1 · 2020 [cited by applicant]
“PacketShader: A GPU-Accelerated Software Router”, Han et al. (Sep. 2010), retrieved from ACM Digital Library (Year: 2010). [cited by examiner]
“GAMT: A Fast and Scalable IP Lookup Engine for GPU-based Software Routers”, Li et al. (Oct. 2013) retrieved from IEEE (Year: 2013). [cited by examiner]
Chakroun I et al: “Combining multi-core and GPU computing for solving combinatoriai optimization probiems”, Journal of Parallel and Distributed Computing, Elsevier, Amsterdam, NL, vol. 73, No. 12, Aug. 22, 2013 (Aug. 22… [cited by applicant]
Fang Juan, et al., “Miss-Aware LLC Buffer Management Strategy Based on Heterogeneous Muiti-Core”, Journal of Supercomputing, Kluwer Academic Publishers, Dordrecht. NL, vol. 75, No. 8, Feb. 20, 2019, pp. 4519-4528, XP036… [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US20/17897 mailed Sep. 30, 2021, 10 pages. [cited by applicant]
International Search Report and the Written Opinion of the International Searching Authority for PCT Application No. PCT/US2020/017897, mailed May 29, 2020, 15 pages. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/355,015 mailed May 13, 2020, 11 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/161,465 mailed Nov. 5, 2021, 13 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/355,015 mailed Sep. 29, 2020, 7 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/161,465 mailed Mar. 28, 2022, 8 pages. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Wenji Wu et al: “G-NetMon: A GPU-accelerated network performance monitoring system for large scale scientific collaborations”, Local Computer Networks (LCN), 2011 IEEE 36th Conference on, IEEE, Oct. 4, 2011 (Oct. 4, 201… [cited by applicant]
Yin-Chi Peng et al: “Cross-layer dynamic prefetching allocation strategies for higll-performance multicores”, VLSI Design, Automation, and Test (VLSI-DAT), 2013 International Symposium on, IEEE, Apr. 22, 2013 (Apr. 22, … [cited by applicant]