IP Library › Granted Patent US 12,561,249
Granted Patent B2
US 12,561,249 · App. 18/388,940 · Granted Feb 24, 2026

Prefetching using a direct memory access engine

Inventors: Vydhyanathan Kalyanasundharam (Santa Clara, CA); Christopher J. Brennan (Boxborough, MA); Joseph L Greathouse (Austin, TX); Mark Fowler (Boxborough, MA)
Assignee: Advanced Micro Devices, Inc.
G06F12/0862G06F13/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,249
App. No.
18/388,940
Granted
Feb 24, 2026
Kind
B2
Abstract

A processing system includes one or more DMA engines that load data from memory or another cache location without storing the data after loading it. As the data propagates past caches located between the memory or other cache location that stores the requested data (“intermediate caches”), the data is selectively copied to the intermediate caches based on a cache replacement policy. Rather than the DMA engine manually storing the data into the intermediate caches, the cache replacement policies of the intermediate caches determine whether the data is copied into each respective cache and a replacement priority of the data. By bypassing storing the data, the DMA engine effectuates prefetching to the intermediate caches without expending unnecessary bandwidth or searching for a memory location to store the data, thus reducing latency and saving energy.

Claims (29)

1 . A method, comprising:

loading data from a memory in response to a prefetch request from a direct memory access (DMA) engine associated with a processor; and

responsive to loading the data, selectively copying the data to one or more caches of the processor located between the DMA engine and the memory.

2 . The method of claim 1 , further comprising:

receiving a command at the DMA engine to send the request to load the data without storing the data.

3 . The method of claim 2 , wherein the command indicates a priority of the request.

4 . The method of claim 1 , wherein the memory is one of system memory or a second cache of the processor.

5 . The method of claim 1 , wherein selectively copying the data at the one or more caches is based on a cache replacement policy.

6 . The method of claim 5 , wherein the cache replacement policy determines the one or more caches into which the data is copied.

7 . The method of claim 5 , wherein the cache replacement policy determines a replacement priority for the data.

8 . A processing system, comprising:

a processor;

a memory;

a direct memory access (DMA engine) configured to send a prefetch request to load data from the memory; and

one or more cache controllers configured to selectively copy the data to one or more caches of the processor located between the DMA engine and the memory responsive to loading the data.

9 . The processing system of claim 8 , wherein the DMA engine is further configured to bypass storing the data in response to receiving a command to send the prefetch request to load the data without storing the data.

10 . The processing system of claim 9 , wherein the command indicates a priority of the command.

11 . The processing system of claim 8 , wherein the memory is one of system memory or a second cache of the processor.

12 . The processing system of claim 8 , wherein selectively copying the data at the one or more caches is based on a cache replacement policy.

13 . The processing system of claim 12 , wherein the cache replacement policy determines the one or more caches into which the data is copied.

14 . The processing system of claim 12 , wherein the cache replacement policy determines a replacement priority for the data.

15 . A processing system, comprising:

a processor;

a memory hierarchy comprising one or more caches and a memory;

a direct memory access (DMA) engine configured to issue a prefetch request to the memory hierarchy; and

one or more cache controllers configured to selectively copy data returned in response to the prefetch request to the one or more caches located between the DMA engine and the memory.

16 . The processing system of claim 15 , wherein the DMA engine is configured to send the prefetch request to load the data without storing the data.

17 . The processing system of claim 16 , wherein the one or more cache controllers are configured to selectively copy the data based on a cache replacement policy.

18 . The processing system of claim 17 , wherein the cache replacement policy determines a replacement priority for the data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 5, 2024
From: KALYANASUNDHARAM, VYDHYANATHAN; BRENNAN, CHRISTOPHER J.; GREATHOUSE, JOSEPH L.; FOWLER, MARK
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 066649/0456 →
Continuity (1)
Related Publication 20250156329A1 · May 15, 2025
References Cited (18)
US 6003106A · Fields, Jr. et al. · 1999 [cited by applicant]
US 7010626B2 · Kahle · 2006 [cited by examiner]
US 8316158B1 · Wright et al. · 2012 [cited by applicant]
US 8627008B2 · Qureshi · 2014 [cited by examiner]
US 11080051B2 · Kerr et al. · 2021 [cited by applicant]
US 11995351B2 · Greathouse · 2024 [cited by examiner]
US 20040193754A1 · Kahle · 2004 [cited by examiner]
US 20050144337A1 · Kahle · 2005 [cited by examiner]
US 20080046657A1 · Eichenberger · 2008 [cited by examiner]
US 20090063777A1 · Usui · 2009 [cited by examiner]
US 20120254576A1 · Dedeoglu · 2012 [cited by examiner]
US 20150339062A1 · Toyoda · 2015 [cited by examiner]
US 20160055107A1 · Ambroladze et al. · 2016 [cited by applicant]
US 20190042436A1 · Klemm · 2019 [cited by examiner]
US 20210255869A1 · Sankaranarayanan et al. · 2021 [cited by applicant]
US 20230359581A1 · Gibb · 2023 [cited by examiner]
US 20240104025A1 · George · 2024 [cited by examiner]
International Search Report and Written Opinion mailed Oct. 2, 2024 for PCT/US2024/034102, 9 pages. [cited by applicant]