IP Library › Granted Patent US 12,743,235
Granted Patent B2
US 12,743,235 · App. 18/665,840 · Granted Sep 22, 2026

DMA engines configured to perform first portion data transfer commands with a first DMA engine and second portion data transfer commands with second DMA engine

Inventors: Joseph L. Greathouse (Santa Clara, CA); Sean Keely (Santa Clara, CA); Alan D. Smith (Santa Clara, CA); Anthony Asaro (Markham, CA); Ling-Ling Wang (Santa Clara, CA); Milind N Nemlekar (Santa Clara, CA); Hari Thangirala (Santa Clara, CA); Felix Kuehling (Markham, CA)
Assignees: Advanced Micro Devices, Inc.; A TI TECHNOLOGIES ULC
G06F3/0659G06F3/061G06F3/0679G06F13/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,235
App. No.
18/665,840
Granted
Sep 22, 2026
Kind
B2
Abstract

A method for hardware management of DMA transfer commands includes accessing, by a first DMA engine, a DMA transfer command and determining a first portion of a data transfer requested by the DMA transfer command. Transfer of a first portion of the data transfer by the first DMA engine is initiated based at least in part on the DMA transfer command. Similarly, a second portion of the data transfer by a second DMA engine is initiated based at least in part on the DMA transfer command. After transferring the first portion and the second portion of the data transfer, an indication is generated that signals completion of the data transfer requested by the DMA transfer command.

Claims (46)

1 . A processor, comprising:

a first direct memory access (DMA) engine configured to initiate, based at least in part on a DMA transfer command, transfer of a first portion of a data transfer; and

a second DMA engine configured to initiate, independent of the first DMA engine and based at least in part on the DMA transfer command, transfer of a second portion of the data transfer.

2 . The processor of claim 1 , wherein:

the first DMA engine is configured to receive a DMA notification indicating that the DMA transfer command is stored at a DMA buffer in system memory; and

the first DMA engine is configured to fetch the DMA transfer command from the DMA buffer.

3 . The processor of claim 2 , wherein the first DMA engine is configured to initiate transfer of the first portion of the data transfer by:

transmitting a cache probe request to a cache memory; and

transferring the first portion of the data transfer based on receiving a return response indicting a cache hit in the cache memory.

4 . The processor of claim 2 , wherein the second DMA engine is configured to initiate transfer of the second portion of the data transfer by:

transmitting a cache probe request to a cache memory; and

transferring the second portion of the data transfer from an owner main memory based on receiving a return response indicting a cache miss in the cache memory.

5 . The processor of claim 4 , wherein the first DMA engine is configured to transfer the first portion of the data transfer further by interleaving a total DMA transfer size between the first DMA engine and the second DMA engine.

6 . The processor of claim 1 , further comprising:

a primary DMA engine configured to receive the DMA transfer command and split the DMA transfer command into a plurality of smaller workloads for independent initiation by respective DMA engines.

7 . The processor of claim 6 , wherein the first DMA engine is configured to:

receive, from the primary DMA engine, one of the plurality of smaller workloads.

8 . The processor of claim 1 , wherein the processor comprises:

a base integrated circuit (IC) die including a plurality of processing stacked die chiplets 3D stacked on top of the base IC die, wherein the base IC die includes an inter-chip data fabric communicably coupling each of the plurality of processing stacked die chiplets together.

9 . The processor of claim 8 wherein:

the first DMA engine and the second DMA engine are stacked on top of the base IC die.

10 . The processor of claim 1 , wherein the first DMA engine includes a single command engine that drives multiple transfer engines.

11 . A system, comprising:

a host processor communicably coupled to a parallel processor multi-chip module, wherein the parallel processor multi-chip module includes:

a first direct memory access (DMA) engine configured to initiate, based at least in part on a DMA transfer command, independent of the first DMA engine and transfer of a first portion of a data transfer; and

a second DMA engine configured to initiate, based at least in part on the DMA transfer command, transfer of a second portion of the data transfer.

12 . The system of claim 11 , wherein:

the first DMA engine is configured to receive a DMA notification indicating that the DMA transfer command is stored at a DMA buffer in system memory; and

the first DMA engine is configured to fetch the DMA transfer command from the DMA buffer.

13 . The system of claim 12 , wherein the first DMA engine is configured to initiate transfer of the first portion of the data transfer by:

transmitting a cache probe request to a cache memory; and

transferring the first portion of the data transfer based on receiving a return response indicting a cache hit in the cache memory.

14 . The system of claim 12 , wherein the second DMA engine is configured to initiate transfer of the second portion of the data transfer by:

transmitting a cache probe request to a cache memory; and

transferring the second portion of the data transfer from an owner main memory based on receiving a return response indicting a cache miss in the cache memory.

15 . The system of claim 14 , wherein the first DMA engine is configured to transfer the first portion of the data transfer further by interleaving a total DMA transfer size between the first DMA engine and the second DMA engine.

16 . The system of claim 11 , further comprising:

a primary DMA engine configured to receive the DMA transfer command and split the DMA transfer command into a plurality of smaller workloads.

17 . The system of claim 16 , wherein the first DMA engine is configured to:

receive, from the primary DMA engine, one of the plurality of smaller workloads.

18 . The system of claim 11 , wherein the host processor comprises:

a base integrated circuit (IC) die including a plurality of processing stacked die chiplets 3D stacked on top of the base IC die, wherein the base IC die includes an inter-chip data fabric communicably coupling each of the plurality of processing stacked die chiplets together.

19 . A method, comprising:

splitting, at a primary direct memory access (DMA) engine, a DMA transfer command into a plurality of smaller workloads; and

submitting a different workload of the plurality of smaller workloads to each of a plurality of DMA engines.

20 . The method of claim 19 , wherein each of the plurality of DMA engines is configured to independently determine a portion of a data transfer by interleaving a total DMA transfer size amongst the plurality of DMA engines.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2024
From: KUEHLING, FELIX; ASARO, ANTHONY
To: ATI TECHNOLOGIES ULC
Reel/Frame 068341/0376 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2024
From: GREATHOUSE, JOSEPH L.; KEELY, SEAN; SMITH, ALAN D.; WANG, LING-LING; NEMLEKAR, MILIND N.; THANGIRALA, HARI
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 068341/0551 →
Continuity (2)
Continuation 17515976 · Nov 1, 2021
Related Publication 20240419358A1 · Dec 19, 2024
References Cited (25)
US 8255593B2 · Siddabathuni · 2012 [cited by examiner]
US 8271700B1 · Annem et al. · 2012 [cited by applicant]
US 8341311B1 · Szewerenko · 2012 [cited by examiner]
US 9424213B2 · Dobbs et al. · 2016 [cited by applicant]
US 9645738B2 · Trachy · 2017 [cited by examiner]
US 10459854B2 · Park · 2019 [cited by examiner]
US 11809953B1 · Jacob · 2023 [cited by examiner]
US 11847507B1 · Borkovic · 2023 [cited by examiner]
US 11995351B2 · Greathouse · 2024 [cited by examiner]
US 12197970B2 · Dobbs · 2025 [cited by examiner]
US 12250163B2 · Kasichainula · 2025 [cited by examiner]
US 12386766B2 · Peng · 2025 [cited by examiner]
US 20080109604A1 · Reilly et al. · 2008 [cited by applicant]
US 20140132611A1 · Chen · 2014 [cited by examiner]
US 20180260343A1 · Park · 2018 [cited by examiner]
US 20200328192A1 · Zaman · 2020 [cited by examiner]
CN 103714027 · 2014 [cited by applicant]
JP 2012039143 · 2012 [cited by applicant]
Extended European Search Report mailed Jul. 2, 2025 for Application No. 22888240.3, 7 pages. [cited by applicant]
OA—International Preliminary Report on Patentability issued in Application No. PCT/US2022/048214, mailed May 16, 2024, 7 pages. [cited by applicant]
Office Action mailed Dec. 2, 2025 for Japanese Application No. 2024-525409, 10 pages. [cited by applicant]
Office Action mailed Mar. 18, 2026 for European Application No. 22888240.3, 7 pages. [cited by applicant]
Office Action mailed Apr. 30, 2026 for Indian Application No. 202417032384, 13 pages. [cited by applicant]
“NPL-A. Tumeo et al. ”Lightweight DMA management mechanisms formultiprocessors on FPGA,“ 2008 International Conference onApplication-Specific Systems, Architectures and Processors, Leuven, Belgium, 2008, pp. 275-280, do… [cited by applicant]
Office Action mailed May 19, 2026 for Japanese Application No. 2024-525409, 4 pages. [cited by applicant]