DMA engines configured to perform first portion data transfer commands with a first DMA engine and second portion data transfer commands with second DMA engine
A method for hardware management of DMA transfer commands includes accessing, by a first DMA engine, a DMA transfer command and determining a first portion of a data transfer requested by the DMA transfer command. Transfer of a first portion of the data transfer by the first DMA engine is initiated based at least in part on the DMA transfer command. Similarly, a second portion of the data transfer by a second DMA engine is initiated based at least in part on the DMA transfer command. After transferring the first portion and the second portion of the data transfer, an indication is generated that signals completion of the data transfer requested by the DMA transfer command.
1 . A processor, comprising:
a first direct memory access (DMA) engine configured to initiate, based at least in part on a DMA transfer command, transfer of a first portion of a data transfer; and
a second DMA engine configured to initiate, independent of the first DMA engine and based at least in part on the DMA transfer command, transfer of a second portion of the data transfer.
2 . The processor of claim 1 , wherein:
the first DMA engine is configured to receive a DMA notification indicating that the DMA transfer command is stored at a DMA buffer in system memory; and
the first DMA engine is configured to fetch the DMA transfer command from the DMA buffer.
3 . The processor of claim 2 , wherein the first DMA engine is configured to initiate transfer of the first portion of the data transfer by:
transmitting a cache probe request to a cache memory; and
transferring the first portion of the data transfer based on receiving a return response indicting a cache hit in the cache memory.
4 . The processor of claim 2 , wherein the second DMA engine is configured to initiate transfer of the second portion of the data transfer by:
transmitting a cache probe request to a cache memory; and
transferring the second portion of the data transfer from an owner main memory based on receiving a return response indicting a cache miss in the cache memory.
5 . The processor of claim 4 , wherein the first DMA engine is configured to transfer the first portion of the data transfer further by interleaving a total DMA transfer size between the first DMA engine and the second DMA engine.
6 . The processor of claim 1 , further comprising:
a primary DMA engine configured to receive the DMA transfer command and split the DMA transfer command into a plurality of smaller workloads for independent initiation by respective DMA engines.
7 . The processor of claim 6 , wherein the first DMA engine is configured to:
receive, from the primary DMA engine, one of the plurality of smaller workloads.
8 . The processor of claim 1 , wherein the processor comprises:
a base integrated circuit (IC) die including a plurality of processing stacked die chiplets 3D stacked on top of the base IC die, wherein the base IC die includes an inter-chip data fabric communicably coupling each of the plurality of processing stacked die chiplets together.
9 . The processor of claim 8 wherein:
the first DMA engine and the second DMA engine are stacked on top of the base IC die.
10 . The processor of claim 1 , wherein the first DMA engine includes a single command engine that drives multiple transfer engines.
11 . A system, comprising:
a host processor communicably coupled to a parallel processor multi-chip module, wherein the parallel processor multi-chip module includes:
a first direct memory access (DMA) engine configured to initiate, based at least in part on a DMA transfer command, independent of the first DMA engine and transfer of a first portion of a data transfer; and
a second DMA engine configured to initiate, based at least in part on the DMA transfer command, transfer of a second portion of the data transfer.
12 . The system of claim 11 , wherein:
the first DMA engine is configured to receive a DMA notification indicating that the DMA transfer command is stored at a DMA buffer in system memory; and
the first DMA engine is configured to fetch the DMA transfer command from the DMA buffer.
13 . The system of claim 12 , wherein the first DMA engine is configured to initiate transfer of the first portion of the data transfer by:
transmitting a cache probe request to a cache memory; and
transferring the first portion of the data transfer based on receiving a return response indicting a cache hit in the cache memory.
14 . The system of claim 12 , wherein the second DMA engine is configured to initiate transfer of the second portion of the data transfer by:
transmitting a cache probe request to a cache memory; and
transferring the second portion of the data transfer from an owner main memory based on receiving a return response indicting a cache miss in the cache memory.
15 . The system of claim 14 , wherein the first DMA engine is configured to transfer the first portion of the data transfer further by interleaving a total DMA transfer size between the first DMA engine and the second DMA engine.
16 . The system of claim 11 , further comprising:
a primary DMA engine configured to receive the DMA transfer command and split the DMA transfer command into a plurality of smaller workloads.
17 . The system of claim 16 , wherein the first DMA engine is configured to:
receive, from the primary DMA engine, one of the plurality of smaller workloads.
18 . The system of claim 11 , wherein the host processor comprises:
a base integrated circuit (IC) die including a plurality of processing stacked die chiplets 3D stacked on top of the base IC die, wherein the base IC die includes an inter-chip data fabric communicably coupling each of the plurality of processing stacked die chiplets together.
19 . A method, comprising:
splitting, at a primary direct memory access (DMA) engine, a DMA transfer command into a plurality of smaller workloads; and
submitting a different workload of the plurality of smaller workloads to each of a plurality of DMA engines.
20 . The method of claim 19 , wherein each of the plurality of DMA engines is configured to independently determine a portion of a data transfer by interleaving a total DMA transfer size amongst the plurality of DMA engines.