IP Library › Granted Patent US 12,645,605
Granted Patent B1
US 12,645,605 · App. 17/449,499 · Granted Jun 2, 2026

DMA coalescing

Inventors: Ron Diamant (San Jose, CA); Yunxuan Yu (Sunnyvale, CA); Taylor Goodhart (Snohomish, WA); Robert Geva (Cupertino, CA)
Assignee: Amazon Technologies, Inc.
G06F12/1081G06F9/30043G06F13/1668G06N3/04G06N3/0464
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,605
App. No.
17/449,499
Granted
Jun 2, 2026
Kind
B1
Abstract

A computer-implemented method includes generating or receiving instruction code for executing by a computing device to implement a neural network model, where the instruction code includes a plurality of direct memory access (DMA) instructions for data transferring between a local memory of an accelerator of the computing device and a system memory of the computing device; modifying the instruction code to arrange sources or destinations of a group of DMA instructions of the plurality of DMA instructions into a contiguous block in the local memory; and replacing the group of DMA instructions with a single DMA instruction, wherein a source address or a destination address of the single DMA instruction is the contiguous block of the local memory.

Claims (60)

1 . A computer-implemented method comprising:

generating, based on a neural network model and configuration of a computing device, instruction code for executing by the computing device to implement the neural network model, the instruction code including a plurality of direct memory access (DMA) instructions for transferring data between a local memory of an accelerator of the computing device and a system memory of the computing device;

identifying, from the plurality of DMA instructions, a group of DMA save instructions that write to a same block in the system memory;

determining, based on the instruction code, a contiguous block in the local memory for storing data to be transferred in the group of DMA save instructions;

modifying the instruction code to place sources of the group of DMA save instructions into the contiguous block in the local memory;

removing the group of DMA save instructions and replacing the group of DMA save instructions in the instruction code implementing the neural network model with a single DMA save instruction at the last DMA save instruction of the group of DMA save instructions in the instruction code, wherein a source address of the single DMA save instruction is the contiguous block of the local memory,

identifying, based on the instruction code, a second contiguous block in the local memory for storing data to be transferred in a group of DMA load instructions of the plurality of DMA instructions;

adding a tensor copy instruction to the instruction code to move data from a region of the second contiguous block to a region of the local memory outside of the second contiguous block to make room for storing the data to be transferred in the group of DMA load instructions; and

replacing the group of DMA load instructions with a single DMA load instruction at the first DMA load instruction of the group of DMA load instructions, wherein a destination address of the single DMA load instruction is the second contiguous block of the local memory.

2 . The computer-implemented method of claim 1 , wherein modifying the instruction code comprises:

adding, before the single DMA save instruction, a first tensor copy instruction to move data from a region of the contiguous block to a region of the local memory outside of the contiguous block to make room for storing the data to be transferred in the group of DMA save instructions to the contiguous block;

changing an output address of a tensor operation instruction of the instruction code to a region in the contiguous block of the local memory, wherein a source of a DMA save instruction of the group of DMA save instructions is an output of the tensor operation instruction;

adding, before the single DMA save instruction, a second tensor copy instruction to move source data of a DMA save instruction of the group of DMA save instructions from a region of the local memory outside of the contiguous block to a region in the contiguous block of the local memory; or

a combination thereof.

3 . The computer-implemented method of claim 2 , wherein the first tensor copy instruction and the second tensor copy instruction are to be performed by a pooling engine of the accelerator, the pooling engine connected to the local memory of the accelerator through a dedicated bus interface.

4 . A computer-implemented method comprising:

generating, based on a neural network model, instruction code for executing by a computing device to implement the neural network model, the instruction code including a plurality of direct memory access (DMA) instructions for transferring data between a local memory of an accelerator of the computing device and a system memory of the computing device;

modifying the instruction code to arrange sources or destinations of a group of DMA instructions of the plurality of DMA instructions into a contiguous block in the local memory; and

removing the group of DMA instructions in the instruction code implementing the neural network model and replacing the group of DMA instructions with a single DMA instruction in the instruction code, wherein a source address or a destination address of the single DMA instruction is the contiguous block of the local memory,

wherein the group of DMA instructions includes a group of DMA save instructions that save data to a same block in the system memory, and

wherein modifying the instruction code includes:

adding a tensor copy instruction to move data in a source of a DMA save instruction of the group of DMA save instructions to a region in the contiguous block of the local memory; or

changing an output address of a tensor operation instruction of the instruction code to the region in the contiguous block of the local memory, wherein the source of a DMA save instruction of the group of DMA save instructions is an output of the tensor operation instruction.

5 . The computer-implemented method of claim 4 , wherein modifying the instruction code includes adding, before the single DMA instruction, a tensor copy instruction to move data from a region of the contiguous block of the local memory to a region of the local memory outside the contiguous block.

6 . The computer-implemented method of claim 5 , wherein the tensor copy instruction is to be performed by a processing engine of the accelerator, the processing engine connected to the local memory of the accelerator through a dedicated bus interface.

7 . The computer-implemented method of claim 6 , wherein the processing engine includes a pooling engine of the accelerator.

8 . The computer-implemented method of claim 6 , wherein the dedicated bus interface is configured to read from or write to each partition of the local memory in parallel.

9 . The computer-implemented method of claim 4 , wherein the contiguous block of the local memory includes equal to or greater than 2 K bytes in a partition of the local memory.

10 . The computer-implemented method of claim 4 , further comprising:

modifying the instruction code to arrange sources or destinations of a second group of DMA instructions of the plurality of DMA instructions into a second contiguous block in the local memory, wherein the second group of DMA instructions includes a group of DMA load instructions;

removing the second group of DMA instructions in the instruction code and replacing the second group of DMA instructions with a second single DMA instruction in the instruction code, wherein a source address or a destination address of the second single DMA instruction is the second contiguous block of the local memory; and

modifying an instruction in the instruction code that uses data to be loaded into a region outside of the second contiguous block of the local memory using a DMA load instruction of the group of DMA load instructions, such that the modified instruction uses data loaded into the second contiguous block of the local memory.

11 . The computer-implemented method of claim 4 , further comprising:

modifying the instruction code to arrange sources or destinations of a second group of DMA instructions of the plurality of DMA instructions into a second contiguous block in the local memory, wherein the second group of DMA instructions includes a group of DMA load instructions;

removing the second group of DMA instructions in the instruction code and replacing the second group of DMA instructions with a second single DMA instruction in the instruction code, wherein a source address or a destination address of the second single DMA instruction is the second contiguous block of the local memory; and

adding, after the second single DMA instruction, a tensor copy instruction to move data from a region of the second contiguous block of the local memory to a second region of the local memory outside of the second contiguous block, the second region of the local memory outside of the second contiguous block being a destination of a DMA load instruction of the group of DMA load instructions.

12 . A non-transitory computer readable medium having stored therein instructions that, when executed by one or more processors, cause the one or more processors to execute a compiler, the compiler performing operations including:

generating, based on a neural network model, instruction code for executing by a computing device to implement the neural network model, the instruction code including a plurality of direct memory access (DMA) instructions for transferring data between a local memory of an accelerator of the computing device and a system memory of the computing device;

modifying the instruction code to arrange sources or destinations of a group of DMA instructions of the plurality of DMA instructions into a contiguous block in the local memory; and

removing the group of DMA instructions in the instruction code implementing the neural network model and replacing the group of DMA instructions with a single DMA instruction in the instruction code, wherein a source address or a destination address of the single DMA instruction is the contiguous block of the local memory,

wherein the group of DMA instructions includes a group of DMA save instructions that save data to a same block in the system memory, and

wherein modifying the instruction code includes:

adding a tensor copy instruction to move data in a source of a DMA save instruction of the group of DMA save instructions to a region in the contiguous block of the local memory; or

changing an output address of a tensor operation instruction of the instruction code to a region in the contiguous block of the local memory, wherein a source of a DMA save instruction of the group of DMA save instructions is an output of the tensor operation instruction.

13 . The non-transitory computer readable medium of claim 12 , wherein modifying the instruction code includes adding, before the single DMA instruction, a tensor copy instruction to move data from a region of the contiguous block of the local memory to a region of the local memory outside the contiguous block.

14 . The non-transitory computer readable medium of claim 13 , wherein the tensor copy instruction is to be performed by a processing engine of the accelerator, the processing engine connected to the local memory of the accelerator through a dedicated bus interface.

15 . The non-transitory computer readable medium of claim 12 , wherein the operations further comprise:

modifying the instruction code to arrange sources or destinations of a second group of DMA instructions of the plurality of DMA instructions into a second contiguous block in the local memory, wherein the second group of DMA instructions includes a group of DMA load instructions;

removing the second group of DMA instructions in the instruction code and replacing the second group of DMA instructions with a second single DMA instruction in the instruction code, wherein a source address or a destination address of the second single DMA instruction is the second contiguous block of the local memory; and

modifying an instruction in the instruction code that uses data to be loaded into a region outside of the second contiguous block of the local memory using a DMA load instruction of the group of DMA load instructions, such that the modified instruction uses data loaded into the second contiguous block of the local memory.

16 . A computer-implemented method comprising:

generating, based on a neural network model, instruction code for executing by a computing device to implement the neural network model, the instruction code including a plurality of direct memory access (DMA) instructions for transferring data between a local memory of an accelerator of the computing device and a system memory of the computing device;

modifying the instruction code to arrange sources or destinations of a group of DMA instructions of the plurality of DMA instructions into a contiguous block in the local memory; and

removing the group of DMA instructions in the instruction code implementing the neural network model and replacing the group of DMA instructions with a single DMA instruction in the instruction code, wherein a source address or a destination address of the single DMA instruction is the contiguous block of the local memory,

wherein the group of DMA instructions includes a group of DMA load instructions, and

wherein the computer-implemented method further comprises adding, after the single DMA instruction, a tensor copy instruction to move data from a region of the contiguous block of the local memory to a region of the local memory outside of the contiguous block, the region of the local memory outside of the contiguous block being a destination of a DMA load instruction of the group of DMA load instructions.

17 . The computer-implemented method of claim 16 , further comprising:

modifying the instruction code to arrange sources or destinations of a second group of DMA instructions of the plurality of DMA instructions into a second contiguous block in the local memory, the second group of DMA instructions including a second group of DMA load instructions;

removing the second group of DMA instructions in the instruction code and replacing the second group of DMA instructions with a second single DMA instruction in the instruction code, wherein a source address or a destination address of the second single DMA instruction is the second contiguous block of the local memory; and

modifying an instruction in the instruction code that uses data to be loaded into a second region outside of the second contiguous block of the local memory using a second DMA load instruction of the second group of DMA load instructions, such that the modified instruction uses data loaded into the second contiguous block of the local memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2021
From: DIAMANT, RON; YU, YUNXUAN; GOODHART, TAYLOR; GEVA, ROBERT
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 057657/0279 →
References Cited (17)
US 7334056B2 · Ellis · 2008 [cited by examiner]
US 7437529B2 · Burugula · 2008 [cited by examiner]
US 10366026B1 · Diamant · 2019 [cited by examiner]
US 10872290B2 · Goulding · 2020 [cited by examiner]
US 11003606B2 · Birsan · 2021 [cited by examiner]
US 11467973B1 · Diamant · 2022 [cited by examiner]
US 11502696B2 · Mathuriya · 2022 [cited by examiner]
US 11513986B1 · Khan · 2022 [cited by examiner]
US 11625269B1 · Geva · 2023 [cited by examiner]
US 11995013B2 · Lim · 2024 [cited by examiner]
US 12131188B1 · Geva · 2024 [cited by examiner]
US 20060288187A1 · Burugula · 2006 [cited by examiner]
US 20140122553A1 · Dehner · 2014 [cited by examiner]
US 20170011288A1 · Brothers · 2017 [cited by examiner]
US 20170200094A1 · Bruestle · 2017 [cited by examiner]
US 20190087708A1 · Goulding · 2019 [cited by examiner]
US 20200401540A1 · Birsan · 2020 [cited by examiner]