IP Library › Granted Patent US 12,204,757
Granted Patent B1
US 12,204,757 · App. 18/067,514 · Granted Jan 21, 2025

Strong ordered transaction for DMA transfers

Inventors: Kun Xu (Austin, TX); Ron Diamant (Santa Clara, CA); Ilya Minkin (Los Altos, CA); Raymond S. Whiteside (Austin, TX)
Assignee: Amazon Technologies, Inc.
G06F3/0611G06F3/0659G06F3/0673
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,204,757
App. No.
18/067,514
Granted
Jan 21, 2025
Kind
B1
Abstract

A technique for processing strong ordered transactions in a direct memory access engine may include retrieving a memory descriptor to perform a strong ordered transaction, and delaying the strong ordered transaction until pending write transactions associated with previous memory descriptors retrieved prior to the memory descriptor are complete. Subsequent transactions associated with memory descriptors following the memory descriptor are allowed to be issued while waiting for the pending write transactions to complete. Upon completion of the pending write transactions, the strong ordered transaction is performed.

Claims (41)

1. A computer-implemented method performed by a direct memory access (DMA) engine, comprising:

processing a first memory descriptor to perform a first DMA write transaction associated with a first set of data;

processing a second memory descriptor retrieved after retrieval of the first memory descriptor to perform a second DMA write transaction to update a semaphore to indicate that the first set of data has been written, the second memory descriptor having a strong ordered indicator bit enabled to indicate that the second DMA write transaction is a strong ordered transaction that is to be performed after pending DMA write transactions associated with memory descriptors retrieved before retrieval of the second memory descriptor are complete;

delaying, based on the strong ordered indicator bit being enabled in the second memory descriptor, performance of the second DMA write transaction until the first set of data has been written;

processing a third memory descriptor retrieved after retrieval of the second memory descriptor to perform a third DMA write transaction to write a second set of data, the third memory descriptor having a strong ordered indicator bit disabled;

performing the third DMA write transaction to write the second set of data before performing the second DMA write transaction; and

upon completion of the first DMA write transaction, performing the second DMA write transaction associated with the second memory descriptor to update the semaphore.

2. The computer-implemented method of claim 1 , further comprising:

tracking, based on the strong ordered indicator bit being set in the second memory descriptor, progress of the first DMA write transaction in a scoreboard.

3. The computer-implemented method of claim 2 , wherein the scoreboard includes a dependency vector of pending DMA write transactions for each outstanding strong ordered transaction.

4. The computer-implemented method of claim 1 , wherein the first DMA write transaction is complete when an acknowledgement for the first DMA write transaction is received.

5. A computer-implemented method comprising:

retrieving a first memory descriptor in a direct memory access (DMA) engine to perform a strong ordered transaction, wherein the strong ordered transaction is to be performed after completion of pending DMA write transactions associated with memory descriptors retrieved prior to retrieval of the first memory descriptor are complete;

delaying performance of the strong ordered transaction associated with the first memory descriptor until the pending DMA write transactions associated with the memory descriptors retrieved prior to retrieval of the first memory descriptor are complete;

allowing subsequent DMA transactions associated with memory descriptors retrieved after retrieval of the first memory descriptor to be issued while waiting for the pending DMA write transactions associated with memory descriptors retrieved prior to retrieval of the first memory descriptor to complete; and

upon completion of the pending DMA write transactions associated with memory descriptors retrieved prior to retrieval of the first memory descriptor, performing the strong ordered transaction associated with the first memory descriptor.

6. The computer-implemented method of claim 5 , further comprising maintaining a scoreboard to track progress of each pending DMA write transaction that the strong ordered transaction is dependent on.

7. The computer-implemented method of claim 6 , wherein the scoreboard includes a dependency vector of pending DMA write transactions for each outstanding strong ordered transaction.

8. The computer-implemented method of claim 5 , wherein the pending DMA write transactions that the strong ordered transaction waits for include DMA write transactions from multiple write descriptor queues in the DMA engine.

9. The computer-implemented method of claim 5 , wherein the first memory descriptor has a strong ordered indicator bit to indicate that the first memory descriptor is for a strong ordered transaction.

10. The computer-implemented method of claim 5 , wherein the strong ordered transaction is a strong ordered write transaction.

11. The computer-implemented method of claim 10 , wherein the strong ordered write transaction is a write transaction to update a semaphore.

12. The computer-implemented method of claim 10 , further comprising:

retrieving an additional memory descriptor to perform an additional strong ordered transaction while the strong ordered write transaction is still pending;

delaying the additional strong ordered transaction until each DMA write transaction pending before the additional memory descriptor, including the strong ordered write transaction, are complete; and

allowing DMA transactions associated with memory descriptors following the additional memory descriptor to be issued while waiting for each DMA write transaction pending before the additional memory descriptor, including the strong ordered write transaction, to complete.

13. The computer-implemented method of claim 5 , wherein the strong ordered transaction is a strong ordered read transaction.

14. The computer-implemented method of claim 5 , wherein a pending DMA write transaction is complete when an acknowledgement is received in response to the pending DMA write transaction.

15. A direct memory access (DMA) engine comprising:

one or more descriptor queues operable to store memory descriptors associated with transactions;

a scoreboard to indicate progress of pending transactions for completion; and

a controller operable to:

retrieve, from a descriptor queue of the one or more descriptor queues, a first memory descriptor to perform a strong ordered transaction, wherein the strong ordered transaction is to be performed after pending DMA write transactions associated with memory descriptors retrieved prior to retrieval of the first memory descriptor are complete;

delay the strong ordered transaction associated with the first memory descriptor until the pending DMA write transactions associated with the memory descriptors retrieved prior to retrieval of the first memory descriptor are complete as indicated by the scoreboard;

allow subsequent DMA transactions associated with memory descriptors retrieved after retrieval of the first memory descriptor to be issued while waiting for the pending DMA write transactions to complete; and

upon completion of the pending DMA write transactions associated with the memory descriptors retrieved prior to retrieval of the first memory descriptor, perform the strong ordered transaction associated with the first memory descriptor.

16. The DMA engine of claim 15 , wherein the first memory descriptor includes a strong ordered indicator bit to indicate the strong ordered transaction.

17. The DMA engine of claim 15 , wherein the pending DMA write transactions belong to more than one descriptor queue.

18. The DMA engine of claim 15 , wherein the strong ordered transaction is a strong ordered write transaction.

19. The DMA engine of claim 15 , wherein the strong ordered transaction is a strong ordered read transaction.

20. The DMA engine of claim 15 , wherein the pending DMA write transactions are associated with writing a tensor into a buffer of a neural network processor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2023
From: XU, KUN; DIAMANT, RON; MINKIN, ILYA; WHITESIDE, RAYMOND S.
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 065324/0885 →
References Cited (27)
US 5392393A · Deering · 1995 [cited by examiner]
US 5459845A · Nguyen · 1995 [cited by examiner]
US 5740409A · Deering · 1998 [cited by examiner]
US 5874969A · Storm · 1999 [cited by examiner]
US 5893165A · Ebrahim · 1999 [cited by examiner]
US 6704831B1 · Avery · 2004 [cited by examiner]
US 7287099B1 · Powderly · 2007 [cited by examiner]
US 9134910B2 · Thompson · 2015 [cited by examiner]
US 11907575B2 · Yoon · 2024 [cited by examiner]
US 20070169042A1 · Janczewski · 2007 [cited by examiner]
US 20080140980A1 · Mei · 2008 [cited by examiner]
US 20110219204A1 · Caspole · 2011 [cited by examiner]
US 20150006834A1 · Dulloor · 2015 [cited by examiner]
US 20150019792A1 · Swanson · 2015 [cited by examiner]
US 20150281126A1 · Regula · 2015 [cited by examiner]
US 20170083326A1 · Burger · 2017 [cited by examiner]
US 20170160929A1 · Ayandeh · 2017 [cited by examiner]
US 20170286113A1 · Shanbhogue · 2017 [cited by examiner]
US 20170351516A1 · Mekkat · 2017 [cited by examiner]
US 20180074827A1 · Mekkat · 2018 [cited by examiner]
US 20200371970A1 · Chachad · 2020 [cited by examiner]
US 20210048991A1 · Tanner · 2021 [cited by examiner]
Defintion of “race condition”; Ben Lutkevich; TechTarget; Jun. 2021; retrieved from https://www.techtarget.com/searchstorage/definition/race-condition on May 10, 2024 (Year: 2021). [cited by examiner]
S. Mitsuno et al., “A High-Performance Out-of-Order Soft Processor Without Register Renaming,” 2020 30th International Conference on Field-Programmable Logic and Applications (FPL), Gothenburg, Sweden, 2020, pp. 73-78, … [cited by examiner]
F. A. Endo et al., “Micro-architectural simulation of in-order and out-of-order ARM microprocessors with gem5,” 2014 International Conference on Embedded Computer Systems: Architectures, Modeling, and Simulation, Agios … [cited by examiner]
R. Ohlendorf et al., “Performance Evaluation of RISC-based SoC Platforms in Network Processing Applications,” 2006 International Conference on Embedded Computer Systems: Architectures, Modeling and Simulation, Samos, Gr… [cited by examiner]
S. Ma, Y. Guo, S. Chen, L. Huang and Z. Wang, “Improving the DRAM Access Efficiency for Matrix Multiplication on Multicore Accelerators,” 2019 Design, Automation & Test in Europe Conference & Exhibition (DATE), Florence… [cited by examiner]
Cited By (1)
US 12,632,401