IP Library Granted Patent US 12,724,734
Granted Patent B2
US 12,724,734 · App. 18/637,305 · Granted Sep 1, 2026

Scatter and gather streaming data through a circular FIFO

Inventors: Marc A. Schaub (Sunnyvale, CA); Roy G. Moss (Palo Alto, CA)
Assignee: Apple Inc.
G06F13/37G06F9/30069G06F9/5022G06F9/544G06F13/1642G06F13/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,734
App. No.
18/637,305
Granted
Sep 1, 2026
Kind
B2
Abstract

Systems, apparatuses, and methods for performing scatter and gather direct memory access (DMA) streaming through a circular buffer are described. A system includes a circular buffer, producer DMA engine, and consumer DMA engine. After the producer DMA engine writes or skips over a given data chunk of a first frame to the buffer, the producer DMA engine sends an updated write pointer to the consumer DMA engine indicating that a data credit has been committed to the buffer and that the data credit is ready to be consumed. After the consumer DMA engine reads or skips over the given data chunk of the first frame from the buffer, the consumer DMA engine sends an updated read pointer to the producer DMA engine indicating that the data credit has been consumed and that space has been freed up in the buffer to be reused by the producer DMA engine.

Claims (68)

1 . An apparatus, comprising:

a buffer to store data from a plurality of producer direct memory access (DMA) engines comprising a first producer DMA engine and second producer DMA engine;

the first producer DMA engine respectively comprising circuitry configured to:

store a respective first portion of data to a respective first location in the buffer; and

send, upon completion of the storing of the first portion of data, a first write pointer identifying the respective first location to indicate the respective first portion of the data is ready to be consumed;

the second producer DMA engine comprising circuitry configured to:

store a second portion of data to a second location in the buffer; and

send, upon completion of the storing of the second portion of data, a second write pointer identifying the second location to indicate the second portion of the data is ready to be consumed;

a companion router comprising routing circuitry configured to merge respective at least the first write pointer and the second write pointer sent from respective ones of the plurality of producer DMA engines into a single updated write pointer; and

a consumer DMA engine comprising circuitry configured to consume the respective first portion of data and the second portion of data from the buffer according to the single updated write pointer.

2 . The apparatus of claim 1 , wherein the plurality of producer DMA engines individually comprise additional circuitry configured to advance a current location of a read pointer by a programmable skip amount of one or more data credits to generate a respective location in the buffer.

3 . The apparatus of claim 1 , wherein the first portion of data and the second portion of data are portions of a superframe of a video sequence and the buffer is a circular buffer with a size that is smaller than a size of the superframe.

4 . The apparatus of claim 3 , the routing circuitry further configured to:

manage initialization and updating of routing tables in a plurality of routers; and

retrieve route descriptors for the superframe from a corresponding route descriptor queue and initialize route entries in the plurality of routers.

5 . The apparatus of claim 1 , wherein:

the consumer DMA engine is one of a plurality of consumer DMA engines respectively comprising circuitry configured to:

consume a respective portion of data from a respective location in the buffer; and

send, upon completion of the consuming, a read pointer identifying the respective location to indicate the respective portion of the data has been consumed;

the routing circuitry is further configured to merge respective read pointers sent from respective ones of the plurality of consumer DMA engines into a single updated read pointer; and

the plurality of producer DMA engines individually comprise additional circuitry configured to store additional portions of data according to the single updated read pointer.

6 . The apparatus of claim 5 , wherein the plurality of consumer DMA engines individually comprise additional circuitry configured to advance a current location of the single updated write pointer by a programmable skip amount of one or more data credits to generate the respective location in the buffer.

7 . The apparatus of claim 5 , wherein the plurality of producer DMA engines individually comprise additional circuitry configured to determine whether space is available to receive the respective portion of data at the respective location in the buffer, and responsive to determining that space is not available, wait for an update of the single updated read pointer prior to storing the respective portion of data to a respective location in the buffer.

8 . A method, comprising:

performing, by circuitry of first producer direct memory access (DMA) engine:

storing a first portion of data to a first location in a buffer, the buffer comprising data received from the plurality of producer DMA engines including the first DMA engine and a second DMA engine; and

sending, upon completion of the storing of the first portion, a first write pointer identifying the first location to indicate the first portion of the data is ready to be consumed;

performing, by circuitry of a second producer direct memory access (DMA) engine:

storing a second portion of data to a second location in a buffer; and

sending, upon completion of the storing of the second portion, a second write pointer identifying the second location to indicate the second portion of the data is ready to be consumed;

merging, by routing circuitry, at least the first write pointer and the second write pointer into a single updated write pointer; and

consuming, by circuitry of a consumer DMA engine, the first portion of data and the second portion of data from the buffer according to the single updated write pointer.

9 . The method of claim 8 , further comprising advancing, by additional circuitry of individual engines of the plurality of producer DMA engines, a current location of a read pointer by a programmable skip amount of one or more data credits to generate a respective location in the buffer.

10 . The method of claim 8 , wherein the first portion of data and the second portion of data are portions of a superframe of a video sequence and the buffer is a circular buffer with a size that is smaller than a size of the superframe.

11 . The method of claim 10 , further comprising performing, by the routing circuitry:

managing initialization and updating of routing tables in a plurality of routers; and

retrieving route descriptors for the superframe from a corresponding route descriptor queue and initialize route entries in the plurality of routers.

12 . The method of claim 8 , further comprising:

performing, by circuitry of a plurality of consumer DMA engines including the consumer DMA engine:

consuming a respective portion of data from a respective location in the buffer; and

sending, upon completion of the consuming, a read pointer identifying the respective location to indicate the respective portion of the data has been consumed;

merging, by the routing circuitry, respective read pointers sent from respective ones of the plurality of consumer DMA engines into a single updated read pointer; and

storing, by additional circuitry of individual engines of the plurality of producer DMA engines, additional portions of data according to the single updated read pointer.

13 . The method of claim 12 , further comprising advancing, by additional circuitry of individual engines of the plurality of consumer DMA engines, a current location of the single updated write pointer by respective programmable skip amounts of one or more data credits to generate the respective locations in the buffer.

14 . The method of claim 12 , further comprising determining, by additional circuitry of individual ones of the plurality of producer DMA engines, whether respective space is available to receive the respective portion of data at the respective location in the buffer, and responsive to determining that space is not available, wait for an update of the single updated read pointer prior to storing the respective portion of data to a respective location in the buffer.

15 . A system, comprising:

one or more processors;

a memory implementing a buffer to provide data to a plurality of consumer direct memory access (DMA) engines comprising a first consumer DMA engine and second consumer DMA engine;

the first consumer DMA engine comprising circuitry configured to:

consume a first portion of data from a first location in the buffer; and

send, upon completion of the consuming of the first portion of data, a first read pointer identifying the first location to indicate the first portion of the data has been consumed;

the second consumer DMA engine comprising circuitry configured to:

consume a second portion of data from a second location in the buffer; and

send, upon completion of the consuming of the second portion of data, a second read pointer identifying the second location to indicate the second portion of the data has been consumed;

routing circuitry configured to merge at least the first read pointer and the second read pointers sent into a single updated read pointer;

and a producer DMA engine comprising circuitry configured to store the respective an additional portion of data in the buffer according to the single updated read pointer.

16 . The system of claim 15 , wherein the plurality of consumer DMA engines individually comprise additional circuitry configured to advance a current location of a write pointer by a programmable skip amount of one or more data credits to generate a respective location in the buffer.

17 . The system of claim 15 , wherein the first portion of data and the second portion of data are portions of a superframe of a video sequence and the buffer is a circular buffer with a size that is smaller than a size of the superframe.

18 . The system of claim 17 , routing circuitry further configured to:

manage initialization and updating of routing tables in a plurality of routers; and

retrieve route descriptors for the superframe from a corresponding route descriptor queue and initialize route entries in the plurality of routers.

19 . The system of claim 15 , wherein:

the producer DMA engine is one of a plurality of producer DMA engines individually comprising circuitry configured to:

store a respective portion of data from a respective location in the buffer; and

send, upon completion of the storing, a write pointer identifying the respective location to indicate the respective portion of the data is ready to be consumed;

the routing circuitry is further configured to merge respective write pointers sent from respective ones of the plurality of producer DMA engines into a single updated write pointer; and

the plurality of consumer DMA engines individually comprise additional circuitry configured to consume the respective portions of data according to the single updated write pointer.

20 . The system of claim 19 , wherein the plurality of producer DMA engines individually comprise additional circuitry configured to advance a current location of the single updated read pointer by a programmable skip amount of one or more data credits to generate the respective location in the buffer.

Continuity (2)
Continuation 16922623 · Jul 7, 2020
Related Publication 20240264963A1 · Aug 8, 2024
References Cited (98)
US 1186802A · Krueger · 1916 [cited by applicant]
US 4864496A · Triolo et al. · 1989 [cited by applicant]
US 5434996A · Bell · 1995 [cited by applicant]
US 5754614A · Wingen · 1998 [cited by applicant]
US 5951635A · Kamgar · 1999 [cited by applicant]
US 6012109A · Schultz · 2000 [cited by applicant]
US 6075833A · Leshay et al. · 2000 [cited by applicant]
US 6108743A · Debs · 2000 [cited by examiner]
US 6480942B1 · Hirairi · 2002 [cited by applicant]
US 6615302B1 · Birns · 2003 [cited by applicant]
US 6680874B1 · Harrison · 2004 [cited by applicant]
US 6717576B1 · Duluk, Jr. · 2004 [cited by applicant]
US 6724683B2 · Liao · 2004 [cited by applicant]
US 6725388B1 · Susnow · 2004 [cited by applicant]
US 6956776B1 · Lowe et al. · 2005 [cited by applicant]
US 6963946B1 · Dwork et al. · 2005 [cited by applicant]
US 7009618B1 · Brunner · 2006 [cited by applicant]
US 7031838B1 · Young · 2006 [cited by applicant]
US 7035983B1 · Fensore · 2006 [cited by applicant]
US 7107393B1 · Sabih · 2006 [cited by applicant]
US 7116601B2 · Fung · 2006 [cited by applicant]
US 7234645B2 · Silverbrook · 2007 [cited by applicant]
US 7237036B2 · Boucher · 2007 [cited by applicant]
US 7337244B2 · Furukawa · 2008 [cited by examiner]
US 7570534B2 · Wang et al. · 2009 [cited by applicant]
US 7630361B2 · Chapman · 2009 [cited by applicant]
US 7672332B1 · Chapman · 2010 [cited by applicant]
US 7860084B2 · Binder · 2010 [cited by applicant]
US 8036214B2 · Elliott · 2011 [cited by applicant]
US 8135878B1 · Jain · 2012 [cited by examiner]
US 8327187B1 · Metcalf · 2012 [cited by applicant]
US 8707370B2 · Carter · 2014 [cited by applicant]
US 8848725B2 · Binder · 2014 [cited by applicant]
US 9294386B2 · Narad · 2016 [cited by applicant]
US 9952991B1 · Bruce · 2018 [cited by examiner]
US 10409524B1 · Branover · 2019 [cited by applicant]
US 10462627B2 · Raleigh · 2019 [cited by applicant]
US 10672098B1 · Chemparathy et al. · 2020 [cited by applicant]
US 10860511B1 · Thompson · 2020 [cited by applicant]
US 11274929B1 · Afrouzi · 2022 [cited by applicant]
US 11669481B2 · Jen · 2023 [cited by examiner]
US 11775452B2 · Jung · 2023 [cited by examiner]
US 12001365B2 · Schaub · 2024 [cited by examiner]
US 12197970B2 · Dobbs · 2025 [cited by examiner]
US 12322068B1 · Kim · 2025 [cited by examiner]
US 20030154341A1 · Asaro · 2003 [cited by applicant]
US 20030165160A1 · Minami · 2003 [cited by applicant]
US 20050125571A1 · Lin et al. · 2005 [cited by applicant]
US 20050128846A1 · Momtaz et al. · 2005 [cited by applicant]
US 20060277329A1 · Paulson et al. · 2006 [cited by applicant]
US 20070011368A1 · Wang · 2007 [cited by applicant]
US 20070220184A1 · Tierno · 2007 [cited by applicant]
US 20070253430A1 · Minami · 2007 [cited by applicant]
US 20080056192A1 · Strong · 2008 [cited by applicant]
US 20080209084A1 · Wang · 2008 [cited by examiner]
US 20090187679A1 · Puri · 2009 [cited by examiner]
US 20100005199A1 · Gadgil · 2010 [cited by applicant]
US 20110187829A1 · Nakajima · 2011 [cited by applicant]
US 20120154375A1 · Zhang · 2012 [cited by applicant]
US 20130010617A1 · Chen · 2013 [cited by applicant]
US 20130054901A1 · Biswas · 2013 [cited by applicant]
US 20130132854A1 · Raleigh · 2013 [cited by applicant]
US 20130282807A1 · Rudy · 2013 [cited by applicant]
US 20140143470A1 · Dobbs et al. · 2014 [cited by applicant]
US 20140344488A1 · Flynn · 2014 [cited by applicant]
US 20150039815A1 · Klein · 2015 [cited by examiner]
US 20150091927A1 · Cote · 2015 [cited by examiner]
US 20150178241A1 · Ajanovic · 2015 [cited by applicant]
US 20150281126A1 · Regula · 2015 [cited by examiner]
US 20150378737A1 · Debbage · 2015 [cited by applicant]
US 20170024568A1 · Pappachan · 2017 [cited by applicant]
US 20170168970A1 · Sajeepa · 2017 [cited by examiner]
US 20180089117A1 · Nicol · 2018 [cited by applicant]
US 20180113826A1 · Li · 2018 [cited by applicant]
US 20180183733A1 · Dcruz · 2018 [cited by applicant]
US 20180376171A1 · Dhandapani · 2018 [cited by applicant]
US 20190004133A1 · Li · 2019 [cited by applicant]
US 20190081903A1 · Kobayashi · 2019 [cited by applicant]
US 20190102859A1 · Hux · 2019 [cited by applicant]
US 20190196745A1 · Persson · 2019 [cited by applicant]
US 20190205153A1 · Niestemski · 2019 [cited by applicant]
US 20190205244A1 · Smith · 2019 [cited by applicant]
US 20190373086A1 · Qi · 2019 [cited by applicant]
US 20200045519A1 · Raleigh · 2020 [cited by applicant]
US 20200125500A1 · Guan · 2020 [cited by applicant]
US 20210243129A1 · Sasu · 2021 [cited by applicant]
US 20220179805A1 · Hu · 2022 [cited by applicant]
US 20230231811A1 · Dalal · 2023 [cited by examiner]
JP 2002254729 · 2002 [cited by applicant]
JP 2017506378 · 2017 [cited by applicant]
KR 1020150095139A · 2015 [cited by applicant]
WO 2013129031 · 2013 [cited by applicant]
Cummings et al., “Simulation and Syntheses Techniques for Asynchronous FIFO Design with Asynchronous Pointer Comparisons”, Sunburst Design, Inc., 2002, 18 pages. [cited by applicant]
Keller, James B., “The PWRficient Processor Family,” PA Semi, Inc., Oct. 2005, 31 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2021/039190, mailed Oct. 28, 2021, 10 pages. [cited by applicant]
Office Action from Japanese Patent Application No. 2023-500423, dated Jan. 19, 2024, pp. 1-11. [cited by applicant]
Notice of Allowance from Korean Patent Application No. 10-2023-7003938, dated Dec. 9, 2025, pp. 1-11. [cited by applicant]
Office Action from Chinese Application No. 2021800486785, dated Jun. 30, 2026, pp. 1-10. [cited by applicant]