IP Library Granted Patent US 12,373,363
Granted Patent B2
US 12,373,363 · App. 17/441,668 · Granted Jul 29, 2025

Adaptive pipeline selection for accelerating memory copy operations

Inventors: Jiayu Hu (Shanghai, CN); Ren Wang (Portland, OR); Cunming Liang (Shanghai, CN)
Assignee: Intel Corporation
G06F13/1673G06F9/3004G06F9/30079G06F13/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,363
App. No.
17/441,668
Granted
Jul 29, 2025
Kind
B2
Abstract

Examples include a computing system having a direct memory access (DMA) engine pipeline, a plurality of processing cores, each processing core including a core pipeline, and a memory coupled to the DMA engine pipeline and the plurality of processing cores. The computing system includes a pipeline selector coupled to the plurality of processing cores and the DMA engine pipeline, the pipeline selector to, during initialization, determine at least one threshold for pipeline selection for the computing system, and during runtime, select one of the core pipelines or the DMA engine pipeline to execute a memory copy operation in the memory based at least in part on the at least one threshold.

Claims (40)

1. A computing system comprising:

a direct memory access (DMA) engine;

a plurality of processing cores, at least one processing core including a core; and

a circuitry coupled to the plurality of processing cores and the DMA engine, the circuitry to select from (1) at least one of the plurality of processing cores or (2) the DMA engine to execute a memory copy operation based at least in part on at least one of: message length, copy length, or buffer length, wherein:

at least one descriptor is associated with the buffer,

for a first memory copy operation, select the at least one of the plurality of processing cores to execute the first memory copy operation based on at least one of: a first message length, a first copy length, or a first buffer length, and

for a second memory copy operation, select the DMA engine to execute the second memory copy operation based on at least one of: a second message length, a second copy length, or a second buffer length.

2. The computing system of claim 1 , wherein the circuitry is to select one of the plurality of processing cores or the DMA engine to execute a memory copy operation based at least in part on a number of copy operations.

3. The computing system of claim 1 , wherein the memory copy operation is for a buffer, the buffer is described by the at least one descriptor, and the at least one descriptor is associated with the buffer length.

4. The computing system of claim 2 , wherein the select from (1) the at least one of the plurality of processing cores or (2) the DMA engine to execute the memory copy operation based at least in part on the number of copy operations comprises:

for the at least one descriptor in a batch of buffers:

based on the at least one descriptor is selected for the DMA engine and a number of descriptors selected for the DMA engine is greater than or equal to a number of copy operations, then select the DMA engine to execute the memory copy operation,

based on the at least one descriptor is selected for the one of the plurality of processing cores, then select the one of the plurality of processing cores to execute the memory copy operation, and

based on the number of descriptors selected for the DMA engine is less than the number of copy operations, then select the one of the plurality of processing cores to execute the memory copy operation.

5. The computing system of claim 4 , wherein a buffer of the batch of buffers is associated with a packet.

6. A method to be performed by a processor in a computing system, the method comprising:

during initialization of the computing system, determining at least one parameter for device selection for the computing system, and during runtime of the computing system, selecting (1) at least one of a plurality of cores or (2) a Direct Memory Access (DMA) engine to execute a memory copy operation based at least in part on the at least one parameter, wherein:

the at least one parameter comprises one or more of: message length, copy length, or buffer length,

for a first memory copy operation, selecting the at least one of the plurality of cores to execute the first memory copy operation based on at least one parameter associated with the first memory copy operation, and

for a second memory copy operation, selecting the DMA engine to execute the second memory copy operation based on at least one parameter associated with the second memory copy operation.

7. The method of claim 6 , wherein the at least one parameter comprises a number of copy operations.

8. The method of claim 6 , wherein the memory copy operation is for a buffer, the buffer is described by a descriptor, and the descriptor is associated with the buffer length.

9. The method of claim 8 , wherein selecting (1) the at least one of the plurality of cores or (2) the DMA engine to execute the memory copy operation based at least in part on the at least one parameter comprises:

for at least one descriptor in a batch of buffers:

based on the at least one descriptor is selected for the DMA engine and based on a number of descriptors selected for the DMA engine is greater than or equal to a number of copy operations, then selecting the DMA engine to execute the memory copy operation,

based on the at least one descriptor is selected for one of the cores, then selecting one of the cores to execute the memory copy operation, and

based on the number of descriptors selected for the DMA engine is less than the number of copy operations, then selecting one of the cores to execute the memory copy operation.

10. The method of claim 9 , wherein the buffer is associated with a packet.

11. At least one tangible non-transitory machine-readable medium comprising a plurality of instructions that in response to being executed by a processor in a computing system, cause the processor to:

during initialization of the computing system, determine at least one parameter for selection for the computing system, and

during runtime of the computing system, select (1) at least one of a plurality of cores or (2) a Direct Memory Access (DMA) engine to execute a memory copy operation based at least in part on the at least one parameter, wherein:

the at least one parameter comprises one or more of: message length, copy length, or buffer length,

for a first memory copy operation, selecting the at least one of the plurality of cores to execute the first memory copy operation based on at least one parameter associated with the first memory copy operation, and

for a second memory copy operation, selecting the DMA engine to execute the second memory copy operation based on at least one parameter associated with the second memory copy operation.

12. The at least one tangible non-transitory machine-readable medium of claim 11 , wherein the at least one parameter comprises a number of copy operations.

13. The at least one tangible non-transitory machine-readable medium of claim 11 , wherein the memory copy operation is for a buffer, the buffer is described by a descriptor, and the descriptor is associated with the buffer length.

14. The at least one tangible non-transitory machine-readable medium of claim 11 , wherein instructions that cause the processor to select one of the cores or the DMA engine to execute the memory copy operation based at least in part on the at least one parameter comprise instructions to:

for at least one descriptor in a batch of buffers: based on the at least one descriptor is selected for the DMA engine and if a number of descriptors selected for the DMA engine is greater than or equal to a number of copy operations, then select the DMA engine to execute the memory copy operation,

based on the at least one descriptor is selected for one of the cores or if the number of descriptors selected for the DMA engine is less than the number of copy operations, then select one of the cores to execute the memory copy operation, and

based on the number of descriptors selected for the DMA engine is less than the number of copy operations, then select one of the cores to execute the memory copy operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2021
From: HU, JIAYU; WANG, REN; LIANG, CUNMING
To: INTEL CORPORATION
Reel/Frame 058023/0628 →
Continuity (1)
Related Publication 20220179805A1 · Jun 9, 2022
References Cited (37)
US 5590313A · Reynolds · 1996 [cited by examiner]
US 6473808B1 · Yeivin · 2002 [cited by examiner]
US 7000244B1 · Adams · 2006 [cited by examiner]
US 7877524B1 · Annem · 2011 [cited by examiner]
US 7937447B1 · Cohen · 2011 [cited by examiner]
US 8458280B2 · Hausauer · 2013 [cited by examiner]
US 10884790B1 · Saidi · 2021 [cited by examiner]
US 11436112B1 · Kumar · 2022 [cited by examiner]
US 11630698B2 · Levin · 2023 [cited by examiner]
US 20060149862A1 · Zaabab · 2006 [cited by examiner]
US 20090006669A1 · Toyama · 2009 [cited by examiner]
US 20090172621A1 · Sathe · 2009 [cited by examiner]
US 20100005199A1 · Gadgil · 2010 [cited by examiner]
US 20100211703A1 · Kasahara · 2010 [cited by examiner]
US 20110004732A1 · Krakirian · 2011 [cited by examiner]
US 20110235869A1 · Lu · 2011 [cited by examiner]
US 20140173600A1 · Nair · 2014 [cited by applicant]
US 20170046179A1 · Teh · 2017 [cited by examiner]
US 20170236053A1 · Lavigueur · 2017 [cited by examiner]
US 20180024951A1 · Edmiston · 2018 [cited by examiner]
US 20180089128A1 · Nicol · 2018 [cited by examiner]
US 20180181503A1 · Nicol · 2018 [cited by examiner]
US 20190007280A1 · Sarangam · 2019 [cited by examiner]
US 20190013965A1 · Sindhu · 2019 [cited by examiner]
US 20190087218A1 · Loftus et al. · 2019 [cited by applicant]
US 20190188148A1 · Dong · 2019 [cited by applicant]
US 20190327173A1 · Gafni · 2019 [cited by examiner]
US 20220179805A1 · Hu · 2022 [cited by examiner]
US 20220407740A1 · Cox · 2022 [cited by examiner]
US 20240403107A1 · Wang · 2024 [cited by examiner]
US 20250028705A1 · Shi · 2025 [cited by examiner]
CN 101556565A · 2009 [cited by applicant]
CN 107168683A · 2017 [cited by applicant]
Kim et al., “Data Cache and Direct Memory Access in Programming Media Processors”, Aug. 2001, IEEE (Year: 2001). [cited by examiner]
Harvey et al., “DMA Fundamentals on Various PC Platforms”, Apr. 1991, National Instrument (Year: 1991). [cited by examiner]
Extended European Search Report for Patent Application No. 19934301.3, Mailed Dec. 19, 2022, 9 pages. [cited by applicant]
International Search Report and Written Opinion for PCT Patent Application No. PCT/CN19/92226, Mailed Mar. 23, 2020, 10 pages. [cited by applicant]