IP Library › Granted Patent US 12,436,705
Granted Patent B2
US 12,436,705 · App. 17/358,914 · Granted Oct 7, 2025

Dynamically scalable and partitioned copy engine

Inventors: Nilay Mistry (Bangalore, IN); David Puffer (Tempe, AZ); Prasoonkumar Surti (Folsom, CA); Hema Chand Nalluri (Bangalore, IN)
Assignee: INTEL CORPORATION
G06F3/065G06F3/0604G06F3/064G06F3/0673G06F9/3887G06F9/3888G06F9/38885G06F13/4027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,436,705
App. No.
17/358,914
Granted
Oct 7, 2025
Kind
B2
Abstract

An apparatus to facilitate a dynamically scalable and partitioned copy engine is disclosed. The apparatus includes a processor comprising copy engine hardware circuitry to facilitate copying surface data in memory and comprising: a plurality of copy front-end hardware circuitry to generate a plurality of surface data sub-blocks, wherein a number of the plurality of copy front-end hardware circuitry corresponds to a number of partitions configured for the processor, with each partition associated with a single copy front-end hardware circuitry; a plurality of copy back-end hardware circuitry to operate in parallel to process the plurality of surface data sub-blocks to perform memory accesses, wherein subsets of the plurality of copy back-end hardware circuitry are each associated with the single copy front-end hardware circuitry associated with each partition; and a connectivity matrix hardware circuitry to communicably connect the plurality of copy front-end hardware circuitry to the plurality of copy back-end hardware circuitry.

Claims (44)

1. A method comprising:

receiving, by a copy engine hardware circuitry of a graphics processor, configuration information for a graphics processor, the configuration information comprising a number of partitions of the processor and a size of each of the partitions;

configuring a number of copy front-ends of the copy engine hardware circuitry based on the number of partitions, wherein each partition is assigned a copy front-end;

configuring a number of copy back-ends of the copy engine hardware circuitry based on the size of each of the partitions, wherein each copy front-end is assigned one or more of the copy back-ends, and wherein each of the copy front-ends and the one or more of the copy back-ends hardware circuitry that is assigned to the single each of the copy front-ends is to form a copy engine building block of the copy engine hardware circuitry;

configuring sub-networks of a connectivity matrix, wherein each set of copy front-end and corresponding one or more copy back-ends is communicably coupled via one of the sub-networks; and

responsive to one of the partitions being reconfigured, resetting the copy engine building block corresponding to the one of the partitions being reconfigured to an unassigned state by resetting the copy front-end hardware circuitry associated with the one of the partitions and draining to an idle state the subset of copy back-end hardware circuitry corresponding to the copy front-end hardware circuitry being reset.

2. The method of claim 1 , wherein the configuration information is received using a register of the graphics processor assigned to the copy engine hardware circuitry.

3. The method of claim 1 , wherein the partitions comprises at least one of hard partitions of the graphics processor, virtual machines (VMs) hosted on the graphics processor, containers hosted on the graphics processor, or a logical resource partition of the processor.

4. The method of claim 1 , wherein the copy front-end comprises copy front-end hardware circuitry to divide surface data from a source location in memory to generate a plurality of surface data sub-blocks.

5. The method of claim 4 , wherein the one or more copy back-ends comprise copy back-end hardware circuitry to operate in parallel to process the plurality of surface data sub-blocks to perform memory accesses.

6. The method of claim 1 , wherein the one or more of the copy back-ends assigned to each copy front-end is determined based on a maximum copy bandwidth supported by the partition associated with the each copy front-end.

7. A processor comprising:

copy engine hardware circuitry to facilitate copying surface data in memory and comprising:

a plurality of copy front-end hardware circuitry to generate a plurality of surface data sub-blocks from the surface data, wherein a number of the plurality of copy front-end hardware circuitry corresponds to a number of partitions configured for the processor, with each partition associated with a single copy front-end hardware circuitry of the plurality of copy front-end hardware circuitry;

a plurality of copy back-end hardware circuitry to operate in parallel to process the plurality of surface data sub-blocks to perform memory accesses, wherein subsets of the plurality of copy back-end hardware circuitry are each associated with the single copy front-end hardware circuitry associated with each partition, and wherein each set of the single copy front-end hardware circuitry and one of the subsets of the copy back-end hardware circuitry that is assigned to the single copy front-end hardware circuitry is to form a copy engine building block of the copy engine hardware circuitry; and

a connectivity matrix hardware circuitry to communicably connect the plurality of copy front-end hardware circuitry to the plurality of copy back-end hardware circuitry;

wherein responsive to one of the partitions being reconfigured, the copy engine building block corresponding to the one of the partitions being reconfigured is reset to an unassigned state by resetting the copy front-end hardware circuitry associated with the one of the partitions and draining to an idle state the subset of copy back-end hardware circuitry corresponding to the copy front-end hardware circuitry being reset.

8. The processor of claim 7 , wherein the partitions comprises at least one of hard partitions of the processor, virtual machines (VMs) hosted on the processor, containers hosted on the processor, or a logical resource partition of the processor.

9. The processor of claim 7 , wherein a number of the copy back-end hardware circuitry in each of the subsets is determined based on a maximum copy bandwidth supported by the partition corresponding to the single copy front-end hardware circuitry.

10. The processor of claim 7 , wherein the connectivity matrix hardware circuitry comprises controller circuitry and a plurality of crossbar circuitry.

11. The processor of claim 7 , wherein the connectivity matrix hardware circuitry comprises a plurality of subnetworks each used to connect each of the copy front-end hardware circuitry to the subset of the copy back-end hardware circuitry that is assigned to the single copy front-end hardware circuitry.

12. The processor of claim 7 , wherein the processor comprises a graphics processing unit (GPU).

13. The processor of claim 7 , wherein the processor is at least one of a single instruction multiple data (SIMD) machine or a single instruction multiple thread (SIMT) machine.

14. A system comprising:

a memory to store surface data in a source location; and

copy engine hardware circuitry to facilitate copying the surface data from the source location in the memory to a destination location in the memory and comprising:

a plurality of copy front-end hardware circuitry to generate a plurality of surface data sub-blocks from the surface data, wherein a number of the plurality of copy front-end hardware circuitry corresponds to a number of partitions configured for a processor, with each partition associated with a single copy front-end hardware circuitry of the plurality of copy front-end hardware circuitry;

a plurality of copy back-end hardware circuitry to operate in parallel to process the plurality of surface data sub-blocks to perform memory accesses, wherein subsets of the plurality of copy back-end hardware circuitry are each associated with the single copy front-end hardware circuitry associated with each partition, and wherein each set of the single copy front-end hardware circuitry and one of the subsets of the copy back-end hardware circuitry that is assigned to the single copy front-end hardware circuitry is to form a copy engine building block of the copy engine hardware circuitry; and

a connectivity matrix hardware circuitry to communicably connect the plurality of copy front-end hardware circuitry to the plurality of copy back-end hardware circuitry;

wherein responsive to one of the partitions being reconfigured, the copy engine building block corresponding to the one of the partitions being reconfigured is reset to an unassigned state by resetting the copy front-end hardware circuitry associated with the one of the partitions and draining to an idle state the subset of copy back-end hardware circuitry corresponding to the copy front-end hardware circuitry being reset.

15. The system of claim 14 , wherein the partitions comprises at least one of hard partitions of the processor, virtual machines (VMs) hosted on the processor, containers hosted on the processor, or a logical resource partition of the processor.

16. The system of claim 14 , wherein a number of the copy back-end hardware circuitry in each of the subsets is determined based on a maximum copy bandwidth supported by the partition corresponding to the single copy front-end hardware circuitry.

17. The system of claim 14 , wherein the connectivity matrix hardware circuitry comprises controller circuitry and a plurality of crossbar circuitry.

18. The system of claim 14 , wherein the connectivity matrix hardware circuitry comprises a plurality of subnetworks each used to connect each of the copy front-end hardware circuitry to the subset of the copy back-end hardware circuitry that is assigned to the single copy front-end hardware circuitry.

19. A non-transitory computer-readable storage medium having stored thereon executable computer program instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving, by a copy engine hardware circuitry of a graphics processor of the one or more processor, configuration information for a graphics processor, the configuration information comprising a number of partitions of the processor and a size of each of the partitions;

configuring a number of copy front-ends of the copy engine hardware circuitry based on the number of partitions, wherein each partition is assigned a copy front-end;

configuring a number of copy back-ends of the copy engine hardware circuitry based on the size of each of the partitions, wherein each copy front-end is assigned one or more of the copy back-ends, and wherein each of the copy front-ends and the one or more of the copy back-ends that is assigned to the each of the copy front-ends is to form a copy engine building block of the copy engine hardware circuitry;

configuring sub-networks of a connectivity matrix, wherein each set of copy front-end and corresponding one or more copy back-ends is communicably coupled via one of the sub-networks; and

responsive to one of the partitions being reconfigured, resetting the copy engine building block corresponding to the one of the partitions being reconfigured to an unassigned state by resetting the copy front-end hardware circuitry associated with the one of the partitions and draining to an idle state the subset of copy back-end hardware circuitry corresponding to the copy front-end hardware circuitry being reset.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the configuration information is received using a register of the graphics processor assigned to the copy engine hardware circuitry.

21. The non-transitory computer-readable storage medium of claim 19 , wherein the partitions comprises at least one of hard partitions of the graphics processor, virtual machines (VMs) hosted on the graphics processor, containers hosted on the graphics processor, or a logical resource partition of the processor.

22. The non-transitory computer-readable storage medium of claim 19 , wherein the copy front-end comprises copy front-end hardware circuitry to divide surface data from a source location in memory to generate a plurality of surface data sub-blocks.

23. The non-transitory computer-readable storage medium of claim 19 , wherein the one or more of the copy back-ends assigned to each copy front-end is determined based on a maximum copy bandwidth supported by the partition associated with the each copy front-end.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2021
From: MISTRY, NILAY; PUFFER, DAVID; SURTI, PRASOONKUMAR; NALLURI, HEMA
To: INTEL CORPORATION
Reel/Frame 057311/0618 →
Continuity (1)
Related Publication 20220413704A1 · Dec 29, 2022
References Cited (33)
US 9094333B1 · Klemin · 2015 [cited by examiner]
US 9582857B1 · Kolesinski · 2017 [cited by examiner]
US 10901647B2 · Surti et al. · 2021 [cited by applicant]
US 20050120173A1 · Minowa · 2005 [cited by applicant]
US 20110276762A1 · Daly et al. · 2011 [cited by applicant]
US 20130262723A1 · Luttenbacher · 2013 [cited by examiner]
US 20140064300A1 · Negishi · 2014 [cited by examiner]
US 20140095825A1 · Moon et al. · 2014 [cited by applicant]
US 20140109102A1 · Duncan et al. · 2014 [cited by applicant]
US 20160231948A1 · Gupta et al. · 2016 [cited by applicant]
US 20160321774A1 · Liang et al. · 2016 [cited by applicant]
US 20180307418A1 · Brown · 2018 [cited by examiner]
US 20200301597A1 · Surti et al. · 2020 [cited by applicant]
US 20210124627A1 · Giroux · 2021 [cited by examiner]
US 20210132816A1 · Jo · 2021 [cited by examiner]
US 20210279229A1 · Parasnis · 2021 [cited by examiner]
US 20210406209A1 · Vishnu · 2021 [cited by examiner]
US 20220207813A1 · Godey · 2022 [cited by examiner]
CN 115525211A · 2022 [cited by applicant]
EP 4109252A1 · 2022 [cited by applicant]
EP 4109252B1 · 2024 [cited by applicant]
Cao et al. Design of HPC Node with Heterogeneous Processors, 2011, IEEE, International Conference on Cluster Computing; pp. 130-137 (Year: 2011). [cited by examiner]
Notification of CN Publication for Application No. 202010101629.0, 59 pages, Oct. 13, 2020. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/358,463, mailed Sep. 23, 2020, 12 pages. [cited by applicant]
Final OA for U.S. Appl. No. 16/358,463, mailed May 19, 2020, 26 pages. [cited by applicant]
Office Action for U.S. Appl. No. 16/358,463, mailed May 6, 2020, 24 pages. [cited by applicant]
Communication under Rule 71(3) Epc, Ep Application No. 22160379.8, Feb. 7, 2024, 8 pages, EPO. [cited by applicant]
Anonymous, “Direct Memory Access—Wikipedia,” Feb. 28, 2018, 8 pages. [cited by applicant]
Anonymous, “Logical Partition—Wikipedia,” Apr. 16, 2020, 4 pages. [cited by applicant]
Extended European Search Report, EP Application No. 22160379.8, Jul. 26, 2022, 16 pages, EPO. [cited by applicant]
Intel Corporation, “Intel Iris XeMAX Graphics Open Source Programmer's Reference Manual,” for the 2020 Discrete GPU formerly named “DG1,” Feb. 2021, vol. 10, Revision 1.0, 48 pages. [cited by applicant]
Zheng Cao et al., “Design of HPC Node with Heterogeneous Processors,” IEEE International Conference on Cluster Computing, Sep. 26, 2011, pp. 130-138, IEEE. [cited by applicant]
Decision to grant a European patent pursuant to Article 97(1) EPC, EP Application No. 22160379.8, Jun. 13, 2024, 2 pages, EPO. [cited by applicant]