IP Library › Granted Patent US 12,639,256
Granted Patent B2
US 12,639,256 · App. 18/781,952 · Granted May 26, 2026

Controllers in data processing engine columns

Inventors: Juan J. Noguera Serra (San Jose, CA); David Patrick Clarke (Dublin, IE); Javier Cabezas Rodriguez (Austin, TX); Mikhail Asiatici (Cambridge, GB); Patrick Schlangen (Cologne, DE)
Assignee: XILINX, INC.
G06F15/80G06F13/28G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,256
App. No.
18/781,952
Granted
May 26, 2026
Kind
B2
Abstract

Embodiments herein describe a hardware accelerator with an array of data processing engines (DPEs) which includes a controller (e.g., a microcontroller) for multiple columns of the array. The controllers can be hardened circuitry that executes software code (or firmware) that controls the hardware accelerator. In one embodiment, the task of the controller is to control and orchestrate the functions performed by the hardware accelerator.

Claims (36)

1 . A hardware accelerator array comprising:

a plurality of data processing engine (DPE) tiles arranged in a plurality of columns; and

a plurality of interface tiles, wherein each of the plurality of columns includes at least one of the plurality of interface tiles, wherein each of the plurality of interface tiles comprises a respective controller configured to control execution of multiple DPE tiles in a respective column of the plurality of columns, wherein each of the controllers comprises:

direct memory access (DMA) circuitry configured to fetch commands, wherein the commands define a task to be performed by the plurality of DPE tiles; and

first circuitry configured to convert the commands into instructions for controlling the execution of the multiple DPE tiles in the respective column.

2 . The hardware accelerator array of claim 1 , wherein the plurality of interface tiles is coupled to a network on chip (NoC) to facilitate data communication into, and out of, the hardware accelerator array.

3 . The hardware accelerator array of claim 1 , wherein

the DMA circuitry is configured to fetch the commands from a binary generated by a compiler

wherein each of the controllers comprises a core comprising the first circuitry that is configured to execute a program that converts the commands into the instructions for controlling the execution of the multiple DPE tiles in the respective column.

4 . The hardware accelerator array of claim 3 , wherein the instructions perform DMA operations using DMA circuitry in at least one of a respective interface tile or the multiple DPE tiles in the respective column.

5 . The hardware accelerator array of claim 3 , wherein each of the plurality of interface tiles comprises DMA circuitry configured to perform DMA operations based on the instructions.

6 . The hardware accelerator array of claim 3 , wherein each of the plurality of interface tiles comprises a memory mapped (MM) switch that is used by the respective controller to transmit the instructions to other tiles in the respective column.

7 . The hardware accelerator array of claim 3 , wherein the core is configured to receive a completion signal from a tile when the instructions are completed, and convert additional commands from the binary into additional instructions for the tile.

8 . The hardware accelerator array of claim 3 , wherein the binary is compiled based on a machine learning (ML) model, wherein the commands comprise ML functions that are converted into DMA operations that enable the plurality of DPE tiles to perform the ML functions.

9 . The hardware accelerator array of claim 1 , further comprising:

a plurality of memory tiles arranged in the plurality of columns and configured to store data processed by the plurality of DPE tiles, wherein the plurality of memory tiles comprises DMA circuitry, but do not have cores.

10 . The hardware accelerator array of claim 9 , wherein the plurality of memory tiles are disposed between the plurality of interface tiles and the plurality of DPE tiles.

11 . A hardware accelerator array comprising:

a plurality of data processing engine (DPE) tiles arranged in a plurality of columns; and

a plurality of controllers arranged in the plurality of columns, the plurality of controllers configured to control execution of only DPE tiles in a same column, wherein each of the plurality of controllers comprises:

direct memory access (DMA) circuitry configured to fetch commands, wherein the commands define a task to be performed by the plurality of DPE tiles; and

a core comprising circuitry configured to convert the commands into instructions for controlling the execution of the DPE tiles in a respective column.

12 . A method for controlling a plurality of data processing engine (DPE) tiles arranged in a plurality of columns in a hardware accelerator array, the method comprising:

fetching, using a plurality of controllers, commands from a binary, wherein the plurality of controllers are disposed within the plurality of columns;

converting, using the plurality of controllers, the commands into DMA and control instructions;

transmitting the DMA and control instructions from the plurality of controllers to DMA circuitry in the hardware accelerator array to move data into the plurality of DPE tiles for processing;

receiving, at the plurality of controllers, a completion signal from a tile in the hardware accelerator array when one of the DMA and control instructions is complete; and

converting, at the plurality of controllers, additional commands from the binary into additional instructions for the tile.

13 . The method of claim 12 , wherein the plurality of controllers is disposed within a plurality of interface tiles arranged in the plurality of columns.

14 . The method of claim 13 , wherein the plurality of interface tiles is coupled to a NoC to facilitate data communication into, and out of, the hardware accelerator array.

15 . The method of claim 13 , wherein the DMA and control instructions are transmitted to DMA circuitry in the plurality of interface tiles or the plurality of DPE tiles.

16 . The method of claim 12 , wherein the completion signal is received on a streaming network in the hardware accelerator array and the DMA and control instructions are transmitted using a MM network in the hardware accelerator array.

17 . The method of claim 13 , wherein the binary is compiled based on a ML model, wherein the commands comprise ML functions that are converted into the DMA and control instructions that enable the plurality of DPE tiles to perform the ML functions.

18 . The method of claim 13 , wherein the hardware accelerator array further comprises:

a plurality of memory tiles arranged in the plurality of columns and configured to store data processed by the plurality of DPE tiles, wherein the plurality of memory tiles comprises DMA circuitry, but do not have cores.

19 . The method of claim 18 , wherein the plurality of memory tiles are disposed between the plurality of controllers and the plurality of DPE tiles in the hardware accelerator array.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2026
From: SILEXICA GMBH
To: XILINX, INC.
Reel/Frame 074528/0263 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2025
From: NOGUERA SERRA, JUAN J.; CLARKE, DAVID PATRICK; CABEZAS RODRIGUEZ, JAVIER, MR.; ASIATICI, MIKHAIL
To: XILINX, INC.
Reel/Frame 070137/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2025
From: SCHLANGEN, PATRICK
To: SILEXICA GMBH
Reel/Frame 070137/0036 →
Continuity (1)
Related Publication 20260030198A1 · Jan 29, 2026
References Cited (11)
US 9218443B1 · Styles · 2015 [cited by examiner]
US 9846660B2 · Styles et al. · 2017 [cited by applicant]
US 10031760B1 · Santan et al. · 2018 [cited by applicant]
US 10790828B1 · Gunter · 2020 [cited by examiner]
US 10802995B2 · Singh et al. · 2020 [cited by applicant]
US 10877766B2 · Soe et al. · 2020 [cited by applicant]
US 10924430B2 · Thyamagondlu et al. · 2021 [cited by applicant]
US 11694066B2 · Ng et al. · 2023 [cited by applicant]
US 12067465B2 · Kalari · 2024 [cited by examiner]
US 20190303311A1 · Bilski · 2019 [cited by examiner]
U.S. Appl. No. 18/781,955, filed Jul. 23, 2024 Entitled “Executing Workloads Across Data Processing Engine Columns”. [cited by applicant]