IP Library › Granted Patent US 11,972,132
Granted Patent B2
US 11,972,132 · App. 18/145,810 · Granted Apr 30, 2024

Data processing engine arrangement in a device

Inventors: Juan J. Noguera Serra (San Jose, CA); Goran H K Bilski (Molndal, SE); Jan Langer (Chemnitz, DE); Baris Ozgul (Dublin, IE); Richard L. Walke (Edinburgh, GB); Ralph D. Wittig (Menlo Park, CA); Kornelis A. Vissers (Sunnyvale, CA); Tim Tuan (San Jose, CA); David Clarke (Dublin, IE)
Assignee: Xilinx, Inc.
G06F3/0647G06F3/061G06F3/0683G06F13/1663G06F15/17331G06F15/7807
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,972,132
App. No.
18/145,810
Granted
Apr 30, 2024
Kind
B2
Abstract

A device includes a data processing engine array having a plurality of data processing engines organized in a grid having a plurality of rows and a plurality of columns. Each data processing engine includes a core, a memory module including a memory and a direct memory access engine. Each data processing engine includes a stream switch connected to the core, the direct memory access engine, and the stream switch of one or more adjacent data processing engines. Each memory module includes a first memory interface directly coupled to the core in the same data processing engine and one or more second memory interfaces directly coupled to the core of each of the one or more adjacent data processing engines.

Claims (38)

1. A device, comprising:

a data processing engine array having a plurality of data processing engines organized in a grid having a plurality of rows and a plurality of columns;

wherein each data processing engine includes a core, a memory module including a memory, and a direct memory access engine;

wherein each data processing engine includes a stream switch, wherein the stream switch of each data processing engine is connected to the core and the direct memory access engine in a same data processing engine, and to the stream switch of one or more adjacent data processing engines;

wherein each memory module includes a plurality of memory interfaces including a first memory interface directly coupled to the core in the same data processing engine and one or more second memory interfaces directly coupled to the core of each of the one or more adjacent data processing engines; and

wherein each core coupled to a selected memory interface of the plurality of memory interfaces of a selected memory module is configured to access the memory of the selected memory module via the selected memory interface independently of the stream switch.

2. The device of claim 1 , wherein a memory module of a first data processing engine is configured to receive data streamed via the stream switch and the direct memory access engine included therein from at least one of a non-adjacent data processing engine or an adjacent data processing engine.

3. The device of claim 1 , wherein a memory module of a first data processing engine is configured to receive data streamed via the stream switch and the direct memory access engine included therein from a non-adjacent data processing engine; and

wherein a core of the first data processing engine is configured to directly access data from the memory module included therein and directly access data from a memory module of an adjacent data processing engine.

4. The device of claim 1 , wherein each data processing engine comprises:

a hardware synchronization circuit including a plurality of hardware locks configured to control access to the memory module in the same data processing engine; and

wherein the hardware synchronization circuit arbitrates access to the memory module between cores of the one or more adjacent data processing engines with direct access and one or more non-adjacent data processing engines coupled via the stream switch and the direct memory access engine.

5. The device of claim 1 , further comprising:

a processor coupled to the data processing engine array, wherein the processor is configured to control operation of the data processing engine array.

6. The device of claim 5 , wherein the processor is configured to reconfigure at least a portion of the data processing engine array in response to a detected condition in the device.

7. The device of claim 5 , wherein the processor, in response to a detected condition in the device, is configured to load new data into one or more selected memory modules of the data processing engine array.

8. The device of claim 5 , wherein the processor is configured to control application parameters stored in the memory module of one or more of the plurality of data processing engines in response to a detected condition in the device.

9. The device of claim 5 , wherein one or more of the plurality of data processing engines is operative as a kernel that uses parameters stored in a selected memory module at runtime; and

wherein the processor, in response to a detected condition in the device, is configured to calculate new parameters dynamically at runtime of the data processing engine array and store the new parameters in the selected memory module.

10. A device, comprising:

a data processing engine array having a plurality of data processing engines organized in a grid having a plurality of rows and a plurality of columns;

wherein each data processing engine includes a core, a memory module including a memory, and a direct memory access engine;

wherein each data processing engine includes a stream switch, wherein the stream switch of each data processing engine is connected to the core and the direct memory access engine in a same data processing engine, and the stream switch of one or more adjacent data processing engines;

wherein each memory module includes a first memory interface directly coupled to the core in the same data processing engine and one or more second memory interfaces directly coupled to the core of each of the one or more adjacent data processing engines;

wherein each data processing engine further comprises a hardware synchronization circuit including a plurality of hardware locks configured to control access to the memory module in the same data processing engine; and

wherein the hardware synchronization circuit is configured to arbitrate access to the memory module between cores of the one or more adjacent data processing engines with direct access and one or more non-adjacent data processing engines coupled via the stream switch and the direct memory access engine.

11. The device of claim 10 , wherein a memory module of a first data processing engine is configured to receive data streamed via the stream switch and the direct memory access engine included therein from a non-adjacent data processing engine.

12. The device of claim 10 , wherein a second data processing engine is configured to directly access data from the memory module included therein and directly access data from a memory module of an adjacent data processing engine.

13. The device of claim 10 , wherein a memory module of a first data processing engine is configured to receive data streamed via the stream switch and the direct memory access engine included therein from a non-adjacent data processing engine; and

wherein a core of the first data processing engine is configured to directly access data from the memory module included therein and directly access data from a memory module of an adjacent data processing engine.

14. The device of claim 10 , further comprising:

a processor coupled to the data processing engine array, wherein the processor is configured to control operation of the data processing engine array.

15. The device of claim 14 , wherein the processor is configured to reconfigure at least a portion of the data processing engine array in response to a detected condition in the device.

16. The device of claim 14 , wherein the processor, in response to a detected condition in the device, is configured to load new data into one or more selected memory modules of the data processing engine array.

17. The device of claim 14 , wherein the processor is configured to control application parameters stored in the memory module of one or more of the plurality of data processing engines in response to a detected condition in the device.

18. The device of claim 14 , wherein one or more of the plurality of data processing engines is operative as a kernel that uses parameters stored in a selected memory module at runtime; and

wherein the processor, in response to a detected condition in the device, is configured to calculate new parameters dynamically at runtime of the data processing engine array and store the new parameters in the selected memory module.

19. The device of claim 1 , wherein, within each memory module, the direct memory access engine, the first memory interface, and the one or more second memory interfaces are coupled to a plurality of arbiter circuits within the memory module, the plurality of arbiter circuits configured to arbitrate among memory accesses received over a plurality of different data paths corresponding to the direct memory access engine, the first memory interface and the one or more second memory interfaces.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2023
From: NOGUERA SERRA, JUAN J.; BILSKI, GORAN HK; LANGER, JAN; OZGUL, BARIS; TUAN, TIM; WALKE, RICHARD L.; WITTIG, RALPH D.; VISSERS, KORNELIS A.; CLARKE, DAVID
To: XILINX, INC.
Reel/Frame 062428/0707 →
Continuity (3)
Continuation 17097917 · Nov 13, 2020
Division 15944160 · Apr 3, 2018
Related Publication 20230131698A1 · Apr 27, 2023
Cited By (1)
US 12,619,371