IP Library › Granted Patent US 12,248,415
Granted Patent B2
US 12,248,415 · App. 17/364,481 · Granted Mar 11, 2025

Automated design of behavioral-based data movers for field programmable gate arrays or other logic devices

Inventors: Stephen R. Reid (Ayer, MA); Sandeep Dutta (Foster City, CA)
Assignee: Raytheon Company
G06F13/28G06F9/30145
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,415
App. No.
17/364,481
Granted
Mar 11, 2025
Kind
B2
Abstract

A method includes obtaining behavioral source code defining logic to be performed using at least one logic device and constraints identifying data movements associated with execution of the logic. The at least one logic device contains multiple components that support at least one of: internal data movements within the at least one logic device and external data movements external to the logic device as defined by the behavioral source code and the constraints. The constraints identify characteristics of at least one of: the internal data movements and the external data movements. The method also includes automatically designing one or more data movers for use within the at least one logic device, where the one or more data movers are configured to perform at least one of the internal and external data movements in accordance with the characteristics.

Claims (64)

1. A method comprising:

obtaining behavioral source code defining logic to be performed using at least one logic device and constraints identifying permitted characteristics of data movements associated with execution of the logic, the at least one logic device containing multiple components that support at least one of: internal data movements within the at least one logic device and external data movements external to the at least one logic device as defined by the behavioral source code and the constraints, the constraints identifying the permitted characteristics of at least one of: the internal data movements and the external data movements, wherein the logic of the behavioral source code comprises a plurality of logic elements each of which is movable to and executable by the at least one logic device, wherein the logic elements comprise instructions and data;

using the constraints to identify data movement logic to be inserted into at least one of one or more data movers for execution by the at least one of the one or more data movers and to identify one or more interfaces of the at least one logic device to be used by the one or more data movers; and

using the constraints and the behavioral source code to automatically design the one or more data movers for use by a run-time scheduler within the at least one logic device, the one or more data movers configured to perform at least one of the internal and external data movements in accordance with the permitted characteristics;

wherein the one or more data movers are optimized to provide a needed latency of the logic of the behavioral source code when performed by the at least one logic device.

2. The method of claim 1 , wherein the one or more data movers comprise at least one of:

one or more remote direct memory access (RDMA) controllers in the at least one logic device, each RDMA controller associated with an internal or external interface of the at least one logic device;

one or more engines or cores in the at least one logic device; and

one or more buffers in the at least one logic device, each buffer configured to temporarily store information transported between one of the one or more engines or cores and one of the one or more RDMA controllers.

3. The method of claim 2 , wherein the one or more data movers further comprise one or more data transformations in the at least one logic device.

4. The method of claim 2 , wherein the one or more RDMA controllers include a memory control function and a sequence random access memory (RAM).

5. The method of claim 1 , wherein the one or more interfaces comprise at least one of:

an external memory interface;

a peripheral component interconnect express (PCI-e) interface;

an Ethernet interface; and

an interface to another logic device.

6. The method of claim 1 , wherein using the constraints and the behavioral source code to automatically design the one or more data movers comprises:

identifying at least one of: a remote direct memory access (RDMA) controller for each of the internal and external data movements and at least one buffer associated with at least one of the RDMA controllers;

identifying flow control, synchronization, or data re-ordering logic for at least one of the one or more data movers; and

identifying source and destination connections to or from the RDMA controllers.

7. The method of claim 1 , wherein the at least one logic device comprises at least one of: a field programmable gate array (FPGA), an adaptive compute accelerator platform (ACAP), an application-specific integrated circuit (ASIC), a very-large-scale integration (VSLI) chip, a memory chip, a data converter, a central processing unit (CPU), and an accelerator chip.

8. An apparatus comprising:

at least one processor configured to:

obtain behavioral source code defining logic to be performed using at least one logic device and constraints identifying permitted characteristics of data movements associated with execution of the logic, the at least one logic device containing multiple components that support at least one of: internal data movements within the at least one logic device and external data movements external to the at least one logic device as defined by the behavioral source code and the constraints, the constraints identifying the permitted characteristics of at least one of: the internal data movements and the external data movements, wherein the logic of the behavioral source code comprises a plurality of logic elements each of which is movable to and executable by the at least one logic device, wherein the logic elements comprise instructions and data;

use the constraints to identify data movement logic to be inserted into at least one of one or more data movers for execution by the at least one of the one or more data movers and to identify one or more interfaces of the at least one logic device to be used by the one or more data movers;

use the constraints and the behavioral source code to automatically generate a design of the one or more data movers for use by a run-time scheduler within the at least one logic device, the one or more data movers configured to perform at least one of the internal and external data movements in accordance with the permitted characteristics; and

configure the at least one logic device based on the design;

wherein the one or more data movers are optimized to provide a needed latency of the logic of the behavioral source code when performed by the at least one logic device.

9. The apparatus of claim 8 , wherein the design of the one or more data movers comprises a design of at least one of:

one or more remote direct memory access (RDMA) controllers in the at least one logic device, each RDMA controller associated with an internal or external interface of the at least one logic device;

one or more engines or cores in the at least one logic device; and

one or more buffers in the at least one logic device, each buffer configured to temporarily store information transported between one of the one or more engines or cores and one of the one or more RDMA controllers.

10. The apparatus of claim 9 , wherein the design of the one or more data movers further comprises a design of one or more data transformations in the at least one logic device.

11. The apparatus of claim 9 , wherein the one or more RDMA controllers include a memory control function and a sequence random access memory (RAM).

12. The apparatus of claim 8 , wherein the one or more interfaces comprise at least one of:

an external memory interface;

a peripheral component interconnect express (PCI-e) interface;

an Ethernet interface; and

an interface to another logic device.

13. The apparatus of claim 8 , wherein, to use the constraints and the behavioral source code to automatically generate the design, the at least one processor is configured to:

identify at least one of: a remote direct memory access (RDMA) controller for each of the internal and external data movements and at least one buffer associated with at least one of the RDMA controllers;

identify flow control, synchronization, or data re-ordering logic for at least one of the one or more data movers; and

identify source and destination connections to or from the RDMA controllers.

14. The apparatus of claim 8 , wherein the at least one logic device comprises at least one of: a field programmable gate array (FPGA), an adaptive compute accelerator platform (ACAP), an application-specific integrated circuit (ASIC), a very-large-scale integration (VSLI) chip, a memory chip, a data converter, a central processing unit (CPU), and an accelerator chip.

15. A non-transitory computer readable medium containing instructions that when executed cause at least one processor to:

obtain behavioral source code defining logic to be performed using at least one logic device and constraints identifying permitted characteristics of data movements associated with execution of the logic, the at least one logic device containing multiple components that support at least one of: internal data movements within the at least one logic device and external data movements external to the at least one logic device as defined by the behavioral source code and the constraints, the constraints identifying the permitted characteristics of at least one of: the internal data movements and the external data movements, wherein the logic of the behavioral source code comprises a plurality of logic elements each of which is movable to and executable by the at least one logic device, wherein the logic elements comprise instructions and data;

use the constraints to identify data movement logic to be inserted into at least one of one or more data movers for execution by the at least one of the one or more data movers and to identify one or more interfaces of the at least one logic device to be used by the one or more data movers; and

use the constraints and the behavioral source code to automatically generate a design of the one or more data movers for use by a run-time scheduler within the at least one logic device, the one or more data movers configured to perform at least one of the internal and external data movements in accordance with the permitted characteristics;

wherein the one or more data movers are optimized to provide a needed latency of the logic of the behavioral source code when performed by the at least one logic device.

16. The non-transitory computer readable medium of claim 15 , wherein the design of the one or more data movers comprises a design of at least one of:

one or more remote direct memory access (RDMA) controllers in the at least one logic device, each RDMA controller associated with an internal or external interface of the at least one logic device;

one or more engines or cores in the at least one logic device; and

one or more buffers in the at least one logic device, each buffer configured to temporarily store information transported between one of the one or more engines or cores and one of the one or more RDMA controllers.

17. The non-transitory computer readable medium of claim 16 , wherein the design of the one or more data movers further comprises a design of one or more data transformations in the at least one logic device.

18. The non-transitory computer readable medium of claim 16 , wherein the one or more RDMA controllers include a memory control function and a sequence random access memory (RAM).

19. The non-transitory computer readable medium of claim 15 , wherein the one or more interfaces comprise at least one of:

an external memory interface;

a peripheral component interconnect express (PCI-e) interface;

an Ethernet interface; and

an interface to another logic device.

20. The non-transitory computer readable medium of claim 15 , wherein the instructions that cause the at least one processor to use the constraints and the behavioral source code to automatically generate the design comprise instructions cause the at least one processor to:

identify at least one of: a remote direct memory access (RDMA) controller for each of the internal and external data movements and at least one buffer associated with at least one of the RDMA controllers;

identify flow control, synchronization, or data re-ordering logic for at least one of the one or more data movers; and

identify source and destination connections to or from the RDMA controllers.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2021
From: REID, STEPHEN R.
To: RAYTHEON COMPANY
Reel/Frame 056724/0756 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2021
From: DUTTA, SANDEEP
To: SYSTEM VIEW, INC.
Reel/Frame 056724/0822 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2021
From: SYSTEM VIEW, INC.
To: RAYTHEON COMPANY
Reel/Frame 056724/0952 →
Continuity (4)
Provisional Application 63117979 · Nov 24, 2020
Provisional Application 63117988 · Nov 24, 2020
Provisional Application 63117998 · Nov 24, 2020
Related Publication 20240028533A1 · Jan 25, 2024
References Cited (67)
US 5949799A · Grivna et al. · 1999 [cited by applicant]
US 6625797B1 · Edwards et al. · 2003 [cited by applicant]
US 7073158B2 · McCubbrey · 2006 [cited by applicant]
US 7571216B1 · McRae et al. · 2009 [cited by applicant]
US 8869121B2 · Vorbach et al. · 2014 [cited by applicant]
US 8972958B1 · Brewer · 2015 [cited by examiner]
US 9098661B1 · Biswas · 2015 [cited by examiner]
US 9652570B1 · Kathail et al. · 2017 [cited by applicant]
US 10084725B2 · Raponi et al. · 2018 [cited by applicant]
US 10949585B1 · Winefeld · 2021 [cited by applicant]
US 11048758B1 · Kim et al. · 2021 [cited by applicant]
US 20010016933A1 · Chang et al. · 2001 [cited by applicant]
US 20020156929A1 · Hekmatpour · 2002 [cited by applicant]
US 20140096119A1 · Vasudevan et al. · 2014 [cited by applicant]
US 20150058832A1 · Gonion · 2015 [cited by applicant]
US 20170262567A1 · Vassiliev · 2017 [cited by applicant]
US 20180268096A1 · Chuang et al. · 2018 [cited by applicant]
US 20190042222A1 · Rong · 2019 [cited by applicant]
US 20210056368A1 · Nudejima · 2021 [cited by examiner]
US 20210081347A1 · Liao · 2021 [cited by examiner]
WO 2020112999A1 · 2020 [cited by applicant]
Non-Final Office Action dated Jul. 11, 2022 in connection with U.S. Appl. No. 17/364,565, 11 pages. [cited by applicant]
“AXI DataMover v5.1—LogiCORE IP Product Guide”, Xilinx, Inc., Vivado Design Suite, Apr. 2017, 59 pages. [cited by applicant]
Patel, “FPGA designs with VHDL”, PythonDSP, Oct. 2017, 229 pages. [cited by applicant]
Garrault et al., “HDL Coding Practices to Accelerate Design Performance”, Xilinx, Inc., White Paper: Virtex-4, Spartan-3/3L, and Spartan-3E FPGAs, Jan. 2006, 22 pages. [cited by applicant]
Shi, “Rapid Prototyping of an FPGA-Based Video Processing System”, Thesis, Virginia Polytechnic Institute and State University, Apr. 2016, 71 pages. [cited by applicant]
“FIFO Generator v13.1—LogiCORE IP Product Guide”, Xilinx, Inc., Vivado Design Suite, Apr. 2017, 218 pages. [cited by applicant]
Russell, “Fine Tune Your Embedded FPGA System with Emulation Tools—DornerWorks”, Dec. 2017, 10 pages. [cited by applicant]
Pagani, “Software support for dynamic partial reconfigurable FPGAs on heterogeneous platforms”, University of Pisa, School of Engineering, 2015/2016, 98 pages. [cited by applicant]
Erusalagandi, “Leveraging Data-Mover IPs for Data Movement in Zynq-7000 AP SoC Systems”, Xilinx, Inc., White Paper: Zynq-7000 AP SoC, Jan. 2015, 27 pages. [cited by applicant]
“A Guide to Vectorization with Intel® C++ Compilers”, Intel Corp., 2010, 39 pages. [cited by applicant]
“Migrating Motor Controller C++ Software from a Microcontroller to a PolarFire FPGA with LegUp High-Level Synthesis—LegUp Computing Blog”, Microchip Technology Inc., 2015, 16 pages. [cited by applicant]
Nabi, “Research Article—Automatic Pipelining and Vectorization of Scientific Code for FPGAs”, Hindawi, International Journal of Reconfigurable Computing, 2019, 13 pages. [cited by applicant]
“How to Cross-Compile Clang/LLVM using Clang/LLVM”, The LLVM Compiler Infrastructure, Documentation—User Guides, Mar. 2019, 9 pages. [cited by applicant]
“Vitis Model Composer User Guide”, Xilinx, Inc., UG1483 (v2021.1), Jun. 2021, 938 pages. [cited by applicant]
“Vivado Design Suite User Guide—High-Level Synthesis”, Xilinx, Inc., UG902 (v2019.2), Jan. 2020, 589 pages. [cited by applicant]
Liang et al., “Vectorization and Parallelization of Loops in C/C++ Code”, International Conference Frontiers in Education: CS and CE, 2017, 4 pages. [cited by applicant]
Wang et al., “DeepBurning: Automatic Generation of FPGA-based Learning Accelerators for the Neural Network Family”, DAC 16, Jun. 2016, 6 pages. [cited by applicant]
Jiang et al., “Hardware/Software Co-Exploration of Neural Architectures”, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Jan. 2020, 10 pages. [cited by applicant]
“AdvMag Optimization”, RIST FX10 and Nagoya University FX100, Jul. 2018, 12 pages. [cited by applicant]
Crain, “Cray Fortran 90 Optimization”, CUG 1996 Fall Proceedings, Cray Research, A Silicon Graphics Company, 1996, 2 pages. [cited by applicant]
Lantz, “Vector Parallelism on Multi-Core Processors”, Cornell University, Jul. 2019, 83 pages. [cited by applicant]
Ren et al., “Exploiting Vector and Multicore Parallelism for Recursive, Data- and Task-Parallel Programs”, Association for Computing Machinery, Feb. 2017, 14 pages. [cited by applicant]
Yanez et al., “Simultaneous multiprocessing in a software-defined heterogeneous FPGA”, Journal of Supercomputing, Apr. 2018, 18 pages. [cited by applicant]
Niu et al., “Reconfiguring Distributed Applications in FPGA Accelerated Cluster With Wireless Networking”, 2011 International Conference on Field Programmable Logic and Applications (FPL), Oct. 2011, 6 pages. [cited by applicant]
Jones et al., “Dynamically Optimizing FPGA Applications by Monitoring Temperature and Workloads”, VLSID 07: Proceedings of the 20th International Conference on VLSI Design, Jan. 2007, 8 pages. [cited by applicant]
Iturbe et al., “Research Article—Runtime Scheduling, Allocation, and Execution of Real-Time Hardware Tasks onto Xilinx FPGAs Subject to Fault Occurrence”, International Journal of Reconfigurable Computing, 2013, 33 page… [cited by applicant]
Ramezani, “A Prefetch-aware Scheduling for FPGA-based Multi-Task Graph Systems”, Journal of Supercomputing, Jan. 2020, 12 pages. [cited by applicant]
Jing et al., Abstract of “Energy-efficient scheduling on multi-FPGA reconfigurable systems”, Microprocessors and Microsystems, Aug.-Oct. 2013, 4 pages. [cited by applicant]
Perng et al., “Energy-Efficient Scheduling on Multi-Context FPGA's”, 2006 IEEE International Symposium on Circuits and Systems, May 2006, 4 pages. [cited by applicant]
Chatarasi et al., “Vyasa: A High-Performance Vectorizing Compiler for Tensor Convolutions on the Xilinx AI Engine”, 2020 IEEE High Performance Extreme Computing Conference (HPEC), Sep. 2020, 12 pages. [cited by applicant]
Non-Final Office Action dated Dec. 19, 2023 in connection with U.S. Appl. No. 17/364,565, 17 pages. [cited by applicant]
Final Office Action dated Oct. 4, 2023 in connection with U.S. Appl. No. 17/364,565, 12 pages. [cited by applicant]
Non-Final Office Action dated Apr. 26, 2023 in connection with U.S. Appl. No. 17/364,565, 10 pages. [cited by applicant]
Hu et al., “Semi-automatic Hardware Design using Ontologies,” ICARCV 2004, 8th Control, Automation, Robotics and Vision Conference, 2004, 6 pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority in connection with International Patent Application No. PCT/US2021/059009 issued Feb. 4, 2022, 16 pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority in connection with International Patent Application No. PCT/US2021 /059018 issued Feb. 10, 2022, 15 pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority in connection with International Patent Application No. PCT/US2021/059013 issued Feb. 11, 2022, 13 pages. [cited by applicant]
Sharma et al., “Run-Time Mitigation of Power Budget Variations and Hardware Faults by Structural Adaptation of FPGA-Based Multi-Modal SoPC”, Computers 2018, vol. 7, No. 4, Oct. 2018, 34 pages. [cited by applicant]
Eckert et al., “Operating System Concepts for Reconfigurable Computing: Review and Survey”, International Journal of Reconfigurable Computing, vol. 2016, Nov. 2016, 12 pages. [cited by applicant]
Sousa et al., “Runtime Adaptation of Application Execution under Thermal and Power Constraints in Massively Parallel Processor Arrays”, Proceedings of the 18th International Workshop on Software and Compilers for Embedd… [cited by applicant]
Pellizzoni et al., “Real-Time Management of Hardware and Software Tasks for FPGA-Based Embedded Systems”, IEEE Transactions on Computers, vol. 56, No. 12, Dec. 2007, 15 pages. [cited by applicant]
Kida et al., “A High Level Synthesis Approach for Application Specific DMA Controllers”, 2019 International Conference on Reconfigurable Computing and FPGAs (Reconfig), IEEE, Dec. 2019, 2 pages. [cited by applicant]
Lo et al., “Model-Based Optimization of High Level Synthesis Directives”, 26th International Conference on Field Programmable Logic and Applications (FPL), EPFL, Aug. 2016, 10 pages. [cited by applicant]
Ferretti et al., “Lattice-Traversing Design Space Exploration for High Level Synthesis”, IEEE 36th International Conference on Computer Design (ICCD), Oct. 2018, 8 pages. [cited by applicant]
Singh et al., “Parallelizing High-Level Synthesis: A Code Transformational Approach to High-Level Synthesis”, EDA for IC System Design, Verification, and Testing, Mar. 2006, 20 pages. [cited by applicant]
Office Action dated Nov. 28, 2024 in connection with European Patent Application No. 21823421.9, 14 pages. [cited by applicant]