Accelerator architecture on a programmable platform
An accelerated processor structure on a programmable integrated circuit device includes a processor and a plurality of configurable digital signal processors (DSPs). Each configurable DSP includes a circuit block, which in turn includes a plurality of multipliers. The accelerated processor structure further includes a first bus to transfer data from the processor to the configurable DSPs, and a second bus to transfer data from the configurable DSPs to the processor.
1 . An integrated circuit comprising:
an accelerator comprising processing nodes arranged in a grid, wherein each of the processing nodes comprises a vector processor that supports a vector data format, wherein the vector processor in each of the processing nodes comprises processing elements and a command processor that performs fetch and dispatch of first instructions for the processing elements, wherein each of the processing elements in the vector processor in each of the processing nodes comprises a multiplier, and wherein each of the processing nodes further comprises local memory;
memory blocks;
a dedicated control subsystem to execute software instructions and to control the processing nodes; and
a mesh network that interconnects the processing nodes and the memory blocks.
2 . The integrated circuit of claim 1 , wherein the mesh network allows for data movement between the processing nodes using mesh routers in the processing nodes.
3 . The integrated circuit of claim 1 , wherein each of the processing elements further comprises a local register file, an arithmetic logic unit, and a load/store unit that interface with a data rotate unit and data memory unit.
4 . The integrated circuit of claim 3 , wherein the accelerator supports data-level parallelism through Single Instruction, Multiple Data (SIMD) vector execution and thread level parallelism.
5 . The integrated circuit of claim 1 , wherein the mesh network comprises a mesh router that connects one of the processing nodes to an external memory via a cache coherency unit such that data in cache units connected to different ones of the processing nodes are consistent.
6 . The integrated circuit of claim 1 , wherein the command processor further comprises an instruction memory that stores instructions for execution by the command processor, a decoder, a scalar register file, an execution module to execute instructions in the instruction memory, and a load/store unit to load or store instructions from data stores in the accelerator.
7 . The integrated circuit of claim 1 , wherein the mesh network comprises a network-on-chip that accesses software instructions from a second memory.
8 . The integrated circuit of claim 1 , wherein the processing nodes provide floating point precision.
9 . The integrated circuit of claim 1 , wherein the command processor in each of the processing nodes coordinates execution on functional units comprising an instruction memory that stores instructions for execution by the command processor, a register file, and an execution module to execute instructions in the instruction memory.
10 . The integrated circuit of claim 1 , wherein the local memory in each of the processing nodes comprises a data rotate unit.
11 . The integrated circuit of claim 1 , wherein each of the processing nodes further comprises a mesh router that routes a connection to other ones of the processing nodes arranged in the grid.
12 . The integrated circuit of claim 1 , wherein the processing nodes in the accelerator are arranged in at least 4 columns and at least 4 rows in the grid.
13 . The integrated circuit of claim 1 , wherein execution of processor instructions is pipelined by a series of pipeline registers, and wherein the processor instructions are processed at the processing elements.
14 . A method for operating an integrated circuit:
receiving, by an accelerator in the integrated circuit that comprises processing nodes arranged in a grid, a first set of software instructions using a mesh network that interconnects the processing nodes;
executing a second set of software instructions using a dedicated control subsystem that controls the processing nodes, wherein the dedicated control subsystem is in the integrated circuit, wherein each of the processing nodes comprises local memory, and wherein each of the processing nodes further comprises a vector processor that supports a vector data format; and
coordinating execution of the first set of software instructions for processing elements using a control processor, wherein the vector processor in each of the processing nodes comprises the processing elements and the control processor, and wherein each of the processing elements in the vector processor in each of the processing nodes comprises a multiplier.
15 . The method of claim 14 further comprising supporting thread level parallelism using the accelerator.
16 . The method of claim 14 further comprising allowing data movement between the processing nodes using mesh routers in the processing nodes in the mesh network.
17 . The method of claim 14 , wherein the processing nodes provide floating point operations.
18 . The method of claim 14 further comprising supporting data-level parallelism through Single Instruction, Multiple Data (SIMD) vector execution using the accelerator.
19 . The method of claim 14 , wherein each of the processing nodes further comprises a data rotate unit.
20 . The method of claim 14 ,
wherein coordinating execution of the first set of software instructions for the processing elements further comprises performing fetch, decode, and dispatch of the first set of software instructions for the processing elements using the control processor.
21 . The method of claim 14 , wherein further comprising processing the first set of software instructions at the processing elements.
22 . An integrated circuit comprising:
an accelerator comprising processing nodes arranged in a grid, wherein each of the processing nodes comprises a vector processor, wherein each of the processing nodes further comprises local memory, wherein each of the processing nodes further comprises a processing element and a control processor that coordinates execution on functional units, and wherein the control processor performs fetch and dispatch of first instructions for the processing element, and wherein the processing element in each of the processing nodes comprises a multiplier;
a network-on-a-chip that accesses software instructions from a second memory;
a dedicated control subsystem that executes the software instructions and that controls the processing nodes; and
a mesh network that interconnects the processing nodes.
23 . The integrated circuit of claim 22 wherein the control processor further comprises an instruction memory that stores instructions for execution by the control processor, a decoder, a scalar register file, an execution module to execute instructions in the instruction memory, and a load/store unit to load or store instructions from data stores in the accelerator.
24 . The integrated circuit of claim 22 , wherein the processing nodes provide floating point precision.
25 . The integrated circuit of claim 22 , wherein the vector processor in each of the processing nodes comprises the control processor and processing elements that comprise the processing element.
26 . The integrated circuit of claim 22 , wherein the processing element in each of the processing nodes further comprises a local register file, an arithmetic logic unit, and a load/store unit that interface with a data rotate unit and data memory unit.