IP Library Granted Patent US 12,613,830
Granted Patent B2
US 12,613,830 · App. 18/368,492 · Granted Apr 28, 2026

Accelerator architecture on a programmable platform

Inventors: David Shippy (Austin, TX); Martin Langhammer (Alderbury, GB); Jeffrey Eastlack (Driftwood, TX)
Assignee: Altera Corporation
G06F15/8023G06F9/3877G06F9/3887G06F9/3888G06F15/7825G06F9/30036G06F13/124G06F13/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,613,830
App. No.
18/368,492
Granted
Apr 28, 2026
Kind
B2
Abstract

An accelerated processor structure on a programmable integrated circuit device includes a processor and a plurality of configurable digital signal processors (DSPs). Each configurable DSP includes a circuit block, which in turn includes a plurality of multipliers. The accelerated processor structure further includes a first bus to transfer data from the processor to the configurable DSPs, and a second bus to transfer data from the configurable DSPs to the processor.

Claims (38)

1 . An integrated circuit comprising:

an accelerator comprising processing nodes arranged in a grid, wherein each of the processing nodes comprises a vector processor that supports a vector data format, wherein the vector processor in each of the processing nodes comprises processing elements and a command processor that performs fetch and dispatch of first instructions for the processing elements, wherein each of the processing elements in the vector processor in each of the processing nodes comprises a multiplier, and wherein each of the processing nodes further comprises local memory;

memory blocks;

a dedicated control subsystem to execute software instructions and to control the processing nodes; and

a mesh network that interconnects the processing nodes and the memory blocks.

2 . The integrated circuit of claim 1 , wherein the mesh network allows for data movement between the processing nodes using mesh routers in the processing nodes.

3 . The integrated circuit of claim 1 , wherein each of the processing elements further comprises a local register file, an arithmetic logic unit, and a load/store unit that interface with a data rotate unit and data memory unit.

4 . The integrated circuit of claim 3 , wherein the accelerator supports data-level parallelism through Single Instruction, Multiple Data (SIMD) vector execution and thread level parallelism.

5 . The integrated circuit of claim 1 , wherein the mesh network comprises a mesh router that connects one of the processing nodes to an external memory via a cache coherency unit such that data in cache units connected to different ones of the processing nodes are consistent.

6 . The integrated circuit of claim 1 , wherein the command processor further comprises an instruction memory that stores instructions for execution by the command processor, a decoder, a scalar register file, an execution module to execute instructions in the instruction memory, and a load/store unit to load or store instructions from data stores in the accelerator.

7 . The integrated circuit of claim 1 , wherein the mesh network comprises a network-on-chip that accesses software instructions from a second memory.

8 . The integrated circuit of claim 1 , wherein the processing nodes provide floating point precision.

9 . The integrated circuit of claim 1 , wherein the command processor in each of the processing nodes coordinates execution on functional units comprising an instruction memory that stores instructions for execution by the command processor, a register file, and an execution module to execute instructions in the instruction memory.

10 . The integrated circuit of claim 1 , wherein the local memory in each of the processing nodes comprises a data rotate unit.

11 . The integrated circuit of claim 1 , wherein each of the processing nodes further comprises a mesh router that routes a connection to other ones of the processing nodes arranged in the grid.

12 . The integrated circuit of claim 1 , wherein the processing nodes in the accelerator are arranged in at least 4 columns and at least 4 rows in the grid.

13 . The integrated circuit of claim 1 , wherein execution of processor instructions is pipelined by a series of pipeline registers, and wherein the processor instructions are processed at the processing elements.

14 . A method for operating an integrated circuit:

receiving, by an accelerator in the integrated circuit that comprises processing nodes arranged in a grid, a first set of software instructions using a mesh network that interconnects the processing nodes;

executing a second set of software instructions using a dedicated control subsystem that controls the processing nodes, wherein the dedicated control subsystem is in the integrated circuit, wherein each of the processing nodes comprises local memory, and wherein each of the processing nodes further comprises a vector processor that supports a vector data format; and

coordinating execution of the first set of software instructions for processing elements using a control processor, wherein the vector processor in each of the processing nodes comprises the processing elements and the control processor, and wherein each of the processing elements in the vector processor in each of the processing nodes comprises a multiplier.

15 . The method of claim 14 further comprising supporting thread level parallelism using the accelerator.

16 . The method of claim 14 further comprising allowing data movement between the processing nodes using mesh routers in the processing nodes in the mesh network.

17 . The method of claim 14 , wherein the processing nodes provide floating point operations.

18 . The method of claim 14 further comprising supporting data-level parallelism through Single Instruction, Multiple Data (SIMD) vector execution using the accelerator.

19 . The method of claim 14 , wherein each of the processing nodes further comprises a data rotate unit.

20 . The method of claim 14 ,

wherein coordinating execution of the first set of software instructions for the processing elements further comprises performing fetch, decode, and dispatch of the first set of software instructions for the processing elements using the control processor.

21 . The method of claim 14 , wherein further comprising processing the first set of software instructions at the processing elements.

22 . An integrated circuit comprising:

an accelerator comprising processing nodes arranged in a grid, wherein each of the processing nodes comprises a vector processor, wherein each of the processing nodes further comprises local memory, wherein each of the processing nodes further comprises a processing element and a control processor that coordinates execution on functional units, and wherein the control processor performs fetch and dispatch of first instructions for the processing element, and wherein the processing element in each of the processing nodes comprises a multiplier;

a network-on-a-chip that accesses software instructions from a second memory;

a dedicated control subsystem that executes the software instructions and that controls the processing nodes; and

a mesh network that interconnects the processing nodes.

23 . The integrated circuit of claim 22 wherein the control processor further comprises an instruction memory that stores instructions for execution by the control processor, a decoder, a scalar register file, an execution module to execute instructions in the instruction memory, and a load/store unit to load or store instructions from data stores in the accelerator.

24 . The integrated circuit of claim 22 , wherein the processing nodes provide floating point precision.

25 . The integrated circuit of claim 22 , wherein the vector processor in each of the processing nodes comprises the control processor and processing elements that comprise the processing element.

26 . The integrated circuit of claim 22 , wherein the processing element in each of the processing nodes further comprises a local register file, an arithmetic logic unit, and a load/store unit that interface with a data rotate unit and data memory unit.

Assignments (1)
SECURITY INTEREST Recorded Sep 12, 2025
From: ALTERA CORPORATION
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 073431/0309 →
Continuity (4)
Continuation 16154517 · Oct 8, 2018
Continuation In Part 14725811 · May 29, 2015
Provisional Application 62004691 · May 29, 2014
Related Publication 20240078211A1 · Mar 7, 2024
References Cited (34)
US 5408677A · Nogi · 1995 [cited by applicant]
US 5752071A · Tubbs et al. · 1998 [cited by applicant]
US 7386704B2 · Schulz et al. · 2008 [cited by applicant]
US 7856545B2 · Casselman · 2010 [cited by applicant]
US 8495122B2 · Simkins et al. · 2013 [cited by applicant]
US 8612725B2 · Tanabe · 2013 [cited by examiner]
US 20020103841A1 · Parviainen · 2002 [cited by applicant]
US 20030055861A1 · Lai et al. · 2003 [cited by applicant]
US 20030088757A1 · Lindner et al. · 2003 [cited by applicant]
US 20040008201A1 · Lewis · 2004 [cited by examiner]
US 20050114565A1 · Gonzalez · 2005 [cited by examiner]
US 20050144210A1 · Simkins · 2005 [cited by examiner]
US 20050171990A1 · Bishop · 2005 [cited by examiner]
US 20050198472A1 · Sih · 2005 [cited by applicant]
US 20050216700A1 · Honary · 2005 [cited by examiner]
US 20070250681A1 · Horvath · 2007 [cited by examiner]
US 20100064115A1 · Hoshi · 2010 [cited by examiner]
US 20120311302A1 · Yang · 2012 [cited by examiner]
US 20130080739A1 · Kyo · 2013 [cited by examiner]
US 20160006471A1 · Pande · 2016 [cited by examiner]
CN 101018055A · 2007 [cited by applicant]
CN 101400178A · 2009 [cited by applicant]
Gustafson, J.L., Hawkinson, S., Scott, K, “The architecture of a homogeneous vector supercomputer”, Springer-Verlag, pp. 62-70 (Year: 1986). [cited by examiner]
Tanqueray, D.A, “The Floating Point Systems T Series”, Springer-Verlag Berlin Heidelberg, pp. 307-317 (Year: 1988). [cited by examiner]
Rodney J. Fazzari and John D. Lynch, “The Second Generation Fps T Series: An Enhanced Parallel Vector Supercomputer”, Jan. 1, pp. 61-70 (Year: 1988). [cited by examiner]
Chinese Office Action for CN Application No. 201580037571.5 mailed Nov. 5, 2018; 7 Pages. [cited by applicant]
EP Office Action for European Application No. 19166751.8 Mailed on Feb. 26, 2021 7 Pages. [cited by applicant]
Wu et al. “A Programming model and a NoC-based architecture for streaming application,” 2010 IEEE, 13th Euromicro Conference on Digital System Design: Architectures, Methods and Tools, Grenoble, France, 5 Pages. [cited by applicant]
CN Office Action for Chinese Application No. 201910218813.0 Mailed on Jan. 4, 2023. [cited by applicant]
EP Office Action for European Application No. 15800257.6-1203 Mailed on Aug. 2, 2023; 7 Pages. [cited by applicant]
Wenhua Fan et al: “Efficient Implementation of OFDM Inner Receiver on a Programmable Multi-Core Processor Platform”, IEICE Transaction on Communication, Communications Society, Tokyo, JP, vol. E95B, No. 4, Apr. 1, 2012 … [cited by applicant]
Gerard J.M. Smit, et al., “Multi-core architectures and streaming applications,” SLIP '08: Proceedings of the 2008 International workshop on System level interconnect prediction, Apr. 2008, pp. 35-42, https://doi.org/10… [cited by applicant]
EP Office Action for European Application No. 19166751.8 Mailed on Oct. 10, 2023 9 Pages. [cited by applicant]
Lertora F et al: “Handling Different Computational Granularity by a Reconfigurable IC Featuring Embedded FPGAs and a Network-on-Chip”, Field-Programmable Custom Computing Machines, 2005. FCCM 2005. 183th an Nual IEEE Sy… [cited by applicant]