IP Library › Granted Patent US 12,541,349
Granted Patent B2
US 12,541,349 · App. 18/553,670 · Granted Feb 3, 2026

Energy-minimal dataflow architecture with programmable on-chip network

Inventors: Brandon Lucia (Pittsburgh, PA); Nathan Beckmann (Pittsburgh, PA); Graham Gobieski (Pittsburgh, PA)
Assignee: CARNEGIE MELLON UNIVERSITY
G06F8/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,349
App. No.
18/553,670
Granted
Feb 3, 2026
Kind
B2
Abstract

Disclosed herein is a co-designed compiler and CGRA architecture that achieves both high programmability and extreme energy efficiency. The architecture includes a rich set of control-flow operators that support arbitrary control flow and memory access on the CGRA fabric. The architecture is able to realize both energy and area savings over prior art implementations by offloading most control operations into a programmable on-chip network where they can re-use existing network switches.

Claims (36)

1 . A system comprising:

a processor; and

compiler software that, when executed by the processor performs the functions of:

compiling a program written in a high-level language to an intermediate representation;

enforcing memory ordering on the intermediate representation to create an ordering graph encoding dependencies between memory operations as ordering arcs;

iteratively pruning the ordering graph to eliminate redundant ordering arcs already enforced by data and control dependencies;

translating the intermediate representation to a dataflow graph;

inserting control-flow operators into the dataflow graph; and

mapping operators in the dataflow graph to a coarse-grained reconfigurable array (CGRA) wherein the control-flow operators utilize logic of an on-chip network of the CGRA to control flow of the program.

2 . The system of claim 1 wherein the dataflow graph is optimized before the operators in the dataflow graph are mapped.

3 . The system of claim 1 wherein path-sensitive transductive reduction is applied to the ordering graph to convert a potentially cyclic ordering graph to an acyclic graph off strongly-connected components.

4 . The system of claim 2 wherein the control-flow operators comprise:

a steer operator;

a carry operator;

an invariant operator;

a merge operator;

an order operator; and

a stream operator.

5 . The system of claim 4 wherein the memory ordering is enforced by inserting order operators in the ordering graph.

6 . The system of claim 4 wherein the steer, carry, invariant and merge operators are inserted into the dataflow graph.

7 . The system, of claim 1 wherein the CGRA comprises:

a fabric of heterogeneous processing elements connected via a bufferless, 2D-torus on-chip network;

a RISC-V scalar core; and

SRAM main memory.

8 . The system of claim 7 wherein the compiler uses the dataflow graph and a topographical description of the CGRA to generate scalar code for execution by the scalar core and a bitstream to configure the fabric.

9 . The system of claim 8 where the processing elements perform arithmetic and memory operations.

10 . The system of claim 7 wherein each processing element comprises:

a functional unit; and

a μcore;

wherein the μcore interfaces with the on-chip network, buffers output values and interfaces with top-level fabric control for configuration of the processing element.

11 . The system of claim 7 wherein the fabric comprises a 6×6 array of processing elements.

12 . The system of claim 7 wherein et processing elements are selected from a group consisting of memory processing elements, arithmetic processing elements, multiplier processing elements, control-flow processing elements and stream processing elements.

13 . The system of claim 7 wherein processing elements buffer values in an output channel to reduce buffering values in the on-chip network.

14 . The system of claim 7 wherein the on-chip network comprises a plurality of routers coupled to the processing elements.

15 . The system of claim 14 wherein the routers include one or more control-flow modules at one of more output ports to implement the control-flow operators.

16 . The system of claim 2 wherein the high level language is C and further wherein the intermediate representation is LLVM produced by clang.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2026
From: LUCIA, BRANDON; BECKMANN, NATHAN; GOBIESKI, GRAHAM
To: CARNEGIE MELLON UNIVERSITY
Reel/Frame 073481/0900 →
Continuity (2)
Provisional Application 63403422 · Sep 2, 2022
Related Publication 20250190189A1 · Jun 12, 2025
References Cited (6)
US 10469397B2 · Fleming · 2019 [cited by examiner]
US 10515049B1 · Fleming · 2019 [cited by applicant]
US 11029958B1 · Zhang · 2021 [cited by examiner]
US 20210373867A1 · Chen · 2021 [cited by applicant]
US 20220164189A1 · Balasubramanian · 2022 [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US23/31348, mailing date Jan. 8, 2024, 16 pages. [cited by applicant]