IP Library › Patent Application 19461086
Patent Application
App. No. 19/461,086

ULTRA-LOW POWER COARSE-GRAINED RECONFIGURABLE ARRAYS WITH PROGRAMMABLE ON-CHIP NETWORK

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/461,086
Abstract

Disclosed herein is a co-designed compiler and CGRA fabric that achieves both high programmability and extreme energy efficiency. The fabric implements a rich set of control-flow operators that support arbitrary control flow and memory access on the fabric. The fabric is able to realize both energy and area savings over prior art implementations by offloading most control operations into a programmable on-chip network that can re-use existing network switches.

Claims (25)

1 . A ultra-low power coarse-grained reconfigurable array fabric comprising:

an array of processing elements;

a scalar core;

a bank of SRAM memory; and

a bufferless, 2D-torus on-chip network connecting the array of processing elements.

2 . The fabric of claim 1 wherein each processing element comprises:

a functional unit; and

a μcore.

3 . The fabric of claim 2 wherein the μcore interfaces the processing unit with the on-chip network.

4 . The fabric of claim 2 wherein the μcore buffers output values of the processing unit.

5 . The fabric of claim 2 wherein the μcore interfaces with top-level fabric control for configuration of the processing element.

6 . The fabric of claim 3 wherein the μcore implements a generic interface using a latency-insensitive ready/valid protocol.

7 . The fabric of claim 2 wherein the μcore handles communication between the on-chip network and the processing unit.

8 . The fabric of claim 1 wherein the array of processing elements is a 6×6 array.

9 . The fabric of claim of claim 1 wherein the processing elements are selected from a group consisting of memory processing elements, arithmetic processing elements, multiplier processing elements, control-flow processing elements and stream processing elements.

10 . The fabric of claim 1 wherein processing elements buffer values in an output channel to reduce buffering values in the on-chip network.

11 . The fabric of claim 1 wherein the fabric implements an instruction set including arithmetic operators, multiplier operators, memory operators, control flow operators, synchronization operators and stream operators.

12 . The fabric of claim 11 wherein the control flow operators include a steer operator, a carry operator and an invariant operator.

13 . The fabric of claim 11 wherein the synchronization operators include a merge operator and an order operator.

14 . The fabric of claim 11 wherein the stream operator generates a sequence of data values in accordance with programming of the fabric.

15 . The fabric of claim 11 wherein a compiler generates scalar code for execution by the scalar core and a bitstream to configure the fabric.

16 . The fabric of claim 11 wherein the processing units implement the arithmetic and memory operators.

17 . The fabric of claim 11 wherein the on-chip network comprises a plurality of routers coupled to the processing elements.

18 . The fabric of claim 17 wherein the on-chip network implements the control flow operators.

19 . The fabric of claim 18 wherein the routers include one or more control-flow modules at one or more output ports to implement the control-flow operators.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2026
From: LUCIA, BRANDON; BECKMANN, NATHAN; GOBIESKI, GRAHAM
To: CARNEGIE MELLON UNIVERSITY
Reel/Frame 073729/0465 →