IP Library Granted Patent US 10,579,390
Granted Patent B2
US 10,579,390 · App. 15/829,924 · Granted Mar 3, 2020

Execution of data-parallel programs on coarse-grained reconfigurable architecture hardware

Inventors: Yoav Etsion (Atlit, IL); Dani Voitsechov (Amirim, IL)
Assignee: TECHNION RESEARCH & DEVELOPMENT FOUNDATION LTD.
G06F9/3869G06F9/38G06F9/3851G06F9/448G06F9/4421G06F9/4436G06F9/4494G06F9/4881
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,579,390
App. No.
15/829,924
Granted
Mar 3, 2020
Kind
B2
Abstract

A GPGPU-compatible architecture combines a coarse-grain reconfigurable fabric (CGRF) with a dynamic dataflow execution model to accelerate execution throughput of massively thread-parallel code. The CGRF distributes computation across a fabric of functional units. The compute operations are statically mapped to functional units, and an interconnect is configured to transfer values between functional units.

Claims (41)

1. A method of computing, comprising the steps of:

providing a coarse grain fabric of processing units having direct interconnects therebetween;

representing a series of computing operations to be processed in the fabric as a control data flow graph having code paths, the computing operations comprising instructions to be executed in the fabric, the instructions having input operands;

establishing a configuration of the fabric by enabling and disabling selected ones of the direct interconnects to match the processing units to the code paths of the control data flow graph; and

executing the instructions of the computing operations in the fabric in multiple simultaneous threads by transmitting input operands of the instructions of different ones of the threads successively to a common receiving set of the processing units via the direct interconnects, wherein the configuration of the fabric does not change while executing the instructions of the computing operations in the fabric,

wherein executing the instructions comprises independently triggering each of the processing units in the common receiving set to execute the instructions of the computing operations upon availability of the input operands thereof.

2. The method according to claim 1 , wherein the threads comprise instructions of the computing operations to be executed in individual processing units, and processing the computing operations comprises dynamically scheduling the instructions of at least a portion of the threads.

3. The method according to claim 2 , further comprising grouping the threads into epochs; wherein dynamically scheduling comprises deferring execution in one of the processing units of a current instruction of one of the epochs until execution in the one processing unit of all preceding instructions belonging to other epochs has completed.

4. The method according to claim 2 , further comprising the steps of:

making a determination that in the control data flow graph one of the code paths is longer than another code path; and

delaying the computing operations in processing units that are matched with the other code path.

5. The method according to claim 2 , wherein the computing operations comprise loops that each of the threads iterate.

6. The method according to claim 5 , wherein different threads perform different numbers of iterations of the loops.

7. The method according to claim 1 , further comprising:

partitioning the series of computing operations into a sequence of smaller series;

executing the instructions of the computing operations in one of the smaller series in the threads;

storing intermediate results of the computing operations; and

iterating the steps of establishing a configuration and executing the instructions of the computing operations by all threads with another of the smaller series.

8. The method according to claim 1 , wherein at least a portion of the processing units are interconnected by switches, further comprising configuring interconnections between the processing units by enabling and disabling the switches.

9. The method according to claim 8 , wherein the switches are crossbar switches.

10. A computing apparatus, comprising:

a coarse grain fabric of processing units; and

direct interconnects between the processing units, wherein the fabric is operative for:

processing a series of computing operations, the computing operations being represented as a control data flow graph having code paths, the computing operations comprising instructions to be executed in the fabric, the instructions having input operands;

establishing a configuration of the fabric by enabling and disabling selected ones of the direct interconnects to match the processing units to the code paths of the control data flow graph;

executing the instructions of the computing operations in the fabric in multiple simultaneous threads by transmitting input operands of the instructions and values relating to a plurality of the threads to a set of the processing units via the direct interconnects, wherein the configuration does not change while executing the instructions of the computing operations in the fabric,

wherein executing the instructions comprises independently triggering each of the processing units in the set to execute the instructions of the computing operations upon availability of the input operands thereof.

11. The apparatus according to claim 10 , wherein the instructions of the computing operations in the threads are executed in individual processing units, and processing the series of computing operations comprises dynamically scheduling the instructions.

12. The apparatus according to claim 11 , wherein the fabric is operative for grouping the threads into epochs; wherein dynamically scheduling comprises deferring execution in one of the processing units of a current instruction of one of the epochs until execution in the one processing unit of all preceding instructions belonging to other epochs has completed.

13. The apparatus according to claim 11 , wherein the fabric is operative for:

making a determination that in the control data flow graph one of the code paths is longer than another code path; and

delaying the computing operations in processing units that are matched with the other code path.

14. The apparatus according to claim 11 , wherein the computing operations comprise loops that each of the threads iterate.

15. The apparatus according to claim 14 , wherein different threads perform different numbers of iterations of the loops.

16. The apparatus according to claim 10 , wherein the fabric is operative for:

partitioning the series of computing operations into a sequence of smaller series;

executing the instructions of the computing operations in one of the smaller series in the plurality of the threads;

storing intermediate results of the computing operations; and

iterating the steps of routing the direct interconnects establishing a configuration and executing the instructions of the computing operations by all threads with another of the smaller series.

17. The apparatus according to claim 10 , further comprising switches that interconnect at least a portion of the processing units, and wherein the fabric is operative for configuring interconnections between the processing units by enabling and disabling the switches.

18. The apparatus according to claim 17 , wherein the switches are crossbar switches.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Jul 3, 2025
From: KREOS CAPITAL VII AGGREGATOR SCSP
To: SPEEDATA LTD
Reel/Frame 071599/0362 →
SECURITY INTEREST Recorded Jul 11, 2023
From: SPEEDATA LTD
To: KREOS CAPITAL VII AGGREGATOR SCSP
Reel/Frame 064205/0454 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2020
From: TECHNION RESEARCH & DEVELOPMENT FOUNDATION LTD.
To: SPEEDATA LTD.
Reel/Frame 054087/0889 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2017
From: ETSION, YOAV; VOITSECHOV, DANI
To: TECHNION RESEARCH & DEVELOPMENT FOUNDATION LTD.
Reel/Frame 044281/0841 →
Continuity (3)
Continuation 14642780 · Mar 10, 2015
Provisional Application 61969184 · Mar 23, 2014
Related Publication 20180101387A1 · Apr 12, 2018