IP Library Granted Patent US 12,386,602
Granted Patent B2
US 12,386,602 · App. 18/130,642 · Granted Aug 12, 2025

Operation fusion in nested meta-pipeline loops

Inventors: Fei Wang (Palo Alto, CA); Weihang Fan (Mountain View, CA); David Alan Koeplinger (Egg Harbor, NJ)
Assignee: SambaNova Systems, Inc.
G06F8/4452G06F8/433G06F8/452
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,386,602
App. No.
18/130,642
Granted
Aug 12, 2025
Kind
B2
Abstract

A method for improving throughput in a reconfigurable computing system includes detecting, in an algebraic representation of a computing task for a reconfigurable dataflow processor, an outer meta-pipeline loop, detecting an inner meta-pipeline loop nested within the outer meta-pipeline loop, and determining that the inner meta-pipeline loop and the outer meta-pipeline loop each conduct a common operation. The method also includes fusing the common operation for the inner meta-pipeline loop and the outer meta-pipeline loop into a single operation within the inner meta-pipeline loop. The instances of the common operation may be fused if the output of a first instance of the common operation is the source for a second instance of the common operation. Examples of the common operation include an accumulator operation, a re-read operation, and a temporal (chip buffer synchronized) operation such as a temporal concatenation operation and a temporal slicing operation.

Claims (39)

1. A system for improving throughput in reconfigurable dataflow processors, the system comprising:

a host computer comprising an optimization module and configured to conduct a method of transforming a high-level program of a computing task into configuration data executable by the reconfigurable dataflow processor (RDP) including one or more arrays of configurable units, the method comprising:

detecting, in an algebraic representation of the computing task for the RDP, an outer meta-pipeline loop;

detecting, in the algebraic representation, an inner meta-pipeline loop nested within the outer meta-pipeline loop;

determining that the inner meta-pipeline loop and the outer meta-pipeline loop each conduct a common operation; and

fusing the common operation for the inner meta-pipeline loop and the outer meta-pipeline loop into a single operation within the inner meta-pipeline loop;

generating the configuration data of the computing task for the RDP processor including placement and routing of configurable units, wherein the configuration data, when loaded onto an instance of the one or more arrays of configurable units of the RDP processor, causes the one or more arrays of configurable units to implement at least the computing task; and

storing the configuration data in a non-transitory computer-readable storage medium.

2. The system of claim 1 , wherein the output of a first instance of the common operation is the source for a second instance of the common operation.

3. The system of claim 1 , wherein the common operation is one of an accumulator operation, a re-read operation and a temporal operation.

4. The system of claim 1 , wherein the algebraic representation comprises a compute graph.

5. The system of claim 4 , wherein nodes in the compute graph correspond to tensor operations.

6. The system of claim 1 , wherein the algebraic representation comprises code blocks.

7. The system of claim 6 , wherein statements in the code blocks correspond to tensor operations.

8. The system of claim 1 , wherein the host computer further comprises an allocation module configured to allocate configurable units based on the configuration data.

9. The system of claim 8 , the host computer further comprises a place and route module configured to place and route the configurable units based on the configuration data.

10. A method for improving throughput in a reconfigurable computing system, the method configured to transform a high-level program of a computing task into configuration data executable by a reconfigurable dataflow processor (RDP) including one or more arrays of configurable units and comprising:

detecting, in an algebraic representation of a computing task for the RDP, an outer meta-pipeline loop;

detecting, in the algebraic representation, an inner meta-pipeline loop nested within the outer meta-pipeline loop;

determining that the inner meta-pipeline loop and the outer meta-pipeline loop each conduct a common operation; and

fusing the common operation for the inner meta-pipeline loop and the outer meta-pipeline loop into a single operation within the inner meta-pipeline loop;

generating the configuration data of the computing task for the RDP processor including placement and routing of configurable units, wherein the configuration data, when loaded onto an instance of the one or more arrays of configurable units of the RDP processor, causes the one or more arrays of configurable units to implement at least the computing task; and

storing the configuration data in a non-transitory computer-readable storage medium.

11. The method of claim 10 , wherein the output of a first instance of the common operation is the source for a second instance of the common operation.

12. The method of claim 10 , wherein the common operation is one of an accumulator operation, a re-read operation and a temporal operation.

13. The method of claim 10 , wherein the algebraic representation comprises a compute graph.

14. The method of claim 13 , wherein nodes in the compute graph correspond to tensor operations.

15. The method of claim 10 , wherein the algebraic representation comprises code blocks.

16. The method of claim 15 , wherein statements in the code blocks correspond to tensor operations.

17. The method of claim 10 , further comprising allocating, placing and routing configurable units and connections based on the configuration data.

18. The method of claim 17 , wherein the configurable units comprise one or more of memory units, compute units and switches.

19. The method of claim 10 , further comprising configuring the RDP to perform the computing task and performing the computing task.

20. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, wherein the computer readable storage medium is not a transitory signal per se, wherein the program instructions are executable by a processor to cause the processor to conduct a method of transforming a high-level program of a computing task into configuration data executable by a reconfigurable dataflow processor (RDP) including one or more arrays of configurable units, the method comprising:

detecting, in an algebraic representation of a computing task for the RDP, an outer meta-pipeline loop;

detecting, in the algebraic representation, an inner meta-pipeline loop nested within the outer meta-pipeline loop;

determining that the inner meta-pipeline loop and the outer meta-pipeline loop each conduct a common operation; and

fusing the common operation for the inner meta-pipeline loop and the outer meta-pipeline loop into a single operation within the inner meta-pipeline loop;

generating the configuration data of the computing task for the RDP processor including placement and routing of configurable units, wherein the configuration data, when loaded onto an instance of the one or more arrays of configurable units of the RDP processor, causes the one or more arrays of configurable units to implement at least the computing task; and

storing the configuration data in a non-transitory computer-readable storage medium.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2023
From: WANG, FEI; KOEPLINGER, DAVID ALAN; FAN, WEIHANG
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 064153/0458 →
Continuity (2)
Provisional Application 63327270 · Apr 4, 2022
Related Publication 20230315411A1 · Oct 5, 2023
References Cited (18)
US 10445451B2 · Fleming · 2019 [cited by examiner]
US 10452452B2 · Hetzel · 2019 [cited by examiner]
US 11809849B1 · Zheng · 2023 [cited by examiner]
US 20070050603A1 · Vorbach · 2007 [cited by examiner]
US 20070294671A1 · Demetriou · 2007 [cited by examiner]
US 20110238948A1 · Vorbach · 2011 [cited by examiner]
US 20160048394A1 · Vorbach · 2016 [cited by examiner]
US 20160055120A1 · Vorbach · 2016 [cited by examiner]
WO 2010142987A1 · 2010 [cited by applicant]
Seffrin, Andre, and Sorin A. Huss. “Ensuring secure information flow in partially reconfigurable architectures by means of process algebra analysis.” 2011IEEE 10th International Conference on Trust, Security and Privacy… [cited by examiner]
Najjar, Walid A., et al. “High-level language abstraction for reconfigurable computing.” Computer 36.8 (2003): pp. 63-69. (Year: 2003). [cited by examiner]
Ayala-Rincón, Mauricio, et al. “Prototyping time-and space-efficient computations of algebraic operations over dynamically reconfigurable systems modeled by rewriting-logic.” ACM Transactions on Design Automation of Ele… [cited by examiner]
Hartenstein, Reiner. “A decade of reconfigurable computing: a visionary retrospective.” Proceedings design, automation and test in Europe. Conference and exhibition 2001. IEEE, 2001. pp. 642-649. (Year: 2001). [cited by examiner]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Podobas et al., A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Zhang et al., “SARA: Scaling a Reconfigurable Dataflow Accelerator,” 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), 2021, pp. 1041-1054. [cited by applicant]