IP Library Granted Patent US 12,430,109
Granted Patent B2
US 12,430,109 · App. 18/115,118 · Granted Sep 30, 2025

Critical stage optimization for reconfigurable architectures

Inventors: Adam Bordelon (Palo Alto, CA); David Alan Koeplinger (Egg Harbor, NJ)
Assignee: SambaNova Systems, Inc.
G06F8/4441G06F15/7867G06F15/825
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,109
App. No.
18/115,118
Granted
Sep 30, 2025
Kind
B2
Abstract

A method for reducing latency and increasing throughput in a reconfigurable computing system includes receiving a user program for execution on a reconfigurable dataflow computing system, comprising a grid of compute units and grid of memory units interconnected with a switching array. The user program includes multiple tensor-based algebraic expressions that are converted to an intermediate representation comprising multiple stages. Each stage includes one or more logical operations executable via dataflow through compute units, and each stage is preceded by and followed by a buffer, each buffer corresponding to one or more memory units. The method includes detecting a memory mapping operation within a critical stage and moving the memory mapping operation to an adjacent stage, wherein the memory mapping operation is executable by memory units within the adjacent stage and dataflow through the buffer is controlled by one or more memory units within the grid of memory units.

Claims (36)

1. A method for reducing latency and increasing throughput in a reconfigurable computing system, the method comprising:

receiving a user program for execution on a reconfigurable dataflow computing system, the reconfigurable dataflow computing system comprising a grid of compute units and a grid of memory units interconnected with a switching array, the user program comprising a plurality of tensor-based algebraic expressions;

converting the plurality of tensor-based algebraic expressions to an intermediate representation comprising a plurality of stages, including a first stage and a second stage which is adjacent to the first stage, each stage comprising one or more logical operations executable via dataflow through one or more compute units of the grid of compute units, each stage preceded by and followed by a buffer, each buffer corresponding to one or more memory units within the grid of memory units;

detecting a memory mapping operation within a critical the first stage; and

moving the memory mapping operation to the second stage;

wherein the memory mapping operation is executable by the one or more memory units within the second stage and wherein dataflow through the buffer is controlled by one or more memory units within the grid of memory units.

2. The method of claim 1 , wherein the memory mapping operation comprises one or more of a transpose operation, a reshape operation, a layout transformation, a roll operation, a permutation operation, a slice operation or a tile operation.

3. The method of claim 1 , wherein the one or more logical operations correspond to one or more template library functions.

4. The method of claim 1 , wherein the first stage comprises a logical operation selected from one or more of matrix multiplication operation, a batch normalization operation, a batch Cholesky operation, or a layer normalization operation.

5. The method of claim 1 , wherein the second stage comprises a logical operation selected from a ReLU operation, a Sigmoid operation, or a Hyperbolic Tangent operation.

6. The method of claim 1 , wherein the one or more logical operations are represented as dataflow statements or compute graph nodes.

7. The method of claim 1 , wherein moving the memory mapping operation the second stage exposes further optimizations.

8. The method of claim 7 , wherein the further optimizations include fusing buffers.

9. The method of claim 1 , wherein the first stage has a highest latency among the plurality of stages.

10. A system for reducing latency and increasing throughput in reconfigurable dataflow processors, the system comprising:

a host computer comprising an optimization module configured to conduct a method comprising:

receiving a user program for execution on a reconfigurable dataflow computing system, the reconfigurable dataflow computing system comprising a grid of compute units and a grid of memory units interconnected with a switching array, the user program comprising a plurality of tensor-based algebraic expressions;

converting the plurality of tensor-based algebraic expressions to an intermediate representation comprising a plurality of stages, including a first stage and a second stage which is adjacent to the first stage, each stage comprising one or more logical operations executable via dataflow through one or more compute units of the grid of compute units, each stage preceded by and followed by a buffer, each buffer corresponding to one or more memory units within the grid of memory units;

detecting a memory mapping operation within the first stage; and

moving the memory mapping operation to the second stage;

wherein the memory mapping operation is executable by the one or more memory units within the second stage and wherein dataflow through the buffer is controlled by one or more memory units within the grid of memory units.

11. The system of claim 10 , wherein the memory mapping operation comprises one or more of a transpose operation, a reshape operation, a layout transformation, a roll operation, a permutation operation, a slice operation or a tile operation.

12. The system of claim 10 , wherein the one or more logical operations correspond to one or more template library functions.

13. The system of claim 10 , wherein the first stage comprises a logical operation selected from one or more of matrix multiplication operation, a batch normalization operation, a batch Cholesky operation, or a layer normalization operation.

14. The system of claim 10 , wherein the second stage comprises a logical operation selected from a ReLU operation, a Sigmoid operation, or a Hyperbolic Tangent operation.

15. The system of claim 10 , wherein the one or more logical operations are represented as dataflow statements or compute graph nodes.

16. The system of claim 10 , wherein moving the memory mapping operation to the second stage exposes further optimizations.

17. The system of claim 16 , wherein the further optimizations include fusing buffers.

18. The system of claim 10 , wherein the first stage has a highest latency among the plurality of stages.

19. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, wherein the computer readable storage medium is not a transitory signal per se, wherein the program instructions are executable by a processor to cause the processor to conduct a method comprising:

receiving a user program for execution on a reconfigurable dataflow computing system, the reconfigurable dataflow computing system comprising a grid of compute units and a grid of memory units interconnected with a switching array, the user program comprising a plurality of tensor-based algebraic expressions;

converting the plurality of tensor-based algebraic expressions to an intermediate representation comprising a plurality of stages, including a first stage and a second stage which is adjacent to the first stage, each stage comprising one or more logical operations executable via dataflow through one or more compute units of the grid of compute units, each stage preceded by and followed by a buffer, each buffer corresponding to one or more memory units within the grid of memory units;

detecting a memory mapping operation within the first stage; and

moving the memory mapping operation to the second stage;

wherein the memory mapping operation is executable by the one or more memory units within the second stage and wherein dataflow through the buffer is controlled by one or more memory units within the grid of memory units.

20. The computer program product of claim 19 , wherein the first stage has a highest latency among the plurality of stages.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2023
From: BORDELON, ADAM; KOEPLINGER, DAVID ALAN
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 062838/0244 →
Continuity (2)
Provisional Application 63314993 · Feb 28, 2022
Related Publication 20230273879A1 · Aug 31, 2023
References Cited (15)
US 11080227B2 · Koeplinger · 2021 [cited by examiner]
US 11204889B1 · Prabhakar et al. · 2021 [cited by applicant]
US 11237971B1 · Brown et al. · 2022 [cited by applicant]
US 11429349B1 · Oklobdzija et al. · 2022 [cited by applicant]
US 11645057B2 · Koeplinger et al. · 2023 [cited by applicant]
US 20200241844A1 · Koeplinger et al. · 2020 [cited by applicant]
US 20210042259A1 · Koeplinger · 2021 [cited by examiner]
US 20210357475A1 · Wang et al. · 2021 [cited by applicant]
US 20210373867A1 · Chen et al. · 2021 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Zhang et al., “SARA: Scaling a Reconfigurable Dataflow Accelerator,” 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), 2021, pp. 1041-1054. [cited by applicant]