IP Library Granted Patent US 12,260,199
Granted Patent B2
US 12,260,199 · App. 18/126,610 · Granted Mar 25, 2025

Merging skip-buffers in a reconfigurable dataflow processor

Inventors: Fei Wang (Palo Alto, CA); David Alan Koeplinger (Egg Harbor, NJ); Kevin Brown (Palo Alto, CA); Weiwei Chen (Palo Alto, CA)
Assignee: SambaNova Systems, Inc.
G06F8/45G06F8/4434
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,199
App. No.
18/126,610
Granted
Mar 25, 2025
Kind
B2
Abstract

A method in a reconfigurable computing system includes connecting a plurality of tensor consumers to their corresponding tensor producers via skip-buffers, which generates a plurality of skip-buffers. The method includes determining that at least one skip-buffer of the plurality of skip-buffers corresponding to a first set of tensor consumers and at least one skip-buffer of the plurality of skip-buffers corresponding to a second set of tensor consumers, are compatible to wholly or partially merge. The method also includes merging, wholly or partially, the compatible skip-buffers to produce a merged skip-buffer having a minimal buffer depth. The described method may reduce memory unit consumption and latency.

Claims (29)

1. A system in a reconfigurable dataflow processor, the system comprising:

a host computer comprising a processor and an optimization module configured to conduct a method comprising:

connecting a plurality of tensor consumers to their corresponding tensor producers via skip-buffers to produce a plurality of skip-buffers;

determining that at least one skip-buffer of the plurality of skip-buffers corresponding to a first set of tensor consumers and at least one skip-buffer of the plurality of skip-buffers corresponding to a second set of tensor consumers, are compatible to wholly or partially merge to produce mergeable skip-buffers; and

merging, wholly or partially, the mergeable skip-buffers to produce a merged skip-buffer having a minimal buffer depth.

2. The system of claim 1 , wherein merging reduces memory unit consumption.

3. The system of claim 1 , wherein connecting, determining, and merging occur in a first intermediate stage of a coarse-grained reconfigurable compiler.

4. The system of claim 3 , wherein additional skip-buffer merging is conducted subsequent to the first intermediate stage of the coarse-grained reconfigurable compiler.

5. The system of claim 4 , wherein resource aware peephole optimization is conducted subsequent to the first intermediate stage of the coarse-grained reconfigurable compiler.

6. The system of claim 1 , wherein the mergeable skip-buffers have a minimal buffer depth.

7. The system of claim 1 , wherein determining that the skip-buffers of the plurality of skip-buffers are mergeable comprises determining that a first set of operations corresponding to the first set of tensor consumers are compatible with a second set of operations corresponding to the second set of tensor consumers.

8. The system of claim 1 , wherein the mergeable skip-buffers are partially mergeable and at least one mergeable skip-buffer of the mergeable skip-buffers has a buffer depth that is greater than the minimal buffer depth.

9. The system of claim 8 , wherein merging the mergeable skip-buffers produces an augmented buffer in addition to the merged skip-buffer.

10. A method in a reconfigurable computing system, the method comprising:

connecting a plurality of tensor consumers to their corresponding tensor producers via skip-buffers to produce a plurality of skip-buffers;

determining that at least one skip-buffer of the plurality of skip-buffers corresponding to a first set of tensor consumers and at least one skip-buffer of the plurality of skip-buffers corresponding to a second set of tensor consumers, are compatible to wholly or partially merge to produce mergeable skip-buffers; and

merging, wholly or partially, the mergeable skip-buffers to produce a merged skip-buffer having a minimal buffer depth.

11. The method of claim 10 , wherein merging reduces memory unit consumption.

12. The method of claim 10 , wherein connecting, determining, and merging occur in a first intermediate stage of a coarse-grained reconfigurable compiler.

13. The method of claim 12 , wherein additional skip-buffer merging is conducted subsequent to the first intermediate stage of the coarse-grained reconfigurable compiler.

14. The method of claim 13 , wherein resource aware peephole optimization is conducted subsequent to the first intermediate stage of the coarse-grained reconfigurable compiler.

15. The method of claim 10 , wherein the mergeable skip-buffers have a minimal buffer depth.

16. The method of claim 10 , wherein determining that the skip-buffers of the plurality of skip-buffers are mergeable comprises determining that a first set of operations corresponding to the first set of tensor consumers are compatible with a second set of operations corresponding to the second set of tensor consumers.

17. The method of claim 10 , wherein the mergeable skip-buffers are partially mergeable and at least one mergeable skip-buffer of the mergeable skip-buffers has a buffer depth that is greater than the minimal buffer depth.

18. The method of claim 17 , wherein merging the mergeable skip-buffers produces an augmented buffer in addition to the merged skip-buffer.

19. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, wherein the computer readable storage medium is not a transitory signal per se, and wherein the program instructions are executable by a processor to cause the processor to conduct a method comprising:

connecting a plurality of tensor consumers to their corresponding tensor producers via skip-buffers to produce a plurality of skip-buffers;

determining that at least one skip-buffer of the plurality of skip-buffers corresponding to a first set of tensor consumers and at least one skip-buffer of the plurality of skip-buffers corresponding to a second set of tensor consumers, are compatible to wholly or partially merge to produce mergeable skip-buffers; and

merging, wholly or partially, the mergeable skip-buffers to produce a merged skip-buffer having a minimal buffer depth.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: WANG, FEI; KOEPLINGER, DAVID ALAN; BROWN, KEVIN; CHEN, WEIWEI
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 063130/0184 →
Continuity (2)
Provisional Application 63324500 · Mar 28, 2022
Related Publication 20230305823A1 · Sep 28, 2023
References Cited (15)
US 11182264B1 · Sivaramakrishnan · 2021 [cited by examiner]
US 11232360B1 · Nama · 2022 [cited by examiner]
US 11250061B1 · Nama · 2022 [cited by examiner]
US 20220197713A1 · Sivaramakrishnan · 2022 [cited by examiner]
US 20220197714A1 · Raumann · 2022 [cited by examiner]
US 20220198117A1 · Raumann · 2022 [cited by examiner]
US 20220261365A1 · Prabhakar · 2022 [cited by examiner]
US 20220309028A1 · Nama · 2022 [cited by examiner]
US 20220309319A1 · Nama · 2022 [cited by examiner]
WO 2010142987A1 · 2010 [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Zhang et al., “SARA: Scaling a Reconfigurable Dataflow Accelerator,” 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), 2021, pp. 1041-1054. [cited by applicant]