IP Library Granted Patent US 12,380,060
Granted Patent B2
US 12,380,060 · App. 18/202,059 · Granted Aug 5, 2025

Graph spatial split

Inventors: Yun Du (Palo Alto, CA); Gao Deng (Palo Alto, CA); Jianding Luo (Palo Alto, CA); Zhengyu Chen (Palo Alto, CA)
Assignee: SambaNova Systems, Inc.
G06F15/825G06F9/3867
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,060
App. No.
18/202,059
Granted
Aug 5, 2025
Kind
B2
Abstract

A method for reducing latency and increasing throughput in a reconfigurable computing system includes receiving a compute graph for execution on a reconfigurable dataflow processor comprising a grid of compute units and grid of memory units interconnected with a switching array. The compute graph includes a node specifying an operation on a tensor. The node may be split into multiple nodes that each specify the operation on a distinctive portion of the tensor to produce a first modified compute graph. The first modified compute graph may be executed. In addition, the multiple nodes may be within a single meta-pipeline stage and may be processed in parallel. Furthermore, the compute graph may further comprise a separate node for gathering the distinctive portions of the tensor into a complete tensor, to produce a second modified compute graph.

Claims (27)

1. A system for reducing latency and increasing throughput in reconfigurable dataflow processors, the system comprising:

a host computer comprising a graph optimization module configured to conduct a method comprising:

receiving a compute graph for execution on a reconfigurable dataflow computing system, the compute graph comprising a node specifying an operation on a tensor;

splitting the node into multiple nodes that each specify the operation on a distinctive portion of the tensor to produce a first modified compute graph; and

a reconfigurable dataflow processor (RDP) configured to execute the first modified compute graph.

2. The system of claim 1 , wherein splitting the node into X nodes improves latency by a factor of X.

3. The system of claim 1 , wherein the multiple nodes that each specify the operation on the distinctive portion of the tensor are within a single meta-pipeline stage.

4. The system of claim 3 , wherein the multiple nodes that each specify the operation on distinctive portions of the tensor are processed in parallel.

5. The system of claim 3 , further comprising adding a separate node for gathering the distinctive portions of the tensor into a complete tensor, to produce a second modified compute graph.

6. The system of claim 5 , wherein the separate node specifies a concatenation operation, a summation operation, or a tensor assembly operation.

7. The system of claim 5 , wherein the multiple nodes and the separate node are within the single meta-pipeline stage.

8. The system of claim 1 , wherein the RDP comprises a grid of compute units and a grid of memory units interconnected with a switching array, each compute unit comprising an array of arithmetic units organized into I lanes and J meta-pipeline stages.

9. A method for reducing latency and increasing throughput in a reconfigurable computing system, the method comprising:

receiving a compute graph for execution on a reconfigurable dataflow processor (RDP), the compute graph comprising a node specifying an operation on a tensor;

splitting the node into multiple nodes that each specify the operation on a distinctive portion of the tensor to produce a first modified compute graph; and

executing the first modified compute graph on the RDP.

10. The method of claim 9 , wherein splitting the node into X nodes improves latency by a factor of X.

11. The method of claim 9 , wherein the multiple nodes that each specify the operation on the distinctive portion of the tensor are within a single meta-pipeline stage.

12. The method of claim 11 , wherein the multiple nodes that each specify the operation on distinctive portions of the tensor are processed in parallel.

13. The method of claim 11 , further comprising adding a separate node for gathering the distinctive portions of the tensor into a complete tensor, to produce a second modified compute graph.

14. The method of claim 13 , wherein the separate node specifies a concatenation operation, a summation operation, or a tensor assembly operation.

15. The method of claim 13 , wherein the multiple nodes and the separate node are within the single meta-pipeline stage.

16. The method of claim 9 , wherein the RDP comprises a grid of compute units and a grid of memory units interconnected with a switching array, each compute unit comprising an array of arithmetic units organized into I lanes and J meta-pipeline stages.

17. A computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, wherein the computer readable storage medium is not a transitory signal per se, wherein the program instructions are executable by a processor to cause the processor to conduct a method comprising:

receiving a compute graph for execution on a reconfigurable dataflow processor (RDP), the compute graph comprising a node specifying an operation on a tensor;

splitting the node into multiple nodes that each specify the operation on a distinctive portion of the tensor to produce a first modified compute graph; and

executing the first modified compute graph on the RDP.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2023
From: DU, YUN; DENG, GAO; LUO, JIANDING; CHEN, ZHENGYU
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 063765/0901 →
Continuity (4)
Provisional Application 63348961 · Jun 3, 2022
Provisional Application 63346234 · May 26, 2022
Provisional Application 63345740 · May 25, 2022
Related Publication 20240168915A1 · May 23, 2024
References Cited (21)
US 10768899B2 · Koeplinger et al. · 2020 [cited by applicant]
US 11204889B1 · Prabhakar et al. · 2021 [cited by applicant]
US 11237971B1 · Brown et al. · 2022 [cited by applicant]
US 11250105B2 · Wang et al. · 2022 [cited by applicant]
US 11429349B1 · Oklobdzija et al. · 2022 [cited by applicant]
US 11443014B1 · Wang et al. · 2022 [cited by applicant]
US 20210157550A1 · Wang et al. · 2021 [cited by applicant]
US 20210248115A1 · Jones · 2021 [cited by examiner]
US 20210373867A1 · Chen et al. · 2021 [cited by applicant]
US 20220092247A1 · Koeplinger et al. · 2022 [cited by applicant]
US 20220124543A1 · Orhan · 2022 [cited by examiner]
US 20220129320A1 · Mohapatra et al. · 2022 [cited by applicant]
US 20220413928A1 · Banuli Nanje Gowda · 2022 [cited by examiner]
US 20230273879A1 · Bordelon et al. · 2023 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
Jingzhi Fang, Yanyan Shen, Yue Wang, and Lei Chen. 2020. Optimizing DNN computation graph using graph substitutions. In VLDB, 13, 12 (2020). (Year: 2020). [cited by examiner]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Zhang et al., “SARA: Scaling a Reconfigurable Dataflow Accelerator,” 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), 2021, pp. 1041-1054. [cited by applicant]
Cited By (1)
US 12,632,269