IP Library Granted Patent US 12,688,148
Granted Patent B2
US 12,688,148 · App. 18/794,143 · Granted Jul 21, 2026

System and method for optimizing data-transfer among multiple compute units in a data-parallel computing system

Inventors: Greg Dykema (Palo Alto, CA); Aarti Lalwani (Palo Alto, CA)
Assignee: SambaNova Systems, Inc.
G06F15/825G06F13/4068
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,148
App. No.
18/794,143
Granted
Jul 21, 2026
Kind
B2
Abstract

System and method for optimizing data-transfer among multiple compute units in a data-parallel computing system. A topological communications configurator (TCC) determines a connections-optimized configuration of processors associated with compute nodes of the computing system. The processors can execute dataflow workers of an application and form intranodal segments of an internodal interconnection topology coupling the intranodal segments. The TCC determines the connections-optimized configuration based on internodal communications costs corresponding to communications routes among the internodal segments via the internodal interconnection fabric.

Claims (42)

1 . A data processing system, comprising:

a first computing system comprising a first processor and a topological communications configurator (TCC) comprising a computing program configured to execute on the first processor;

a second computing system comprising a plurality of compute nodes, compute nodes among the plurality of compute nodes comprising at least one dataflow processor, each of the at least one dataflow processor configured to execute a compute worker among a plurality of compute workers, the plurality of compute workers configured to execute a computing application of the second computing system; and

an internodal fabric configured to communicatively couple compute nodes among the plurality of compute nodes,

wherein the TCC is configured to:

determine an internodal interconnection topology of the plurality of compute nodes, the internodal interconnection topology comprising the internodal fabric;

determine, based on the internodal interconnection topology, internodal communications routes communicatively interconnecting, via the internodal fabric, a set of intranodal segments among a plurality of intranodal segments, each of the plurality of intra nodal segments comprising an intranodal interconnection of dataflow processors included in respective nodes among the plurality of compute nodes, each of the plurality of intranodal segments comprising a set of compute workers, among the plurality of compute workers, executed by the dataflow processors included in a node among the respective nodes, the set of compute workers included in a worker logical topology;

determine internodal communications costs corresponding to communications routes among the internodal communications routes; and

determine, based on the internodal communications costs, a connections-optimized configuration of interconnected segments among the set of intranodal segments.

2 . The data processing system of claim 1 , wherein the TCC is further configured to determine, based on the internodal communications costs, a cost-optimized interconnection topology of the interconnected segments.

3 . The data processing system of claim 2 , wherein the cost-optimized interconnection topology of the interconnected segments is a ring topology.

4 . The data processing system of claim 1 , wherein a first segment, among the plurality of intranodal segments, comprises a head dataflow processor;

wherein a second segment, among the plurality of intranodal segments, comprises a tail dataflow processor; and

wherein the TCC is further configured to determine, based on the internodal communications costs, a cost-optimized interconnection of the head dataflow processor, included in the first segment, and the tail dataflow processor included in the second segment.

5 . The data processing system of claim 1 , wherein the TCC is further configured to determine a connections-optimized configuration of a set of dataflow processors of a node among the plurality of compute nodes, the set of dataflow processors included in a segment, among the plurality of intranodal segments included in the node.

6 . The data processing system of claim 1 , wherein the TCC is further configured to determine the connections-optimized configuration of dataflow processors included in the segment by

determining an intranodal interconnection topology of the node, the intranodal interconnection topology comprising interconnections of the set of dataflow processors via an intranodal fabric;

determining, based on the intranodal interconnection topology, a set of intranodal communications routes communicatively interconnecting, via the intranodal fabric, the set of dataflow processors;

determining intranodal communications costs corresponding to communications routes among the set of intranodal communications routes; and

determining based on the intranodal communications costs, a connections-optimized configuration of the set of dataflow processors.

7 . The data processing system of claim 1 , wherein the set of compute workers is configured to execute operations of the computing application as a pipeline of compute workers.

8 . The data processing system of claim 1 , wherein the TCC configured to determine the internodal communications costs comprises based on performance characteristics selected from a group consisting of performance characteristics of the internodal fabric, and performance characteristics of an interconnect coupling a first segment among the plurality of intranodal segments.

9 . The data processing system of claim 8 , wherein a performance characteristic among the performance characteristics of the internodal fabric is selected from a group consisting of performance characteristics of the internodal fabric, performance characteristics of the interconnect coupling the first segment and the internodal fabric, and performance characteristics associated with a physical locality of the internodal fabric within the second computing system.

10 . The data processing system of claim 8 , wherein a performance characteristic among the performance characteristics of the interconnect is selected from a group consisting of a utilization of the interconnect, a throughput of the interconnect, a data rate of the interconnect, a communications latency of the interconnect, and a physical locality of the interconnect within the second computing system.

11 . A computer-implemented method comprising:

determining, by a topological communications configurator (TCC) of a first computing system, an internodal interconnection topology of a plurality of compute nodes of a second computing system, the internodal interconnection topology comprising an internodal fabric;

determining, by the TCC, based on the internodal interconnection topology, internodal communications routes interconnecting, via the internodal fabric, a set of intranodal segments among a plurality of intranodal segments, each of the plurality of intranodal segments comprising an intranodal interconnection of dataflow processors included in respective nodes among the plurality of compute nodes, each of the plurality of intranodal segments corresponding to a respective portion of a worker logical topology comprising compute workers configured to execute an application of the second computing system;

determining, by the TCC, internodal communications costs corresponding to communications routes among the internodal communications routes; and

determining a connections-optimized configuration of interconnected segments.

12 . The method of claim 11 , wherein the method of determining the connections-optimized configuration of integrated segments comprises the TCC determining, based on the internodal communications costs, a cost-optimized interconnection of the interconnected segments to form a ring topology among the interconnected segments.

13 . The method of claim 11 , wherein a first segment and a second segment, among the plurality of intranodal segments, each comprise a head and a tail dataflow processor; and

wherein the method determining, by the TCC, the connections-optimized configuration of integrated segments comprises the TCC determining, further based on the internodal communications costs, a cost-optimized interconnection of the tail dataflow processor of the first segment and a head dataflow processor of the second segment.

14 . The method of claim 11 , the method further comprising the TCC determining a connections-optimized configuration of dataflow processors included in a first segment of a first node, the first segment among the plurality of intranodal segments, the first node among the plurality of compute nodes.

15 . The method of claim 11 , wherein the compute workers comprise data-parallel workers configured to execute operations of the application, on dataflow processors among the dataflow processors included in the respective nodes among the plurality of compute nodes, as a pipeline.

16 . The method of claim 11 , wherein the internodal communications costs are based on performance characteristics selected from a group consisting of performance characteristics of the internodal fabric and performance characteristics of an interconnect coupling a first segment among the plurality of intranodal segments and the internodal fabric.

17 . The method of claim 16 , wherein a performance characteristic among the performance characteristics of the internodal fabric is selected from a group consisting of a utilization of the internodal fabric; a throughput of a communications route through the internodal fabric; a latency of a communications route through the internodal fabric; and a physical locality of the internodal fabric within the second computing system.

18 . The method of claim 16 , wherein a performance characteristic among the performance characteristics of the interconnect is selected from a group consisting of a utilization of the interconnect; a throughput of the interconnect; a data rate of the interconnect; a communications latency of the interconnect; and a physical locality of the interconnect within the second computing system.

19 . A computer program product comprising a computer readable storage medium having first program instructions embodied therewith, wherein the first program instructions are executable by at least one processor to cause the at least one processor to:

determine an internodal interconnection topology of a plurality of compute nodes of a computing system, the internodal interconnection topology comprising an internodal fabric;

determine, based on the internodal interconnection topology, internodal communications routes communicatively interconnecting, via the internodal fabric, a set of intranodal segments among a plurality of intranodal segments, each of the plurality of intranodal segments comprising an intranodal interconnection of processors included in respective nodes among the plurality of compute nodes, each of the plurality of intranodal segments corresponding to a respective portion of a worker logical topology comprising compute workers configured to execute an application of the computing system;

determine internodal communications costs corresponding to communications routes among the internodal communications routes; and

determine, based on the internodal communications costs, a connections-optimized configuration of interconnected segments.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2024
From: DYKEMA, GREG; LALWANI, AARTI
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 069141/0643 →
Continuity (3)
Continuation 18096253 · Jan 12, 2023
Provisional Application 63301464 · Jan 20, 2022
Related Publication 20240394218A1 · Nov 28, 2024
References Cited (31)
US 5754543A · Seid · 1998 [cited by applicant]
US 7231638B2 · Blackmore et al. · 2007 [cited by applicant]
US 7414978B2 · Lun et al. · 2008 [cited by applicant]
US 9489202B2 · Stark · 2016 [cited by examiner]
US 10698853B1 · Grohoski et al. · 2020 [cited by applicant]
US 11574253B2 · Santos et al. · 2023 [cited by applicant]
US 11853244B2 · Sankaralingam · 2023 [cited by examiner]
US 12175252B2 · Ould-Ahmed-Vall · 2024 [cited by examiner]
US 20040114569A1 · Naden et al. · 2004 [cited by applicant]
US 20070198752A1 · Danz · 2007 [cited by examiner]
US 20100217949A1 · Schopp · 2010 [cited by examiner]
US 20120120803A1 · Farkas et al. · 2012 [cited by applicant]
US 20140281210A1 · Rehm · 2014 [cited by examiner]
US 20150271236A1 · Chen et al. · 2015 [cited by applicant]
US 20150277990A1 · Xiong et al. · 2015 [cited by applicant]
US 20170177517A1 · Nicol · 2017 [cited by examiner]
US 20180210730A1 · Sankaralingam · 2018 [cited by examiner]
US 20200319324A1 · Au et al. · 2020 [cited by applicant]
US 20210035027A1 · Santos et al. · 2021 [cited by applicant]
US 20220012077A1 · Kumar et al. · 2022 [cited by applicant]
US 20230024785A1 · Dutta · 2023 [cited by applicant]
US 20230229624A1 · Dykema et al. · 2023 [cited by applicant]
US 20240038086A1 · Yang et al. · 2024 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
Dong et al., Low-Cost Datacenter Load Balancing with Multipath Transport and Top-of-Rack Switches, IEEE Transactions on Parallel and Distributed Systems, vol. 31, Issue 10, Apr. 22, 2020, 16 pages. [cited by applicant]
Heydari, et al, Efficient network structures with separable heterogeneous connection costs, Economic Letters, vol. 134, Sep. 2015, pp. 82-85. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Xue et al., ROTOS: A Reconfigurable and Cost-Effective Architecture for High-Performance Optical Data Center Networks, Journal Lightwave Technology, vol. 38, dated Jun. 16, 2020, pp. 3485-3494. [cited by applicant]
Yang, Complexity analysis of new task allocation problem using network flow method on multicore clusters, Mathematical Problems in Engineering 2014, 7 pages. [cited by applicant]