IP Library Granted Patent US 12,236,220
Granted Patent B2
US 12,236,220 · App. 18/206,829 · Granted Feb 25, 2025

Flow control for reconfigurable processors

Inventors: Weiwei Chen (Palo Alto, CA); Raghu Prabhakar (San Jose, CA); David Alan Koeplinger (Egg Harbor, NJ); Sitanshu Gupta (Palo Alto, CA); Ruddhi Chaphekar (Palo Alto, CA); Ajit Punj (Palo Alto, CA); Sumti Jairath (Palo Alto, CA)
Assignee: SambaNova Systems, Inc.
G06F8/452G06F8/41G06F15/7867G06F15/825
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,220
App. No.
18/206,829
Granted
Feb 25, 2025
Kind
B2
Abstract

The technology disclosed relates to storing a dataflow graph with a plurality of compute nodes that transmit data along data connections, and controlling data transmission between compute nodes in the plurality of compute nodes along the data connections by using control connections to control writing of data.

Claims (35)

1. A computer-implemented method, including:

executing a dataflow graph on a processing system having a plurality of compute nodes, each having a ready-to-read credit counter and a write credit counter, that transmit data along data connections; and

controlling data transmission between compute nodes in the plurality of compute nodes along the data connections by using dedicated control connections between the compute nodes to selectively control writing of data based on both the ready-to-read credit counter and the write credit counter in a particular compute node of the plurality of compute nodes that provides the data for a particular data transmission, wherein the control connections are distinct from the data connections and are configured to transmit control signals that manage flow of the data by selectively incrementing the ready-to-read credit counter and the write credit counter in the particular compute node.

2. The computer-implemented method of claim 1 , further including initializing the ready-to-read credit counter of the particular compute node with as many read credits as a buffer depth of a corresponding compute node of the plurality of compute nodes that reads data from the particular compute node.

3. The computer-implemented method of claim 2 , further including decrementing the ready-to-read credit counter in the particular compute node when the particular compute node begins writing a buffer data unit into the corresponding compute node along a data connection of the data connections.

4. The computer-implemented method of claim 3 , further including incrementing the ready-to-read credit counter in the particular compute node when the particular compute node receives from the corresponding compute node a read ready token along a control connection of the control connections, wherein the read ready token indicates to the particular compute node that the corresponding compute node has freed a buffer data unit and is ready to receive an additional buffer data unit.

5. The computer-implemented method of claim 4 , wherein the particular compute node stops writing data into the corresponding compute node when the ready-to-read credit counter in the particular compute node has zero read credits.

6. The computer-implemented method of claim 5 , wherein the particular compute node resumes writing data into the corresponding compute node when the particular compute node receives the read ready token from the corresponding compute node.

7. The computer-implemented method of claim 1 , further including initializing the write credit counter in the particular compute node with one or more write credits.

8. The computer-implemented method of claim 7 , further including decrementing the write credit counter in the particular compute node when the particular compute node begins writing a buffer data unit of the data into a corresponding compute node of the plurality of compute nodes along a data connection of the data connections.

9. The computer-implemented method of claim 8 , further including incrementing the write credit counter in the particular compute node when the particular compute node receives from the corresponding compute node a write done token along a control connection of the control connections, wherein the write done token indicates to the particular compute node that the writing of the buffer data unit into the corresponding compute node has completed.

10. The computer-implemented method of claim 9 , wherein the particular compute node stops writing data into the corresponding compute node when the write credit counter in the particular compute node has zero write credits.

11. The computer-implemented method of claim 10 , wherein the particular compute node resumes writing data into the corresponding compute node when the particular compute node receives the write done token from the corresponding compute node.

12. A computer-implemented method comprising:

executing a dataflow graph on a processing system having a plurality of compute nodes coupled by data connections and control connections, including a first node coupled to a second node by a first data connection and a first control connection, the first node including a write credit counter and a ready-to-read credit counter;

sending data from the first node to the second node over the first data connection while both the write credit counter and the ready-to-read credit counter are greater than zero while pausing transmission of the data from the first node to the second node over the first data connection while either the write credit counter or the ready-to-read credit counter are equal to zero;

decrementing both the write credit counter and the ready-to-read credit counter upon sending a buffer unit of the data from the first node to the second node over the first data connection;

incrementing the write credit counter upon receipt of a write done token from the second node over the first control connection; and

incrementing the ready-to-read credit counter upon receipt of a read ready token from the second node over the first control connection.

13. The computer-implemented method of claim 12 , further comprising:

initializing the write credit counter to a first predetermined value; and

initializing the ready-to-read credit counter to a second predetermined value that is based on a size of a buffer in the second node.

14. The computer-implemented method of claim 12 , further including:

sending the write done token from the second node to the first node over the first control connection in response to storing the buffer unit of the data received from the first node over the first data connection into a buffer of the second node.

15. The computer-implemented method of claim 12 , further including:

sending the read ready token from the second node to the first node over the first control connection in response to removing the buffer unit of the data from a buffer of the second node.

16. A processing system comprising:

a plurality of compute nodes coupled by data connections and control connections, including a first node coupled to a second node by a first data connection and a first control connection, wherein the first node is configured to send data to the second node over the first data connection;

a write credit counter in the first node that is incremented in response to receipt of a write done token from the second node over the first control connection and decremented in response to receipt of a write done token from the second node over the first control connection;

a ready-to-read credit counter in the first node that is incremented in response to receipt of a read ready token from the second node over the first control connection and decremented in response to receipt of a write done token from the second node over the first control connection; and

dataflow control circuitry in the first node to pause transmission of the data from the first node to the second node over the first data connection while either the write credit counter or the ready-to-read credit counter are equal to zero.

17. The processing system of claim 16 , further comprising a buffer in the second node, wherein the ready-to-read credit counter is initialized to a predetermined value that is based on a size of the buffer in the second node.

18. The processing system of claim 17 , further comprising buffer management circuitry configured to:

send the write done token from the second node to the first node over the first control connection in response to storing a buffer unit of the data received from the first node over the first data connection into the buffer of the second node; and

send the read ready token from the second node to the first node over the first control connection in response to removing the buffer unit of the data from the buffer of the second node.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2023
From: CHEN, WEIWEI; PRABHAKAR, RAGHU; KOEPLINGER, DAVID ALAN; GUPTA, SITANSHU; CHAPHEKAR, RUDDHI ARUN; PUNJ, AJIT; JAIRATH, SUMTI
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 063899/0962 →
Continuity (2)
Continuation 16890841 · Jun 2, 2020
Related Publication 20230325163A1 · Oct 12, 2023
References Cited (59)
US 6636223B1 · Morein · 2003 [cited by applicant]
US 7369637B1 · Mauer · 2008 [cited by applicant]
US 9063668B1 · Jung et al. · 2015 [cited by applicant]
US 10970217B1 · Dastidar et al. · 2021 [cited by applicant]
US 11036546B1 · Bhandari · 2021 [cited by examiner]
US 11709664B2 · Chen · 2023 [cited by examiner]
US 20020010852A1 · Arnold · 2002 [cited by examiner]
US 20040068612A1 · Stolowitz · 2004 [cited by applicant]
US 20040208173A1 · Gregorio · 2004 [cited by applicant]
US 20070011396A1 · Singh et al. · 2007 [cited by applicant]
US 20100161948A1 · Abdallah · 2010 [cited by applicant]
US 20100191878A1 · Nandagopalan et al. · 2010 [cited by applicant]
US 20110246170A1 · Oh et al. · 2011 [cited by applicant]
US 20120066690A1 · Gupta · 2012 [cited by examiner]
US 20130061028A1 · Ishebabi · 2013 [cited by examiner]
US 20130110784A1 · Guo et al. · 2013 [cited by applicant]
US 20140085318A1 · Nadar et al. · 2014 [cited by applicant]
US 20180129624A1 · Krutsch et al. · 2018 [cited by applicant]
US 20180157464A1 · Lutz et al. · 2018 [cited by applicant]
US 20180204117A1 · Brevdo · 2018 [cited by applicant]
US 20180210730A1 · Sankaralingam et al. · 2018 [cited by applicant]
US 20180212894A1 · Nicol et al. · 2018 [cited by applicant]
US 20180300181A1 · Hetzel et al. · 2018 [cited by applicant]
US 20190057053A1 · Tsuchida et al. · 2019 [cited by applicant]
US 20190130269A1 · Nicol · 2019 [cited by applicant]
US 20190229996A1 · ChoFleming, Jr. et al. · 2019 [cited by applicant]
US 20190392002A1 · Lavasani · 2019 [cited by examiner]
US 20190392296A1 · Brady et al. · 2019 [cited by applicant]
US 20200005155A1 · Datta et al. · 2020 [cited by applicant]
US 20200034306A1 · Luo et al. · 2020 [cited by applicant]
US 20200142743A1 · Zhang et al. · 2020 [cited by applicant]
US 20200241797A1 · Kanno et al. · 2020 [cited by applicant]
US 20200387397A1 · Ohta et al. · 2020 [cited by applicant]
US 20210042259A1 · Koeplinger et al. · 2021 [cited by applicant]
US 20210081769A1 · Chen et al. · 2021 [cited by applicant]
US 20210182021A1 · Wang et al. · 2021 [cited by applicant]
US 20210192314A1 · Aarts et al. · 2021 [cited by applicant]
US 20220057958A1 · Patriarca et al. · 2022 [cited by applicant]
US 20220066739A1 · Croxford et al. · 2022 [cited by applicant]
TW 202230129A · 2022 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
WO 2019202216A2 · 2019 [cited by applicant]
IEEE search results, Year: 2021, 7 pages. [cited by applicant]
Jung, Optimization of the Memory Subsystem of a Coarse Grained Reconfigurable Hardware Accelerator, Technical University at Darmstadt, dated 2019, 184 pages. [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
Lopez, DA, A Memory Layout for Dynamically Routed Capsule Layers, 16th International Conference on Information Technology—New Generations (ITNG), pp. 317-324, Published in May 2019. [doi:http://dx.doi.org/10.1007/978-3-… [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Pawłowski et al., High performance tensor-vector multiplies on shared memory systems, published in 2019, 24 pages. [cited by applicant]
PCT/US2021/035305—International Search Report and Written Opinion dated Sep. 1, 2021,17 pages. [cited by applicant]
PCT/US2021/050586—International Search Report and Written Opinion, dated Jan. 18, 2022, 16 pages. [cited by applicant]
PCT/US2021/051305—International Search Report and Written Opinion, dated Jan. 11, 2022, 16 pages. [cited by applicant]
Podobas et al., A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Prabhakar, Design of Programmable, Energy-Efficient Reconfigurable Accelerators, Stanford University, dated Aug. J018, 104 pages. [cited by applicant]
Rotem et al., Glow: Graph Lowering Compiler Techniques for Neural Networks, Cornell University Library, New York, dated May 2, 2018, 12 pages. [cited by applicant]
U.S. Appl. No. 17/023,015—Notice of Allowance dated Sep. 30, 2021, 9 pages. [cited by applicant]
U.S. Appl. No. 17/031,679—Notice of Allowance, dated Dec. 22, 2022, 13 pages. [cited by applicant]
Zhao et al. “Serving Recurrent Neural Networks Efficiently with a Spatial Accelerator”, 2019, Proceedings of the 2nd sysML Conference. (Year: 2019). [cited by applicant]
Vasiljevic et al., OpenCL Library of Stream Memory Components Targeting FPGAs, Dec. 2015, IEEE (Year: 2015). [cited by applicant]