IP Library Granted Patent US 11,182,221
Granted Patent B1
US 11,182,221 · App. 17/127,929 · Granted Nov 23, 2021

Inter-node buffer-based streaming for reconfigurable processor-as-a-service (RPaaS)

Inventors: Ram Sivaramakrishnan (San Jose, CA); Sumti Jairath (Santa Clara, CA); Emre Ali Burhan (Sunnyvale, CA); Manish K. Shah (Austin, TX); Raghu Prabhakar (San Jose, CA); Ravinder Kumar (Fremont, CA); Arnav Goel (San Jose, CA); Ranen Chatterjee (Fremont, CA); Gregory Frederick Grohoski (Bee Cave, TX); Kin Hing Leung (Cupertino, CA); Dawei Huang (San Diego, CA); Manoj Unnikrishnan (Saratoga, CA); Martin Russell Raumann (San Leandro, CA); Bandish B. Shah (San Francisco, CA)
Assignee: SambaNova Systems, Inc.
G06F9/5077G06F9/45558G06F9/5027G06F2009/4557
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,182,221
App. No.
17/127,929
Granted
Nov 23, 2021
Kind
B1
Abstract

The technology disclosed relates to buffer-based inter-node streaming of configuration data over a network fabric. In particular, the technology disclosed relates to a runtime processor configured to load and execute a first subset of configuration files in a set of configuration files on a first reconfigurable processor operatively coupled to a first processing node, load and execute a second subset of configuration files in the set of configuration files on a second reconfigurable processor operatively coupled to a second processing node, and use a first plurality of buffers operatively coupled to the first processing node, and a second plurality of buffers operatively coupled to the second processing node to stream data between the first reconfigurable processor and the second reconfigurable processor to load and execute the first subset of configuration files and the second subset of configuration files.

Claims (30)

1. A data processing system, comprising:

a pool of reconfigurable dataflow resources including a plurality of processing nodes, respective processing nodes in the plurality of processing nodes operatively coupled to respective pluralities of reconfigurable processors and respective pluralities of buffers; and

a runtime processor operatively coupled to the pool of reconfigurable dataflow resources, and configured to:

receive a plurality of configuration files for applications, configuration files in the plurality of configuration files specifying configurations of virtual dataflow resources required to execute the configuration files, and the virtual dataflow resources including a first virtual reconfigurable processor in a first virtual processing node, a second virtual reconfigurable processor in a second virtual processing node, and virtual buffers that stream data between the first virtual reconfigurable processor and the second virtual reconfigurable processor;

allocate reconfigurable dataflow resources in the pool of reconfigurable dataflow resources to the virtual dataflow resources, the allocated reconfigurable dataflow resources including

a first processing node in the respective processing nodes allocated to the first virtual processing node,

a second processing node in the respective processing nodes allocated to the second virtual processing node,

a first reconfigurable processor, operatively coupled to the first processing node, allocated to the first virtual reconfigurable processor,

a second reconfigurable processor operatively coupled to the second processing node allocated to the second virtual reconfigurable processor, and

a first plurality of send and receive network interface card (NIC) buffers, operatively coupled to the first processing node, and a second plurality of send and receive network interface card (NIC) buffers, operatively coupled to the second processing node, allocated to the virtual buffers, wherein the first and second pluralities of send and receive NIC buffers use routing tables to specify local reconfigurable processors and destination reconfigurable processors; and

execute the configuration files and process data for the applications using the allocated reconfigurable dataflow resources.

2. The data processing system of claim 1 , wherein the first plurality of buffers includes a first set of sender buffers configured to receive data from the first reconfigurable processor and provide the data to a second set of receiver buffers in the second plurality of buffers, the second set of receiver buffers configured to provide the data to the second reconfigurable processor.

3. The data processing system of claim 2 , wherein the second plurality of buffers includes a second set of sender buffers configured to receive data from the second reconfigurable processor and provide the data to a first set of receiver buffers in the first plurality of buffers, the first set of receiver buffers configured to provide the data to the first reconfigurable processor.

4. The data processing system of claim 3 , wherein the respective processing nodes are operatively coupled to respective host processors.

5. The data processing system of claim 4 , wherein the first plurality of buffers operates in a memory of a first host processor operatively coupled to the first processing node, and the second plurality of buffers operates in a memory of a second host processor operatively coupled to the second processing node.

6. The data processing system of claim 3 , wherein the respective processing nodes are operatively coupled to respective pluralities of Smart Network Interface Controllers (SmartNICs), and wherein the first and second pluralities of send and receive NIC buffers are located on the SmartNICs.

7. The data processing system of claim 6 , wherein the first plurality of buffers operates in a memory of a first SmartNIC operatively coupled to the first processing node.

8. The data processing system of claim 7 , wherein the runtime logic is further configured to configure the first SmartNIC with a routing table that specifies the first reconfigurable processor as a local reconfigurable processor, and the second reconfigurable processor as a destination reconfigurable processor.

9. The data processing system of claim 6 , wherein the second plurality of buffers operates in a memory of a second SmartNIC operatively coupled to the second processing node.

10. The data processing system of claim 9 , wherein the runtime logic is further configured to configure the second SmartNIC with a routing table that specifies the second reconfigurable processor as a local reconfigurable processor, and the first reconfigurable processor as a destination reconfigurable processor.

11. The data processing system of claim 3 , wherein the first plurality of buffers operates in a memory of the first reconfigurable processor, and the second plurality of buffers operates in a memory of the second reconfigurable processor.

12. The data processing system of claim 1 , wherein at least one of the applications is a dataflow graph with a set of processing modules.

13. The data processing system of claim 12 , wherein the runtime logic is further configured to partition the set of processing modules into a first subset of processing modules and a second subset of processing modules.

14. The data processing system of claim 13 , wherein the runtime logic is further configured to execute configuration files for the first subset of processing modules and data therefor on the first reconfigurable processor.

15. The data processing system of claim 14 , wherein the runtime logic is further configured to execute configuration files for the second subset of processing modules and data therefor on the second reconfigurable processor.

16. The data processing system of claim 15 , wherein the runtime logic is further configured to use the first plurality of buffers and the second plurality of buffers to stream data between the first subset of processing modules and the second subset of processing modules, wherein the data includes feature maps and/or activations generated during a forward pass, and loss gradients generated during a backward pass.

17. The data processing system of claim 12 , wherein the runtime logic is further configured to initialize a first instance of the dataflow graph and a second instance of the dataflow graph.

18. The data processing system of claim 17 , wherein the runtime logic is further configured to execute configuration files for the first instance of the dataflow graph and data therefor on the first reconfigurable processor.

19. The data processing system of claim 18 , wherein the runtime logic is further configured to execute configuration files for the second instance of the dataflow graph and data therefor on the second reconfigurable processor.

20. The data processing system of claim 19 , wherein the runtime logic is further configured to use the first plurality of buffers and the second plurality of buffers to stream data between the first instance of the dataflow graph and the second instance of the dataflow graph, wherein the data includes gradients generated during a backward pass.

Assignments (3)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ORIGINAL ASSIGNMENT DOCUMENT PREVIOUSLY RECORDED ON REEL 057531 FRAME 0072. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded Oct 2, 2021
From: SIVARAMAKRISHNAN, RAM; JAIRATH, SUMTI; BURHAN, EMRE ALI; SHAH, MANISH K.; PRABHAKAR, RAGHU; KUMAR, RAVINDER; GOEL, ARNAV; CHATTERJEE, RANEN; GROHOSKI, GREGORY FREDERICK; LEUNG, KIN HING; HUANG, DAWEI; UNNIKRISHNAN, MANOJ; RAUMANN, MARTIN RUSSELL; SHAH, BANDISH B.
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 057689/0186 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2021
From: SIVARAMAKRISHNAN, RAM; JAIRATH, SUMTI; BURHAN, EMRE ALI; SHAH, MANISH K.; PRABHAKAR, RAGHU; KUMAR, RAVINDER; GOEL, ARNAV; CHATTERJEE, RANEN; GROHOSKI, GREGORY FREDERICK; LEUNG, KIN HING; HUANG, DAWEI; UNNIKRISHNAN, MANOJ; RAUMANN, MARTIN RUSSELL; SHAH, BANDISH B.
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 057531/0072 →
Cited By (14)
US 12,210,468 US 12,229,057 US 12,242,403 US 12,271,322 US 12,340,190 US 12,367,022 US 12,380,041 US 12,413,530 US 12,430,272 US 12,474,897 US 12,602,349 US 12,645,693 US 12,681,806 US 12,705,205