IP Library Granted Patent US 12,461,889
Granted Patent B2
US 12,461,889 · App. 18/243,994 · Granted Nov 4, 2025

Intelligent graph execution and orchestration engine for a reconfigurable data processor

Inventors: Arnav Goel (San Jose, CA); Ravinder Kumar (Fremont, CA); Arjun Sabnis (San Francisco, CA); Qi Zheng (Fremont, CA); Neal Sanghvi (Palo Alto, CA)
Assignee: SambaNova Systems, Inc.
G06F15/80G06F15/825
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,461,889
App. No.
18/243,994
Granted
Nov 4, 2025
Kind
B2
Abstract

A data processing system including an array of reconfigurable units and a compiler configured to generate to execute a dataflow graph of a user application is disclosed. The dataflow graph includes a sequence of temporal partitions, each temporal partition including a sequence of graph control operations. Also disclosed is an intelligent graph orchestration and execution engine (IGOEE) configured to receive an optimization objective from the complier. The optimization objective can be for minimizing execution time of the reconfigurable processor or maximizing computing resource utilization of the reconfigurable processor. The IGOEE can reorganize the sequence of temporal partitions and the sequence of graph control operations within each temporal partition to satisfy the optimization objective; and execute the reorganized dataflow graph on the reconfigurable processor. A corresponding method is also disclosed herein.

Claims (48)

1 . A system comprising:

a processor comprising an array of reconfigurable units, configured to execute a dataflow graph of a user application from a compiler, wherein the dataflow graph includes a sequence of temporal partitions, and wherein each temporal partition includes a sequence of graph control operations;

an intelligent graph orchestration and execution engine (IGOEE) configured to

receive at least one optimization objective from the complier, wherein the at least one optimization objective specifies at least one of: minimizing an execution time of the reconfigurable processor and maximizing a computing resource utilization of the reconfigurable processor;

reorganize the sequence of temporal partitions and the sequence of graph control operations within each temporal partition to satisfy the at least one optimization objective;

generate by a finite state machine (FSM), a plurality of hardware states, wherein each hardware state is coupled to unroll a single graph control operation or a plurality of graph control operations to a runtime; and

execute the reorganized dataflow graph on the reconfigurable processor.

2 . The system of claim 1 , wherein the sequence of graph control operations includes two or more of the following:

loading a configuration file;

loading an argument file;

loading an address translation file; and

executing the configuration file.

3 . The system of claim 1 , wherein the IGOEE is configured to reorganize the sequence of graph control operations by combining a subset of graph control operations in the sequence of graph control operations into a single operation.

4 . The system of claim 1 , wherein the IGOEE is configured to reorganize the sequence of temporal partitions by pipelining consecutive temporal partitions.

5 . The system of claim 1 , wherein the IGOEE is configured to execute the reorganized dataflow graph on the reconfigurable processor by allocating a subset of reconfigurable processing units within the reconfigurable processor to the reorganized dataflow graph; and loading the reorganized dataflow graph into the allocated subset of reconfigurable processing units.

6 . The system of claim 1 , wherein each graph control operation includes a software (SW) operation having a SW setup latency equal to a time required for iterating & updating through the array of reconfigurable units to start a HW operation.

7 . The system of claim 6 , wherein minimizing for execution time of the reconfigurable processor includes reorganizing the sequence of graph control operations to have a minimum possible SW setup latency.

8 . The system of claim 1 , wherein each graph control operation includes a HW operation having a HW execution latency equal to an execution time including a time required to push operation-related data to or pull operation-related data from a memory and a total time required by the processor to start and complete the HW operation.

9 . The system of claim 8 , wherein minimizing for execution time of the reconfigurable processor includes reorganizing the sequence of graph control operations to have a minimum possible HW execution latency.

10 . A method of managing executing a dataflow graph of a user application on a reconfigurable processor comprising an array of reconfigurable units, the method comprising:

receiving a dataflow graph of a user application from a complier, wherein the dataflow graph includes a sequence of temporal partitions, and wherein each temporal partition includes a sequence of graph control operations;

receiving at least one optimization objective from the complier, wherein the at least one optimization objective specifies at least one of: minimizing an execution time of the reconfigurable processor and maximizing a computing resource utilization of the reconfigurable processor;

reorganizing the sequence of temporal partitions and the sequence of graph control operations within each temporal partition to satisfy the at least one optimization objective;

generating by a finite state machine (FSM), a plurality of hardware states and unrolling by each hardware state, a single graph control operation or a plurality of graph control operations to a runtime; and executing the reorganized dataflow graph on the reconfigurable processor.

11 . The method of claim 10 , wherein the sequence of graph control operations includes two or more of the following:

loading a configuration file;

loading an argument file;

loading an address translation file; and

executing the configuration file.

12 . The method of claim 10 , wherein reorganizing the sequence of graph control operations includes combining a subset of graph control operations in the sequence of graph control operations into a single operation.

13 . The method of claim 10 , wherein reorganizing the sequence of temporal partitions includes pipelining consecutive temporal partitions.

14 . The method of claim 10 , wherein executing the reorganized dataflow graph on the reconfigurable processor further comprises:

allocating a subset of reconfigurable processing units within the reconfigurable processor to the reorganized dataflow graph; and

loading the reorganized dataflow graph into the allocated subset of reconfigurable processing units.

15 . The method of claim 10 , wherein each graph control operation includes a software (SW) operation having a SW setup latency equal to a time required for iterating & updating through the array of reconfigurable units to start a HW operation.

16 . The method of claim 15 , wherein minimizing for execution time of the reconfigurable processor includes reorganizing the sequence of graph control operations to have a minimum possible SW setup latency.

17 . The method of claim 10 , wherein each graph control operation includes a HW operation having a HW execution latency equal to an execution time including a time required to push operation-related data to or pull operation-related data from a memory and a total time required by the processor to start and complete the HW operation.

18 . The method of claim 17 , wherein minimizing for execution time of the reconfigurable processor includes reorganizing the sequence of graph control operations to have a minimum possible HW execution latency.

19 . A non-transitory computer readable medium having instructions encoded thereon for a data processing system comprising a coarse-grained reconfigurable (CGR) processor including an array of CGR unit reconfigurable units, the instructions configured to cause the processor to conduct a method comprising:

receiving a dataflow graph of a user application from a complier, wherein the dataflow graph includes a sequence of temporal partitions, and wherein each temporal partition includes a sequence of graph control operations;

receiving at least one optimization objective from the complier, wherein the at least one optimization objective specifies at least one of: minimizing an execution time of the reconfigurable processor and maximizing a computing resource utilization of the reconfigurable processor;

reorganizing the sequence of temporal partitions and the sequence of graph control operations within each temporal partition to satisfy the at least one optimization objective;

generating by a finite state machine (FSM), a plurality of hardware states and unrolling by each hardware state, a single graph control operation or a plurality of graph control operations to a runtime; and executing the reorganized dataflow graph on the reconfigurable processor.

20 . The non-transitory computer readable medium of claim 19 , wherein the sequence of graph control operations includes two or more of the following:

loading a configuration file;

loading an argument file;

loading an address translation file; and

executing the configuration file.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2024
From: GOEL, ARNAV; KUMAR, RAVINDER; SABNIS, ARJUN; ZHENG, QI; SANGHVI, NEAL
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 066278/0510 →
Continuity (2)
Provisional Application 63458315 · Apr 10, 2023
Related Publication 20240338340A1 · Oct 10, 2024
References Cited (24)
US 7895586B2 · Ozone · 2011 [cited by examiner]
US 7953956B2 · Okada · 2011 [cited by examiner]
US 10007746B1 · Wolfovitz · 2018 [cited by examiner]
US 10025566B1 · Ahmed · 2018 [cited by examiner]
US 10698853B1 · Grohoski et al. · 2020 [cited by applicant]
US 10831507B2 · Shah et al. · 2020 [cited by applicant]
US 11003429B1 · Zejda · 2021 [cited by examiner]
US 11386038B2 · Prabhakar et al. · 2022 [cited by applicant]
US 11593112B2 · Shah · 2023 [cited by examiner]
US 11709664B2 · Chen et al. · 2023 [cited by applicant]
US 11809908B2 · Kumar et al. · 2023 [cited by applicant]
US 20170039048A1 · Gschwind · 2017 [cited by examiner]
US 20170068764A1 · Takashina · 2017 [cited by examiner]
US 20220345535A1 · Enrici · 2022 [cited by examiner]
WO 2010142987A1 · 2010 [cited by applicant]
Yin, C et al., A Rescheduable Dataflow-SIMD Execution for Increased Utilization in CGRA Cross-Domain Acceleration, Jun. 2022, IEEE pp. 874-886. (Year: 2022). [cited by examiner]
Wijerathne, D et al., HiMap: Fast and Scalable High Quality Mapping on CGRA via Hierarchical Abstraction, Dec. 2021, IEEE, pp. 3290-3303. (Year: 2021). [cited by examiner]
Sankaralingam, K et al., The Mozart Resue Exposed Dataflow Processor for AI and Beyond, 2022, ACM, pp. 978-992 (Year: 2022). [cited by examiner]
Zhang, Y, et al., SARA: Scaling a Reconfigurable Dataflow Accelerator , 2021, IEEE, pp. 1041-1054. (Year: 2021). [cited by examiner]
Zhao, Z et al., Towards Higher Performance and Robust Compilation for CGRA Modulo Scheduling, 2020, IEEE, pp. 2201-2219. (Year: 2020). [cited by examiner]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]