IP Library Granted Patent US 12,413,530
Granted Patent B2
US 12,413,530 · App. 18/740,240 · Granted Sep 9, 2025

Data processing system with link-based resource allocation for reconfigurable processors

Inventors: Raghunath Shenbagam (San Jose, CA); Ravinder Kumar (Fremont, CA)
Assignee: SambaNova Systems, Inc.
H04L47/28G06F9/44505G06F9/45558G06F9/5044G06F9/5077G06F15/7892H04L41/0896H04L49/10G06F2009/4557G06F2009/45595G06F2209/5011G06F2213/0062G06F2213/0064H04L41/0895H04L45/745
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,413,530
App. No.
18/740,240
Granted
Sep 9, 2025
Kind
B2
Abstract

The disclosed technology relates to link-based resource allocation for a pool of reconfigurable processors. Resource allocation is provided for reconfigurable processors based on link bandwidths and link latencies. Runtime logic receives target link bandwidth and target link latency and rated link bandwidth and rated link latency. In response, the runtime logic allocates configuration files for an application, reconfigurable processors, and links between the processors. The runtime logic executes the allocated configuration files using the allocated processors and the allocated links. In another embodiment, the pool of reconfigurable processors comprise a cluster of processing nodes connected through a network.

Claims (44)

1. A data processing system, comprising:

a pool of reconfigurable processors operatively coupled by links, the links having rated link bandwidths and rated link latencies;

runtime logic operatively coupled to the pool of reconfigurable processors, the runtime logic configured to receive, for a first application, a first plurality of configuration files containing configuration data;

a first configuration of a first plurality of virtual reconfigurable processors required to execute the first application;

virtual links between the virtual reconfigurable processors in the first plurality of virtual reconfigurable processors; and

a first specification of target link bandwidths and target link latencies of the virtual links between the virtual reconfigurable processors in the first plurality of virtual reconfigurable processors;

using the runtime logic to allocate the reconfigurable processors in the plurality of reconfigurable processors to the virtual reconfigurable processors in the first plurality of virtual reconfigurable processors;

using the runtime logic to allocate links between the reconfigurable processors to the virtual links between the virtual reconfigurable processors in the first plurality of virtual reconfigurable processors based on a link bandwidth comparison that compares the target link bandwidths, specified by the first specification, against the rated link bandwidths, and a link latency comparison that compares the target link latencies, specified by the first specification, against the rated link latencies; and

using the runtime logic to allocate the reconfigurable processors and the links with configuration data in the first plurality of configuration files, forming configured reconfigurable processors and configured links; and

executing the first application using the configured reconfigurable processors and the configured links.

2. The data processing system of claim 1 , wherein the runtime logic is further configured to receive, for a second application, a second plurality of configuration files that contain configuration data;

a second configuration of a second plurality of virtual reconfigurable processors required to execute the second application;

virtual links between virtual reconfigurable processors in the second plurality of virtual reconfigurable processors; and

a second specification of target link bandwidths and target link latencies of the virtual links between the virtual reconfigurable processors in the second plurality of virtual reconfigurable processors.

3. The data processing system of claim 1 , wherein the pool of reconfigurable processors comprise a cluster of processing nodes connected through a network.

4. The data processing system of claim 1 , wherein one or more reconfigurable processors are host processors for controlling dataflow through a network.

5. The data processing system of claim 3 , wherein the pool of reconfigurable processors is associated with a single processing node in a network.

6. The data processing system of claim 3 , wherein the pool of reconfigurable processors is associated with multiple processing nodes in a network.

7. The data processing system of claim 3 , where the allocated reconfigurable processors are on a same processing node.

8. The data processing system of claim 3 , where the allocated reconfigurable processors are on a different processing node.

9. The data processing system of claim 1 , wherein the pool of reconfigurable processors is dynamically scalable to meet the performance requirements of applications requesting execution.

10. The system of claim 1 , wherein the processor elements are respective arrays of configurable units.

11. The system of claim 5 , wherein the reconfigurable processors are pattern compute units (PCUs) and pattern memory units (PMUs).

12. A computer-implemented method, comprising:

providing a pool of reconfigurable processors operatively coupled by links, the links having rated link bandwidths and rated link latencies;

providing runtime logic operatively coupled to the pool of reconfigurable processors, the runtime logic configured to receive, for a first application, a first plurality of configuration files containing configuration data;

providing a first configuration of a first plurality of virtual reconfigurable processors required to execute the first application;

providing virtual links between the virtual reconfigurable processors in the first plurality of virtual reconfigurable processors; and

providing a first specification of target link bandwidths and target link latencies of the virtual links between the virtual reconfigurable processors in the first plurality of virtual reconfigurable processors;

using the runtime logic, allocating reconfigurable processors in the plurality of reconfigurable processors to the virtual reconfigurable processors in the first plurality of virtual reconfigurable processors;

using the runtime logic, allocating links between the reconfigurable processors to the virtual links between the virtual reconfigurable processors in the first plurality of virtual reconfigurable processors based on a link bandwidth comparison that compares the target link bandwidths, specified by the first specification, against the rated link bandwidths, and a link latency comparison that compares the target link latencies, specified by the first specification, against the rated link latencies;

using the runtime logic, allocating the reconfigurable processors and the links with configuration data in the first plurality of configuration files, forming configured reconfigurable processors and configured links; and

executing the first application using the configured reconfigurable processors and the configured links.

13. The computer-implemented method of claim 12 , wherein the runtime logic is further configured to receive, for a second application, a second plurality of configuration files that contain configuration data;

a second configuration of a second plurality of virtual reconfigurable processors required to execute the second application;

virtual links between virtual reconfigurable processors in the second plurality of virtual reconfigurable processors; and

a second specification of target link bandwidths and target link latencies of the virtual links between the virtual reconfigurable processors in the second plurality of virtual reconfigurable processors.

14. The computer-implemented method of claim 12 , wherein the bandwidth and latency are achievable bandwidth and achievable latency.

15. The computer-implemented method of claim 12 , wherein the pool of reconfigurable processors comprise a cluster of processing nodes connected through a network.

16. The computer-implemented method of claim 12 , wherein one or more reconfigurable processors are host processors for controlling dataflow through a network.

17. The computer-implemented method of claim 15 , wherein the pool of reconfigurable processors is a single processing node or multiple processing nodes coupled to a plurality of reconfigurable processors.

18. The computer-implemented method of claim 15 , wherein the allocated reconfigurable processors are on a same processing node.

19. The computer-implemented method of claim 15 , where the allocated reconfigurable processors are on a different processing node.

20. The computer-implemented method of claim 12 , wherein the pool of reconfigurable processors is dynamically scalable to meet the performance requirements of applications requesting execution.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2024
From: SHENBAGAM, RAGHUNATH; KUMAR, RAVINDER
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 068596/0699 →
Continuity (3)
Continuation 17528081 · Nov 16, 2021
Continuation 17214768 · Mar 26, 2021
Related Publication 20240330074A1 · Oct 3, 2024
References Cited (200)
US 5684980A · Casselman · 1997 [cited by applicant]
US 6470485B1 · Cote et al. · 2002 [cited by applicant]
US 6539438B1 · Ledzius et al. · 2003 [cited by applicant]
US 6667983B1 · Lo et al. · 2003 [cited by applicant]
US 8745626B1 · Sandstrom · 2014 [cited by examiner]
US 9009723B2 · Degenaro et al. · 2015 [cited by applicant]
US 9501325B2 · Pell et al. · 2016 [cited by applicant]
US 10511479B2 · Xie et al. · 2019 [cited by applicant]
US 10621138B2 · Hu et al. · 2020 [cited by applicant]
US 10698853B1 · Grohoski et al. · 2020 [cited by applicant]
US 10768899B2 · Koeplinger et al. · 2020 [cited by applicant]
US 10802870B2 · Lu · 2020 [cited by applicant]
US 10831507B2 · Shah et al. · 2020 [cited by applicant]
US 10831523B2 · Kochevar-Cureton et al. · 2020 [cited by applicant]
US 10877822B1 · Wang et al. · 2020 [cited by applicant]
US 11055141B2 · Prabhakar et al. · 2021 [cited by applicant]
US 11068780B2 · Mellempudi et al. · 2021 [cited by applicant]
US 11080227B2 · Koeplinger et al. · 2021 [cited by applicant]
US 11126574B1 · Prabhakar et al. · 2021 [cited by applicant]
US 11150872B2 · Wang et al. · 2021 [cited by applicant]
US 11182221B1 · Sivaramakrishnan et al. · 2021 [cited by applicant]
US 11182264B1 · Sivaramakrishnan et al. · 2021 [cited by applicant]
US 11184439B2 · Eran et al. · 2021 [cited by applicant]
US 11188497B2 · Shah et al. · 2021 [cited by applicant]
US 11200096B1 · Shenbagam et al. · 2021 [cited by applicant]
US 11237880B1 · Raumann et al. · 2022 [cited by applicant]
US 11237971B1 · Brown et al. · 2022 [cited by applicant]
US 11250105B2 · Wang et al. · 2022 [cited by applicant]
US 11327713B2 · Wang et al. · 2022 [cited by applicant]
US 11327717B2 · Wang et al. · 2022 [cited by applicant]
US 11327923B2 · Wang et al. · 2022 [cited by applicant]
US 11328038B2 · Wang et al. · 2022 [cited by applicant]
US 11328207B2 · Lauterbach et al. · 2022 [cited by applicant]
US 11328208B2 · Lie et al. · 2022 [cited by applicant]
US 11347965B2 · Dutta et al. · 2022 [cited by applicant]
US 11360800B2 · Kochevar-Cureton et al. · 2022 [cited by applicant]
US 11386038B2 · Prabhakar et al. · 2022 [cited by applicant]
US 11392740B2 · Raumann et al. · 2022 [cited by applicant]
US 11410027B2 · Chen et al. · 2022 [cited by applicant]
US 11436429B2 · Jaganathan et al. · 2022 [cited by applicant]
US 11645057B2 · Koeplinger et al. · 2023 [cited by applicant]
US 11709664B2 · Chen et al. · 2023 [cited by applicant]
US 11782729B2 · Grohoski et al. · 2023 [cited by applicant]
US 11809908B2 · Kumar et al. · 2023 [cited by applicant]
US 11836629B2 · Liu · 2023 [cited by applicant]
US 11886930B2 · Sivaramakrishnan et al. · 2024 [cited by applicant]
US 20020156998A1 · Casselman · 2002 [cited by applicant]
US 20030108119A1 · Mohebbi et al. · 2003 [cited by applicant]
US 20060012395A1 · Huppenthal et al. · 2006 [cited by applicant]
US 20060015712A1 · Ang et al. · 2006 [cited by applicant]
US 20070186126A1 · Smith et al. · 2007 [cited by applicant]
US 20070220522A1 · Coene et al. · 2007 [cited by applicant]
US 20080013448A1 · Horie et al. · 2008 [cited by applicant]
US 20090089475A1 · Chitlur · 2009 [cited by applicant]
US 20090172351A1 · Vorbach et al. · 2009 [cited by applicant]
US 20090300209A1 · Elzur · 2009 [cited by applicant]
US 20140137123A1 · Hartmann et al. · 2014 [cited by applicant]
US 20140258438A1 · Ayoub et al. · 2014 [cited by applicant]
US 20150058614A1 · Degenaro et al. · 2015 [cited by applicant]
US 20150100971A1 · Dube et al. · 2015 [cited by applicant]
US 20150106823A1 · Canoy et al. · 2015 [cited by applicant]
US 20160308719A1 · Putnam et al. · 2016 [cited by applicant]
US 20160314025A1 · Mcgarry et al. · 2016 [cited by applicant]
US 20160378550A1 · Monfort et al. · 2016 [cited by applicant]
US 20170220499A1 · Gray · 2017 [cited by applicant]
US 20170289060A1 · Aftab · 2017 [cited by examiner]
US 20170315815A1 · Smith et al. · 2017 [cited by applicant]
US 20170317679A1 · Suh et al. · 2017 [cited by applicant]
US 20180285295A1 · Abel et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180308200A1 · Surti et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20190089616A1 · Chabbi et al. · 2019 [cited by applicant]
US 20190138890A1 · Liang et al. · 2019 [cited by applicant]
US 20190171604A1 · Brewer · 2019 [cited by applicant]
US 20190171612A1 · Shahar et al. · 2019 [cited by applicant]
US 20190180176A1 · Yudanov et al. · 2019 [cited by applicant]
US 20190258921A1 · Lie et al. · 2019 [cited by applicant]
US 20190286973A1 · Kovvuri et al. · 2019 [cited by applicant]
US 20190347136A1 · Miyoshi · 2019 [cited by examiner]
US 20190384642A1 · Bolkhovitin et al. · 2019 [cited by applicant]
US 20200090313A1 · Bugdary et al. · 2020 [cited by applicant]
US 20200142753A1 · Harwood et al. · 2020 [cited by applicant]
US 20200142857A1 · Catiller et al. · 2020 [cited by applicant]
US 20200151573A1 · Das et al. · 2020 [cited by applicant]
US 20200174840A1 · Zhao et al. · 2020 [cited by applicant]
US 20200183745A1 · Ernst et al. · 2020 [cited by applicant]
US 20200226444A1 · Sharma et al. · 2020 [cited by applicant]
US 20200264876A1 · Lo et al. · 2020 [cited by applicant]
US 20200301898A1 · Samynathan et al. · 2020 [cited by applicant]
US 20200314181A1 · Eran · 2020 [cited by applicant]
US 20200326992A1 · Jin et al. · 2020 [cited by applicant]
US 20200341930A1 · Cannata et al. · 2020 [cited by applicant]
US 20210011770A1 · Prabhakar et al. · 2021 [cited by applicant]
US 20210081691A1 · Chen et al. · 2021 [cited by applicant]
US 20210089343A1 · Hyoudou · 2021 [cited by applicant]
US 20210097366A1 · Wagner et al. · 2021 [cited by applicant]
US 20210097379A1 · Yang et al. · 2021 [cited by applicant]
US 20210103820A1 · Ghosh · 2021 [cited by applicant]
US 20210125058A1 · Chowdhury et al. · 2021 [cited by applicant]
US 20210192357A1 · Sinha et al. · 2021 [cited by applicant]
US 20210192358A1 · Song et al. · 2021 [cited by applicant]
US 20210200610A1 · Chu et al. · 2021 [cited by applicant]
US 20210241093A1 · Byrne et al. · 2021 [cited by applicant]
US 20210374503A1 · Kim et al. · 2021 [cited by applicant]
US 20220058034A1 · Grohoski et al. · 2022 [cited by applicant]
US 20220121928A1 · Dong · 2022 [cited by examiner]
US 20220197712A1 · Sivaramakrishnan et al. · 2022 [cited by applicant]
US 20220197714A1 · Raumann et al. · 2022 [cited by applicant]
US 20220198117A1 · Raumann et al. · 2022 [cited by applicant]
US 20220269534A1 · Misra et al. · 2022 [cited by applicant]
US 20220308935A1 · Shenbagam et al. · 2022 [cited by applicant]
EP 1372084A2 · 2003 [cited by applicant]
JP 2020112901A · 2020 [cited by applicant]
TW 201606508A · 2016 [cited by applicant]
TW 202240386A · 2022 [cited by applicant]
TW 202240394A · 2022 [cited by applicant]
TW 202248853A · 2022 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
WO 2022133047A1 · 2022 [cited by applicant]
WO 2022182573A1 · 2022 [cited by applicant]
WO 2022203925A1 · 2022 [cited by applicant]
Accelerated Computing with a Reconfigurable Dataflow Architecture, SambaNova Systems Whitepaper, 10 pages. [cited by applicant]
Bae et al., Auto-Tuning CNNs for Coarse-Grained Reconfigurable Array-based Accelerators, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 37, Issue: 11, Nov. 2018, 10 pages. [cited by applicant]
Busa et al., A Run-Time Word-Level Reconfigurable Coarse-Grain Functional Unit for a VLIW processor; ACM 2002, pp. 44-49. (Year: 2002). [cited by applicant]
Dettmers, How to Parallelize Deep Learning on GPUs Part 1 of 2: Data Parallelism, dated Oct. 9, 2014, 19 pages. Retrieved on Sep. 3, 2021 Retrieved from [URL: https://timdettmers.com/2014/10/09/deep-learning-data-parall… [cited by applicant]
Dettmers, How to Parallelize Deep Learning on GPUs Part 2 of 2: Model Parallelism, dated Nov. 9, 2014, 19 pages. Retrieved on Sep. 3, 2021. Retrieved from [URL: https://timdettmers.com/2014/11/09/model-parallelism-deep-… [cited by applicant]
Donges, Gradient Descent: An Introduction to Machine Learning's Most Popular Algorithms, dated Jun. 16, 2019, 10 pages. Retrieved on Mar. 24, 2021, retrieved from [URL: https://builtin.com/data-science/gradient-descent … [cited by applicant]
Ekanayake, Model Parallelism in Deep Learning is NOT What you think, dated Nov. 10, 2018, 4 pages. Retrieved an Sep. 3, 2021. Retrieved from [ URL: https://medium.com/@esaliya/model-parallelism-in-deep-learning-is-not-w… [cited by applicant]
Éricles Sousa, A reconfigurable memory architecture for system integration of coarse-grained reconfigurable arrays, Published in: 2017 International Conference on ReConFigurable Computing and FPGAs (ReConFig) Dec. 4-6, … [cited by applicant]
Fazlali et al., Efficient task scheduling for runtime reconfigurable systems, Journal of Systems Architecture, vol. 56, dated Jul. 26, 2010, pp. 623-632, 10 pages. [cited by applicant]
Galanis et al., A design flow for speeding-up dsp applications in heterogeneous reconfigurable systems, Microelectronics Journal, vol. 37, dated 2006, pp. 554-564, 11 pages. [cited by applicant]
Galanis et al., Accelerating Applications by Mapping Critical Kernels on Coarse-Grain Reconfigurable Hardware in Hybrid Systems, Field-Programmable Custom Computing Machines, 2005, 13th Annual IEEE Symposium on Napa, CA… [cited by applicant]
Galanis et al., Partitioning Methodology for Heterogeneous Reconfigurable Functional Units, The Journal of Supercomputing, vol. 38, No. 1, dated Oct. 1, 2006, 18 pages. [cited by applicant]
Goodfellow et. al., Deep Learning Book Chapter 6 Deep Feedforward Networks, 2016, 60 pages. [cited by applicant]
Insujang, GPU Architecture Overview, Better Tomorrow with Computer Science, published Apr. 27, 2017, retrieved on Jun. 17, 2021, retrieved from the Internet [ URL: https://insujang.github.io/2017-04-17/gpu-architecture-… [cited by applicant]
Iqbal et al., Reconfigurable Processor Architecture for High Speed Applications, IEEE, dated 2009, pp. 624-629. [cited by applicant]
Jackson et al., PCI Express Technology Comprehensive Guide to Generation 1.x, 2.x and 3.0, dated Jun. 2020, 1057 pages. [cited by applicant]
Jafri et al., NeuroCGRA: A CGRAs with Support for Neural Networks, 2014 International Conference on High Performance Computing & Simulation (HPCS), 8 pages. [cited by applicant]
Jin et. al., How to scale distributed deep learning, dated Nov. 14, 2016, 16 pages. [cited by applicant]
Kachris et al.; “A Survey on Reconfigurable Accelerators for Cloud Computing”, IEEE 2016, Aug. 29, 2016, pp. 1-11. [cited by applicant]
Knodel, Oliver, et. al., “RC3E: Reconfigurable Accelerators in Data Centers and their Provision by Adapted Service Models”, IEEE 9th International Converence on Cloud Computing, 2016, pp. 1-8. [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
Lecture 11: Distributed Training and Communication Protocols, CSE599W: Spring 2018, UW Paul G. Allen School of Computer Science and Engineering, 41 pages. [cited by applicant]
Li, Ang, et. al., “Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect”, Mar. 11, 2019, 15 pages. [cited by applicant]
Li, et al., “CATERPILLAR: Coarse Grain Reconfigurable Architecture for Accelerating the Training of Deep Neural Networks,” arXiv: 1706.00517v2 [cs.DC], Jun. 8, 2017, 10 pages. [cited by applicant]
Liang et al., Dynamic Coarse Grain Dataflow Reconfiguration Technique for Real-Time Systems, IEEE, dated 2005, pp. 3511-3514. [cited by applicant]
Liu et. al., Offloading distributed Applications onto SmartNICs using iPipe, ACM 2019, pp. 1-16. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Ma et al., DeepGauge: Multi-Granularity Testing Criteria for Deep Learning Systems; ACM 2018, pp. 1-12. [cited by applicant]
Mao, Data Parallelism vs Model Parallelism in Distributed Deep Learning Training, dated Mar. 23, 2019, 4 pages, retrieved on Mar. 30, 2021, Retrieved from the internet [ URL: https://leimao.github.io]. [cited by applicant]
Marshall, Dave, “Remote Procedure Calls (RPC)”, Jan. 5, 1999, 15 pages, Retreived from URL. [cited by applicant]
Mazur, A step by step backpropagation example, dated Mar. 17, 2015, 26 pages. Retrieved on Sep. 3, 2021. Retrieved from [URL: https://mattmazur.com/2015/03/17/a-step-by-step-backpropagation-example/ ]. [cited by applicant]
NVIDIA, “NVIDIA DGX-1 System Architecture”, WP-08437-001_v02, 2017, 33 pages. [cited by applicant]
NVIDIA, “NVIDIA DGX-1 With Tesla V100 System Architecture”, WP-08437-002_v01, 2017, 43 pages. [cited by applicant]
NVIDIA, “NVIDIA Tesla P100”, WP-08019-001 v01.1, 2016, 45 pages. [cited by applicant]
NVIDIA, “NVIDIA Turing GPU Architecture”, WP-09183-001_v01, 2018, 86 pages. [cited by applicant]
Padole et al., Configuration Memory Based Dynamic Coarse Grained Reconfigurable Multiscore Architecture, IEEE 2013, pp. 3511-3514, 5 pages. [cited by applicant]
Paek et al., “Binary Acceleration Using Coarse-Grained Reconfigurable Architecture,” ACM SIGARCH Computer Architecture News, vol. 38, No. 4, Sep. 2010, 7 pages. [cited by applicant]
PCT/US2021/063728—International Search Report and Written Opinion, dated Apr. 4, 2022, 15 pages. [cited by applicant]
PCT/US2021/063733—International Search Report and Written Opinion, dated Apr. 4, 2022, 17 pages. [cited by applicant]
PCT/US2022/016871—International Search Report and Written Opinion, dated Jun. 1, 2022, 14 pages. [cited by applicant]
PCT/US2022/020638—International Search Report and Written Opinion, dated Jun. 21, 2022, 17 pages. [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Rubattu et al., Dataflow-Functional High-Level Synthesis for Coarse-Grained Reconfigurable Accelerators, IEEE 2019, pp. 69-72 (Year: 2019). [cited by applicant]
Ruder, An overview of gradient descent optimization algorithms, NUI Galway Aylien Lyd, dated Jun. 15, 2017, 14 pages. [cited by applicant]
Souissi et al., Optimization of Run-time Mapping on Heterogeneous CPU/FPGA Architecture, 9th International Conference of Modeling, Optimization and Simulation—MOSIM'12, Jun. 6-8, 2012, Bordeaux, France, 9 pages. [cited by applicant]
Strom, Scalable Distributed DNN Training Using Commodity GPU Cloud Computing, Amazon.com, 5 pages. [cited by applicant]
Tanaka et. al., Distributed Deep Learning with GPU-FPGA heterogenous computing, IEEE 2021, 9 pages. [cited by applicant]
TW 110147198, Allowance Decision, Jul. 27, 2022, 2 pages. [cited by applicant]
TW 110147198, Search Report, Jul. 27, 2022, 1 page. [cited by applicant]
TW 111110601, Non-final office action, dated Oct. 18, 2022, 3 pages. [cited by applicant]
TW 111110601, Search Report, Oct. 18, 2022, 1 page. [cited by applicant]
U.S. Appl. No. 17/127,818 Notice of Allowance, dated Jul. 21, 2021, 10 pages. [cited by applicant]
U.S. Appl. No. 17/127,818—Office Action dated Apr. 1, 2021, 15 pages. [cited by applicant]
U.S. Appl. No. 17/127,929 Notice of Allowance, dated Jul. 21, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/127,929—Office Action dated Apr. 1, 2021, 26 pages. [cited by applicant]
U.S. Appl. No. 17/127,818—Response to Office Action dated Apr. 1, 2021, filed Jul. 1, 2021, 15 pages. [cited by applicant]
U.S. Appl. No. 17/127,929—Response to Office Action dated Apr. 1, 2021, filed Jul. 1, 2021, 10 pages. [cited by applicant]
U.S. Appl. No. 17/185,264—Non-Final Office Action, dated Jan. 26, 2023, 9 pages. [cited by applicant]
U.S. Appl. No. 17/185,264—Notice of Allowance, dated May 30, 2023, 7 pages. [cited by applicant]
U.S. Appl. No. 17/214,768—Notice of Allowance, dated Aug. 11, 2021, 26 pages. [cited by applicant]
U.S. Appl. No. 17/214,768—Supplemental Notice of Allowance, dated Aug. 25, 2021, 10 pages. [cited by applicant]
U.S. Appl. No. 17/379,921—Notice of Allowance, dated Nov. 26, 2021, 21 pages. [cited by applicant]
U.S. Appl. No. 17/379,921 Notice of Allowance dated Mar. 21, 2022, 25 pages. [cited by applicant]
U.S. Appl. No. 17/379,924—Notice of Allowance, dated Sep. 16, 2021, 23 pages. [cited by applicant]
U.S. Appl. No. 17/522,655—Notice of Allowance, dated Nov. 16, 2022, 21 pages. [cited by applicant]
U.S. Appl. No. 17/522,658—Notice of Allowance, dated Dec. 14, 2022, 22 pages. [cited by applicant]
U.S. Appl. No. 17/522,682—Notice of Allowance, dated Jan. 11, 2023, 23 pages. [cited by applicant]
U.S. Appl. No. 17/522,682—Supplemental Notice of Allowance, dated Jan. 24, 2023, 2 pages. [cited by applicant]
U.S. Appl. No. 17/522,694—Non-Final Office Action, dated Mar. 31, 2023, 29 pages. [cited by applicant]
U.S. Appl. No. 17/522,694—Notice of Allowance, dated Sep. 6, 2023, 62 pages. [cited by applicant]
U.S. Appl. No. 17/528,081—Non-Final Office Action, dated Jun. 7, 2023, 29 pages. [cited by applicant]
Vucha et al., Dynamic Task Distribution Model for On-Chip Reconfigurable High Speed Computing System, Hindawi, dated Jun. 30, 2015, 13 pages. [cited by applicant]
What is the difference between model paralellism and data paralellism, Quora, 27 pages. Retrieved on Sep. 3, 2021. Retrieved from [URL: https://www.quora.com/What-is-the-difference-between-model-parallelism-and-data-par… [cited by applicant]
Woolloy, NCCL: Accelerated Multi-GPU Collective Communications, NVIDIA, 56 pages. [cited by applicant]
Xiandong Qi, Introduction to Distributed Deep Learning, dated May 13, 2017, 13 pages. [cited by applicant]
Zhang et. al., Dive into Deep Learning, Release 0.16.2, dated Mar. 20, 2021, 1027 pages. [cited by applicant]