IP Library Granted Patent US 12,346,729
Granted Patent B2
US 12,346,729 · App. 18/211,962 · Granted Jul 1, 2025

Runtime virtualization of reconfigurable data flow resources

Inventors: Ravinder Kumar (Palo Alto, CA); Conrad Alexander Turlik (Palo Alto, CA); Arnav Goel (Palo Alto, CA); Qi Zheng (Palo Alto, CA); Raghunath Shenbagam (San Jose, CA); Anand Misra (Palo Alto, CA); Ananda Reddy Vayyala (San Jose, CA); Pushkar Shridhar Nandkar (Hayward, CA)
Assignee: SambaNova Systems, Inc.
G06F9/5011G06F9/5016G06F9/5077G06F15/7867G06F15/7871G06F2209/501G06F2209/5011
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,346,729
App. No.
18/211,962
Granted
Jul 1, 2025
Kind
B2
Abstract

A data processing system comprises a pool of reconfigurable data flow resources and a runtime processor. The pool of reconfigurable data flow resources includes arrays of physical configurable units and memory. The runtime processor includes logic to receive a plurality of configuration files for user applications. The configuration files include configurations of virtual data flow resources required to execute the user applications. The runtime processor also includes logic to allocate physical configurable units and memory in the pool of reconfigurable data flow resources to the virtual data flow resources and load the configuration files to the allocated physical configurable units. The runtime processor further includes logic to execute the user applications using the allocated physical configurable units and memory.

Claims (13)

1. A system, comprising:

a plurality of reconfigurable data flow resources, reconfigurable data flow resources in the plurality of reconfigurable data flow resources including a plurality of reconfigurable processors, reconfigurable processors in the plurality of reconfigurable processors including an array of configurable units, and the array of configurable units partitionable into a plurality of subarrays of configurable units;

a plurality of transfer resources usable by the reconfigurable data flow resources to receive and send to a plurality of storage resources usable by the reconfigurable data flow resources to store data; and

a runtime processor configured with logic to:

present a unified interface to the plurality of reconfigurable data flow resources, the plurality of transfer resources, and the plurality of storage resources, wherein the unified interface enables attachment of one of the configurable units to every other one of the configurable units;

control execution of a plurality of application graphs based on an execution file wherein the application graphs are representations of how the configurable units interact to exchange data to provide data flow with each other through the unified interface, the execution file including configuration files for application graphs in the plurality of application graphs, topologies for subarrays of configurable units in the plurality of subarrays of configurable units with the topologies indicating how to load the configuration files so that the configurable units interact to exchange data according to the application graphs, and resource requests for transfer resources in the plurality of transfer resources and storage resources in the plurality of storage resources required to satisfy data flow for the topologies according to the application graphs;

allocate the subarrays of configurable units to the application graphs based on the topologies;

allocate the transfer resources and the storage resources to the application graphs based on the resource requests; and

load and execute the configuration files using the allocated subarrays of configurable units, transfer resources, and storage resources.

2. The system of claim 1 , wherein the topologies specify a set of two or more subarrays of configurable units of a single reconfigurable processor along a vertical and horizontal orientation of the set of subarrays of configurable units.

3. The system of claim 1 , wherein the topologies specify a set of subarrays of configurable units spanning two or more reconfigurable processors.

4. The system of claim 1 , wherein the runtime processor allocates one or more subarrays of configurable units of a single reconfigurable processor to two or more configuration files of two or more application graphs based on the topologies according to the application graphs, and wherein a device driver concurrently loads and executes the two or more configuration files on the subarrays of the single reconfigurable processor.

5. The system of claim 1 , wherein the runtime processor allocates subarrays of two or more reconfigurable processors to a single configuration file of a single application graph based on the topologies, and wherein a device driver concurrently loads and executes the single configuration file on the subarrays of the two or more reconfigurable processors.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 2, 2026
From: NANDKAR, PUSHKAR SHRIDHAR
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 073350/0060 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2023
From: KUMAR, RAVINDER; TURLIK, CONRAD ALEXANDER; GOEL, ARNAV; ZHENG, QI; SHENBAGAM, RAGHUNATH; MISRA, ANAND; VAYYALA, ANANDA REDDY
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 064004/0313 →
Continuity (2)
Division 16922975 · Jul 7, 2020
Related Publication 20230409395A1 · Dec 21, 2023
References Cited (270)
US 4769790A · Yamashita · 1988 [cited by applicant]
US 5506797A · Koshiba · 1996 [cited by applicant]
US 5560029A · Papadopoulos et al. · 1996 [cited by applicant]
US 5794033A · Aldebert et al. · 1998 [cited by applicant]
US 5963746A · Barker et al. · 1999 [cited by applicant]
US 6105119A · Kerr et al. · 2000 [cited by applicant]
US 6119181A · Vorbach et al. · 2000 [cited by applicant]
US 6256653B1 · Juffa et al. · 2001 [cited by applicant]
US 6470485B1 · Cote et al. · 2002 [cited by applicant]
US 6539438B1 · Ledzius · 2003 [cited by examiner]
US 6667983B1 · Lo et al. · 2003 [cited by applicant]
US 6728871B1 · Vorbach et al. · 2004 [cited by applicant]
US 7015921B1 · Trivedi et al. · 2006 [cited by applicant]
US 7472149B2 · Endo · 2008 [cited by applicant]
US 7734895B1 · Agarwal et al. · 2010 [cited by applicant]
US 7797258B1 · Bowman et al. · 2010 [cited by applicant]
US 7952387B1 · Frazer · 2011 [cited by applicant]
US 7996684B2 · Wasson et al. · 2011 [cited by applicant]
US 8006021B1 · Li et al. · 2011 [cited by applicant]
US 8045546B1 · Bao et al. · 2011 [cited by applicant]
US 8184317B2 · Okamoto · 2012 [cited by applicant]
US 8261042B2 · Kanstein et al. · 2012 [cited by applicant]
US 8271557B1 · Lysaght · 2012 [cited by examiner]
US 9009723B2 · Degenaro et al. · 2015 [cited by applicant]
US 9201899B2 · Nishimura et al. · 2015 [cited by applicant]
US 9335977B2 · Wang et al. · 2016 [cited by applicant]
US 9411532B2 · Vorbach et al. · 2016 [cited by applicant]
US 9411756B2 · Nogueira et al. · 2016 [cited by applicant]
US 9501325B2 · Pell et al. · 2016 [cited by applicant]
US 9569214B2 · Govindu et al. · 2017 [cited by applicant]
US 9690747B2 · Vorbach et al. · 2017 [cited by applicant]
US 9697318B2 · Hutton et al. · 2017 [cited by applicant]
US 9875105B2 · Rozas et al. · 2018 [cited by applicant]
US 9952831B1 · Ross et al. · 2018 [cited by applicant]
US 10037227B2 · Therien et al. · 2018 [cited by applicant]
US 10067911B2 · Gholaminejad et al. · 2018 [cited by applicant]
US 10186011B2 · Nurvitadhi et al. · 2019 [cited by applicant]
US 10331836B1 · Hosangadi et al. · 2019 [cited by applicant]
US 10621138B2 · Hu et al. · 2020 [cited by applicant]
US 10698853B1 · Grohoski et al. · 2020 [cited by applicant]
US 10831507B2 · Shah et al. · 2020 [cited by applicant]
US 11080227B2 · Koeplinger et al. · 2021 [cited by applicant]
US 20030068097A1 · Wilson et al. · 2003 [cited by applicant]
US 20030108119A1 · Mohebbi et al. · 2003 [cited by applicant]
US 20040049672A1 · Nollet · 2004 [cited by examiner]
US 20040088666A1 · Poznanovic et al. · 2004 [cited by applicant]
US 20040153608A1 · Vorbach et al. · 2004 [cited by applicant]
US 20050108503A1 · Sandon et al. · 2005 [cited by applicant]
US 20050160129A1 · Endo · 2005 [cited by applicant]
US 20070180172A1 · Schmidt et al. · 2007 [cited by applicant]
US 20070220522A1 · Coene et al. · 2007 [cited by applicant]
US 20080013448A1 · Horie et al. · 2008 [cited by applicant]
US 20090135739A1 · Hoover et al. · 2009 [cited by applicant]
US 20090172351A1 · Vorbach et al. · 2009 [cited by applicant]
US 20090187756A1 · Nollet et al. · 2009 [cited by applicant]
US 20090300209A1 · Elzur · 2009 [cited by applicant]
US 20120079498A1 · Kim · 2012 [cited by examiner]
US 20120126851A1 · Kelem et al. · 2012 [cited by applicant]
US 20120131257A1 · Rudosky et al. · 2012 [cited by applicant]
US 20120303930A1 · Coker · 2012 [cited by examiner]
US 20130151576A1 · Lutz et al. · 2013 [cited by applicant]
US 20130339564A1 · Nogueira et al. · 2013 [cited by applicant]
US 20140040334A1 · Burgess et al. · 2014 [cited by applicant]
US 20140137123A1 · Hartmann et al. · 2014 [cited by applicant]
US 20140201642A1 · Vicat-Blanc · 2014 [cited by applicant]
US 20140237227A1 · Aizawa et al. · 2014 [cited by applicant]
US 20140258438A1 · Ayoub et al. · 2014 [cited by applicant]
US 20140317628A1 · Kim · 2014 [cited by applicant]
US 20150058614A1 · Degenaro et al. · 2015 [cited by applicant]
US 20150100971A1 · Dube et al. · 2015 [cited by applicant]
US 20150106823A1 · Canoy et al. · 2015 [cited by applicant]
US 20150347192A1 · Blaine et al. · 2015 [cited by applicant]
US 20160012012A1 · Yen et al. · 2016 [cited by applicant]
US 20160134702A1 · Gertner · 2016 [cited by applicant]
US 20160210167A1 · Bolic · 2016 [cited by examiner]
US 20160308719A1 · Putnam et al. · 2016 [cited by applicant]
US 20160314025A1 · Mcgarry et al. · 2016 [cited by applicant]
US 20160335120A1 · Gupta · 2016 [cited by examiner]
US 20170054449A1 · Mani et al. · 2017 [cited by applicant]
US 20170083313A1 · Sankaralingam · 2017 [cited by examiner]
US 20170123794A1 · Chen et al. · 2017 [cited by applicant]
US 20170185564A1 · Toichi et al. · 2017 [cited by applicant]
US 20170195173A1 · Izenberg et al. · 2017 [cited by applicant]
US 20170244982A1 · Fuldseth et al. · 2017 [cited by applicant]
US 20170317678A1 · Coole et al. · 2017 [cited by applicant]
US 20170317679A1 · Suh · 2017 [cited by examiner]
US 20170322774A1 · Zhang · 2017 [cited by applicant]
US 20170322805A1 · Zohar et al. · 2017 [cited by applicant]
US 20180121121A1 · Mehra et al. · 2018 [cited by applicant]
US 20180157465A1 · Bittner et al. · 2018 [cited by applicant]
US 20180157825A1 · Eksten et al. · 2018 [cited by applicant]
US 20180174022A1 · Young · 2018 [cited by applicant]
US 20180189231A1 · Fleming, Jr. et al. · 2018 [cited by applicant]
US 20180220144A1 · Su et al. · 2018 [cited by applicant]
US 20180246834A1 · Catiller · 2018 [cited by applicant]
US 20180275193A1 · Rouge et al. · 2018 [cited by applicant]
US 20180293185A1 · Vembu et al. · 2018 [cited by applicant]
US 20180300181A1 · Hetzel et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180308200A1 · Surti et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180329681A1 · Zhang et al. · 2018 [cited by applicant]
US 20180336052A1 · Iyer · 2018 [cited by examiner]
US 20180349098A1 · Manohararajah · 2018 [cited by applicant]
US 20190004878A1 · Adler · 2019 [cited by examiner]
US 20190042513A1 · Fleming, Jr. et al. · 2019 [cited by applicant]
US 20190042924A1 · Pasca et al. · 2019 [cited by applicant]
US 20190087606A1 · Subhaschandra · 2019 [cited by examiner]
US 20190089616A1 · Chabbi et al. · 2019 [cited by applicant]
US 20190114139A1 · Zhang · 2019 [cited by applicant]
US 20190138890A1 · Liang et al. · 2019 [cited by applicant]
US 20190147323A1 · Li et al. · 2019 [cited by applicant]
US 20190171612A1 · Shahar et al. · 2019 [cited by applicant]
US 20190180176A1 · Yudanov et al. · 2019 [cited by applicant]
US 20190197655A1 · Sun et al. · 2019 [cited by applicant]
US 20190205734A1 · Guntoro · 2019 [cited by applicant]
US 20190213153A1 · Pan et al. · 2019 [cited by applicant]
US 20190258921A1 · Lie et al. · 2019 [cited by applicant]
US 20190279075A1 · Liu et al. · 2019 [cited by applicant]
US 20190286973A1 · Kovvuri et al. · 2019 [cited by applicant]
US 20190317770A1 · Sankaralingam et al. · 2019 [cited by applicant]
US 20200090313A1 · Bugdary et al. · 2020 [cited by applicant]
US 20200125396A1 · Chynoweth et al. · 2020 [cited by applicant]
US 20200151573A1 · Das et al. · 2020 [cited by applicant]
US 20200159544A1 · Shah et al. · 2020 [cited by applicant]
US 20200159692A1 · Shah et al. · 2020 [cited by applicant]
US 20200167309A1 · Nicol · 2020 [cited by applicant]
US 20200174840A1 · Zhao et al. · 2020 [cited by applicant]
US 20200225996A1 · Sharma et al. · 2020 [cited by applicant]
US 20200226444A1 · Sharma et al. · 2020 [cited by applicant]
US 20200241844A1 · Koeplinger et al. · 2020 [cited by applicant]
US 20200241899A1 · Al-Aghbari et al. · 2020 [cited by applicant]
US 20200264876A1 · Lo et al. · 2020 [cited by applicant]
US 20200272882A1 · Lo · 2020 [cited by applicant]
US 20200310994A1 · ChoFleming et al. · 2020 [cited by applicant]
US 20200326992A1 · Jin et al. · 2020 [cited by applicant]
US 20200341812A1 · McClure · 2020 [cited by examiner]
US 20200341930A1 · Cannata et al. · 2020 [cited by applicant]
US 20200356523A1 · Prabhakar et al. · 2020 [cited by applicant]
US 20200371805A1 · Lutz et al. · 2020 [cited by applicant]
US 20210011770A1 · Prabhakar et al. · 2021 [cited by applicant]
US 20210034982A1 · Sather et al. · 2021 [cited by applicant]
US 20210042259A1 · Koeplinger et al. · 2021 [cited by applicant]
US 20210064341A1 · Kuo et al. · 2021 [cited by applicant]
US 20210064372A1 · Sun et al. · 2021 [cited by applicant]
US 20210064568A1 · Wang et al. · 2021 [cited by applicant]
US 20210072955A1 · Mellempudi et al. · 2021 [cited by applicant]
US 20210081691A1 · Chen et al. · 2021 [cited by applicant]
US 20210081769A1 · Chen et al. · 2021 [cited by applicant]
US 20210089343A1 · Hyoudou · 2021 [cited by applicant]
US 20210096816A1 · Wang et al. · 2021 [cited by applicant]
US 20210097366A1 · Wagner et al. · 2021 [cited by applicant]
US 20210097379A1 · Yang et al. · 2021 [cited by applicant]
US 20210103820A1 · Ghosh · 2021 [cited by applicant]
US 20210110066A1 · Liu · 2021 [cited by examiner]
US 20210125058A1 · Chowdhury et al. · 2021 [cited by applicant]
US 20210149634A1 · Wang et al. · 2021 [cited by applicant]
US 20210157550A1 · Wang et al. · 2021 [cited by applicant]
US 20210182021A1 · Wang et al. · 2021 [cited by applicant]
US 20210192357A1 · Sinha et al. · 2021 [cited by applicant]
US 20210192358A1 · Song et al. · 2021 [cited by applicant]
US 20210200610A1 · Chu et al. · 2021 [cited by applicant]
US 20210241093A1 · Byrne et al. · 2021 [cited by applicant]
US 20210263853A1 · Waters et al. · 2021 [cited by applicant]
US 20210265015A1 · Parnaby et al. · 2021 [cited by applicant]
US 20220100680A1 · Chrysos et al. · 2022 [cited by applicant]
US 20220188028A1 · Mesnier et al. · 2022 [cited by applicant]
EP 0733234A · 1995 [cited by applicant]
EP 1372084A2 · 2003 [cited by applicant]
JP 2020112901A · 2020 [cited by applicant]
TW 200736953A · 2007 [cited by applicant]
TW 200801964A · 2008 [cited by applicant]
TW 200928736A · 2009 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
WO 2018100920A1 · 2018 [cited by applicant]
WO 2021067318A1 · 2021 [cited by applicant]
WO 2021108328A1 · 2021 [cited by applicant]
NVIDIA, “NVIDIA Turing GPU Architecture”, WP-09183-001_v01, 2018, 86 pages. [cited by applicant]
Olukotun, Designing Computer Sytems for Software 2.0, ISCA 2018 keynote, Jun. 2018, 49 pages. [cited by applicant]
Paek et al., “Binary Acceleration Using Coarse-Grained Reconfigurable Architecture,” ACM SIGARCH Computer Architecture News, vol. 38, No. 4, Sep. 2010, 7 pages. [cited by applicant]
PCT/US/2021/040382—International Search Report and Written Opinion, dated Nov. 29, 2021, 22 pages. [cited by applicant]
Petersen, “Softmax with cross-entropy,” https://mattpetersen.github.io/softmax-with-cross-entropy, Jun. 25, 2017, 19 pages. [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Rubattu et al., Dataflow-Functional High-Level Synthesis for Coarse-Grained Reconfigurable Accelerators, IEEE 2019, pp. 69-72 (Year: 2019). [cited by applicant]
Ruder, An overview of gradient descent optimization algorithms, NUI Galway Aylien Lyd, dated Jun. 15, 2017, 14 pages. [cited by applicant]
Strom, Scalable Distributed DNN Training Using Commodity GPU Cloud Computing, Amazon.com, 5 pages. [cited by applicant]
Tanaka et al., Distributed Deep Learning with GPU-FPGA heterogenous computing, IEEE 2021, 9 pages. [cited by applicant]
Tanomoto et al., “A CGRA-based Approach for Accelerating Convolutional Neural Networks,” 2015 IEEE 9th International Symposium on Embedded Multicore/Many-core Systems-on-Chip, 2015, pp. 73-80. [cited by applicant]
Tobuschat, et al., “IDAMC: A NoC for mixed criticality systems,” 2013 IEEE 19th International Conference on Embedded and Real-Time Computing Systems and Applications, Taipei, Aug. 19-21, 2013, pp. 149-156. [cited by applicant]
Turkson et al. “Artificial neural network applications in the calibration of spark-ignition engines: An overview,” Engineering Science and Technology, an International Journal, vol. 19, Issue 3, Sep. 2016, 1346-1359. [cited by applicant]
TW 110124802—First Office Action and Search Report dated May 24, 2022, 17 pages. [cited by applicant]
U.S. Appl. No. 16/922,975—Final Office Action, dated Mar. 9, 2023, 23 pages. [cited by applicant]
U.S. Appl. No. 16/922,975—Non-Final Office Action, dated Oct. 27, 2022, 26 pages. [cited by applicant]
U.S. Appl. No. 16/922,975—Notice of Allowance, dated Jul. 3, 2023, 11 pages. [cited by applicant]
U.S. Appl. No. 17/127,818—Notice of Allowance, dated Jul. 21, 2021, 10 pages. [cited by applicant]
U.S. Appl. No. 17/127,818—Office Action dated Apr. 1, 2021, 15 pages. [cited by applicant]
U.S. Appl. No. 17/127,818—Response to Office Action dated Apr. 1, 2021, filed Jul. 1, 2021, 15 pages. [cited by applicant]
U.S. Appl. No. 17/127,929—Notice of Allowance dated Jul. 21, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/127,929—Office Action dated Apr. 1, 2021, 26 pages. [cited by applicant]
U.S. Appl. No. 17/127,929—Response to Office Action dated Apr. 1, 2021, filed Jul. 1, 2021, 10 pages. [cited by applicant]
U.S. Appl. No. 17/214,768—Notice of Allowance, dated Aug. 11, 2021, 26 pages. [cited by applicant]
U.S. Appl. No. 17/214,768—Supplemental Notice of Allowance, dated Aug. 25, 2021, 10 pages. [cited by applicant]
U.S. Appl. No. 17/379,921—Notice of Allowance, dated Nov. 26, 2021, 21 pages. [cited by applicant]
U.S. Appl. No. 17/379,924—Notice of Allowance, dated Sep. 16, 2021, 23 pages. [cited by applicant]
Vadivel et al., “Loop Overhead Reduction Techniques for Coarse Grained Reconfigurable Architectures,” ResearchGate, Conference Paper, Aug. 2017, https://www.researchgate.net/publication/319416458, 9 pages. [cited by applicant]
Vranjkovic et al., “Coarse-Grained Reconfigurable Hardware Accelerator of Machine Learning Classifiers,” IWSSIP 2016, The 23rd International Conference on Systems, Signals and Image Processing, May 23-25, 2016, Bratisla… [cited by applicant]
Wang, et al., “Reconfigurable Hardware Accelerators: Opportunities, Trends and Challenges,” Cornell University, Dec. 13, 2017, 25 pages. [cited by applicant]
Wentzlaff et al: “On-Chip Interconnection Architecture of the Tile Processor”, IEEE Micro, IEEE Service Center, Los Alamitos, CA, US, vol. 27, No. 5, Sep. 1, 2007 Sep. 1, 2007), pp. 15-31, XP011196754. [cited by applicant]
What is the difference between model paralellism and data paralellism, Quora, 27 pages. Retrieved on Sep. 3, 2021. Retrieved from [URL: https://www.quora.com/What-is-the-difference-between-model-parallelism-and-data-par… [cited by applicant]
Wijtvliet et al., “Coarse Grained Reconfigurable Architectures in the Past 25 Years: Overview and Classification,” IEEE 2016, pp. 235-244. [cited by applicant]
Wijtvliet, Course Syllabus for “Accelerators and Coarse Grained Reconfigurable Architectures,” Advanced School for Computing and Imaging, 2017, 2 pages. [cited by applicant]
Wikipedia “bfloat16 floating-point format,” downloaded Aug. 19, 2019, 2 pages. [cited by applicant]
Wikipedia, Batch normalization, downloaded Feb. 25, 2021, 10 pages. [cited by applicant]
Wikipedia, Floor and ceiling functions, downloaded Aug. 12, 2019, 5 pages. [cited by applicant]
Woolloy, NCCL: Accelerated Multi-GPU Collective Communications, NVIDIA, 56 pages. [cited by applicant]
Xiandong Qi, Introduction to Distributed Deep Learning, dated May 13, 2017, 13 pages. [cited by applicant]
Zhang et al., Dive into Deep Learning, Release 0.16.2, dated Mar. 20, 2021, 1027 pages. [cited by applicant]
Zhang, “Design of Coarse-Grained Reconfigurable Architecture for Digital Signal Processing,” Implementation Aspects, Master of Science Thesis, Feb. 2009, 110 pages. [cited by applicant]
80.192.25.230: “Producer-consumer problem”, Feb. 7, 2013 {Feb. 7, 2013), XP055530821, Retrieved from the nternet: URL:https://en.wikipedia.org/w/index.php?t>ille=Producer/oE2%80%93consumer_problem&old d=537111527 [retri… [cited by applicant]
Accelerated Computing with a Reconfigurable Dataflow Architecture, SambaNova Systems Whitepaper, 10 pages. [cited by applicant]
Amba Axi and ACE Protocol Specification, ARM, as early as Jan. 2003, 440 pages. [cited by applicant]
Ando et al., “A Multithreaded CGRA for Convolutional Neural Network Processing,” Scientific Research Publishing, Circuits and Systems, Jun. 2017, pp. 149-170. [cited by applicant]
Anonymous, Activation Function, Wikipedia, Retrieved on Aug. 16, 2019, 3 pages. Retrieved from [ URL: https://en.wikipedia.org/wiki/Activation_function ]. [cited by applicant]
Arvind, A., Dataflow: Passing the Token, Jun. 6, 2005, 42 pages. [cited by applicant]
Bae et al., Auto-Tuning CNNs for Coarse-Grained Reconfigurable Array-based Accelerators, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 37, Issue: 11, Nov. 2018, 10 pages. [cited by applicant]
Basterretxea et al., “Approximation of sigmoid function and the derivative for hardware implementation of artificial neurons,” IEE Proceedings—Circuits, Devices and Systems, vol. 151, Issue 1, Feb. 5, 2004, 7 pages. [cited by applicant]
Bendersky, The Soflmax function and its derivative, Oct. 18, 2016, 11 pages. URL: https://eli.thegreenplace.net/2016/the-softmax-function-and-its-derivative. [cited by applicant]
Benoit et al., Automatic Task Scheduling/ Loop Unrolling using Dedicated RTR Controllers in Coarse Grain Reconfigurable Architectures, Parallel and Distributed Processing Symposium, 2005. Proceedings. 19th IEEE Internat… [cited by applicant]
Busa et al., A Run-Time Word-Level Reconfigurable Coarse-Grain Functional Unit for a VLIW processor; ACM 2002, pp. 44-49. (Year: 2002). [cited by applicant]
Cook, Comparing bfloat16 range and precision to other 16-bit numbers, dated Nov. 15, 2018, 8 pages. Retrieved on Dec. 3, 2021. Retrieved from the internet [URL: www.johndcook.com/blog/2018/11/15/bfloat16 ]. [cited by applicant]
De Sutter et al., Coarse-Grained Reconfigurable Array Architectures, 2010 Handbook of Signal Processing Systems, 37 pages. [cited by applicant]
Dettmers, How to Parallelize Deep Learning on GPUs Part 1 of 2: Data Parallelism, dated Oct. 9, 2014, 19 pages. Retrieved on Sep. 3, 2021 Retrieved from [URL: https://timdettmers.com/2014/10/09/deep-learning-data-parall… [cited by applicant]
Dettmers, How to Parallelize Deep Learning on GPUs Part 2 of 2: Model Parallelism, dated Nov. 9, 2014, 19 pages. Retrieved on Sep. 3, 2021. Retrieved from [URL: https://timdettmers.com/2014/11/09/model-parallelism-deep-… [cited by applicant]
Donges, Gradient Descent: An Introduction to Machine Learning's Most Popular Algorithms, dated Jun. 16, 2019, 10 pages. Retrieved on Mar. 24, 2021, retrieved from [URL: https://builtin.com/data-science/gradient-descent … [cited by applicant]
Ekanayake, Model Parallelism in Deep Learning is NOT What you think, dated Nov. 10, 2018, 4 pages. Retrieved an Sep. 3, 2021. Retrieved from [ URL: https://medium.com/@esaliya/model-parallelism-in-deep-learning-is-not-w… [cited by applicant]
Eppler et al., High speed neural network chip for trigger purposes in high energy physics, IEEE, Proc. of the conference on design, automation and test in Europe, Feb. 1998, 8 pages. [cited by applicant]
Éricles Sousa, A reconfigurable memory architecture for system integration of coarse-grained reconfigurable arrays, Published in: 2017 International Conference on ReConFigurable Computing and FPGAs (ReConFig) Dec. 4-6, … [cited by applicant]
Fiolhais et al., “Overlay Architectures for Space Applications,” SpacE FPGA Users Workshop, Apr. 9-11, 2018, pp. 1-20. [cited by applicant]
Gomar et al. “Precise digital implementations of hyperbolic tanh and sigmoid function,” 2016 50th Asilomar Conference on Signals, Systems and Computers, Nov. 6-9, 2016, 4 pages. [cited by applicant]
Goodfellow et. al., Deep Learning Book Chapter 6 Deep Feedforward Networks, 2016, 60 pages. [cited by applicant]
Harris et al., Architectures and Algorithms for User Customization of CNNs, ASP-DAC 2018, 32 pages. [cited by applicant]
Hartenstein, Coarse Grain Reconfigurable Architectures, IEEE, 2001, 6 pages. [cited by applicant]
Iannucci, “Toward a dataflow/von Neumann hybrid architecture,” ISCA '88 Proc. of the 15th Annual ISCA, May 30-Jun. 2, 1988, 10 pages. [cited by applicant]
Insujang, GPU Architecture Overview, Better Tomorrow with Computer Science, published Apr. 27, 2017, retrieved on Jun. 17, 2021, retrieved from the Internet [ URL: https://insujang.github.io/2017-04-17/gpu-architecture-… [cited by applicant]
Intel BLOAT16—Hardware Numerics Definition White Paper, Rev. 1.0, Nov. 2018, 7 pages. [cited by applicant]
Ioffe, et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” Cornell University, available at https://arxiv.org/abs/1502.03167, Mar. 2, 2015, 11 pages. [cited by applicant]
Iqbal et al., Reconfigurable Processor Architecture for High Speed Applications, IEEE, dated 2009, pp. 624-629. [cited by applicant]
Jafri et al., NeuroCGRA: A CGRAs with Support for Neural Networks, 2014 International Conference on High Performance Computing & Simulation (HPCS), 8 pages. [cited by applicant]
Jin et al., How to scale distributed deep learning, dated Nov. 14, 2016, 16 pages. [cited by applicant]
Kachris et al.; “A Survey on Reconfigurable Accelerators for Cloud Computing”, IEEE 2016, Aug. 29, 2016, pp. 1-11. [cited by applicant]
Knodel, Oliver, et al., “RC3E: Reconfigurable Accelerators in Data Centers and their Provision by Adapted Service Models”, IEEE 9th International Converence on Cloud Computing, 2016, pp. 1-8. [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
Lecture 11: Distributed Training and Communication Protocols, CSE599W: Spring 2018, UW Paul G. Allen School of Computer Science and Engineering, 41 pages. [cited by applicant]
Li, Ang, et al., “Evaluating Modern GPU Interconnect: PCle, NVLink, NV-SLI, NVSwitch and GPUDirect”, Mar. 11, 2019, 15 pages. [cited by applicant]
Li, et al., “CATERPILLAR: Coarse Grain Reconfigurable Architecture for Accelerating the Training of Deep Neural Networks,” arXiv: 1706.00517v2 [cs.DC], Jun. 8, 2017, 10 pages. [cited by applicant]
Lin et al., “A Digital Circuit Design of Hyperbolic Tangent Sigmoid Function for Neural Networks,” 2018 IEEE Int'l Symp. on Circuits and Systems, May 18-21, 2018, 4 pages. [cited by applicant]
Liu et al., Offloading distributed Applications onto SmartNICs using iPipe, ACM 2019, pp. 1-16. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Ma et al., DeepGauge: Multi-Granularity Testing Criteria for Deep Learning Systems; ACM 2018, pp. 1-12. [cited by applicant]
Mao, Data Parallelism vs Model Parallelism in Distributed Deep Learning Training, dated Mar. 23, 2019, 4 pages, retrieved on Mar. 30, 2021, Retrieved from the internet [ URL: https://leimao.github.io]. [cited by applicant]
Marshall, Dave, “Remote Procedure Calls (RPC)”, Jan. 5, 1999, 15 pages, Retreived from URL. [cited by applicant]
Mazur, A step by step backpropagation example, dated Mar. 17, 2015, 26 pages. Retrieved on Sep. 3, 2021. Retrieved from [URL: https://mattmazur.com/2015/03/17/a-step-by-step-backpropagation-example/ ]. [cited by applicant]
MISB ST 1201.4, “Floating Point to Integer Mapping,” Feb. 28, 2019, pp. 1-21. [cited by applicant]
Nicol, “A Course Grain Reconfigurable Array (CGRA) for Statically Scheduled Data Flow Computing,” Wave Computing, May 3, 2017, 9 pages. [cited by applicant]
Nicol, “Wave Computing: A Dataflow Processing Chip for Training Deep Neural Networks,” 2017, 25 pages. [cited by applicant]
NVIDIA, “NVIDIA DGX-1 System Architecture”, WP-08437-001_v02, 2017, 33 pages. [cited by applicant]
NVIDIA, “NVIDIA DGX-1 With Tesla V100 System Architecture”, WP-08437-002_v01, 2017, 43 pages. [cited by applicant]
NVIDIA, “NVIDIA Tesla P100”, WP-08019-001 v01.1, 2016, 45 pages. [cited by applicant]