US 4769790A
· Yamashita
· 1988
[cited by applicant]
US 5506797A
· Koshiba
· 1996
[cited by applicant]
US 5560029A
· Papadopoulos et al.
· 1996
[cited by applicant]
US 5794033A
· Aldebert et al.
· 1998
[cited by applicant]
US 5963746A
· Barker et al.
· 1999
[cited by applicant]
US 6105119A
· Kerr et al.
· 2000
[cited by applicant]
US 6119181A
· Vorbach et al.
· 2000
[cited by applicant]
US 6256653B1
· Juffa et al.
· 2001
[cited by applicant]
US 6470485B1
· Cote et al.
· 2002
[cited by applicant]
US 6539438B1
· Ledzius
· 2003
[cited by examiner]
US 6667983B1
· Lo et al.
· 2003
[cited by applicant]
US 6728871B1
· Vorbach et al.
· 2004
[cited by applicant]
US 7015921B1
· Trivedi et al.
· 2006
[cited by applicant]
US 7472149B2
· Endo
· 2008
[cited by applicant]
US 7734895B1
· Agarwal et al.
· 2010
[cited by applicant]
US 7797258B1
· Bowman et al.
· 2010
[cited by applicant]
US 7952387B1
· Frazer
· 2011
[cited by applicant]
US 7996684B2
· Wasson et al.
· 2011
[cited by applicant]
US 8006021B1
· Li et al.
· 2011
[cited by applicant]
US 8045546B1
· Bao et al.
· 2011
[cited by applicant]
US 8184317B2
· Okamoto
· 2012
[cited by applicant]
US 8261042B2
· Kanstein et al.
· 2012
[cited by applicant]
US 8271557B1
· Lysaght
· 2012
[cited by examiner]
US 9009723B2
· Degenaro et al.
· 2015
[cited by applicant]
US 9201899B2
· Nishimura et al.
· 2015
[cited by applicant]
US 9335977B2
· Wang et al.
· 2016
[cited by applicant]
US 9411532B2
· Vorbach et al.
· 2016
[cited by applicant]
US 9411756B2
· Nogueira et al.
· 2016
[cited by applicant]
US 9501325B2
· Pell et al.
· 2016
[cited by applicant]
US 9569214B2
· Govindu et al.
· 2017
[cited by applicant]
US 9690747B2
· Vorbach et al.
· 2017
[cited by applicant]
US 9697318B2
· Hutton et al.
· 2017
[cited by applicant]
US 9875105B2
· Rozas et al.
· 2018
[cited by applicant]
US 9952831B1
· Ross et al.
· 2018
[cited by applicant]
US 20030068097A1
· Wilson et al.
· 2003
[cited by applicant]
US 20030108119A1
· Mohebbi et al.
· 2003
[cited by applicant]
US 20040049672A1
· Nollet
· 2004
[cited by examiner]
US 20040088666A1
· Poznanovic et al.
· 2004
[cited by applicant]
US 20040153608A1
· Vorbach et al.
· 2004
[cited by applicant]
US 20090300209A1
· Elzur
· 2009
[cited by applicant]
US 20140137123A1
· Hartmann et al.
· 2014
[cited by applicant]
US 20140201642A1
· Vicat-Blanc
· 2014
[cited by applicant]
US 20140317628A1
· Kim
· 2014
[cited by applicant]
US 20170083313A1
· Sankaralingam
· 2017
[cited by examiner]
US 20170322774A1
· Zhang
· 2017
[cited by applicant]
US 20170322805A1
· Zohar et al.
· 2017
[cited by applicant]
US 20190004878A1
· Adler
· 2019
[cited by examiner]
US 20190114139A1
· Zhang
· 2019
[cited by applicant]
US 20190138890A1
· Liang et al.
· 2019
[cited by applicant]
US 20190279075A1
· Liu et al.
· 2019
[cited by applicant]
US 20190286973A1
· Kovvuri et al.
· 2019
[cited by applicant]
US 20200167309A1
· Nicol
· 2020
[cited by applicant]
US 20200226444A1
· Sharma et al.
· 2020
[cited by applicant]
US 20210081691A1
· Chen et al.
· 2021
[cited by applicant]
US 20210089343A1
· Hyoudou
· 2021
[cited by applicant]
US 20210103820A1
· Ghosh
· 2021
[cited by applicant]
US 20210192358A1
· Song et al.
· 2021
[cited by applicant]
EP 0733234A
· 1995
[cited by applicant]
EP 1372084A2
· 2003
[cited by applicant]
JP 2020112901A
· 2020
[cited by applicant]
TW 200736953A
· 2007
[cited by applicant]
TW 200801964A
· 2008
[cited by applicant]
TW 200928736A
· 2009
[cited by applicant]
WO 2010142987A1
· 2010
[cited by applicant]
WO 2018100920A1
· 2018
[cited by applicant]
WO 2021067318A1
· 2021
[cited by applicant]
WO 2021108328A1
· 2021
[cited by applicant]
NVIDIA, “NVIDIA Turing GPU Architecture”, WP-09183-001_v01, 2018, 86 pages.
[cited by applicant]
Olukotun, Designing Computer Sytems for Software 2.0, ISCA 2018 keynote, Jun. 2018, 49 pages.
[cited by applicant]
Paek et al., “Binary Acceleration Using Coarse-Grained Reconfigurable Architecture,” ACM SIGARCH Computer Architecture News, vol. 38, No. 4, Sep. 2010, 7 pages.
[cited by applicant]
PCT/US/2021/040382—International Search Report and Written Opinion, dated Nov. 29, 2021, 22 pages.
[cited by applicant]
Petersen, “Softmax with cross-entropy,” https://mattpetersen.github.io/softmax-with-cross-entropy, Jun. 25, 2017, 19 pages.
[cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages.
[cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages.
[cited by applicant]
Rubattu et al., Dataflow-Functional High-Level Synthesis for Coarse-Grained Reconfigurable Accelerators, IEEE 2019, pp. 69-72 (Year: 2019).
[cited by applicant]
Ruder, An overview of gradient descent optimization algorithms, NUI Galway Aylien Lyd, dated Jun. 15, 2017, 14 pages.
[cited by applicant]
Strom, Scalable Distributed DNN Training Using Commodity GPU Cloud Computing, Amazon.com, 5 pages.
[cited by applicant]
Tanaka et al., Distributed Deep Learning with GPU-FPGA heterogenous computing, IEEE 2021, 9 pages.
[cited by applicant]
Tanomoto et al., “A CGRA-based Approach for Accelerating Convolutional Neural Networks,” 2015 IEEE 9th International Symposium on Embedded Multicore/Many-core Systems-on-Chip, 2015, pp. 73-80.
[cited by applicant]
Tobuschat, et al., “IDAMC: A NoC for mixed criticality systems,” 2013 IEEE 19th International Conference on Embedded and Real-Time Computing Systems and Applications, Taipei, Aug. 19-21, 2013, pp. 149-156.
[cited by applicant]
Turkson et al. “Artificial neural network applications in the calibration of spark-ignition engines: An overview,” Engineering Science and Technology, an International Journal, vol. 19, Issue 3, Sep. 2016, 1346-1359.
[cited by applicant]
TW 110124802—First Office Action and Search Report dated May 24, 2022, 17 pages.
[cited by applicant]
U.S. Appl. No. 16/922,975—Final Office Action, dated Mar. 9, 2023, 23 pages.
[cited by applicant]
U.S. Appl. No. 16/922,975—Non-Final Office Action, dated Oct. 27, 2022, 26 pages.
[cited by applicant]
U.S. Appl. No. 16/922,975—Notice of Allowance, dated Jul. 3, 2023, 11 pages.
[cited by applicant]
U.S. Appl. No. 17/127,818—Notice of Allowance, dated Jul. 21, 2021, 10 pages.
[cited by applicant]
U.S. Appl. No. 17/127,818—Office Action dated Apr. 1, 2021, 15 pages.
[cited by applicant]
U.S. Appl. No. 17/127,818—Response to Office Action dated Apr. 1, 2021, filed Jul. 1, 2021, 15 pages.
[cited by applicant]
U.S. Appl. No. 17/127,929—Notice of Allowance dated Jul. 21, 2021, 14 pages.
[cited by applicant]
U.S. Appl. No. 17/127,929—Office Action dated Apr. 1, 2021, 26 pages.
[cited by applicant]
U.S. Appl. No. 17/127,929—Response to Office Action dated Apr. 1, 2021, filed Jul. 1, 2021, 10 pages.
[cited by applicant]
U.S. Appl. No. 17/214,768—Notice of Allowance, dated Aug. 11, 2021, 26 pages.
[cited by applicant]
U.S. Appl. No. 17/214,768—Supplemental Notice of Allowance, dated Aug. 25, 2021, 10 pages.
[cited by applicant]
U.S. Appl. No. 17/379,921—Notice of Allowance, dated Nov. 26, 2021, 21 pages.
[cited by applicant]
U.S. Appl. No. 17/379,924—Notice of Allowance, dated Sep. 16, 2021, 23 pages.
[cited by applicant]
Vadivel et al., “Loop Overhead Reduction Techniques for Coarse Grained Reconfigurable Architectures,” ResearchGate, Conference Paper, Aug. 2017, https://www.researchgate.net/publication/319416458, 9 pages.
[cited by applicant]
Vranjkovic et al., “Coarse-Grained Reconfigurable Hardware Accelerator of Machine Learning Classifiers,” IWSSIP 2016, The 23rd International Conference on Systems, Signals and Image Processing, May 23-25, 2016, Bratisla…
[cited by applicant]
Wang, et al., “Reconfigurable Hardware Accelerators: Opportunities, Trends and Challenges,” Cornell University, Dec. 13, 2017, 25 pages.
[cited by applicant]
Wentzlaff et al: “On-Chip Interconnection Architecture of the Tile Processor”, IEEE Micro, IEEE Service Center, Los Alamitos, CA, US, vol. 27, No. 5, Sep. 1, 2007 Sep. 1, 2007), pp. 15-31, XP011196754.
[cited by applicant]
What is the difference between model paralellism and data paralellism, Quora, 27 pages. Retrieved on Sep. 3, 2021. Retrieved from [URL: https://www.quora.com/What-is-the-difference-between-model-parallelism-and-data-par…
[cited by applicant]
Wijtvliet et al., “Coarse Grained Reconfigurable Architectures in the Past 25 Years: Overview and Classification,” IEEE 2016, pp. 235-244.
[cited by applicant]
Wijtvliet, Course Syllabus for “Accelerators and Coarse Grained Reconfigurable Architectures,” Advanced School for Computing and Imaging, 2017, 2 pages.
[cited by applicant]
Wikipedia “bfloat16 floating-point format,” downloaded Aug. 19, 2019, 2 pages.
[cited by applicant]
Wikipedia, Batch normalization, downloaded Feb. 25, 2021, 10 pages.
[cited by applicant]
Wikipedia, Floor and ceiling functions, downloaded Aug. 12, 2019, 5 pages.
[cited by applicant]
Woolloy, NCCL: Accelerated Multi-GPU Collective Communications, NVIDIA, 56 pages.
[cited by applicant]
Xiandong Qi, Introduction to Distributed Deep Learning, dated May 13, 2017, 13 pages.
[cited by applicant]
Zhang et al., Dive into Deep Learning, Release 0.16.2, dated Mar. 20, 2021, 1027 pages.
[cited by applicant]
Zhang, “Design of Coarse-Grained Reconfigurable Architecture for Digital Signal Processing,” Implementation Aspects, Master of Science Thesis, Feb. 2009, 110 pages.
[cited by applicant]
80.192.25.230: “Producer-consumer problem”, Feb. 7, 2013 {Feb. 7, 2013), XP055530821, Retrieved from the nternet: URL:https://en.wikipedia.org/w/index.php?t>ille=Producer/oE2%80%93consumer_problem&old d=537111527 [retri…
[cited by applicant]
Accelerated Computing with a Reconfigurable Dataflow Architecture, SambaNova Systems Whitepaper, 10 pages.
[cited by applicant]
Amba Axi and ACE Protocol Specification, ARM, as early as Jan. 2003, 440 pages.
[cited by applicant]
Ando et al., “A Multithreaded CGRA for Convolutional Neural Network Processing,” Scientific Research Publishing, Circuits and Systems, Jun. 2017, pp. 149-170.
[cited by applicant]
Anonymous, Activation Function, Wikipedia, Retrieved on Aug. 16, 2019, 3 pages. Retrieved from [ URL: https://en.wikipedia.org/wiki/Activation_function ].
[cited by applicant]
Arvind, A., Dataflow: Passing the Token, Jun. 6, 2005, 42 pages.
[cited by applicant]
Bae et al., Auto-Tuning CNNs for Coarse-Grained Reconfigurable Array-based Accelerators, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 37, Issue: 11, Nov. 2018, 10 pages.
[cited by applicant]
Basterretxea et al., “Approximation of sigmoid function and the derivative for hardware implementation of artificial neurons,” IEE Proceedings—Circuits, Devices and Systems, vol. 151, Issue 1, Feb. 5, 2004, 7 pages.
[cited by applicant]
Bendersky, The Soflmax function and its derivative, Oct. 18, 2016, 11 pages. URL: https://eli.thegreenplace.net/2016/the-softmax-function-and-its-derivative.
[cited by applicant]
Benoit et al., Automatic Task Scheduling/ Loop Unrolling using Dedicated RTR Controllers in Coarse Grain Reconfigurable Architectures, Parallel and Distributed Processing Symposium, 2005. Proceedings. 19th IEEE Internat…
[cited by applicant]
Busa et al., A Run-Time Word-Level Reconfigurable Coarse-Grain Functional Unit for a VLIW processor; ACM 2002, pp. 44-49. (Year: 2002).
[cited by applicant]
Cook, Comparing bfloat16 range and precision to other 16-bit numbers, dated Nov. 15, 2018, 8 pages. Retrieved on Dec. 3, 2021. Retrieved from the internet [URL: www.johndcook.com/blog/2018/11/15/bfloat16 ].
[cited by applicant]
De Sutter et al., Coarse-Grained Reconfigurable Array Architectures, 2010 Handbook of Signal Processing Systems, 37 pages.
[cited by applicant]
Dettmers, How to Parallelize Deep Learning on GPUs Part 1 of 2: Data Parallelism, dated Oct. 9, 2014, 19 pages. Retrieved on Sep. 3, 2021 Retrieved from [URL: https://timdettmers.com/2014/10/09/deep-learning-data-parall…
[cited by applicant]
Dettmers, How to Parallelize Deep Learning on GPUs Part 2 of 2: Model Parallelism, dated Nov. 9, 2014, 19 pages. Retrieved on Sep. 3, 2021. Retrieved from [URL: https://timdettmers.com/2014/11/09/model-parallelism-deep-…
[cited by applicant]
Donges, Gradient Descent: An Introduction to Machine Learning's Most Popular Algorithms, dated Jun. 16, 2019, 10 pages. Retrieved on Mar. 24, 2021, retrieved from [URL: https://builtin.com/data-science/gradient-descent …
[cited by applicant]
Ekanayake, Model Parallelism in Deep Learning is NOT What you think, dated Nov. 10, 2018, 4 pages. Retrieved an Sep. 3, 2021. Retrieved from [ URL: https://medium.com/@esaliya/model-parallelism-in-deep-learning-is-not-w…
[cited by applicant]
Eppler et al., High speed neural network chip for trigger purposes in high energy physics, IEEE, Proc. of the conference on design, automation and test in Europe, Feb. 1998, 8 pages.
[cited by applicant]
Éricles Sousa, A reconfigurable memory architecture for system integration of coarse-grained reconfigurable arrays, Published in: 2017 International Conference on ReConFigurable Computing and FPGAs (ReConFig) Dec. 4-6, …
[cited by applicant]
Fiolhais et al., “Overlay Architectures for Space Applications,” SpacE FPGA Users Workshop, Apr. 9-11, 2018, pp. 1-20.
[cited by applicant]
Gomar et al. “Precise digital implementations of hyperbolic tanh and sigmoid function,” 2016 50th Asilomar Conference on Signals, Systems and Computers, Nov. 6-9, 2016, 4 pages.
[cited by applicant]
Goodfellow et. al., Deep Learning Book Chapter 6 Deep Feedforward Networks, 2016, 60 pages.
[cited by applicant]
Harris et al., Architectures and Algorithms for User Customization of CNNs, ASP-DAC 2018, 32 pages.
[cited by applicant]
Hartenstein, Coarse Grain Reconfigurable Architectures, IEEE, 2001, 6 pages.
[cited by applicant]
Iannucci, “Toward a dataflow/von Neumann hybrid architecture,” ISCA '88 Proc. of the 15th Annual ISCA, May 30-Jun. 2, 1988, 10 pages.
[cited by applicant]
Insujang, GPU Architecture Overview, Better Tomorrow with Computer Science, published Apr. 27, 2017, retrieved on Jun. 17, 2021, retrieved from the Internet [ URL: https://insujang.github.io/2017-04-17/gpu-architecture-…
[cited by applicant]
Intel BLOAT16—Hardware Numerics Definition White Paper, Rev. 1.0, Nov. 2018, 7 pages.
[cited by applicant]
Ioffe, et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” Cornell University, available at https://arxiv.org/abs/1502.03167, Mar. 2, 2015, 11 pages.
[cited by applicant]
Iqbal et al., Reconfigurable Processor Architecture for High Speed Applications, IEEE, dated 2009, pp. 624-629.
[cited by applicant]
Jafri et al., NeuroCGRA: A CGRAs with Support for Neural Networks, 2014 International Conference on High Performance Computing & Simulation (HPCS), 8 pages.
[cited by applicant]
Jin et al., How to scale distributed deep learning, dated Nov. 14, 2016, 16 pages.
[cited by applicant]
Kachris et al.; “A Survey on Reconfigurable Accelerators for Cloud Computing”, IEEE 2016, Aug. 29, 2016, pp. 1-11.
[cited by applicant]
Knodel, Oliver, et al., “RC3E: Reconfigurable Accelerators in Data Centers and their Provision by Adapted Service Models”, IEEE 9th International Converence on Cloud Computing, 2016, pp. 1-8.
[cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages.
[cited by applicant]
Lecture 11: Distributed Training and Communication Protocols, CSE599W: Spring 2018, UW Paul G. Allen School of Computer Science and Engineering, 41 pages.
[cited by applicant]
Li, Ang, et al., “Evaluating Modern GPU Interconnect: PCle, NVLink, NV-SLI, NVSwitch and GPUDirect”, Mar. 11, 2019, 15 pages.
[cited by applicant]
Li, et al., “CATERPILLAR: Coarse Grain Reconfigurable Architecture for Accelerating the Training of Deep Neural Networks,” arXiv: 1706.00517v2 [cs.DC], Jun. 8, 2017, 10 pages.
[cited by applicant]
Lin et al., “A Digital Circuit Design of Hyperbolic Tangent Sigmoid Function for Neural Networks,” 2018 IEEE Int'l Symp. on Circuits and Systems, May 18-21, 2018, 4 pages.
[cited by applicant]
Liu et al., Offloading distributed Applications onto SmartNICs using iPipe, ACM 2019, pp. 1-16.
[cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572…
[cited by applicant]
Ma et al., DeepGauge: Multi-Granularity Testing Criteria for Deep Learning Systems; ACM 2018, pp. 1-12.
[cited by applicant]
Mao, Data Parallelism vs Model Parallelism in Distributed Deep Learning Training, dated Mar. 23, 2019, 4 pages, retrieved on Mar. 30, 2021, Retrieved from the internet [ URL: https://leimao.github.io].
[cited by applicant]
Marshall, Dave, “Remote Procedure Calls (RPC)”, Jan. 5, 1999, 15 pages, Retreived from URL.
[cited by applicant]
Mazur, A step by step backpropagation example, dated Mar. 17, 2015, 26 pages. Retrieved on Sep. 3, 2021. Retrieved from [URL: https://mattmazur.com/2015/03/17/a-step-by-step-backpropagation-example/ ].
[cited by applicant]
MISB ST 1201.4, “Floating Point to Integer Mapping,” Feb. 28, 2019, pp. 1-21.
[cited by applicant]
Nicol, “A Course Grain Reconfigurable Array (CGRA) for Statically Scheduled Data Flow Computing,” Wave Computing, May 3, 2017, 9 pages.
[cited by applicant]
Nicol, “Wave Computing: A Dataflow Processing Chip for Training Deep Neural Networks,” 2017, 25 pages.
[cited by applicant]
NVIDIA, “NVIDIA DGX-1 System Architecture”, WP-08437-001_v02, 2017, 33 pages.
[cited by applicant]
NVIDIA, “NVIDIA DGX-1 With Tesla V100 System Architecture”, WP-08437-002_v01, 2017, 43 pages.
[cited by applicant]
NVIDIA, “NVIDIA Tesla P100”, WP-08019-001 v01.1, 2016, 45 pages.
[cited by applicant]