IP Library Granted Patent US 12,210,468
Granted Patent B2
US 12,210,468 · App. 18/099,014 · Granted Jan 28, 2025

Data transfer between accessible memories of multiple processors incorporated in coarse-grained reconfigurable (CGR) architecture within heterogeneous processing system using one memory to memory transfer operation

Inventors: Arnav Goel (Palo Alto, CA); Neal Sanghvi (Palo Alto, CA); Jiayu Bai (Palo Alto, CA); Qi Zheng (Palo Alto, CA); Ravinder Kumar (Palo Alto, CA)
Assignee: SambaNova Systems, Inc.
G06F13/28G06F12/0238G06F13/1642
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,468
App. No.
18/099,014
Granted
Jan 28, 2025
Kind
B2
Abstract

A heterogeneous processing system including a host processor, a first processor with a first memory and a first data transfer resource, a second processor with a second memory, and switch and bus circuitry that communicatively couples the processors and the data transfer resource. The host processor is programmed to map virtual addresses of the second memory to physical addresses of the switch and bus circuitry and to configure the first processor to perform one memory to memory transfer operation between the first and second memories using the data transfer resource. The first processor may be configured to program the first data transfer resource. A method including mapping virtual addresses of the second memory to physical addresses of the switch and bus circuitry, and configuring the first processor to perform one memory to memory transfer operation between the first and second memories using the first data transfer resource.

Claims (42)

1. A heterogeneous processing system programmed to execute a computation graph for machine learning and/or inference, the system comprising:

a host processor;

a first processor, having a coarse-grained reconfigurable architecture, coupled to a first memory and configured to execute a first node of the computation graph to generate and store first data into the first memory;

a second processor coupled to a second memory and configured to execute a second node of the computation graph using the first data and store second data into the second memory;

a first data transfer resource, incorporated into the first processor, and a second data transfer resource; and

switch and bus circuitry that communicatively couples the host processor, the first processor, the second processor, the first data transfer resource, and the second data transfer resource;

wherein the host processor is configured to map virtual addresses of the second memory to physical addresses of the switch and bus circuitry, to configure the first data transfer resource to transfer the first data from the first memory into the second memory using the mapped physical addresses, and to configure the second data transfer resource to transfer the second data from the second memory into host memory of the host processor.

2. The heterogeneous processing system of claim 1 , wherein the first processor with the coarse-grained reconfigurable architecture comprises a reconfigurable dataflow unit including:

an array of configurable units, including a plurality of configurable compute units, a plurality of configurable memory units, and a plurality of address generation units, coupled together by an array-level network;

a top-level network coupled to the plurality of address generation units of the array of configurable units; and

the first data transfer resource.

3. The heterogeneous processing system of claim 1 , wherein the second processor comprises a compute engine incorporating the second data transfer resource.

4. The heterogeneous processing system of claim 1 , wherein the first data transfer resource comprises a direct memory access (DMA) engine.

5. The heterogeneous processing system of claim 1 , wherein the first data transfer resource is configured by the host processor to directly transfer the first data from the first memory into the second memory using the mapped physical addresses.

6. A method of transferring data in a heterogeneous system programmed to execute a computation graph for machine learning and/or inference, wherein the heterogeneous system includes a host processor coupled to a host memory, a first processor having a coarse-grained reconfigurable architecture coupled to a first memory, a second processor coupled to a second memory, a first data transfer resource incorporated into the first processor, and switch and bus circuitry that communicatively couples the host processor, the first processor, the second processor, and the first data transfer resource, the method comprising:

mapping, by the host processor, virtual addresses of the second memory to physical addresses of the switch and bus circuitry;

executing, by the first processor, a first node of the computation graph to generate and store first data into the first memory;

configuring, by the host processor, the first data transfer resource to transfer the first data from the first memory into the second memory using the mapped physical addresses;

prompting the first data transfer resource to transfer the first data from the first memory into the second memory; and

executing, by the second processor, a second node of the computation graph using the first data to generate and store second data into the second memory.

7. The method of claim 6 , wherein the first processor comprises a reconfigurable dataflow unit including:

an array of configurable units, including a plurality of configurable compute units, a plurality of configurable memory units, and a plurality of address generation units, coupled together by an array-level network;

a top-level network coupled to the plurality of address generation units of the array of configurable units; and

the first data transfer resource.

8. The method of claim 6 , wherein the configuring comprises configuring the first data transfer resource to directly transfer the first data from the first memory into the second memory using the mapped physical addresses.

9. The method of claim 6 , wherein the heterogeneous system further includes a second data transfer resource coupled to the switch and bus circuitry, the method further comprising:

configuring, by the host processor, the second data transfer resource to transfer the second data from the second memory into host memory using the mapped physical addresses; and

prompting the second data transfer resource to transfer the second data from the second memory into the host memory.

10. A method of transferring data in a heterogeneous system programmed to execute a computation graph for machine learning and/or inference, wherein the heterogeneous system includes a host processor coupled to a host memory, a first processor having a coarse-grained reconfigurable architecture coupled to a first memory, a second processor coupled to a second memory, a first data transfer resource incorporated into the first processor, and switch and bus circuitry that communicatively couples the host processor, the first processor, the second processor, and the first data transfer resource, the method comprising:

mapping, by the host processor, virtual addresses of the second memory to physical addresses of the switch and bus circuitry;

executing, by the second processor, a first node of the computation graph to generate and store first data into the second memory;

configuring, by the host processor, the first data transfer resource to transfer the first data from the second memory into the first memory using the mapped physical addresses;

prompting the first data transfer resource to transfer the first data from the second memory into the first memory; and

executing, by the first processor, a second node of the computation graph using the first data to generate and store second data into the first memory.

11. The method of claim 10 , wherein the first processor comprises a reconfigurable dataflow unit including:

an array of configurable units, including a plurality of configurable compute units, a plurality of configurable memory units, and a plurality of address generation units, coupled together by an array-level network;

a top-level network coupled to the plurality of address generation units of the array of configurable units; and

the first data transfer resource.

12. The method of claim 10 , the method further comprising:

configuring, by the host processor, the first data transfer resource to transfer the second data from the first memory into host memory; and

prompting the first data transfer resource to transfer the second data from the first memory into the host memory.

13. The method of claim 10 , wherein the configuring comprises configuring the first data transfer resource to directly transfer the first data from the second memory into the first memory using the mapped physical addresses.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2023
From: GOEL, ARNAV; SANGHVI, NEAL; BAI, JIAYU; ZHENG, QI; KUMAR, RAVINDER
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 062444/0507 →
Continuity (1)
Related Publication 20240248863A1 · Jul 25, 2024
References Cited (165)
US 5301287A · Herrell · 1994 [cited by examiner]
US 5684980A · Casselman · 1997 [cited by applicant]
US 6470485B1 · Cote et al. · 2002 [cited by applicant]
US 6539438B1 · Ledzius et al. · 2003 [cited by applicant]
US 6557156B1 · Guccione · 2003 [cited by applicant]
US 6667983B1 · Lo et al. · 2003 [cited by applicant]
US 7707392B2 · Maeda · 2010 [cited by examiner]
US 8244930B1 · Dykema et al. · 2012 [cited by applicant]
US 9009723B2 · Degenaro et al. · 2015 [cited by applicant]
US 9288101B1 · Dalal et al. · 2016 [cited by applicant]
US 9300574B2 · Ditya · 2016 [cited by applicant]
US 9501325B2 · Pell et al. · 2016 [cited by applicant]
US 9715475B2 · Lavasani · 2017 [cited by applicant]
US 9875167B1 · Norrie et al. · 2018 [cited by applicant]
US 10031857B2 · Menachem · 2018 [cited by examiner]
US 10318475B2 · Kaimalettu et al. · 2019 [cited by applicant]
US 10511479B2 · Xie et al. · 2019 [cited by applicant]
US 10558496B2 · Apodaca · 2020 [cited by applicant]
US 10613881B2 · Lee · 2020 [cited by examiner]
US 10621138B2 · Hu et al. · 2020 [cited by applicant]
US 10802870B2 · Lu · 2020 [cited by applicant]
US 10831507B2 · Shah et al. · 2020 [cited by applicant]
US 10831523B2 · Kochevar-Cureton et al. · 2020 [cited by applicant]
US 10838895B2 · Yu et al. · 2020 [cited by applicant]
US 10841243B2 · Levi et al. · 2020 [cited by applicant]
US 10877822B1 · Wang et al. · 2020 [cited by applicant]
US 10936533B2 · LeBeane et al. · 2021 [cited by applicant]
US 11068780B2 · Mellempudi et al. · 2021 [cited by applicant]
US 11080227B2 · Koeplinger et al. · 2021 [cited by applicant]
US 11182221B1 · Sivaramakrishnan et al. · 2021 [cited by applicant]
US 11182264B1 · Sivaramakrishnan et al. · 2021 [cited by applicant]
US 11184439B2 · Eran et al. · 2021 [cited by applicant]
US 11200096B1 · Shenbagam et al. · 2021 [cited by applicant]
US 11237880B1 · Raumann et al. · 2022 [cited by applicant]
US 11328207B2 · Lauterbach et al. · 2022 [cited by applicant]
US 11328208B2 · Lie et al. · 2022 [cited by applicant]
US 11347965B2 · Dutta et al. · 2022 [cited by applicant]
US 11360800B2 · Kochevar-Cureton et al. · 2022 [cited by applicant]
US 11392740B2 · Raumann et al. · 2022 [cited by applicant]
US 11436400B2 · Liao et al. · 2022 [cited by applicant]
US 11436429B2 · Jaganathan et al. · 2022 [cited by applicant]
US 11487694B1 · Misra et al. · 2022 [cited by applicant]
US 11568218B2 · Ng et al. · 2023 [cited by applicant]
US 11609798B2 · Sivaramakrishnan et al. · 2023 [cited by applicant]
US 11625283B2 · Sivaramakrishnan et al. · 2023 [cited by applicant]
US 11625284B2 · Sivaramakrishnan et al. · 2023 [cited by applicant]
US 11687264B2 · Chang et al. · 2023 [cited by applicant]
US 11847395B2 · Raumann et al. · 2023 [cited by applicant]
US 11886930B2 · Sivaramakrishnan et al. · 2024 [cited by applicant]
US 11886931B2 · Sivaramakrishnan et al. · 2024 [cited by applicant]
US 11893424B2 · Raumann et al. · 2024 [cited by applicant]
US 20020156998A1 · Casselman · 2002 [cited by applicant]
US 20030108119A1 · Mohebbi et al. · 2003 [cited by applicant]
US 20040039980A1 · Zak et al. · 2004 [cited by applicant]
US 20060012395A1 · Huppenthal et al. · 2006 [cited by applicant]
US 20070186126A1 · Smith et al. · 2007 [cited by applicant]
US 20070220522A1 · Coene et al. · 2007 [cited by applicant]
US 20080013448A1 · Horie et al. · 2008 [cited by applicant]
US 20090089475A1 · Chitlur · 2009 [cited by applicant]
US 20090172351A1 · Vorbach et al. · 2009 [cited by applicant]
US 20090300209A1 · Elzur · 2009 [cited by applicant]
US 20120303932A1 · Farabet et al. · 2012 [cited by applicant]
US 20140137123A1 · Hartmann et al. · 2014 [cited by applicant]
US 20140258438A1 · Ayoub et al. · 2014 [cited by applicant]
US 20150058614A1 · Degenaro et al. · 2015 [cited by applicant]
US 20150100971A1 · Dube et al. · 2015 [cited by applicant]
US 20160077976A1 · Raikin et al. · 2016 [cited by applicant]
US 20160308719A1 · Putnam et al. · 2016 [cited by applicant]
US 20160314025A1 · Mcgarry et al. · 2016 [cited by applicant]
US 20170147228A1 · Yudanov et al. · 2017 [cited by applicant]
US 20170262291A1 · Lai et al. · 2017 [cited by applicant]
US 20170317679A1 · Suh et al. · 2017 [cited by applicant]
US 20180026916A1 · Banerjee et al. · 2018 [cited by applicant]
US 20180225403A1 · Nicol et al. · 2018 [cited by applicant]
US 20180285295A1 · Abel et al. · 2018 [cited by applicant]
US 20180307950A1 · Nealis et al. · 2018 [cited by applicant]
US 20180308200A1 · Surti et al. · 2018 [cited by applicant]
US 20180314941A1 · Lie et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20190089616A1 · Chabbi et al. · 2019 [cited by applicant]
US 20190138890A1 · Liang et al. · 2019 [cited by applicant]
US 20190171604A1 · Brewer · 2019 [cited by applicant]
US 20190171612A1 · Shahar et al. · 2019 [cited by applicant]
US 20190180176A1 · Yudanov et al. · 2019 [cited by applicant]
US 20190243571A1 · Narayanan et al. · 2019 [cited by applicant]
US 20190258921A1 · Lie et al. · 2019 [cited by applicant]
US 20190286973A1 · Kovvuri et al. · 2019 [cited by applicant]
US 20190384642A1 · Bolkhovitin et al. · 2019 [cited by applicant]
US 20200090313A1 · Bugdary et al. · 2020 [cited by applicant]
US 20200142753A1 · Harwood et al. · 2020 [cited by applicant]
US 20200142857A1 · Catiller et al. · 2020 [cited by applicant]
US 20200151573A1 · Das et al. · 2020 [cited by applicant]
US 20200174840A1 · Zhao et al. · 2020 [cited by applicant]
US 20200226444A1 · Sharma et al. · 2020 [cited by applicant]
US 20200264876A1 · Lo et al. · 2020 [cited by applicant]
US 20200314181A1 · Eran et al. · 2020 [cited by applicant]
US 20200326992A1 · Jin et al. · 2020 [cited by applicant]
US 20210011770A1 · Prabhakar et al. · 2021 [cited by applicant]
US 20210073625A1 · Cai · 2021 [cited by applicant]
US 20210089343A1 · Hyoudou · 2021 [cited by applicant]
US 20210097366A1 · Wagner et al. · 2021 [cited by applicant]
US 20210097379A1 · Yang et al. · 2021 [cited by applicant]
US 20210103820A1 · Ghosh · 2021 [cited by applicant]
US 20210125058A1 · Chowdhury et al. · 2021 [cited by applicant]
US 20210192287A1 · Dwivedi et al. · 2021 [cited by applicant]
US 20210192357A1 · Sinha et al. · 2021 [cited by applicant]
US 20210192358A1 · Song et al. · 2021 [cited by applicant]
US 20210200610A1 · Chu et al. · 2021 [cited by applicant]
US 20210241093A1 · Byrne et al. · 2021 [cited by applicant]
US 20210306142A1 · Willis et al. · 2021 [cited by applicant]
US 20210373867A1 · Chen et al. · 2021 [cited by applicant]
US 20210374503A1 · Kim et al. · 2021 [cited by applicant]
US 20220058034A1 · Grohoski et al. · 2022 [cited by applicant]
US 20220197712A1 · Sivaramakrishnan et al. · 2022 [cited by applicant]
US 20220197713A1 · Sivaramakrishnan et al. · 2022 [cited by applicant]
US 20220198117A1 · Raumann et al. · 2022 [cited by applicant]
US 20220269534A1 · Misra et al. · 2022 [cited by applicant]
US 20230128529A1 · Zeng et al. · 2023 [cited by applicant]
US 20230195478A1 · Brot et al. · 2023 [cited by applicant]
US 20230205585A1 · Chatterjee et al. · 2023 [cited by applicant]
US 20240248853A1 · Goel et al. · 2024 [cited by applicant]
US 20240248855A1 · Goel et al. · 2024 [cited by applicant]
US 20240248860A1 · Goel et al. · 2024 [cited by applicant]
EP 1372084A2 · 2003 [cited by applicant]
JP 2020112901A · 2020 [cited by applicant]
TW 202240386A · 2022 [cited by applicant]
TW 202240394A · 2022 [cited by applicant]
TW 202248853A · 2022 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
U.S. Appl. No. 18/099,006—Non-Final Rejection, dated May 24, 2024, 9 pages. [cited by applicant]
U.S. Appl. No. 18/099,021—Non-Final Rejection, dated Mar. 7, 2024, 12 pages. [cited by applicant]
U.S. Appl. No. 18/099,021—Notice of Allowance, dated Jun. 6, 2024, 9 pages. [cited by applicant]
U.S. Appl. No. 18/099,032—Non-Final Rejection, dated Jun. 24, 2024, 15 pages. [cited by applicant]
Accelerated Computing with a Reconfigurable Dataflow Architecture, SambaNova Systems Whitepaper, 10 pages. [cited by applicant]
Bae et al., Auto-Tuning CNNs for Coarse-Grained Reconfigurable Array-based Accelerators, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 37, Issue: 11, Nov. 2018, 10 pages. [cited by applicant]
Busa et al., A Run-Time Word-Level Reconfigurable Coarse-Grain Functional Unit for a VLIW processor; ACM 2002, pp. 44-49. (Year: 2002). [cited by applicant]
Donges, Gradient Descent: An Introduction to Machine Learning's Most Popular Algorithms, dated Jun. 16, 2019, 10 pages. Retrieved on Mar. 24, 2021, retrieved from [URL: https://builtin.com/data-science/gradient-descent … [cited by applicant]
Éricles Sousa, A reconfigurable memory architecture for system integration of coarse-grained reconfigurable arrays, Published in: 2017 International Conference on ReConFigurable Computing and FPGAs (ReConFig) Dec. 4-6, … [cited by applicant]
Galanis et al., Accelerating Applications by Mapping Critical Kernels on Coarse-Grain Reconfigurable Hardware in Hybrid Systems, Field-Programmable Custom Computing Machines, 2005, 13th Annual IEEE Symposium on Napa, CA… [cited by applicant]
Galanis et al., Partitioning Methodology for Heterogeneous Reconfigurable Functional Units, The Journal of Supercomputing, vol. 38, No. 1, dated Oct. 1, 2006, 18 pages. [cited by applicant]
Iqbal et al., Reconfigurable Processor Architecture for High Speed Applications, IEEE, dated 2009, pp. 624-629. [cited by applicant]
Zhang et. al., Dive into Deep Learning, Release 0.16.2, dated Mar. 20, 2021, 1027 pages. [cited by applicant]
Jafri et al., NeuroCGRA: A CGRAs with Support for Neural Networks, 2014 International Conference on High Performance Computing & Simulation (HPCS), 8 pages. [cited by applicant]
Jin et al., How to scale distributed deep learning, dated Nov. 14, 2016, 16 pages. [cited by applicant]
Kachris et al.; “A Survey on Reconfigurable Accelerators for Cloud Computing”, IEEE 2016, Aug. 29, 2016, pp. 1-11. [cited by applicant]
Knodel, Oliver, et al., “RC3E: Reconfigurable Accelerators in Data Centers and their Provision by Adapted Service Models”, IEEE 9th International Converence on Cloud Computing, 2016, pp. 1-8. [cited by applicant]
Lecture 11: Distributed Training and Communication Protocols, CSE599W: Spring 2018, UW Paul G. Allen School of Computer Science and Engineering, 41 pages. [cited by applicant]
Li, et al., “CATERPILLAR: Coarse Grain Reconfigurable Architecture for Accelerating the Training of Deep Neural Networks,” arXiv: 1706.00517v2 [cs. DC], Jun. 8, 2017, 10 pages. [cited by applicant]
Liang et al., Dynamic Coarse Grain Dataflow Reconfiguration Technique for Real-Time Systems, IEEE, dated 2005, pp. 3511-3514. [cited by applicant]
Liu et al., Offloading distributed Applications onto SmartNICs using iPipe, ACM 2019, pp. 1-16 (SBNV 1029-2). [cited by applicant]
Ma et al., DeepGauge: Multi-Granularity Testing Criteria for Deep Learning Systems; ACM 2018, pp. 1-12. [cited by applicant]
Mao, Data Parallelism vs Model Parallelism in Distributed Deep Learning Training, dated Mar. 23, 2019, 4 pages, retrieved on Mar. 30, 2021, Retrieved from the internet [ URL: https://leimao.github.io]. [cited by applicant]
Marshall, Dave, “Remote Procedure Calls (RPC)”, Jan. 5, 1999, 15 pages, Retreived from URL. [cited by applicant]
Padole et al., Configuration Memory Based Dynamic Coarse Grained Reconfigurable Multiscore Architecture, IEEE 2013, pp. 3511-3514, 5 pages. [cited by applicant]
Paek et al., “Binary Acceleration Using Coarse-Grained Reconfigurable Architecture,” ACM SIGARCH Computer Architecture News, vol. 38, No. 4, Sep. 2010, 7 pages. [cited by applicant]
Rubattu et al., Dataflow-Functional High-Level Synthesis for Coarse-Grained Reconfigurable Accelerators, IEEE 2019, pp. 69-72 (Year: 2019). [cited by applicant]
Ruder, An overview of gradient descent optimization algorithms, NUI Galway Aylien Lyd, dated Jun. 15, 2017, 14 pages. [cited by applicant]
Strom, Scalable Distributed DNN Training Using Commodity GPU Cloud Computing, Amazon.com, 5 pages. [cited by applicant]
Tanaka et al., Distributed Deep Learning with GPU-FPGA heterogenous computing, IEEE 2021, 9 pages. [cited by applicant]
Woolloy, NCCL: Accelerated Multi-GPU Collective Communications, NVIDIA, 56 pages. [cited by applicant]
Xiandong Qi, Introduction to Distributed Deep Learning, dated May 13, 2017, 13 pages. [cited by applicant]
Cited By (4)
US 12,380,041 US 12,474,897 US 12,639,242 US 12,639,255