IP Library Granted Patent US 12,619,370
Granted Patent B2
US 12,619,370 · App. 18/888,015 · Granted May 5, 2026

Memory unit partitioning for reconfigurable dataflow computers

Inventors: Yaqi Zhang (Foster City, CA); Matthew Feldman (Palo Alto, CA)
Assignee: SambaNova Systems, Inc.
G06F3/0644G06F3/0604G06F3/0683G06F9/5077
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,370
App. No.
18/888,015
Granted
May 5, 2026
Kind
B2
Abstract

A system and method for memory unit partitioning for reconfigurable dataflow computing systems includes a parser that receives and parses source code for a reconfigurable dataflow processor, a tensor expression extractor that extracts tensor indexing expressions from the source code, a logical memory constraint generator that converts the tensor indexing expressions to logical memory indexing constraints, a grouping module that groups the logical memory indexing constraints into concurrent access groups, and a memory partitioning module that determines a memory unit partitioning solution for each concurrent access group.

Claims (35)

1 . A system for controlling memory unit partitioning for reconfigurable dataflow computing systems, comprising:

a parser configured to receive and parse source code for a reconfigurable dataflow processor that comprises an array of compute units and an array of memory units interconnected with a switching fabric, the source code comprising a plurality of tensor indexing expressions;

a tensor expression extractor configured to extract the plurality of tensor indexing expressions from the source code;

a logical memory constraint generator configured to convert the plurality of tensor indexing expressions to a plurality of logical memory indexing constraints;

a grouping module configured to group the plurality of logical memory indexing constraints into concurrent access groups; and

a memory partitioning module configured to determine a memory unit partitioning solution for each concurrent access group that supports the plurality of logical memory indexing constraints without concurrent usage conflicts including memory unit and memory port conflicts,

wherein the reconfigurable dataflow processor is configured to execute the plurality of tensor indexing expressions and access the array of memory units according to the memory unit partitioning solution,

wherein memory units in the array of memory units comprise address generators that generate, for each memory cycle, a physical address comprising a bank identifier and a bank offset,

wherein said memory units in the array of memory units are configured to respond to a specific bank identifier, and wherein the memory partitioning module is further configured to determine the memory unit partitioning solution by selecting a set of logical-to-physical mapping parameters, and

wherein the set of logical-to-physical mapping parameters comprise a logical memory unit count N, a blocking parameter B, a scaling vector alpha and a packing vector P.

2 . A system for providing memory unit partitioning solutions for reconfigurable dataflow computing systems, the system comprising:

a parser configured to receive and parse source code for a reconfigurable dataflow processor that comprises an array of compute units and an array of memory units interconnected with a switching fabric, the source code comprising a plurality of tensor indexing expressions;

a tensor expression extractor configured to extract the plurality of tensor indexing expressions from the source code;

a logical memory constraint generator configured to convert the plurality of tensor indexing expressions to a plurality of logical memory indexing constraints;

a grouping module configured to group the plurality of logical memory indexing constraints into concurrent access groups; and

a memory partitioning module configured to determine a memory unit partitioning solution for each concurrent access group that supports the plurality of logical memory indexing constraints without concurrent usage conflicts including memory unit and memory port conflicts,

wherein the dataflow processor is configured to execute the plurality of tensor indexing expressions and access the array of memory units according to the memory unit partitioning solution,

wherein memory units in the array of memory units comprise address generators that generate, for each memory cycle, a physical address comprising a bank identifier and a bank offset, and wherein said memory units in the array of memory units are configured to respond to a specific bank identifier, and

wherein the memory partitioning module is further configured to determine the memory unit partitioning solution by selecting a set of logical-to-physical mapping parameters, and wherein the set of logical-to-physical mapping parameters comprise a logical memory unit count N, a blocking parameter B, a scaling vector alpha and a packing vector P.

3 . The system of claim 2 , wherein the selecting comprises testing legal combinations of N, B and alpha.

4 . The system of claim 2 , further comprising a capacity modification module configured to perform a capacity modification to legalize the memory unit partitioning solution.

5 . The system of claim 4 , wherein the capacity modification comprises scaling packing vector P or increasing a logical memory unit count N of a set of logical-to-physical mapping parameters.

6 . A method for controlling memory unit partitioning solutions for reconfigurable dataflow computing systems, the method comprising:

receiving source code for a reconfigurable dataflow processor that comprises an array of compute units and an array of memory units interconnected with a switching fabric, the source code comprising a plurality of tensor indexing expressions;

converting the plurality of tensor indexing expressions to a plurality of logical memory indexing constraints;

grouping the plurality of logical memory indexing constraints into concurrent access groups;

determining a memory unit partitioning solution for each concurrent access group that supports the plurality of logical memory indexing constraints without concurrent usage conflicts including memory unit and memory port conflicts; and

accessing the array of memory units according to the memory unit partitioning solution in conjunction with executing the plurality of tensor indexing expressions with the reconfigurable dataflow processor,

wherein determining the memory unit partitioning solution comprises selecting a set of logical-to-physical mapping parameters, and

wherein the set of logical-to-physical mapping parameters comprise a logical memory unit count N, a blocking parameter B, a scaling vector alpha and a packing vector P.

7 . The method of claim 6 , wherein selecting comprises testing legal combinations of N, B and alpha.

8 . The method of claim 6 , wherein the capacity modification comprises scaling a packing vector P or increasing a logical memory unit count N of a set of logical-to-physical mapping parameters.

9 . The method of claim 6 , wherein the selecting comprises testing legal combinations of N, B and alpha.

10 . The method of claim 8 , wherein the set of logical-to-physical mapping parameters define a hyperplane partitioning or a parallel-piped partitioning.

11 . The method of claim 6 , wherein the array of compute units operate on vectors and the memory unit partitioning solution is vectorized.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 18, 2024
From: ZHANG, YAQI; FELDMAN, MATTHEW S.
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 069611/0229 →
Continuity (4)
Continuation 18208343 · Jun 12, 2023
Continuation 17878504 · Aug 1, 2022
Provisional Application 63271906 · Oct 26, 2021
Related Publication 20250013375A1 · Jan 9, 2025
References Cited (96)
US 4204636A · Hayman · 1980 [cited by applicant]
US 5220545A · Tomimitsu · 1993 [cited by applicant]
US 7565465B2 · Secatch · 2009 [cited by applicant]
US 9063668B1 · Jung et al. · 2015 [cited by applicant]
US 10698853B1 · Grohoski et al. · 2020 [cited by applicant]
US 10831507B2 · Shah et al. · 2020 [cited by applicant]
US 10970217B1 · Dastidar et al. · 2021 [cited by applicant]
US 11204889B1 · Prabhakar et al. · 2021 [cited by applicant]
US 11366783B1 · Prabhakar et al. · 2022 [cited by applicant]
US 11467827B1 · Zuckerman et al. · 2022 [cited by applicant]
US 20040100963A1 · Guo · 2004 [cited by applicant]
US 20060190517A1 · Guerrero · 2006 [cited by applicant]
US 20070028076A1 · Wezelenburg et al. · 2007 [cited by applicant]
US 20070177677A1 · Thomsen et al. · 2007 [cited by applicant]
US 20090031089A1 · Tuominen · 2009 [cited by applicant]
US 20100313060A1 · Bjoerklund et al. · 2010 [cited by applicant]
US 20140085318A1 · Nadar et al. · 2014 [cited by applicant]
US 20140201642A1 · Vicat-Blanc · 2014 [cited by applicant]
US 20140281417A1 · Peters · 2014 [cited by applicant]
US 20150039851A1 · Uliel et al. · 2015 [cited by applicant]
US 20160012012A1 · Yen et al. · 2016 [cited by applicant]
US 20160055004A1 · Grochowski et al. · 2016 [cited by applicant]
US 20160134702A1 · Gertner · 2016 [cited by applicant]
US 20170031865A1 · Eyole et al. · 2017 [cited by applicant]
US 20170083313A1 · Sankaralingam et al. · 2017 [cited by applicant]
US 20170123794A1 · Chen et al. · 2017 [cited by applicant]
US 20170317678A1 · Coole et al. · 2017 [cited by applicant]
US 20180173668A1 · Rusten et al. · 2018 [cited by applicant]
US 20180212894A1 · Nicol et al. · 2018 [cited by applicant]
US 20180293185A1 · Vembu et al. · 2018 [cited by applicant]
US 20180300181A1 · Hetzel et al. · 2018 [cited by applicant]
US 20180324112A1 · Nicol · 2018 [cited by applicant]
US 20180341487A1 · Kim et al. · 2018 [cited by applicant]
US 20190042241A1 · Akin et al. · 2019 [cited by applicant]
US 20190108423A1 · Jones · 2019 [cited by applicant]
US 20190130270A1 · Nicol et al. · 2019 [cited by applicant]
US 20190197655A1 · Sun et al. · 2019 [cited by applicant]
US 20190229996A1 · ChoFleming, Jr. et al. · 2019 [cited by applicant]
US 20190286973A1 · Kovvuri et al. · 2019 [cited by applicant]
US 20190324888A1 · Evans et al. · 2019 [cited by applicant]
US 20190385046A1 · Cassidy et al. · 2019 [cited by applicant]
US 20190391811A1 · Garegrat et al. · 2019 [cited by applicant]
US 20190392296A1 · Brady et al. · 2019 [cited by applicant]
US 20200020156A1 · Howson · 2020 [cited by applicant]
US 20200042856A1 · Datta et al. · 2020 [cited by applicant]
US 20200097289A1 · Eapen et al. · 2020 [cited by applicant]
US 20200142743A1 · Zhang et al. · 2020 [cited by applicant]
US 20200159692A1 · Shah et al. · 2020 [cited by applicant]
US 20200202198A1 · Lee et al. · 2020 [cited by applicant]
US 20200211262A1 · Vaidyanathan et al. · 2020 [cited by applicant]
US 20200225996A1 · Sharma et al. · 2020 [cited by applicant]
US 20200302305A1 · Oki et al. · 2020 [cited by applicant]
US 20200310994A1 · ChoFleming et al. · 2020 [cited by applicant]
US 20200349420A1 · Ovsiannikov et al. · 2020 [cited by applicant]
US 20200410327A1 · Chinya et al. · 2020 [cited by applicant]
US 20210011785A1 · Sanghi et al. · 2021 [cited by applicant]
US 20210192314A1 · Aarts et al. · 2021 [cited by applicant]
US 20210263853A1 · Waters et al. · 2021 [cited by applicant]
US 20210265015A1 · Parnaby et al. · 2021 [cited by applicant]
US 20220012058A1 · Hanrahan et al. · 2022 [cited by applicant]
US 20220100680A1 · Chrysos et al. · 2022 [cited by applicant]
US 20220188028A1 · Mesnier et al. · 2022 [cited by applicant]
TW 200736953A · 2007 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
WO 2019202216A2 · 2019 [cited by applicant]
WO 2021247614A1 · 2021 [cited by applicant]
WO 2022010810A1 · 2022 [cited by applicant]
WO 2022066639A1 · 2022 [cited by applicant]
Cong et al., Automatic Memory Partitioning and Scheduling for Throughput and Power Optimization, ACM Transactions on Design Automation of Electronic Systems, vol. 16, No. 2, Article 15, dated Mar. 2011, 25 pages. [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
List of Related Cases for U.S. Appl. No. 18/208,343, filed Jun. 12, 2023, 2 pages. [cited by applicant]
Lopez, DA, A Memory Layout for Dynamically Routed Capsule Layers, 16th International Conference on Information Technology—New Generations (ITNG), pp. 317-324, Published in May 2019. [doi:http://dx.doi.org/10.1007/978-3-… [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Pawłowski et al., High performance tensor-vector multiplies on shared memory systems, published in 2019, 24 pages. [cited by applicant]
PCT/US/2021/040382—International Search Report and Written Opinion, dated Nov. 29, 2021, 22 pages. [cited by applicant]
PCT/US2021/035305—International Search Report and Written Opinion dated Sep. 1, 2021, 17 pages. [cited by applicant]
PCT/US2021/051305—International Search Report and Written Opinion, dated Jan. 11, 2022, 16 pages. [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Rotem et al., Glow: Graph Lowering Compiler Techniques for Neural Networks, Cornell University Library, New York, dated May 2, 2018, 12 pages. [cited by applicant]
TW 110124802—First Office Action and Search Report dated May 24, 2022, 17 pages. [cited by applicant]
U.S. Appl. No. 16/922,975—Final Office Action, dated Mar. 9, 2023, 23 pages. [cited by applicant]
U.S. Appl. No. 16/922,975—Non-Final Office Action, dated Oct. 27, 2022, 26 pages. [cited by applicant]
U.S. Appl. No. 17/031,679—Notice of Allowance, dated Dec. 22, 2022, 13 pages. [cited by applicant]
U.S. Appl. No. 17/216,647—Notice of Allowance dated Aug. 13, 2021, 13 pages. [cited by applicant]
U.S. Appl. No. 17/216,647—Office Action dated Jun. 18, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,647—Response to Office Action dated Jun. 18, 2021, filed Jul. 29, 2021, 15 pages. [cited by applicant]
U.S. Appl. No. 17/216,650—Office Action dated Aug. 18, 2021, 25 pages. [cited by applicant]
U.S. Appl. No. 17/216,650—Office Action dated Jun. 30, 2021, 15 pages. [cited by applicant]
U.S. Appl. No. 17/216,650 Notice of Allowance, dated Feb. 16, 2022, 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,650 Response to Office Action dated Aug. 18, 2021, filed Sep. 9, 2021, 13 pages. [cited by applicant]
U.S. Appl. No. 17/216,650 Response to Office Action dated Jun. 30, 2021, filed Jul. 29, 2021, 17 pages. [cited by applicant]
U.S. Appl. No. 17/878,504—Non-Final Office Action, dated Nov. 16, 2022, 11 pages. [cited by applicant]
U.S. Appl. No. 17/878,504—Notice of Allowance, dated Mar. 3, 2023, 8 pages. [cited by applicant]
Wang et al., Memory Partitioning for Multidimensional Arrays in High-level Synthesis, DAC 2013, dated May 29-Jun. 7, 2013, 8 pages. [cited by applicant]
Wang et al., Theory and Algorithm for Generalized Memory Partitioning in High-Level Synthesis, FPGA 2014, dated Feb. 26-28, 2014, 10 pages. [cited by applicant]