IP Library Granted Patent US 12,321,843
Granted Patent B2
US 12,321,843 · App. 17/700,336 · Granted Jun 3, 2025

Lossless tiling in convolution networks—data flow logic

Inventors: Tejas Nagendra Babu Nama (Sunnyvale, CA); Ruddhi Chaphekar (Santa Clara, CA); Ram Sivaramakrishnan (San Jose, CA); Raghu Prabhakar (San Jose, CA); Sumti Jairath (Santa Clara, CA); Junjue Wang (San Mateo, CA); Kaizhao Liang (Palo Alto, CA); Adi Fuchs (West Windsor, NJ); Matheen Musaddiq (Austin, TX); Arvind Krishna Sujeeth (San Francisco, CA)
Assignee: SambaNova Systems, Inc.
G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,843
App. No.
17/700,336
Granted
Jun 3, 2025
Kind
B2
Abstract

A data processing system includes memory and reconfigurable processors, operatively coupled to the memory, configured to execute a sequence of subgraphs of a graph. The sequence of subgraphs includes a preceding subgraph and a succeeding subgraph. The data processing system also includes data flow logic, operatively coupled to the reconfigurable processors and the memory, configured to store a tiled output of the preceding subgraph as a composed input in the memory and make available parts of the composed input for processing by the succeeding subgraph.

Claims (43)

1. A data processing system, comprising:

one or more reconfigurable processors, operatively coupled to a memory, configured to execute a sequence of subgraphs of a processing graph to process an input tensor for machine learning and/or inference, the processing graph including two or more processing nodes connected by one or more edges representing dataflow between the two or more processing nodes, the sequence of subgraphs including a preceding subgraph and a succeeding subgraph each comprising a subsequence of at least one processing node of the two or more processing nodes; and

data flow logic configured to:

store, in the memory, a composed tensor generated from output tiles of the preceding subgraph having a first tiling configuration and comprising a respective portion of the composed tensor; and

make input tiles available for processing by the succeeding subgraph, the input tiles each generated from the composed tensor and having a second tiling configuration that is different from the first tiling configuration.

2. The data processing system of claim 1 , further comprising a host system, including the memory, operatively coupled to the one or more reconfigurable processors.

3. The data processing system of claim 1 , the first tiling configuration comprising non-overlapping tiles and the second tiling configuration comprising overlapping tiles, wherein a first tile of the input tiles overlaps with a second tile of the input tiles.

4. The data processing system of claim 1 , the data flow logic further configured to:

fill a portion of the memory with zeros, the portion of memory having a size larger than a sum of sizes of the output tiles; and

store the output tiles into the portion of the memory to form the composed tensor in the portion of the memory having a zero padding area around its perimeter;

wherein an edge tile of the input tiles includes some of the zero padding area.

5. The data processing system of claim 4 , the first tiling configuration comprising unpadded non-overlapping tiles and the second tiling configuration comprising overlapping tiles, wherein a first tile of the input tiles overlaps with a second tile of the input tiles.

6. The data processing system of claim 5 , the second tiling configuration having larger tiles than the first tiling configuration.

7. The data processing system of claim 1 , wherein the data flow logic is configured to store the output tiles in the memory by serially writing individual output tiles of the output tiles to the memory.

8. The data processing system of claim 1 , wherein the one or more reconfigurable processors each have a coarse-grained reconfigurable architecture comprising:

an array of configurable units, including a plurality of configurable compute units, a plurality of configurable memory units, and a plurality of address generation units, coupled together by an array-level network;

a top-level network coupled to the plurality of address generation units of the array of configurable units;

a first interface coupled between the top-level network and a first physical link that is coupled to a host system; and

a second interface coupled between the top-level network and an external memory interface coupled to at least a portion of the memory;

wherein the array of configurable units comprises the data flow logic.

9. A computer-implemented method to execute a sequence of subgraphs of a processing graph to process an input tensor for machine learning and/or inference using one or more reconfigurable processors, the processing graph including two or more processing nodes connected by one or more edges representing dataflow between the two or more processing nodes, the sequence of subgraphs including a preceding subgraph and a succeeding subgraph each comprising a subsequence of at least one processing node of the two or more processing nodes, the computer-implemented method comprising:

storing, in a memory, a composed tensor generated from output tiles of the preceding subgraph having a first tiling configuration and comprising a respective portion of the composed tensor; and

making input tiles generated from the composed tensor available for processing by the succeeding subgraph, the input tiles each generated from the composed tensor and having a second tiling configuration that is different from the first tiling configuration.

10. The computer-implemented method of claim 9 , the second tiling configuration comprising overlapping tiles, wherein a first tile of the input tiles overlaps with a second tile of the input tiles.

11. The computer-implemented method of claim 9 , further comprising:

filling a portion of the memory with zeros, the portion of memory having a size larger than a sum of sizes of the output tiles; and

storing the output tiles into the portion of the memory to form the composed tensor in the portion of the memory having a zero padding area around its perimeter;

wherein an edge tile of the input tiles includes some of the zero padding area.

12. The computer-implemented method of claim 11 , the first tiling configuration comprising unpadded non-overlapping tiles and the second tiling configuration comprising overlapping tiles, wherein a first tile of the input tiles overlaps with a second tile of the input tiles.

13. The computer-implemented method of claim 12 , the second tiling configuration having larger tiles than the first tiling configuration.

14. The computer-implemented method of claim 9 , further comprising storing the output tiles in the memory by serially writing individual output tiles of the output tiles to the memory.

15. A computer program product, the computer program product comprising a non-transitory computer readable storage medium having one or more configuration files embodied therewith, wherein the one or more configuration files, when loaded into one or more reconfigurable processors, cause the one or more reconfigurable processors to perform a method comprising:

executing a sequence of subgraphs of a processing graph to process an input tensor for machine learning and/or inference using one or more reconfigurable processors, the processing graph including two or more processing nodes connected by one or more edges representing dataflow between the two or more processing nodes, the sequence of subgraphs including a preceding subgraph and a succeeding subgraph each comprising a subsequence of at least one processing node of the two or more processing nodes;

storing, in a memory, a composed tensor generated from output tiles of the preceding subgraph having a first tiling configuration and comprising a respective portion of the composed tensor; and

reading input tiles from the composed tensor for processing by the succeeding subgraph, the input tiles having a second tiling configuration that is different from the first tiling configuration.

16. The computer program product of claim 15 , the second tiling configuration comprising overlapping tiles, wherein a first tile of the input tiles overlaps with a second tile of the input tiles.

17. The computer program product of claim 15 , the method further comprising:

filling a portion of the memory with zeros, the portion of memory having a size larger than a sum of sizes of the output tiles; and

storing the output tiles into the portion of the memory to form the composed tensor in the portion of the memory having a zero padding area around its perimeter;

wherein an edge tile of the input tiles includes some of the zero padding area.

18. The computer program product of claim 17 , the first tiling configuration comprising unpadded non-overlapping tiles and the second tiling configuration comprising overlapping tiles, wherein a first tile of the input tiles overlaps with a second tile of the input tiles.

19. The computer program product of claim 18 , the second tiling configuration having larger tiles than the first tiling configuration.

20. The computer program product of claim 15 , the method further comprising storing the output tiles in the memory by serially writing individual output tiles of the output tiles to the memory.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2022
From: NAMA, TEJAS NAGENDRA BABU; CHAPHEKAR, RUDDHI; SIVARAMAKRISHNAN, RAM; PRABHAKAR, RAGHU; JAIRATH, SUMTI; WANG, JUNJUE; LIANG, KAIZHAO; FUCHS, ADI; MUSADDIQ, MATHEEN; SUJEETH, ARVIND KRISHNA
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 059509/0704 →
Continuity (3)
Continuation 17364110 · Jun 30, 2021
Division 17216651 · Mar 29, 2021
Related Publication 20220309323A1 · Sep 29, 2022
References Cited (105)
US 9646243B1 · Gokmen · 2017 [cited by applicant]
US 10310768B1 · Gauria et al. · 2019 [cited by applicant]
US 10607331B1 · Tandia et al. · 2020 [cited by applicant]
US 10812828B2 · Han et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 10990650B1 · Vantrease et al. · 2021 [cited by applicant]
US 20140177941A1 · Docherty et al. · 2014 [cited by applicant]
US 20140201450A1 · Haugen · 2014 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342888A1 · Yang et al. · 2016 [cited by applicant]
US 20160350645A1 · Brothers et al. · 2016 [cited by applicant]
US 20170102901A1 · Burke · 2017 [cited by applicant]
US 20170103309A1 · Chang et al. · 2017 [cited by applicant]
US 20190138898A1 · Song et al. · 2019 [cited by applicant]
US 20190146497A1 · Urtasun et al. · 2019 [cited by applicant]
US 20190220742A1 · Kuo · 2019 [cited by examiner]
US 20190266485A1 · Singh · 2019 [cited by examiner]
US 20190273536A1 · McCallister · 2019 [cited by applicant]
US 20190286973A1 · Kovvuri · 2019 [cited by examiner]
US 20190294108A1 · Ozcan et al. · 2019 [cited by applicant]
US 20190295228A1 · Liu et al. · 2019 [cited by applicant]
US 20190347549A1 · Phanishayee et al. · 2019 [cited by applicant]
US 20190370631A1 · Fais · 2019 [cited by examiner]
US 20200082215A1 · Aliabadi et al. · 2020 [cited by applicant]
US 20200082243A1 · Jin et al. · 2020 [cited by applicant]
US 20200159809A1 · Catthoor et al. · 2020 [cited by applicant]
US 20200160226A1 · Ross et al. · 2020 [cited by applicant]
US 20200193267A1 · Aydonat et al. · 2020 [cited by applicant]
US 20200234129A1 · Fagerholm et al. · 2020 [cited by applicant]
US 20200272892A1 · Desappan et al. · 2020 [cited by applicant]
US 20200279157A1 · Gao et al. · 2020 [cited by applicant]
US 20200302223A1 · Dutta et al. · 2020 [cited by applicant]
US 20210049463A1 · Ruff · 2021 [cited by applicant]
US 20210049804A1 · Sarel et al. · 2021 [cited by applicant]
US 20210097347A1 · Kwon et al. · 2021 [cited by applicant]
US 20210141571A1 · Lew et al. · 2021 [cited by applicant]
US 20210158167A1 · Sharma et al. · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210182676A1 · Zlateski et al. · 2021 [cited by applicant]
US 20210201124A1 · Gelashvili · 2021 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jun. 29, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jun. 18, 2021, 12 pages. [cited by applicant]
Sinha, Sudipta N., et al., “Feature tracking and matching in video using programmable graphics hardware”, Jul. 16, 2006, 11 pages. [cited by applicant]
Li, Peng, et. al., “Deep Convolutional Computation Model for Feature Learning on Big Data in Internet of Things”, IEEE Transactions on Industrial Informatics, vol. 14, No. 2, Feb. 2018, 9 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA 2017, Jun. 24-28, 2017, Toronto, ON, Canada, 14 pages. [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, Proceedings of the 39th ACM SIGPLAN Conference On Programming Language Design And Embodiment (PLDI), Proceedings of the 43rd Internationa… [cited by applicant]
Luo, Yandong, et al., “AILC: Accelerate On-chip Incremental Learning with Compute-in-Memory Technology”, DOI 10.1109/TC.2021.3053199, IEEE Transactions on Computers, Jan. 20, 2021, 15 pages. [cited by applicant]
Jangda et al., Model-Based Warp overlapped Tiling for Image Processing Programs on GPUs, dated Sep. 8, 2020, 14 pages. [cited by applicant]
Dumoulin et. al., A guide to convolution arithmetic for deep learning, dated Jan. 12, 2018, 31 pages. [cited by applicant]
Chen et. al., Training Deep Nets with Sublinear Memory Cost, dated Apr. 22, 2016, 12 pages. [cited by applicant]
Azarkhish et. al., Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes, IEEE, dated Sep. 24, 2017, 15 pages. [cited by applicant]
Kong et. al., Take it in your stride: Do we need striding in CNNs, Carnegie Mellon University, dated Dec. 7, 2017, 9 pages. [cited by applicant]
Pinckaers et. al., Training Convolutional Neural Networks with Megapixel Images, dated Apr. 16, 2018, 3 pages. [cited by applicant]
Gui et. al., A Survey on Graph Processing Accelerators: Challenges and Opportunities, dated Feb. 26, 2019, 41 pages. [cited by applicant]
Pinckaers et. al., Streaming convolutional neural networks for end-to-end learning with multi-megapixel images, dated Nov. 11, 2019, 10 pages. [cited by applicant]
Kim et. al., DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks, IEEE Transactions on Computer Aided Design of Integrated Circuits and Systems, dated Nov. 11, 2018, 11 pages. [cited by applicant]
Schuiki et. al., A scalable near-memory architecture for Training Deep Neural Networks on Large In-Memory Datasets, IEEE Transactions on Computers, vol. 68, No. 4, Apr. 2019, 14 pages. [cited by applicant]
Liu et. al., An FPGA-Based CNN Accelerator Integrating Depthwise Separable Convolution, Electronics, published Mar. 3, 2019, 18 pages. [cited by applicant]
Park et. al., HetPipe: enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data Parallelism, Proceedings of the USENIX Annual Technical Conference, J… [cited by applicant]
Anonymous, BDS-GCN: Efficient full-graph training of graph convolutional nets with partition parallelism and boundary sampling, ICLR, 2021, 11 pages. [cited by applicant]
Venkataramanaiah et. al., Automatic Compiler Based FPGA Accelerator for CNN Training, dated Aug. 15, 2019, 7 pages. [cited by applicant]
Podlozhnyuk, Image Convolution with CUDA, NVIDIA, Jun. 2007, 21 pages. [cited by applicant]
NVIDIA, CSE 599 I Accelerated Computing—Programming GPUs, GPU teaching kit, 61 pages. [cited by applicant]
Hua et. al., Reverse Engineering Convolutional Neural Networks Through side-channel Information Leaks, DAC 2018, Jun. 24-29, 2018, San Francisco, USA, 6 pages. [cited by applicant]
Clemons et. al., A Patch memory system for Image Processing and Computer Vision, IEEE 2016, 13 pages. [cited by applicant]
Anonymous, ECE Exam 1 study guide, Spring 2019, 25 pages. [cited by applicant]
Zlateski et. al., FFt Convolutions are faster than Winograd on Modern CPUs, Here's Why, dated Sep. 20, 2018, 17 pages. [cited by applicant]
Zhou et. al., Hierarchical Overlapped Tiling, 12 pages. [cited by applicant]
Boris Ginzburg, Lecture 3: CNN: Backpropagation, Intel, 18 pages. [cited by applicant]
Ma, Hardware Acceleration of Deep Convolutional Neural Networks on FPGA, Arizona State University, Dec. 2018, 169 pages. [cited by applicant]
Versamopoulos et. al., Decoding surface code with a distributed neural network based decoder, published Feb. 6, 2019, 12 pages. Retrieved from the internet [URL: https://researchgate.net/publication/330751526 ]. [cited by applicant]
Liu et. al., Memory-Efficient Architecture for Accelerating Generative Networks on FPGA, EPSRC and European Union Horizon 2020, 8 pages. [cited by applicant]
Werkhoven et. al., Optimizing Convolution Operations in CUDA and Adaptive Tiling, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,651—Non-Final Rejection, dated Jul. 13, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action, dated Jun. 29, 2021, filed Jul. 12, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Response to Office Action, dated Jun. 18, 2021, filed Jul. 12, 2021, 12 pages. [cited by applicant]
M. Emani et al., “Accelerating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture,” in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 1-Apr. 2021, doi: 10.1109/MCSE.2021.3… [cited by applicant]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
U.S. Appl. No. 17/216,651—Response to First Office Action, dated Jul. 13, 2021, filed Jul. 23, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/216,651—Notice of Allowance, dated Aug. 5, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Office Action, dated Aug. 2, 2021, 24 pages. [cited by applicant]
Ren et. al., SBNet: Sparse Blocks Network for Fast Inference, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 8711-8720. [cited by applicant]
Dathathri et. al., CHET: Compiler and Runtime for Homomorphic Evaluation of Tensor Programs, arXiv:1810.00845v1, dated Oct. 1, 2018. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jul. 28, 2021, 28 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jul. 28, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Office Action dated Sep. 1, 2021, 81 pages. [cited by applicant]
Zeng, Hanging, et., al, “GraphACT: Accelerating GCN Training on CPU FPGA Heterogeneous Platforms” Feb. 23-25, 2020 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Sep. 22, 2021, 22 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Notice of Allowance dated Aug. 25, 2021, 9 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action dated Sep. 22, 2021, filed Sep. 27, 2021, 13 pages. [cited by applicant]
U.S. Appl. No. 17/216,657 Notice of Allowance, dated Oct. 20, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Response to Office Action dated Aug. 2, filed Aug. 13, 2021, 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Notice of Allowance, dated Aug. 23, 2021, 9 pages. [cited by applicant]
Mittal et. al., A Survey on Hardware Accelerators and Optimization on Techniques for RNNs, Journal of Systems Architecture, dated Jul. 2020, 57 pages. [cited by applicant]
Wang et. al., FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters, Transactions on Computers, vol. 14, No. 8, dated Aug. 2020, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Response to Office Action dated Sep. 1, 2021, filed Sep. 8, 2021, 18 pages. [cited by applicant]
U.S. Appl. No. 17/216,654 Notice of Allowance, dated Oct. 8, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 9 pages. [cited by applicant]
List of Related Cases, Aug. 8, 2023, 2 pages. [cited by applicant]
TW 111111672—Office Action dated Sep. 14, 2022, 5 pgs. [cited by applicant]
U.S. Appl. No. 17/364,110 Non Final Office Action, dated Jun. 30, 2023, 58 pages. [cited by applicant]
U.S. Appl. No. 17/364,129—Office Action dated Aug. 9, 2023, 62 pages. [cited by applicant]
U.S. Appl. No. 17/384,507—Non-final Rejection, dated May 5, 2023, 64 pages. [cited by applicant]
Cited By (1)
US 12,639,255