IP Library Granted Patent US 12,511,252
Granted Patent B2
US 12,511,252 · App. 18/518,695 · Granted Dec 30, 2025

Lossless tiling in convolution networks—tiling configuration between two sections

Inventors: Tejas Nagendra Babu Nama (Palo Alto, CA); Ruddhi Chaphekar (Palo Alto, CA); Ram Sivaramakrishnan (San Jose, CA); Raghu Prabhakar (San Jose, CA); Sumti Jairath (Palo Alto, CA); Junjue Wang (Newark, CA); Kaizhao Liang (Palo Alto, CA); Adi Fuchs (Palo Alto, CA); Matheen Musaddiq (Austin, TX); Arvind Krishna Sujeeth (Palo Alto, CA)
Assignee: SambaNova Systems, Inc.
G06F15/7885G06F15/7839G06F16/9024G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,252
App. No.
18/518,695
Granted
Dec 30, 2025
Kind
B2
Abstract

Disclosed is a method that includes sectioning a graph into a sequence of sections, the sequence of sections including at least a first section followed by a second section. The first section is configured to generate a first output in a first target tiling configuration in response to processing a first input in a first input tiling configuration. The graph is configured to reconfigure the first output in the first target tiling configuration to a second input in a second input tiling configuration. The second section is configured to generate a second output in a second target tiling configuration in response to processing the second input in the second input tiling configuration.

Claims (71)

1 . A computer-implemented method comprising:

sectioning, using a processor, a processing graph for an application into a sequence of sections, the sequence of sections including at least a first section followed by a second section, the first section and the second section both respectively comprising one or more computational operations;

configuring the first section to generate, using a first reconfigurable processor as a part of executing the application, a first output in a first target tiling configuration in response to processing a first input in a first input tiling configuration;

configuring the processing graph to reconfigure the first output in the first target tiling configuration to a second input in a second input tiling configuration different than the first target tiling configuration; and

configuring the second section to generate, using a second reconfigurable processor as a part of executing the application, a second output in a second target tiling configuration in response to processing the second input in the second input tiling configuration.

2 . The method of claim 1 , further comprising:

saving, by the processor to a non-transitory computer-readable medium, configuration information for executing the application by the first section and the second section;

executing the first section of the application using the first reconfigurable processor; and

executing the second section of the application using the second reconfigurable processor.

3 . The method of claim 2 , further comprising:

reconfiguring the first output in the first target tiling configuration to the second input in the second input tiling configuration by:

storing a plurality of output tiles of the first output in the first target tiling configuration, with padding around the plurality of output tiles to generate a padded output, and

retiling the padded output to generate a plurality of input tiles of the second input in the second input tiling configuration.

4 . The method of claim 3 , further comprising storing the plurality of output tiles of the first output by:

storing the plurality of output tiles of the first output in a non-overlapping configuration.

5 . The method of claim 4 , further comprising retiling the padded output by:

retiling the padded output to generate the plurality of input tiles in an overlapping configuration.

6 . The method of claim 3 , wherein one or more input tiles of the plurality of input tiles of the second input have padding on one or more corresponding edges.

7 . The method of claim 3 , wherein:

a first input tile of the plurality of input tiles of the second input has padding only on a top edge and a left edge;

a second input tile of the plurality of input tiles of the second input has padding only on a top edge and a right edge;

a third input tile of the plurality of input tiles of the second input has padding only on a bottom edge and a left edge; and

a fourth input tile of the plurality of input tiles of the second input has padding only on a bottom edge and a right edge.

8 . The method of claim 7 , wherein:

a fifth input tile of the plurality of input tiles of the second input has padding only on a top edge;

a sixth input tile of the plurality of input tiles of the second input has padding only on a right edge;

a seventh input tile of the plurality of input tiles of the second input has padding only on a bottom edge;

an eighth input tile of the plurality of input tiles of the second input has padding only on a left edge; and

a ninth input tile of the plurality of input tiles of the second input does not have any padding of any of its edges.

9 . The method of claim 1 , wherein:

the first target tiling configuration tiles the first output into a first output set of non-overlapping tiles; and

the first input tiling configuration tiles the first input into a first input set of overlapping tiles.

10 . The method of claim 9 , wherein:

the first output set of non-overlapping tiles is generated by using tiles in the first input set of overlapping tiles as effective receptive fields.

11 . The method of claim 9 , wherein:

the second input tiling configuration tiles the second input into a second input set of overlapping tiles; and

the second target tiling configuration tiles the second output into a second output set of non-overlapping tiles.

12 . A non-transitory computer readable storage medium impressed with computer program instructions, the computer program instructions, when executed on a processor, implement a method comprising:

sectioning a processing graph for an application into a sequence of sections, the sequence of sections including at least a first section followed by a second section, the first section and the second section both respectively comprising one or more computational operations;

configuring the first section, using a first reconfigurable processor as a part of executing the application, to generate a first output in a first target tiling configuration in response to processing a first input in a first input tiling configuration;

configuring the processing graph to reconfigure the first output in the first target tiling configuration to a second input in a second input tiling configuration different than the first target tiling configuration; and

configuring the second section to generate, using a second reconfigurable processor as a part of executing the application, a second output in a second target tiling configuration in response to processing the second input in the second input tiling configuration.

13 . The non-transitory computer readable storage medium of claim 12 , wherein:

the first target tiling configuration tiles the first output into a first output set of non-overlapping tiles; and

the first input tiling configuration tiles the first input into a first input set of overlapping tiles.

14 . The non-transitory computer readable storage medium of claim 13 , wherein:

the first output set of non-overlapping tiles is generated by using tiles in the first input set of overlapping tiles as effective receptive fields.

15 . The non-transitory computer readable storage medium of claim 13 , wherein:

the second input tiling configuration tiles the second input into a second input set of overlapping tiles; and

the second target tiling configuration tiles the second output into a second output set of non-overlapping tiles.

16 . The non-transitory computer readable storage medium of claim 12 , the method further comprising:

reconfiguring the first output in the first target tiling configuration to the second input in the second input tiling configuration by:

storing a plurality of output tiles of the first output in the first target tiling configuration, with padding around the plurality of output tiles to generate a padded output, and

retiling the padded output to generate a plurality of input tiles of the second input in the second input tiling configuration.

17 . The non-transitory computer readable storage medium of claim 16 , the method further comprising:

storing the plurality of output tiles of the first output in a non-overlapping configuration; and

retiling the padded output to generate the plurality of input tiles in an overlapping configuration;

wherein one or more input tiles of the plurality of input tiles of the second input have padding on one or more corresponding edges.

18 . A data processing system, comprising:

a processor; and

a compiler configured to execute on the processor and to:

section a processing graph for an application into a sequence of sections, the sequence of sections including at least a first section and a second section, the first section and the second section both respectively comprising one or more computational operations;

configure the first section to generate, using a first reconfigurable processor as a part of executing the application, a first output in a first target tiling configuration in response to processing a first input in a first input tiling configuration,

configure the processing graph to reconfigure the first output in the first target tiling configuration to a second input in a second input tiling configuration different than the first target tiling configuration, and

configure the second section, using a second reconfigurable processor as a part of executing the application, to generate a second output in a second target tiling configuration in response to processing the second input in the second input tiling configuration.

19 . The data processing system of claim 18 , wherein:

the first target tiling configuration tiles the first output into a first output set of non-overlapping tiles; and

the first input tiling configuration tiles the first input into a first input set of overlapping tiles.

20 . The data processing system of claim 18 , the compiler further configured to:

store a plurality of output tiles of the first output in the first target tiling configuration, with padding around the plurality of output tiles to generate a padded output, and

retile the padded output to generate a plurality of input tiles of the second input in the second input tiling configuration.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2024
From: NAMA, TEJAS NAGENDRA BABU, MR.; CHAPHEKAR, RUDDHI; SIVARAMAKRISHNAN,, RAM; PRABHAKAR,, RAGHU; JAIRATH, SUMTI; WANG, JUNJUE; LIANG, KAIZHAO; FUCHS, ADI; MUSADDIQ, MATHEEN; SUJEETH, ARVIND KRISHNA
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 066498/0668 →
Continuity (3)
Continuation 17384515 · Jul 23, 2021
Continuation 17216657 · Mar 29, 2021
Related Publication 20240168913A1 · May 23, 2024
References Cited (108)
US 9646243B1 · Gokmen · 2017 [cited by applicant]
US 10310768B1 · Gauria et al. · 2019 [cited by applicant]
US 10607331B1 · Tandia et al. · 2020 [cited by applicant]
US 10812828B2 · Han et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 10990650B1 · Vantrease et al. · 2021 [cited by applicant]
US 20120303932A1 · Farabet · 2012 [cited by examiner]
US 20140007065A1 · Schmidt · 2014 [cited by examiner]
US 20140177941A1 · Docherty et al. · 2014 [cited by applicant]
US 20140201450A1 · Haugen et al. · 2014 [cited by applicant]
US 20150278642A1 · Chertok · 2015 [cited by examiner]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342888A1 · Yang et al. · 2016 [cited by applicant]
US 20160350645A1 · Brothers et al. · 2016 [cited by applicant]
US 20170083313A1 · Sankaralingam · 2017 [cited by examiner]
US 20170102901A1 · Burke · 2017 [cited by applicant]
US 20170103309A1 · Chang et al. · 2017 [cited by applicant]
US 20190138898A1 · Song et al. · 2019 [cited by applicant]
US 20190146497A1 · Urtasun et al. · 2019 [cited by applicant]
US 20190220742A1 · Kuo et al. · 2019 [cited by applicant]
US 20190266485A1 · Singh et al. · 2019 [cited by applicant]
US 20190273536A1 · Mccallister · 2019 [cited by applicant]
US 20190286973A1 · Kovvuri et al. · 2019 [cited by applicant]
US 20190294108A1 · Ozcan et al. · 2019 [cited by applicant]
US 20190295228A1 · Liu et al. · 2019 [cited by applicant]
US 20190347549A1 · Phanishayee et al. · 2019 [cited by applicant]
US 20190370631A1 · Fais et al. · 2019 [cited by applicant]
US 20200082215A1 · Aliabadi et al. · 2020 [cited by applicant]
US 20200082243A1 · Jin et al. · 2020 [cited by applicant]
US 20200159809A1 · Catthoor et al. · 2020 [cited by applicant]
US 20200160226A1 · Ross et al. · 2020 [cited by applicant]
US 20200193267A1 · Aydonat et al. · 2020 [cited by applicant]
US 20200234129A1 · Fagerholm et al. · 2020 [cited by applicant]
US 20200272892A1 · Desappan et al. · 2020 [cited by applicant]
US 20200279157A1 · Gao et al. · 2020 [cited by applicant]
US 20200302223A1 · Dutta et al. · 2020 [cited by applicant]
US 20210049463A1 · Ruff · 2021 [cited by applicant]
US 20210049804A1 · Sarel et al. · 2021 [cited by applicant]
US 20210097347A1 · Kwon et al. · 2021 [cited by applicant]
US 20210141571A1 · Lew et al. · 2021 [cited by applicant]
US 20210158167A1 · Sharma et al. · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210182676A1 · Zlateski et al. · 2021 [cited by applicant]
US 20210201124A1 · Gelashvili · 2021 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,657 Notice of Allowance, dated Oct. 20, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,657 Response to Office Action, dated Jun. 29, 2021, filed Jul. 12, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/364,110 Non Final Office Action, dated Jun. 30, 2023, 58 pages. [cited by applicant]
U.S. Appl. No. 17/364,129—Office Action dated Aug. 9, 2023, 62 pages. [cited by applicant]
U.S. Appl. No. 17/384,507—Non-final Rejection, dated May 5, 2023, 64 pages. [cited by applicant]
Venkataramanaiah et al., Automatic Compiler Based FPGA Accelerator for CNN Training, dated Aug. 15, 2019, 7 pages. [cited by applicant]
Versamopoulos et al., Decoding surface code with a distributed neural network based decoder, published Feb. 6, 2019, 12 pages. Retrieved from the internet [URL: https://researchgate.net/publication/330751526 ]. [cited by applicant]
Wang et al., FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters, Transactions on Computers, vol. 14, No. 8, dated Aug. 2020, 16 pages. [cited by applicant]
Werkhoven et al., Optimizing Convolution Operations in CUDA and Adaptive Tiling, published 2011, 12 pages. Retrieved from the internet. URL {https://api.semanticscholar.org/CorpusID:26235621}. [cited by applicant]
Zeng, Hanging, et., al, “GraphACT: Accelerating GCN Training on CPU FPGA Heterogeneous Platforms” Feb. 23-25, 2020 11 pages. [cited by applicant]
Zhou et al., Hierarchical Overlapped Tiling, published dt. Apr. 1, 2012 12 pages. [cited by applicant]
Zlateski et al., FFt Convolutions are faster than Winograd on Modern CPUs, Here's Why, dated Sep. 20, 2018, 17 pages. [cited by applicant]
Anonymous, BDS-GCN: Efficient full-graph training of graph convolutional nets with partition parallelism and boundary sampling, ICLR, 2021, 11 pages. [cited by applicant]
Anonymous, ECE Exam 1 study guide, Spring 2019, 25 pages. [cited by applicant]
Azarkhish et al., Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes, IEEE, dated Sep. 24, 2017, 15 pages. [cited by applicant]
Boris Ginzburg, Lecture 3: CNN: Backpropagation, Intel IMEC 2006, 18 pages. [cited by applicant]
Chen et al., Training Deep Nets with Sublinear Memory Cost, dated Apr. 22, 2016, 12 pages. [cited by applicant]
Clemons et al., A Patch memory system for Image Processing and Computer Vision, IEEE 2016, 13 pages. [cited by applicant]
Dathathri et al., CHET: Compiler and Runtime for Homomorphic Evaluation of Tensor Programs, arXiv:1810.00845v1, dated Oct. 1, 2018, 12 pgs. [cited by applicant]
Dumoulin et al., A guide to convolution arithmetic for deep learning, dated Jan. 12, 2018, 31 pages. [cited by applicant]
Gui et al., A Survey on Graph Processing Accelerators: Challenges and Opportunities, dated Feb. 26, 2019, 41 pages. [cited by applicant]
Hua et al., Reverse Engineering Convolutional Neural Networks Through side-channel Information Leaks, DAC 2018, Jun. 24-29, 2018, San Francisco, USA, 6 pages. [cited by applicant]
Jangda et al., Model-Based Warp overlapped Tiling for Image Processing Programs on GPUs, dated Sep. 8, 2020, 14 pages. [cited by applicant]
Kim et al., DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks, IEEE Transactions an Computer Aided Design of Integrated Circuits and Systems, dated Nov. 11, 2018, 11 pages. [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, PLDI '18, Jun. 18-22, 2018, Association for Computng Machinery, 16 pages. [cited by applicant]
Kong et al., Take it in your stride: Do we need striding in CNNs, Carnegie Mellon University, dated Dec. 7, 2017, 9 pages. [cited by applicant]
Li, Peng, et al., “Deep Convolutional Computation Model for Feature Learning on Big Data in Internet of Things”, IEEE Transactions on Industrial Informatics, vol. 14, No. 2, Feb. 2018, 9 pages. [cited by applicant]
List of Related Cases, Aug. 8, 2023, 2 pages. [cited by applicant]
Liu et al., An FPGA-Based CNN Accelerator Integrating Depthwise Separable Convolution, Electronics, published Mar. 3, 2019, 18 pages. [cited by applicant]
Liu et al., Memory-Efficient Architecture for Accelerating Generative Networks on FPGA, EPSRC and European Union Horizon 2020, 8 pages. [cited by applicant]
Luo, Yandong, et al., “AILC: Accelerate On-chip Incremental Learning with Compute-in-Memory Technology”, DOI 10.1109/TC.2021.3053199, IEEE Transactions on Computers, Jan. 20, 2021, 15 pages. [cited by applicant]
M. Emani et al., Accelerating Scientific Applications With Sambanova Reconfigurable Dataflow Architecture, in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 26, 2021, [doi: 10.1109/MCSE.2021.30572… [cited by applicant]
Ma, Hardware Acceleration of Deep Convolutional Neural Networks on FPGA, Arizona State University, Dec. 2018, 169 pages. [cited by applicant]
Mittal et al., A Survey on Hardware Accelerators and Optimization on Techniques for RNNs, Journal of Systems Architecture, dated Jul. 2020, 57 pages. [cited by applicant]
NVIDIA, CSE 599 I Accelerated Computing—Programming GPUs, Lecture 7, GPU teaching kit, Spring 2017, 61 pages. [cited by applicant]
Park et al., HetPipe: enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data Parallelism, Proceedings of the USENIX Annual Technical Conference, Ju… [cited by applicant]
Pinckaers et al., Streaming convolutional neural networks for end-to-end learning with multi-megapixel images, dated Nov. 11, 2019, 10 pages. [cited by applicant]
Pinckaers et al., Training Convolutional Neural Networks with Megapixel Images, dated Apr. 16, 2018, 3 pages. [cited by applicant]
Podlozhnyuk, Image Convolution with CUDA, NVIDIA, Jun. 2007, 21 pages. [cited by applicant]
Podobas et al., A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA, Jun. 24-28, 2017, 14 pages. [cited by applicant]
Ren et al., SBNet: Sparse Blocks Network for Fast Inference, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 8711-8720. [cited by applicant]
Schuiki et al., A scalable near-memory architecture for Training Deep Neural Networks on Large In-Memory Datasets, IEEE Transactions on Computers, vol. 68, No. 4, Apr. 2019, 14 pages. [cited by applicant]
Sinha, Sudipta N., et al., “Feature tracking and matching in video using programmable graphics hardware”, Jul. 16, 2006, 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,651 Non-Final Rejection, dated Jul. 13, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,651 Notice of Allowance, dated Aug. 5, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/216,651 Response to First Office Action, dated Jul. 13, 2021, filed Jul. 23, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Notice of Allowance, dated Aug. 23, 2021, 9 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Office Action, dated Aug. 2, 2021, 24 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Response to Office Action dated Aug. 2, filed Aug. 13, 2021, 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Office Action dated Sep. 1, 2021, 81 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Response to Office Action dated Sep. 1, 2021, filed Sep. 8, 2021, 18 pages. [cited by applicant]
U.S. Appl. No. 17/216,654 Notice of Allowance, dated Oct. 8, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Notice of Allowance dated Aug. 25, 2021, 9 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jun. 18, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jul. 28, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 9 pages. [cited by applicant]
U.S. Appl. No. 17/216,655 Response to Office Action, dated Jun. 18, 2021, filed Jul. 12, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Sep. 22, 2021, 22 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jul. 28, 2021, 21 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jun. 29, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action dated Sep. 22, 2021, filed Sep. 27, 2021, 13 pages. [cited by applicant]