IP Library Granted Patent US 12,488,218
Granted Patent B2
US 12,488,218 · App. 17/364,141 · Granted Dec 2, 2025

Lossless tiling in convolution networks—padding and re-tilling at section boundaries

Inventors: Tejas Nagendra Babu Nama (Sunnyvale, CA); Ruddhi Chaphekar (Santa Clara, CA); Ram Sivaramakrishnan (San Jose, CA); Raghu Prabhakar (San Jose, CA); Sumti Jairath (Santa Clara, CA); Junjue Wang (San Mateo, CA); Kaizhao Liang (Palo Alto, CA); Adi Fuchs (West Windsor, NJ); Matheen Musaddiq (Austin, TX); Arvind Krishna Sujeeth (San Francisco, CA)
Assignee: SambaNova Systems, Inc.
G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,218
App. No.
17/364,141
Granted
Dec 2, 2025
Kind
B2
Abstract

Disclosed is a method that includes generating by an output processing node of a first section of a processing graph, a plurality of output tiles of an output tensor. The plurality of output tiles of the output tensor is written in a memory, where the writing includes zero-padding the plurality of output tiles of the output tensor in the memory. The zero-padded plurality of output tiles of the output tensor are tiled, to generate a plurality of input tiles of an input tensor. The plurality of input tiles of the input tensor is processed in a second section of the processing graph.

Claims (74)

1 . A non-transitory computer readable storage medium impressed with computer program instructions that, when executed on a processor, implement a method comprising:

generating by an output processing node of a first section of a processing graph, a plurality of output tiles of an output tensor, individual output tiles in the plurality of output tiles having a first size;

writing the plurality of output tiles of the output tensor in a memory, wherein the writing comprises zero-padding the plurality of output tiles of the output tensor in the memory to create a zero-padded plurality of output tiles;

tiling the zero-padded plurality of output tiles of the output tensor to generate a plurality of input tiles of an input tensor, individual input files in the plurality of input tiles having a second size that is larger than the first size; and

processing the plurality of input tiles of the input tensor in a second section of the processing graph.

2 . The non-transitory computer readable storage medium of claim 1 , further comprising:

initializing a plurality of memory locations to zero, the plurality of memory locations including (i) a first subset of memory locations, and (ii) a second subset of memory locations surrounding the first subset of memory locations,

wherein writing the plurality of output tiles comprises writing the plurality of output tiles of the output tensor in the first subset of memory locations in the memory, and

wherein the plurality of output tiles in the first subset of memory locations is surrounded by zeros in the second subset of memory locations.

3 . The non-transitory computer readable storage medium of claim 2 , wherein the zero-padded plurality of output tiles of the output tensor comprises an aggregate of (i) the plurality of output tiles in the first subset of memory locations and (ii) the zeros in the second subset of memory locations.

4 . The non-transitory computer readable storage medium of claim 2 , wherein writing the plurality of output tiles further comprises:

sequentially writing the plurality of output tiles of the output tensor in the first subset of memory locations in the memory, such that a first output tile of the plurality of output tiles of the output tensor is written to a first section of the first subset of memory locations, followed by writing of a second output tile of the plurality of output tiles of the output tensor to a second section of the first subset of memory locations.

5 . The non-transitory computer readable storage medium of claim 2 , wherein writing the plurality of output tiles further comprises:

at least in part parallelly writing the plurality of output tiles of the output tensor in the first subset of memory locations in the memory, such that a first output tile of the plurality of output tiles of the output tensor is written to a first section of the first subset of memory locations at least in part simultaneously with writing of a second output tile of the plurality of output tiles of the output tensor to a second section of the first subset of memory locations.

6 . The non-transitory computer readable storage medium of claim 2 , wherein tiling the zero-padded plurality of output tiles of the output tensor comprises:

tiling a combination of (i) the plurality of output tiles of the output tensor in the first subset of memory locations and (ii) the zeros in the second subset of memory locations surrounding the plurality of output tiles of the output tensor.

7 . The non-transitory computer readable storage medium of claim 1 , wherein:

one or more first input tiles of the plurality of input tiles of the input tensor have zero padding along one or more edges, and one or more second input tiles of the plurality of input tiles of the input tensor do not have zero padding along any edge.

8 . The non-transitory computer readable storage medium of claim 7 , wherein:

the one or more first input tiles of the plurality of input tiles of the input tensor have zero padding along those edges that coincide with edges of the input tensor.

9 . The non-transitory computer readable storage medium of claim 1 , wherein:

a first input tile of the plurality of input tiles of the input tensor has zero padding only on a top edge and a left edge;

a second input tile of the plurality of input tiles of the input tensor has zero padding only on a top edge and a right edge;

a third input tile of the plurality of input tiles of the input tensor has zero padding only on a bottom edge and a left edge; and

a fourth input tile of the plurality of input tiles of the input tensor has zero padding only on a bottom edge and a right edge.

10 . The non-transitory computer readable storage medium of claim 9 , wherein:

a fifth input tile of the plurality of input tiles of the input tensor has zero padding only on a top edge;

a sixth input tile of the plurality of input tiles of the input tensor has zero padding only on a right edge;

a seventh input tile of the plurality of input tiles of the input tensor has zero padding only on a bottom edge;

an eighth input tile of the plurality of input tiles of the input tensor has zero padding only on a left edge; and

a ninth input tile of the plurality of input tiles of the input tensor does not have any zero padding of any of its edges.

11 . The non-transitory computer readable storage medium of claim 1 , wherein:

the plurality of output tiles of the output tensor is non-overlapping tiles; and

the plurality of input tiles of the input tensor is overlapping tiles.

12 . A computer implemented method, comprising:

generating by an output processing node of a first section of a processing graph, a plurality of output tiles of an output tensor, individual output tiles in the plurality of output tiles having a first size;

writing the plurality of output tiles of the output tensor in a memory, wherein the writing comprises zero-padding the plurality of output tiles of the output tensor in the memory to create a zero-padded plurality of output tiles;

tiling the zero-padded plurality of output tiles of the output tensor to generate a plurality of input tiles of an input tensor, individual input files in the plurality of input tiles having a second size that is larger than the first size; and

processing the plurality of input tiles of the input tensor in a second section of the processing graph.

13 . The method of claim 12 , further comprising:

initializing a plurality of memory locations to zero, the plurality of memory locations including (i) a first subset of memory locations, and (ii) a second subset of memory locations surrounding the first subset of memory locations,

wherein writing the plurality of output tiles comprises writing the plurality of output tiles of the output tensor in the first subset of memory locations in the memory, and

wherein the plurality of output tiles in the first subset of memory locations is surrounded by zeros in the second subset of memory locations.

14 . The method of claim 12 , wherein:

one or more first input tiles of the plurality of input tiles of the input tensor have zero padding along one or more edges that coincide with edges of the input tensor, and one or more second input tiles of the plurality of input tiles of the input tensor do not have zero padding along any edge.

15 . The method of claim 12 , wherein:

a first input tile of the plurality of input tiles of the input tensor has zero padding only on a top edge and a left edge;

a second input tile of the plurality of input tiles of the input tensor has zero padding only on a top edge and a right edge;

a third input tile of the plurality of input tiles of the input tensor has zero padding only on a bottom edge and a left edge;

a fourth input tile of the plurality of input tiles of the input tensor has zero padding only on a bottom edge and a right edge;

a fifth input tile of the plurality of input tiles of the input tensor has zero padding only on a top edge;

a sixth input tile of the plurality of input tiles of the input tensor has zero padding only on a right edge;

a seventh input tile of the plurality of input tiles of the input tensor has zero padding only on a bottom edge;

an eighth input tile of the plurality of input tiles of the input tensor has zero padding only on a left edge; and

a ninth input tile of the plurality of input tiles of the input tensor does not have any zero padding of any of its edges.

16 . The method of claim 12 , wherein:

the plurality of output tiles of the output tensor is non-overlapping tiles; and

the plurality of input tiles of the input tensor is overlapping tiles.

17 . A data processing system, comprising:

runtime logic configured to:

generate, at an output processing node of a first section of a processing graph, a plurality of output tiles of an output tensor, individual output tiles in the plurality of output tiles having a first size,

write the plurality of output tiles of the output tensor in a memory, wherein the writing comprises zero-padding the plurality of output tiles of the output tensor in the memory to create a zero-padded plurality of output tiles,

tile the zero-padded plurality of output tiles of the output tensor to generate a plurality of input tiles of an input tensor, individual input files in the plurality of input tiles having a second size that is larger than the first size, and

process the plurality of input tiles of the input tensor in a second section of the processing graph.

18 . The data processing system of claim 17 , wherein the runtime logic is further configured to:

initialize a plurality of memory locations to zero, the plurality of memory locations including (i) a first subset of memory locations, and (ii) a second subset of memory locations surrounding the first subset of memory locations; and

write the plurality of output tiles of the output tensor in the first subset of memory locations in the memory,

wherein the plurality of output tiles in the first subset of memory locations is surrounded by zeros in the second subset of memory locations.

19 . The data processing system of claim 18 , wherein:

the first section and the second section of the processing graph are executed in one or more processing units that are in an Integrated Circuit (IC) chip; and

the memory is external to the IC chip.

20 . The data processing system of claim 17 , wherein:

the plurality of output tiles of the output tensor is non-overlapping tiles; and

the plurality of input tiles of the input tensor is overlapping tiles.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2022
From: NAMA, TEJAS NAGENDRA BABU; CHAPHEKAR, RUDDHI; SIVARAMAKRISHNAN, RAM; PRABHAKAR, RAGHU; JAIRATH, SUMTI; WANG, JUNJUE; LIANG, KAIZHAO; FUCHS, ADI; MUSADDIQ, MATHEEN; SUJEETH, ARVIND KRISHNA
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 059697/0670 →
Continuity (2)
Division 17216652 · Mar 29, 2021
Related Publication 20220309318A1 · Sep 29, 2022
References Cited (107)
US 9646243B1 · Gokmen · 2017 [cited by applicant]
US 10310768B1 · Gauria et al. · 2019 [cited by applicant]
US 10607331B1 · Tandia et al. · 2020 [cited by applicant]
US 10812828B2 · Han et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 10990650B1 · Vantrease et al. · 2021 [cited by applicant]
US 20140177941A1 · Docherty et al. · 2014 [cited by applicant]
US 20140201450A1 · Haugen · 2014 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342888A1 · Yang et al. · 2016 [cited by applicant]
US 20160350645A1 · Brothers et al. · 2016 [cited by applicant]
US 20170102901A1 · Burke · 2017 [cited by applicant]
US 20170103309A1 · Chang et al. · 2017 [cited by applicant]
US 20190138898A1 · Song et al. · 2019 [cited by applicant]
US 20190146497A1 · Urtasun et al. · 2019 [cited by applicant]
US 20190220742A1 · Kuo et al. · 2019 [cited by applicant]
US 20190273536A1 · McCallister · 2019 [cited by applicant]
US 20190294108A1 · Ozcan et al. · 2019 [cited by applicant]
US 20190295228A1 · Liu et al. · 2019 [cited by applicant]
US 20190347549A1 · Phanishayee et al. · 2019 [cited by applicant]
US 20190370631A1 · Fais et al. · 2019 [cited by applicant]
US 20200082215A1 · Aliabadi et al. · 2020 [cited by applicant]
US 20200082243A1 · Jin et al. · 2020 [cited by applicant]
US 20200159809A1 · Catthoor et al. · 2020 [cited by applicant]
US 20200160226A1 · Ross et al. · 2020 [cited by applicant]
US 20200193267A1 · Aydonat et al. · 2020 [cited by applicant]
US 20200234129A1 · Fagerholm et al. · 2020 [cited by applicant]
US 20200272892A1 · Desappan et al. · 2020 [cited by applicant]
US 20200279157A1 · Gao et al. · 2020 [cited by applicant]
US 20200302223A1 · Dutta et al. · 2020 [cited by applicant]
US 20210049463A1 · Ruff · 2021 [cited by applicant]
US 20210049804A1 · Sarel et al. · 2021 [cited by applicant]
US 20210097347A1 · Kwon et al. · 2021 [cited by applicant]
US 20210141571A1 · Lew et al. · 2021 [cited by applicant]
US 20210158167A1 · Sharma et al. · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210182676A1 · Zlateski et al. · 2021 [cited by applicant]
US 20210201124A1 · Gelashvili · 2021 [cited by applicant]
US 20220309336A1 · Minkin · 2022 [cited by examiner]
EP 3282398A1 · 2018 [cited by examiner]
WO 2010142987A1 · 2010 [cited by applicant]
Le et al. “Tiled convolutional neural networks.” Advances in neural information processing systems 23 (2010); Total pp. 6 (Year: 2010). [cited by examiner]
Lin et al., GrateTile: Efficient Sparse Tensor Tiling for CNN Processing, arXiv:2009.08685v1 [cs.LG] Sep. 18, 2020; Total pp. 6 (Year: 2020). [cited by examiner]
Podobas et al, A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
U.S. Appl. No. 17/216,651 Response to First Office Action, dated Jul. 13, 2021, filed Jul. 23, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/216,651 Notice of Allowance, dated Aug. 5, 2021, 14 pages. [cited by applicant]
Ren et. al., SBNet: Sparse Blocks Network for Fast Inference, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 8711-8720. [cited by applicant]
Dathathri et. al., CHET: Compiler and Runtime for Homomorphic Evaluation of Tensor Programs, arXiv:1810.00845v1, dated Oct. 1, 2018. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jul. 28, 2021, 28 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jul. 28, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Office Action dated Sep. 1, 2021, 81 pages. [cited by applicant]
Zeng, Hanqing, et., al, “GraphACT: Accelerating GCN Training on CPU FPGA Heterogeneous Platforms” Feb. 23-25, 2020 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Sep. 22, 2021, 22 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Notice of Allowance dated Aug. 25, 2021, 9 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action dated Sep. 22, 2021, filed Sep. 27, 2021, 13 pages. [cited by applicant]
U.S. Appl. No. 17/216,657 Notice of Allowance, dated Oct. 20, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Response to Office Action dated Aug. 2, filed Aug. 13, 2021, 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Notice of Allowance, dated Aug. 23, 2021, 9 pages. [cited by applicant]
Mittal et. al., A Survey on Hardware Accelerators and Optimization on Techniques for RNNs, Journal of Systems Architecture, dated Jul. 2020, 57 pages. [cited by applicant]
Wang et. al., FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters, Transactions on Computers, vol. 14, No. 8, dated Aug. 2020, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Response to Office Action dated Sep. 1, 2021, filed Sep. 8, 2021, 18 pages. [cited by applicant]
U.S. Appl. No. 17/216,654 Notice of Allowance, dated Oct. 8, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 9 pages. [cited by applicant]
List of Related Cases, Aug. 8, 2023, 2 pages. [cited by applicant]
TW 111111672—Office Action dated Sep. 14, 2022, 5 pgs. [cited by applicant]
U.S. Appl. No. 17/364,110 Non Final Office Action, dated Jun. 30, 2023, 58 pages. [cited by applicant]
U.S. Appl. No. 17/364,129—Office Action dated Aug. 9, 2023, 62 pages. [cited by applicant]
U.S. Appl. No. 17/384,507—Non-final Rejection, dated May 5, 2023, 64 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jun. 29, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jun. 18, 2021, 12 pages. [cited by applicant]
Sinha, Sudipta N., et. al., “Feature tracking and matching in video using programmable graphics hardware”, Jul. 16, 2006, 11 pages. [cited by applicant]
Li, Peng, et. al., “Deep Convolutional Computation Model for Feature Learning on Big Data in Internet of Things”, IEEE Transactions on Industrial Informatics, vol. 14, No. 2, Feb. 2018, 9 pages. [cited by applicant]
Prabhakar et. al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA 2017, Jun. 24-28, 2017, Toronto, ON, Canada. [cited by applicant]
Koeplinger et. al., Spatial: A Language and Compiler for Application Accelerators, Proceedings of the 39th ACM SIGPLAN Conference On Programming Language Design And Embodiment (PLDI), Proceedings of the 43rd Internation… [cited by applicant]
Luo, Yandong, et. al., “AILC: Accelerate On-chip Incremental Learning with Compute-in-Memory Technology”, DOI 10.1109/TC.2021.3053199, IEEE Transactions on Computers, Jan. 20, 2021, 15 pages. [cited by applicant]
Jangda et. al., Model-Based Warp overlapped Tiling for Image Processing Programs on GPUs, dated Sep. 8, 2020, 14 pages. [cited by applicant]
Dumoulin et. al., A guide to convolution arithmetic for deep learning, dated Jan. 12, 2018, 31 pages. [cited by applicant]
Chen et. al., Training Deep Nets with Sublinear Memory Cost, dated Apr. 22, 2016, 12 pages. [cited by applicant]
Azarkhish et. al., Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes, IEEE, dated Sep. 24, 2017, 15 pages. [cited by applicant]
Kong et. al., Take it in your stride: Do we need striding in CNNs, Carnegie Mellon University, dated Dec. 7, 2017, 9 pages. [cited by applicant]
Pinckaers et. al., Training Convolutional Neural Networks with Megapixel Images, dated Apr. 16, 2018, 3 pages. [cited by applicant]
Gui et. al., A Survey on Graph Processing Accelerators: Challenges and Opportunities, dated Feb. 26, 2019, 41 pages. [cited by applicant]
Pinckaers et. al., Streaming convolutional neural networks for end-to-end learning with multi-megapixel images, dated Nov. 11, 2019, 10 pages. [cited by applicant]
Kim et. al., DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks, IEEE Transactions on Computer Aided Design of Integrated Circuits and Systems, dated Nov. 11, 2018, 11 pages. [cited by applicant]
Schuiki et. al., A scalable near-memory architecture for Training Deep Neural Networks on Large In-Memory Datasets, IEEE Transactions on Computers, vol. 68, No. 4, Apr. 2019, 14 pages. [cited by applicant]
Liu et. al., An FPGA-Based CNN Accelerator Integrating Depthwise Separable Convolution, Electronics, published Mar. 3, 2019, 18 pages. [cited by applicant]
Park et. al., HetPipe: enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data Parallelism, Proceedings of the USENIX Annual Technical Conference, J… [cited by applicant]
Anonymous, BDS-GCN: Efficient full-graph training of graph convolutional nets with partition parallelism and boundary sampling, ICLR, 2021, 11 pages. [cited by applicant]
Venkataramanaiah et. al., Automatic Compiler Based FPGA Accelerator for CNN Training, dated Aug. 15, 2019, 7 pages. [cited by applicant]
Podlozhnyuk, Image Convolution with Cuda, NVIDIA, Jun. 2007, 21 pages. [cited by applicant]
NVIDIA, CSE 599 I Accelerated Computing—Programming GPUs, GPU teaching kit, 61 pages. [cited by applicant]
Hua et. al., Reverse Engineering Convolutional Neural Networks Through side-channel Information Leaks, DAC 2018, Jun. 24-29, 2018, San Francisco, USA, 6 pages. [cited by applicant]
Clemons et. al., A Patch memory system for Image Processing and Computer Vision, IEEE 2016, 13 pages. [cited by applicant]
Anonymous, ECE Exam 1 study guide, Spring 2019, 25 pages. [cited by applicant]
Zlateski et. al., FFt Convolutions are faster than Winograd on Modern CPUs, Here's Why, dated Sep. 20, 2018, 17 pages. [cited by applicant]
Zhou et. al., Hierarchical Overlapped Tiling, 12 pages. [cited by applicant]
Boris Ginzburg, Lecture 3: CNN: Backpropagation, Intel, 18 pages. [cited by applicant]
Ma, Hardware Acceleration of Deep Convolutional Neural Networks on FPGA, Arizona State University, Dec. 2018, 169 pages. [cited by applicant]
Versamopoulos et. al., Decoding surface code with a distributed neural network based decoder, published Feb. 6, 2019, 12 pages. Retrieved from the internet [URL: https://researchgate.net/publication/330751526 ]. [cited by applicant]
Liu et. al., Memory-Efficient Architecture for Accelerating Generative Networks on FPGA, EPSRC and European Union Horizon 2020, 8 pages. [cited by applicant]
Werkhoven et. al., Optimizing Convolution Operations in CUDA and Adaptive Tiling, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,651 Non-Final Rejection, dated Jul. 13, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,657 Response to Office Action, dated Jun. 29, 2021, filed Jul. 12, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,655 Response to Office Action, dated Jun. 18, 2021, filed Jul. 12, 2021, 12 pages. [cited by applicant]
M. Emani et al., “Accelerating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture,” in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 1-Apr. 2021, doi: 10.1109/MCSE.2021.3… [cited by applicant]
U.S. Appl. No. 17/216,652 Non-Final Rejection, dated Aug. 2, 2021, 24 pages. [cited by applicant]