IP Library Granted Patent US 12,541,669
Granted Patent B2
US 12,541,669 · App. 17/477,409 · Granted Feb 3, 2026

Lossless tiling in convolution networks—section boundaries

Inventors: Tejas Nagendra Babu Nama (Sunnyvale, CA); Ruddhi Chaphekar (Santa Clara, CA); Ram Sivaramakrishnan (San Jose, CA); Raghu Prabhakar (San Jose, CA); Sumti Jairath (Santa Clara, CA); Junjue Wang (San Mateo, CA); Kaizhao Liang (Palo Alto, CA); Adi Fuchs (West Windsor, NJ); Matheen Musaddiq (Austin, TX); Arvind Krishna Sujeeth (San Francisco, CA)
Assignee: SambaNova Systems, Inc.
G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,669
App. No.
17/477,409
Granted
Feb 3, 2026
Kind
B2
Abstract

Disclosed is a data processing system that includes compile time logic to section a graph into a sequence of sections, configure a first section to generate a first set of output tiles in a first target tiling configuration in response to processing a first set of input tiles in a first input tiling configuration, and configure a second section to generate a second set of output tiles in a second target tiling configuration in response to processing the first set of output tiles in a second input tiling configuration. Runtime logic is configured to pad a first input into a first padded input, read the first set of input tiles from the first padded input in the first input tiling configuration, and process the first set of input tiles through the first section to generate the first set of output tiles in the first target tiling configuration.

Claims (59)

1 . A computer implemented method comprising:

processing, in a first section of a processing graph, a first set of input tiles in a first input tiling configuration, to generate a first set of output tiles in a first target tiling configuration;

generating, from the first set of output tiles in the first target tiling configuration, a second set of input tiles in a second input tiling configuration;

reading, by a second section of the processing graph, the second set of input tiles in the second input tiling configuration that is different from the first target tiling configuration; and

processing, in the second section of the processing graph, the second set of input tiles in the second input tiling configuration, to generate a second set of output tiles in a second target tiling configuration;

wherein individual output tiles in the first set of output tiles have a first size; and

individual input tiles in the second set of input tiles have a second size that is larger than the first size.

2 . The method of claim 1 , wherein generating, from the first set of output tiles in the first target tiling configuration, the second set of input tiles in the second input tiling configuration comprises:

writing the first set of output tiles in a memory, with zero-padding applied to the first set of output tiles, to create a zero-padded first set of output tiles; and

retiling the zero-padded first set of output tiles in the memory, to generate the second set of input tiles in the second input tiling configuration.

3 . The method of claim 2 , wherein:

writing the first set of output tiles in the memory comprises writing the first set of output tiles in the memory in a non-overlapping configuration; and

retiling the zero-padded first set of output tiles comprises retiling the zero-padded first set of output tiles in an overlapping configuration.

4 . The method of claim 2 , wherein writing the first set of output tiles in the memory comprises:

initializing an area of the memory with zeros; and

writing individual tiles of the first set of output tiles in the area of the memory that was initialized with the zeros.

5 . The method of claim 4 , wherein the area of the memory that was initialized with the zeros is larger than a combined size of individual tiles of the first set of output tiles, such that a periphery of the area of the memory that was initialized with the zeros forms the zero-padding applied around the first set of output tiles in the memory.

6 . The method of claim 1 , further comprising:

compiling the processing graph by sectioning the processing graph into a sequence of sections, the sequence of sections including at least the first section and the second section.

7 . The method of claim 1 , wherein:

the first target tiling configuration configures neighboring tiles in the first set of output tiles to be non-overlapping with each other.

8 . The method of claim 1 , wherein:

the second input tiling configuration configures at least a first tile and a neighboring second tile in the second set of input tiles to be overlapping with each other.

9 . The method of claim 1 , wherein:

the first input tiling configuration configures at least a first tile and a second tile of the first set of input tiles to be overlapping with each other.

10 . The method of claim 1 , wherein:

the second target tiling configuration configures neighboring tiles in the second set of output tiles to be non-overlapping with each other.

11 . A data processing system, comprising:

storage medium storing runtime logic; and

one or more processors coupled to the storage medium, the one or more processors executing the runtime logic,

wherein the runtime logic executed by the one or more processors is configured to:

cause a first section of a processing graph to process a first set of input tiles in a first input tiling configuration, to generate a first set of output tiles in a first target tiling configuration,

cause to generate, from the first set of output tiles in the first target tiling configuration, a second set of input tiles in a second input tiling configuration, the second input tiling configuration different from the first target tiling configuration,

cause a second section of the processing graph to read the second set of input tiles in the second input tiling configuration, and

cause the second section of the processing graph to process the second set of input tiles in the second input tiling configuration, to generate a second set of output tiles in a second target tiling configuration;

wherein individual output tiles in the first set of output tiles have a first size; and

individual input tiles in the second set of input tiles have a second size that is larger than the first size.

12 . The data processing system of claim 11 , wherein to generate the second set of input tiles from the first set of output tiles, the runtime logic is configured to:

write the first set of output tiles in a memory, with zero-padding applied to the first set of output tiles, to create a zero-padded first set of output tiles; and

retile the zero-padded first set of output tiles in the memory, to generate the second set of input tiles in the second input tiling configuration.

13 . The data processing system of claim 12 , wherein:

to write the first set of output tiles in the memory, the runtime logic is configured to write the first set of output tiles in the memory in a non-overlapping configuration; and

to retile the zero-padded first set of output tiles, the runtime logic is configured to retile the zero-padded first set of output tiles in an overlapping configuration, such that a tile in the second set of input tiles overlaps with one or more neighboring tiles in the second set of input tiles.

14 . The data processing system of claim 11 , wherein:

the first target tiling configuration configures neighboring tiles in the first set of output tiles to be non-overlapping with each other; and

the second input tiling configuration configures at least a first tile and a neighboring second tile in the second set of input tiles to be overlapping with each other.

15 . The data processing system of claim 11 , wherein the storage medium stores compile time logic, wherein the one or more processors execute the compile time logic, and wherein the compile time logic executed by the one or more processors is configured to:

compile the processing graph by sectioning the processing graph into a sequence of sections, the sequence of sections including at least the first section and the second section.

16 . A non-transitory computer readable storage medium impressed with computer program instructions, the instructions, when executed on a processor, implement a method comprising:

generating, by a first section of a processing graph, a first set of output tiles of an output tensor in a first target tiling configuration;

zero-padding and retiling the first set of output tiles, to generate a set of input tiles of an input tensor in an input tiling configuration, the input tiling configuration different from the first target tiling configuration;

reading, by a second section of the processing graph, the set of input tiles in the input tiling configuration; and

processing, in the second section of the processing graph, the set of input tiles in the input tiling configuration, to generate a second set of output tiles in a second target tiling configuration.

17 . The non-transitory computer readable storage medium of claim 16 , wherein:

the first target tiling configuration configures neighboring tiles in the first set of output tiles to be non-overlapping with each other; and

the input tiling configuration configures at least a first tile and a neighboring second tile in the set of input tiles to be overlapping with each other.

18 . The non-transitory computer readable storage medium of claim 16 , wherein:

individual tiles in the first set of output tiles have a first size; and

individual tiles in the set of input tiles have a second size that is larger than the first size.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2026
From: NAMA, TEJAS NAGENDRA BABU; CHAPHEKAR, RUDDHI; SIVARAMAKRISHNAN, RAM; PRABHAKAR, RAGHU; JAIRATH, SUMTI; WANG, JUNJUE; LIANG, KAIZHAO; FUCHS, ADI; MUSADDIQ, MATHEEN; SUJEETH, ARVIND KRISHNA
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 073381/0769 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
Continuity (2)
Continuation 17216652 · Mar 29, 2021
Related Publication 20220309319A1 · Sep 29, 2022
References Cited (109)
US 9646243B1 · Gokmen · 2017 [cited by applicant]
US 10310768B1 · Gauria et al. · 2019 [cited by applicant]
US 10607331B1 · Tandia et al. · 2020 [cited by applicant]
US 10812828B2 · Han et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 10990650B1 · Vantrease et al. · 2021 [cited by applicant]
US 20140177941A1 · Docherty et al. · 2014 [cited by applicant]
US 20140201450A1 · Haugen · 2014 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342888A1 · Yang et al. · 2016 [cited by applicant]
US 20160350645A1 · Brothers et al. · 2016 [cited by applicant]
US 20170102901A1 · Burke · 2017 [cited by applicant]
US 20170103309A1 · Chang et al. · 2017 [cited by applicant]
US 20190138898A1 · Song et al. · 2019 [cited by applicant]
US 20190146497A1 · Urtasun et al. · 2019 [cited by applicant]
US 20190220742A1 · Kuo et al. · 2019 [cited by applicant]
US 20190266485A1 · Singh et al. · 2019 [cited by applicant]
US 20190273536A1 · McCallister · 2019 [cited by applicant]
US 20190286973A1 · Kovvuri et al. · 2019 [cited by applicant]
US 20190294108A1 · Ozcan et al. · 2019 [cited by applicant]
US 20190295228A1 · Liu et al. · 2019 [cited by applicant]
US 20190347549A1 · Phanishayee et al. · 2019 [cited by applicant]
US 20190370631A1 · Fais et al. · 2019 [cited by applicant]
US 20200082215A1 · Aliabadi et al. · 2020 [cited by applicant]
US 20200082243A1 · Jin et al. · 2020 [cited by applicant]
US 20200159809A1 · Catthoor et al. · 2020 [cited by applicant]
US 20200160226A1 · Ross et al. · 2020 [cited by applicant]
US 20200193267A1 · Aydonat et al. · 2020 [cited by applicant]
US 20200234129A1 · Fagerholm et al. · 2020 [cited by applicant]
US 20200272892A1 · Desappan et al. · 2020 [cited by applicant]
US 20200279157A1 · Gao et al. · 2020 [cited by applicant]
US 20200302223A1 · Dutta et al. · 2020 [cited by applicant]
US 20210049463A1 · Ruff · 2021 [cited by applicant]
US 20210049804A1 · Sarel et al. · 2021 [cited by applicant]
US 20210097347A1 · Kwon et al. · 2021 [cited by applicant]
US 20210141571A1 · Lew et al. · 2021 [cited by applicant]
US 20210158167A1 · Sharma et al. · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210182676A1 · Zlateski et al. · 2021 [cited by applicant]
US 20210201124A1 · Gelashvili · 2021 [cited by applicant]
US 20220309336A1 · Minkin · 2022 [cited by examiner]
EP 3282398A1 · 2018 [cited by examiner]
WO 2010142987A1 · 2010 [cited by applicant]
Le et al., Tiled convolutional neural networks, Advances in neural information processing systems 23 (2010); Total pp. 1-9 (Year: 2010). [cited by examiner]
Lin et al., GrateTile: Efficient Sparse Tensor Tiling for CNN Processing, arXiv:2009.08685v1 [cs.LG] Sep. 18, 2020; Total Pages: (Year: 2020). [cited by examiner]
U.S. Appl. No. 17/216,657—Response to Office Action dated Sep. 22, 2021, filed Sep. 27, 2021, 13 pages. [cited by applicant]
U.S. Appl. No. 17/216,657 Notice of Allowance, dated Oct. 20, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Response to Office Action dated Aug. 2, filed Aug. 13, 2021, 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Notice of Allowance, dated Aug. 23, 2021, 9 pages. [cited by applicant]
Mittal et. al., A Survey on Hardware Accelerators and Optimization on Techniques for RNNs, Journal of Systems Architecture, dated Jul. 2020, 57 pages. [cited by applicant]
Wang et. al., FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters, Transactions on Computers, vol. 14, No. 8, dated Aug. 2020, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Response to Office Action dated Sep. 1, 2021, filed Sep. 8, 2021, 18 pages. [cited by applicant]
U.S. Appl. No. 17/216,654 Notice of Allowance, dated Oct. 8, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 9 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jun. 29, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jun. 18, 2021, 12 pages. [cited by applicant]
Sinha, Sudipta N., et al., “Feature tracking and matching in video using programmable graphics hardware”, Jul. 16, 2006, 11 pages. [cited by applicant]
Li, Peng, et. al., “Deep Convolutional Computation Model for Feature Learning on Big Data in Internet of Things”, IEEE Transactions on Industrial Informatics, vol. 14, No. 2, Feb. 2018, 9 pages. [cited by applicant]
Prabhakar et. al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA 2017, Jun. 24-28, 2017, Toronto, ON, Canada. [cited by applicant]
Koeplinger et. al., Spatial: A Language and Compiler for Application Accelerators, Proceedings of the 39th ACM SIGPLAN Conference On Programming Language Design And Embodiment (PLDI), Proceedings of the 43rd Internation… [cited by applicant]
Luo, Yandong, et. al., “AILC: Accelerate On-chip Incremental Learning with Compute-in-Memory Technology”, DOI 10.1109/TC.2021.3053199, IEEE Transactions on Computers, Jan. 20, 2021, 15 pages. [cited by applicant]
Jangda et. al., Model-Based Warp overlapped Tiling for Image Processing Programs on GPUs, dated Sep. 8, 2020, 14 pages. [cited by applicant]
Dumoulin et. al., A guide to convolution arithmetic for deep learning, dated Jan. 12, 2018, 31 pages. [cited by applicant]
Chen et. al., Training Deep Nets with Sublinear Memory Cost, dated Apr. 22, 2016, 12 pages. [cited by applicant]
Azarkhish et. al., Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes, IEEE, dated Sep. 24, 2017, 15 pages. [cited by applicant]
Kong et. al., Take it in your stride: Do we need striding in CNNs, Carnegie Mellon University, dated Dec. 7, 2017, 9 pages. [cited by applicant]
Pinckaers et. al., Training Convolutional Neural Networks with Megapixel Images, dated Apr. 16, 2018, 3 pages. [cited by applicant]
Gui et. al., A Survey on Graph Processing Accelerators: Challenges and Opportunities, dated Feb. 26, 2019, 41 pages. [cited by applicant]
Pinckaers et. al., Streaming convolutional neural networks for end-to-end learning with multi-megapixel images, dated Nov. 11, 2019, 10 pages. [cited by applicant]
Kim et. al., DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks, IEEE Transactions on Computer Aided Design of Integrated Circuits and Systems, dated Nov. 11, 2018, 11 pages. [cited by applicant]
Schuiki et. al., A scalable near-memory architecture for Training Deep Neural Networks on Large In-Memory Datasets, IEEE Transactions on Computers, vol. 68, No. 4, Apr. 2019, 14 pages. [cited by applicant]
Liu et. al., An FPGA-Based CNN Accelerator Integrating Depthwise Separable Convolution, Electronics, published Mar. 3, 2019, 18 pages. [cited by applicant]
Park et. al., HetPipe: enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data Parallelism, Proceedings of the USENIX Annual Technical Conference, J… [cited by applicant]
Anonymous, BDS-GCN: Efficient full-graph training of graph convolutional nets with partition parallelism and boundary sampling, ICLR, 2021, 11 pages. [cited by applicant]
Venkataramanaiah et. al., Automatic Compiler Based FPGA Accelerator for CNN Training, dated Aug. 15, 2019, 7 pages. [cited by applicant]
Podlozhnyuk, Image Convolution with CUDA, NVIDIA, Jun. 2007, 21 pages. [cited by applicant]
NVIDIA, CSE 599 I Accelerated Computing—Programming GPUs, GPU teaching kit, 61 pages. [cited by applicant]
Hua et. al., Reverse Engineering Convolutional Neural Networks Through side-channel Information Leaks, DAC 2018, Jun. 24-29, 2018, San Francisco, USA, 6 pages. [cited by applicant]
Clemons et. al., A Patch memory system for Image Processing and Computer Vision, IEEE 2016, 13 pages. [cited by applicant]
Anonymous, ECE Exam 1 study guide, Spring 2019, 25 pages. [cited by applicant]
Zlateski et. al., FFt Convolutions are faster than Winograd on Modern CPUs, Here's Why, dated Sep. 20, 2018, 17 pages. [cited by applicant]
Zhou et. al., Hierarchical Overlapped Tiling, 12 pages. [cited by applicant]
Boris Ginzburg, Lecture 3: CNN: Backpropagation, Intel, 18 pages. [cited by applicant]
Ma, Hardware Acceleration of Deep Convolutional Neural Networks on FPGA, Arizona State University, Dec. 2018, 169 pages. [cited by applicant]
Versamopoulos et. al., Decoding surface code with a distributed neural network based decoder, published Feb. 6, 2019, 12 pages. Retrieved from the internet [URL: https://researchgate.net/publication/330751526 ]. [cited by applicant]
Liu et. al., Memory-Efficient Architecture for Accelerating Generative Networks on FPGA, EPSRC and European Union Horizon 2020, 8 pages. [cited by applicant]
Werkhoven et. al., Optimizing Convolution Operations in CUDA and Adaptive Tiling, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,651—Non-Final Rejection, dated Jul. 13, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action, dated Jun. 29, 2021, filed Jul. 12, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Response to Office Action, dated Jun. 18, 2021, filed Jul. 12, 2021, 12 pages. [cited by applicant]
M. Emani et al., “Accelerating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture,” in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 1-Apr. 2021, doi: 10.1109/MCSE.2021.3… [cited by applicant]
Podobas et al., A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
U.S. Appl. No. 17/216,651—Response to First Office Action, dated Jul. 13, 2021, filed Jul. 23, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/216,651—Notice of Allowance, dated Aug. 5, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Office Action, dated Aug. 2, 2021, 24 pages. [cited by applicant]
Ren et. al., SBNet: Sparse Blocks Network for Fast Inference, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 8711-8720. [cited by applicant]
Dathathri et. al., CHET: Compiler and Runtime for Homomorphic Evaluation of Tensor Programs, arXiv:1810.00845v1, dated Oct. 1, 2018. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jul. 28, 2021, 28 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jul. 28, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Office Action dated Sep. 1, 2021, 81 pages. [cited by applicant]
Zeng, Hanging, et., al, “GraphACT: Accelerating GCN Training on CPU FPGA Heterogeneous Platforms” Feb. 23-25, 2020 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Sep. 22, 2021, 22 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Notice of Allowance dated Aug. 25, 2021, 9 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 12 pages. [cited by applicant]
List of Related Cases, Aug. 8, 2023, 2 pages. [cited by applicant]
U.S. Appl. No. 17/384,507 Non-final Rejection, dated May 5, 2023, 64 pages. [cited by applicant]
TW 111111672 Office Action dated Sep. 14, 2022, 5 pgs. [cited by applicant]
U.S. Appl. No. 17/364,129—Office Action dated Aug. 9, 2023, 62 pages. [cited by applicant]
U.S. Appl. No. 17/364,110 Non Final Office Action, dated Jun. 30, 2023, 58 pages. [cited by applicant]