IP Library Granted Patent US 12,210,953
Granted Patent B2
US 12,210,953 · App. 17/687,516 · Granted Jan 28, 2025

Lossless tiling in convolution networks—section cuts

Inventors: Tejas Nagendra Babu Nama (Sunnyvale, CA); Ruddhi Chaphekar (Santa Clara, CA); Ram Sivaramakrishnan (San Jose, CA); Raghu Prabhakar (San Jose, CA); Sumti Jairath (Santa Clara, CA); Junjue Wang (San Mateo, CA); Kaizhao Liang (Palo Alto, CA); Adi Fuchs (West Windsor, NJ); Matheen Musaddiq (Austin, TX); Arvind Krishna Sujeeth (San Francisco, CA)
Assignee: SambaNova Systems, Inc.
G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,953
App. No.
17/687,516
Granted
Jan 28, 2025
Kind
B2
Abstract

A data processing system receives a graph that includes a sequence of layers and executes graph cuts between a preceding layer in the graph and a succeeding layer in the graph that succeeds the preceding layer. The preceding layer generates a set of tiles on a tile-by-tile basis and the succeeding layer processes a tensor that includes multiple tiles in the set of tiles. Thus the graph is partitioned into a sequence of subgraphs, and a subgraph in the sequence of subgraphs including a sub-sequence of layers in the sequence of layers. One or more configuration files is generated to configure runtime logic to execute the sequence of subgraphs and the one or more configuration files are stored on a computer-readable media.

Claims (48)

1. One or more non-transitory computer-readable media storing instructions, that if executed by a processor, perform a method comprising:

receiving a processing graph to process a tensor for machine learning and/or inference, the processing graph including a sequence of processing nodes connected by one or more edges representing dataflow between the sequence of processing nodes;

executing graph cuts to partition the processing graph into a sequence of subgraphs, a subgraph in the sequence of subgraphs including a sub-sequence of one or more processing nodes in the sequence of processing nodes, wherein a graph cut of the graph cuts is executed between a preceding processing node in the processing graph, the preceding processing node configured to generate a set of tiles on a tile-by-tile basis, and a succeeding processing node in the processing graph that succeeds the preceding processing node, the succeeding processing node is configured to process the tensor that includes multiple tiles in the set of tiles;

generating one or more configuration files to configure runtime logic to execute the sequence of subgraphs; and

storing the one or more configuration files on a computer-readable media.

2. The one or more non-transitory computer-readable media of claim 1 , the method further comprising:

retrieving the one or more configuration files from the computer-readable media; and

sending the one or more configuration files to the runtime logic.

3. The one or more non-transitory computer-readable media of claim 1 , wherein the succeeding processing node implements a batch normalization operation or a reduction operation, wherein the reduction operation comprises a pooling operation or a convolution.

4. The one or more non-transitory computer-readable media of claim 1 , wherein the tensor includes padding.

5. The one or more non-transitory computer-readable media of claim 4 , wherein the tensor is provided as tiles in a second set of tiles to the succeeding processing node, and only tiles of the second set of tiles that coincide with edges of the tensor include a portion of the padding.

6. The one or more non-transitory computer-readable media of claim 1 , wherein the tensor represents an image and the tiles in the set of tiles comprise pixels.

7. The one or more non-transitory computer-readable media of claim 1 , wherein the tensor comprises a feature map or a gradient map.

8. The one or more non-transitory computer-readable media of claim 7 , wherein gradients of the gradient map are input gradients.

9. The one or more non-transitory computer-readable media of claim 1 , wherein the preceding processing node is configured as a final processing node of a preceding subgraph in the sequence of subgraphs.

10. The one or more non-transitory computer-readable media of claim 9 , wherein the succeeding processing node is configured as a first processing node of a succeeding subgraph in the sequence of subgraphs that succeeds the preceding subgraph.

11. The one or more non-transitory computer-readable media of claim 1 , wherein the processing graph comprises a convolutional neural network.

12. The one or more non-transitory computer-readable media of claim 1 , wherein the subgraphs include forward pass subgraphs.

13. The one or more non-transitory computer-readable media of claim 1 , wherein the subgraphs include backward pass subgraphs.

14. The one or more non-transitory computer-readable media of claim 1 , wherein processing nodes in the sequence of processing nodes include one or more of a convolution processing node, a max pooling processing node, a min pooling processing node, an average pooling processing node, a non-linearity processing node, a normalization processing node, a dropout processing node, a concatenation processing node, a transpose convolution processing node, a fully connected processing node, a softmax processing node, or a loss processing node.

15. The one or more non-transitory computer-readable media of claim 1 , wherein the runtime logic comprises a Coarse-Grained Reconfigurable Architecture system.

16. The one or more non-transitory computer-readable media of claim 1 , wherein the preceding processing node is further configured to provide the set of tiles on a tile-by-tile basis to data flow logic which aggregates the set of tiles into the tensor in a memory.

17. A computer-implemented method comprising:

receiving a processing graph to process a tensor for machine learning and/or inference, the processing graph including a sequence of processing nodes connected by one or more edges representing dataflow between the sequence of processing nodes;

executing graph cuts to partition the processing graph into a sequence of subgraphs, a subgraph in the sequence of subgraphs including a sub-sequence of one or more processing nodes in the sequence of processing nodes, wherein a graph cut of the graph cuts is executed between a preceding processing node in the processing graph, the preceding processing node configured to generate a set of tiles on a tile-by-tile basis, and a succeeding processing node in the processing graph that succeeds the preceding processing node, the succeeding processing node is configured to process the tensor that includes multiple tiles in the set of tiles;

generating one or more configuration files to configure runtime logic to execute the sequence of subgraphs; and

storing the one or more configuration files on a computer-readable media.

18. The method of claim 17 , further comprising:

retrieving the one or more configuration files from the computer-readable media; and

sending the one or more configuration files to the runtime logic.

19. The method of claim 17 , wherein the tensor includes padding.

20. The method of claim 17 , wherein the tensor represents an image and the tiles in the set of tiles comprise pixels.

21. The method of claim 17 , wherein the subgraphs include backward pass subgraphs.

22. The method of claim 17 , wherein the preceding processing node is further configured to provide the set of tiles on a tile-by-tile basis to data flow logic which aggregates the set of tiles into the tensor in a memory.

23. A data processing system comprising:

a reconfigurable processor comprising an array of configurable units that is partitionable into a plurality of subarrays of configurable units;

compile time logic configured to receive a processing graph to process a tensor for machine learning and/or inference, the processing graph including a sequence of processing nodes connected by one or more edges representing dataflow between the sequence of processing nodes, execute graph cuts to partition the processing graph into a sequence of subgraphs, and generate one or more configuration files for the sequence of subgraphs; and

runtime logic to use the one or more configuration files to execute the sequence of subgraphs; wherein

a subgraph in the sequence of subgraphs includes a sub-sequence of one or more processing nodes in the sequence of processing nodes,

a graph cut is executed between a preceding processing node in the processing graph and a succeeding processing node in the processing graph that succeeds the preceding processing node,

the preceding processing node is configured to generate a set of tiles on a tile-by-tile basis, and

the succeeding processing node is implemented in at least one subarray of the plurality of subarrays of configurable units and configured to process the tensor that includes multiple tiles in the set of tiles.

24. The data processing system of claim 23 , further comprising:

a memory device to store the set of tiles; and

an integrated circuit that includes the runtime logic, coupled to, but separate from, the memory device.

25. The data processing system of claim 24 , wherein the preceding processing node is further configured to provide the set of tiles on a tile-by-tile basis to data flow logic which aggregates the set of tiles into the tensor in the memory device.

26. The data processing system of claim 23 , further comprising an integrated circuit that includes the runtime logic and memory to store the tensor.

27. The data processing system of claim 23 , wherein the reconfigurable processor further comprises a Coarse-Grained Reconfigurable Architecture system comprising the runtime logic.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2022
From: NAMA, TEJAS NAGENDRA BABU; CHAPHEKAR, RUDDHI; SIVARAMAKRISHNAN, RAM; PRABHAKAR, RAGHU; JAIRATH, SUMTI; WANG, JUNJUE; LIANG, KAIZHAO; FUCHS, ADI; MUSADDIQ, MATHEEN; SUJEETH, ARVIND KRISHNA
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 059178/0472 →
Continuity (3)
Continuation 17364110 · Jun 30, 2021
Division 17216651 · Mar 29, 2021
Related Publication 20220309322A1 · Sep 29, 2022
References Cited (110)
US 9646243B1 · Gokmen · 2017 [cited by applicant]
US 10310768B1 · Gauria et al. · 2019 [cited by applicant]
US 10607331B1 · Tandia et al. · 2020 [cited by applicant]
US 10812828B2 · Han et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 10990650B1 · Vantrease et al. · 2021 [cited by applicant]
US 20140177941A1 · Docherty et al. · 2014 [cited by applicant]
US 20140201450A1 · Haugen · 2014 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160342888A1 · Yang et al. · 2016 [cited by applicant]
US 20160350645A1 · Brothers et al. · 2016 [cited by applicant]
US 20170102901A1 · Burke · 2017 [cited by applicant]
US 20170103309A1 · Chang et al. · 2017 [cited by applicant]
US 20190138898A1 · Song et al. · 2019 [cited by applicant]
US 20190146497A1 · Urtasun et al. · 2019 [cited by applicant]
US 20190213473A1 · Dutta · 2019 [cited by examiner]
US 20190220742A1 · Kuo · 2019 [cited by examiner]
US 20190266218A1 · Scott · 2019 [cited by examiner]
US 20190266485A1 · Singh · 2019 [cited by examiner]
US 20190273536A1 · McCallister · 2019 [cited by applicant]
US 20190286973A1 · Kovvuri · 2019 [cited by examiner]
US 20190294108A1 · Ozcan et al. · 2019 [cited by applicant]
US 20190295228A1 · Liu et al. · 2019 [cited by applicant]
US 20190347549A1 · Phanishayee et al. · 2019 [cited by applicant]
US 20190370631A1 · Fais · 2019 [cited by examiner]
US 20200082215A1 · Aliabadi et al. · 2020 [cited by applicant]
US 20200082243A1 · Jin et al. · 2020 [cited by applicant]
US 20200134778A1 · He · 2020 [cited by examiner]
US 20200159809A1 · Catthoor et al. · 2020 [cited by applicant]
US 20200160226A1 · Ross et al. · 2020 [cited by applicant]
US 20200193267A1 · Aydonat et al. · 2020 [cited by applicant]
US 20200234129A1 · Fagerholm et al. · 2020 [cited by applicant]
US 20200272892A1 · Desappan et al. · 2020 [cited by applicant]
US 20200279157A1 · Gao et al. · 2020 [cited by applicant]
US 20200302223A1 · Dutta et al. · 2020 [cited by applicant]
US 20210049463A1 · Ruff · 2021 [cited by applicant]
US 20210049804A1 · Sarel et al. · 2021 [cited by applicant]
US 20210097347A1 · Kwon et al. · 2021 [cited by applicant]
US 20210141571A1 · Lew et al. · 2021 [cited by applicant]
US 20210158167A1 · Sharma et al. · 2021 [cited by applicant]
US 20210173787A1 · Nagy et al. · 2021 [cited by applicant]
US 20210182676A1 · Zlateski · 2021 [cited by examiner]
US 20210201124A1 · Gelashvili · 2021 [cited by applicant]
WO 2010142987 · 2010 [cited by applicant]
WO 2010142987A1 · 2010 [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jun. 29, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jun. 18, 2021, 12 pages. [cited by applicant]
Sinha, Sudipta N., et al., “Feature tracking and matching in video using programmable graphics hardware”, Jul. 16, 2006, 11 pages. [cited by applicant]
Li, Peng, et al., “Deep Convolutional Computation Model for Feature Learning on Big Data in Internet of Things”, IEEE Transactions on Industrial Informatics, vol. 14, No. 2, Feb. 2018, 9 pages. [cited by applicant]
Prabhakar et al., Plasticine: A Reconfigurable Architecture for Parallel Patterns, ISCA 2017, Jun. 24-28, 2017, Toronto, ON, Canada, 14 pages. [cited by applicant]
Koeplinger et al., Spatial: A Language and Compiler for Application Accelerators, Proceedings of the 39th ACM SIGPLAN Conference On Programming Language Design And Embodiment (PLDI), Proceedings of the 43rd Internationa… [cited by applicant]
Luo, Yandong, et al., “AILC: Accelerate On-chip Incremental Learning with Compute-in-Memory Technology”, DOI 10.1109/TC.2021.3053199, IEEE Transactions on Computers, Jan. 20, 2021, 15 pages. [cited by applicant]
Jangda et al., Model-Based Warp overlapped Tiling for Image Processing Programs on GPUs, dated Sep. 8, 2020, 14 pages. [cited by applicant]
Dumoulin et al., A guide to convolution arithmetic for deep learning, dated Jan. 12, 2018, 31 pages. [cited by applicant]
Chen et. al., Training Deep Nets with Sublinear Memory Cost, dated Apr. 22, 2016, 12 pages. [cited by applicant]
Azarkhish et al., Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes, IEEE, dated Sep. 24, 2017, 15 pages. [cited by applicant]
Kong et. al., Take it in your stride: Do we need striding in CNNs, Carnegie Mellon University, dated Dec. 7, 2017, 9 pages. [cited by applicant]
Pinckaers et. al., Training Convolutional Neural Networks with Megapixel Images, dated Apr. 16, 2018, 3 pages. [cited by applicant]
Gui et. al., A Survey on Graph Processing Accelerators: Challenges and Opportunities, dated Feb. 26, 2019, 41 pages. [cited by applicant]
Pinckaers et. al., Streaming convolutional neural networks for end-to-end learning with multi-megapixel images, dated Nov. 11, 2019, 10 pages. [cited by applicant]
Kim et al., DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks, IEEE Transactions on Computer Aided Design of Integrated Circuits and Systems, dated Nov. 11, 2018, 11 pages. [cited by applicant]
Schuiki et. al., A scalable near-memory architecture for Training Deep Neural Networks on Large In-Memory Datasets, IEEE Transactions on Computers, vol. 68, No. 4, Apr. 2019, 14 pages. [cited by applicant]
Liu et. al., An FPGA-Based CNN Accelerator Integrating Depthwise Separable Convolution, Electronics, published Mar. 3, 2019, 18 pages. [cited by applicant]
Park et. al., HetPipe: enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data Parallelism, Proceedings of the USENIX Annual Technical Conference, J… [cited by applicant]
Anonymous, BDS-GCN: Efficient full-graph training of graph convolutional nets with partition parallelism and boundary sampling, ICLR, 2021, 11 pages. [cited by applicant]
Venkataramanaiah et. al., Automatic Compiler Based FPGA Accelerator for CNN Training, dated Aug. 15, 2019, 7 pages. [cited by applicant]
Podlozhnyuk, Image Convolution with CUDA, NVIDIA, Jun. 2007, 21 pages. [cited by applicant]
NVIDIA, CSE 599 I Accelerated Computing—Programming GPUs, GPU teaching kit, 61 pages. [cited by applicant]
Hua et. al., Reverse Engineering Convolutional Neural Networks Through side-channel Information Leaks, DAC 2018, Jun. 24-29, 2018, San Francisco, USA, 6 pages. [cited by applicant]
Clemons et. al., A Patch memory system for Image Processing and Computer Vision, IEEE 2016, 13 pages. [cited by applicant]
Anonymous, ECE Exam 1 study guide, Spring 2019, 25 pages. [cited by applicant]
Zlateski et. al., FFt Convolutions are faster than Winograd on Modern CPUs, Here's Why, dated Sep. 20, 2018, 17 pages. [cited by applicant]
Zhou et. al., Hierarchical Overlapped Tiling, 12 pages. [cited by applicant]
Boris Ginzburg, Lecture 3: CNN: Backpropagation, Intel, 18 pages. [cited by applicant]
Ma, Hardware Acceleration of Deep Convolutional Neural Networks on FPGA, Arizona State University, Dec. 2018, 169 pages. [cited by applicant]
Versamopoulos et al., Decoding surface code with a distributed neural network based decoder, published Feb. 6, 2019, 12 pages. Retrieved from the internet [URL: https://researchgate.net/publication/330751526]. [cited by applicant]
Liu et. al., Memory-Efficient Architecture for Accelerating Generative Networks on FPGA, EPSRC and European Union Horizon 2020, 8 pages. [cited by applicant]
Werkhoven et al., Optimizing Convolution Operations in CUDA and Adaptive Tiling, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,651—Non-Final Rejection, dated Jul. 13, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action, dated Jun. 29, 2021, filed Jul. 12, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Response to Office Action, dated Jun. 18, 2021, filed Jul. 12, 2021, 12 pages. [cited by applicant]
M. Emani et al., “Accelerating Scientific Applications With SambaNova Reconfigurable Dataflow Architecture,” in Computing in Science & Engineering, vol. 23, No. 2, pp. 114-119, Mar. 1-Apr. 2021, doi: 10.1109/MCSE.2021.3… [cited by applicant]
Podobas et al., A Survey on Coarse-Grained Reconfigurable Architectures From a Performance Perspective, IEEEAccess, vol. 2020.3012084, Jul. 27, 2020, 25 pages. [cited by applicant]
U.S. Appl. No. 17/216,651—Response to First Office Action, dated Jul. 13, 2021, filed Jul. 23, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/216,651—Notice of Allowance, dated Aug. 5, 2021, 14 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Office Action, dated Aug. 2, 2021, 24 pages. [cited by applicant]
Ren et. al., SBNet: Sparse Blocks Network for Fast Inference, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 8711-8720. [cited by applicant]
Dathathri et. al., CHET: Compiler and Runtime for Homomorphic Evaluation of Tensor Programs, arXiv:1810.00845v1, dated Oct. 1, 2018. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Jul. 28, 2021, 28 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Office Action dated Jul. 28, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Office Action dated Sep. 1, 2021, 81 pages. [cited by applicant]
Zeng, Hanqing, et., al, “GraphACT: Accelerating GCN Training on CPU FPGA Heterogeneous Platforms” Feb. 23-25, 2020 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Office Action dated Sep. 22, 2021, 22 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Notice of Allowance dated Aug. 25, 2021, 9 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 12 pages. [cited by applicant]
U.S. Appl. No. 17/216,657—Response to Office Action dated Sep. 22, 2021, filed Sep. 27, 2021, 13 pages. [cited by applicant]
U.S. Appl. No. 17/216,657 Notice of Allowance, dated Oct. 20, 2021, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Response to Office Action dated Aug. 2, filed Aug. 13, 2021, 11 pages. [cited by applicant]
U.S. Appl. No. 17/216,652—Notice of Allowance, dated Aug. 23, 2021, 9 pages. [cited by applicant]
Mittal et. al., A Survey on Hardware Accelerators and Optimization on Techniques for RNNs, Journal of Systems Architecture, dated Jul. 2020, 57 pages. [cited by applicant]
Wang et. al., FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters, Transactions on Computers, vol. 14, No. 8, dated Aug. 2020, 16 pages. [cited by applicant]
U.S. Appl. No. 17/216,654—Response to Office Action dated Sep. 1, 2021, filed Sep. 8, 2021, 18 pages. [cited by applicant]
U.S. Appl. No. 17/216,654 Notice of Allowance, dated Oct. 8, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 17/216,655—Response to Office Action dated Jul. 28, 2021, filed Aug. 11, 2021, 9 pages. [cited by applicant]
Zeng, Hanging, et., al, “GraphACT: Accelerating GCN Training on CPU FPGA Heterogeneous Platforms” Feb. 23-25, 2020 11 pages. [cited by applicant]
List of Related Cases, Aug. 8, 2023, 2 pages. [cited by applicant]
TW 111111672—Office Action dated Sep. 14, 2022, 5 pgs. [cited by applicant]
U.S. Appl. No. 17/364,110 Non Final Office Action, dated Jun. 30, 2023, 58 pages. [cited by applicant]
U.S. Appl. No. 17/364,129—Office Action dated Aug. 9, 2023, 62 pages. [cited by applicant]
U.S. Appl. No. 17/384,507—Non-final Rejection, dated May 5, 2023, 64 pages. [cited by applicant]