IP Library Granted Patent US 12,443,833
Granted Patent B2
US 12,443,833 · App. 17/271,326 · Granted Oct 14, 2025

Systems and methods for neural network convolutional layer matrix multiplication using cache memory

Inventor: Rati Gelashvili (Cambridge, MA)
Assignee: RED HAT, INC.
G06N3/063G06F17/16G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,833
App. No.
17/271,326
Granted
Oct 14, 2025
Kind
B2
Abstract

A computer processor may include a number of cores, a shared cache shared among the cores, and a local cache associated with each core and used by that core only. Input data for a neural network (NN) layer may be partitioned into a set of tiles of size T×T, and the tile set may be partitioned into blocks of R tiles. For each block, a core may perform a transform operation on the tiles to produce transformed data matrices fitting in a local cache, and a set of multiply operations, each multiply operation using a transformed data matrix and a transformed kernel matrix from a set of transformed kernel matrices. The set of transformed kernel matrices may fit in the shared cache. The result of at least one of the multiply operations may be stored in a location used to store a transformed data matrix.

Claims (34)

1. A method of executing matrix multiply operations for a neural network (NN), the method comprising, using a computer processor comprising a plurality of cores and a shared cache shared among the cores, each core associated with a local cache used by that core only:

partitioning input data for a NN layer into a set of tiles using a parameter T, each tile being of size T×T;

using a parameter R, partitioning the set of tiles into blocks of R tiles each; and

for each block of R tiles, performing by a single core:

a transform operation on the R tiles to produce a set of transformed data matrices, the set of transformed data matrices stored in the local cache of the single core; and

a set of multiply operations, each multiply operation using a transformed data matrix of the set of transformed data matrices and a transformed kernel matrix from a set of transformed kernel matrices, the set of transformed kernel matrices stored in the shared cache, wherein a result of at least one of the multiply operations is stored in a location used to store a transformed data matrix.

2. The method of claim 1 , wherein the partitioning of the input data and the partitioning of the tile set increase a likelihood that the transformed matrices are stored in a local cache.

3. The method of claim 1 , wherein the location used to store the result of at least one of the multiply operations is a location used to store a transformed data matrix used in a previous multiply operation.

4. The method of claim 1 , wherein the set of transformed kernel matrices is an entire set of transformed kernel matrices for the NN layer.

5. The method of claim 1 , comprising performing by the core an inverse transform operation on the results of the multiply operations.

6. The method of claim 5 , comprising, for a second NN layer:

partitioning a portion of the results of the inverse transform operations into a second set of tiles using a parameter T′, each tile being of size T′×T′;

using a parameter R′, partitioning the second set of tiles into blocks of R′ tiles each; and

for each block of R′ tiles, performing by a core a set of transform-multiply-inverse transform operations.

7. The method of claim 1 , wherein the transformed data matrices for the NN layer fit within the shared cache.

8. A system for of executing matrix multiply operations for a neural network (NN), the system comprising a computer processor comprising:

a plurality of cores each core associated with a local cache used by that core only; and

a shared cache shared among the cores, wherein, for input data partitioned for a NN layer into a set of tiles using a parameter T, each tile being of size T×T and the set of tiles partitioned using a parameter R into blocks of R tiles each:

each single core configured to, for a block of R tiles perform:

a transform operation on the R tiles to produce a set of transformed data matrices, the set of transformed data matrices stored in the local cache of the single core; and

a set of multiply operations, each multiply operation using a transformed data matrix of the set of transformed data matrices and a transformed kernel matrix from a set of transformed kernel matrices, the set of transformed kernel matrices stored in the shared cache, wherein a result of at least one of the multiply operations is stored in a location used to store a transformed data matrix.

9. The system of claim 8 , wherein the partitioning of the input data and the partitioning of the tile set increase a likelihood that the transformed matrices are stored in a local cache.

10. The system of claim 8 , wherein the location used to store the result of at least one of the multiply operations is a location used to store a transformed data matrix used in a previous multiply operation.

11. The system of claim 8 , wherein the set of transformed kernel matrices is an entire set of transformed kernel matrices for the NN layer.

12. The system of claim 8 , wherein the core is configured to perform an inverse transform operation on the results of the multiply operations.

13. The system of claim 12 , wherein for a second NN layer a portion of the results of the inverse transform operations is partitioned into a second set of tiles using a parameter T′, each tile being of size T′×T′, and the second set of tiles is partitioned using a parameter R′ into blocks of R′ tiles each; and

each single core is configured to, for a block of R′ tiles, perform a set of transform-multiply-inverse transform operations.

14. The system of claim 8 , wherein the transformed data matrices for the NN layer fit within the shared cache.

15. A method of performing inference for a neural network (NN), the method comprising, in a computer processor comprising a plurality of cores, a shared cache, and a private cache used by each core, the method comprising:

using input data divided into a set of tiles using a parameter T, each tile being of size T×T, the set of tiles divided into blocks of tiles, for each block of tiles, a single core executing a transform-multiply-inverse transform operation on the block of tiles, such that a set of transformed data matrices is stored in the local cache of the single core, wherein the transform-multiply-inverse transform operation operates on a transformed kernel matrix, the transformed kernel matrix stored in the shared cache, and wherein results of a multiply operation is stored in a location used to store a transformed data matrix.

16. The method of claim 15 , wherein the division of the input data and the division of the tile set increase a likelihood that the transformed matrices are stored in a local cache.

17. The method of claim 15 , wherein the location used to store the result is a location used to store a transformed data matrix used in a previous multiply operation.

18. The method of claim 15 , wherein, for a second NN layer a portion of the results of the inverse transform operations are partitioned into a second set of tiles using a parameter T′, each tile being of size T′×T′, and wherein the second set of tiles is partitioned using a parameter R′ into blocks of R′ tiles each, wherein the method comprises for each block of R′ tiles, performing by a core a set of transform-multiply-inverse transform operations.

19. The method of claim 15 wherein in the transform-multiply-inverse operation a set of transformed kernel matrices used in the transform-multiply-inverse operation fits in the shared cache.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2025
From: GELASHVILI, RATI
To: NEURALMAGIC INC.
Reel/Frame 072444/0301 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2025
From: NEURALMAGIC, INC.
To: RED HAT, INC.
Reel/Frame 072278/0309 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2025
From: GELASHVILI, RATI
To: NEURALMAGIC INC.
Reel/Frame 071425/0649 →
Continuity (2)
Provisional Application 62723350 · Aug 27, 2018
Related Publication 20210201124A1 · Jul 1, 2021
References Cited (107)
US 5577166A · Mizuno · 1996 [cited by applicant]
US 9558156B1 · Bekas et al. · 2017 [cited by applicant]
US 9811775B2 · Krizhevsky et al. · 2017 [cited by applicant]
US 9818059B1 · Woo et al. · 2017 [cited by applicant]
US 10157045B2 · Venkataramani et al. · 2018 [cited by applicant]
US 10223333B2 · Chetlur et al. · 2019 [cited by applicant]
US 10521488B1 · Ross · 2019 [cited by examiner]
US 10572568B2 · Narayanamoorthy et al. · 2020 [cited by applicant]
US 10685082B2 · Bekas et al. · 2020 [cited by applicant]
US 10719323B2 · Baum et al. · 2020 [cited by applicant]
US 20100076915A1 · Xu et al. · 2010 [cited by applicant]
US 20110119467A1 · Cadambi et al. · 2011 [cited by applicant]
US 20130138589A1 · Yu et al. · 2013 [cited by applicant]
US 20160224465A1 · Morad et al. · 2016 [cited by applicant]
US 20160239706A1 · Dijkman et al. · 2016 [cited by applicant]
US 20160328643A1 · Liu et al. · 2016 [cited by applicant]
US 20160358070A1 · Brothers et al. · 2016 [cited by applicant]
US 20160379109A1 · Chung et al. · 2016 [cited by applicant]
US 20170032487A1 · Ashari et al. · 2017 [cited by applicant]
US 20170103313A1 · Ross et al. · 2017 [cited by applicant]
US 20170103317A1 · Young · 2017 [cited by applicant]
US 20170132496A1 · Shoaib et al. · 2017 [cited by applicant]
US 20170169567A1 · Chefd'Hotel et al. · 2017 [cited by applicant]
US 20170193361A1 · Chilimbi et al. · 2017 [cited by applicant]
US 20170200094A1 · Bruestle et al. · 2017 [cited by applicant]
US 20170220524A1 · Herrero Abellanas et al. · 2017 [cited by applicant]
US 20170270073A1 · Badin · 2017 [cited by examiner]
US 20170316311A1 · Pilly et al. · 2017 [cited by applicant]
US 20170316312A1 · Goyal et al. · 2017 [cited by applicant]
US 20170372202A1 · Ginsburg et al. · 2017 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180096226A1 · Aliabadi et al. · 2018 [cited by applicant]
US 20180173571A1 · Huang et al. · 2018 [cited by applicant]
US 20180253402A1 · Redfern et al. · 2018 [cited by applicant]
US 20180315159A1 · Ould-Ahmed-Vall et al. · 2018 [cited by applicant]
US 20180322390A1 · Das et al. · 2018 [cited by applicant]
US 20180336468A1 · Kadav et al. · 2018 [cited by applicant]
US 20190042250A1 · Anders et al. · 2019 [cited by applicant]
US 20190042542A1 · Narayanamoorthy et al. · 2019 [cited by applicant]
US 20190056916A1 · Varma et al. · 2019 [cited by applicant]
US 20190138902A1 · Matveev et al. · 2019 [cited by applicant]
US 20190156206A1 · Graham et al. · 2019 [cited by applicant]
US 20190156214A1 · Matveev et al. · 2019 [cited by applicant]
US 20190156215A1 · Matveev et al. · 2019 [cited by applicant]
US 20190179818A1 · Lee · 2019 [cited by applicant]
US 20190179869A1 · Park · 2019 [cited by examiner]
US 20190205746A1 · Nurvitadhi · 2019 [cited by examiner]
US 20190212982A1 · Yoda et al. · 2019 [cited by applicant]
US 20190278600A1 · Frumkin · 2019 [cited by examiner]
US 20190303743A1 · Venkataramani et al. · 2019 [cited by applicant]
US 20190311242A1 · Chen · 2019 [cited by examiner]
US 20190354894A1 · Lazovich et al. · 2019 [cited by applicant]
US 20190370071A1 · Matveev et al. · 2019 [cited by applicant]
US 20190370644A1 · Kenney et al. · 2019 [cited by applicant]
US 20200034710A1 · Sidhu et al. · 2020 [cited by applicant]
US 20200097826A1 · Du et al. · 2020 [cited by applicant]
US 20200104717A1 · Alistarh · 2020 [cited by applicant]
US 20200117701A1 · Ohno · 2020 [cited by examiner]
US 20200160181A1 · Zlateski et al. · 2020 [cited by applicant]
US 20200160182A1 · Matveev et al. · 2020 [cited by applicant]
US 20200193274A1 · Darvish Rouhani et al. · 2020 [cited by applicant]
US 20200218978A1 · Kopinsky · 2020 [cited by applicant]
US 20200293783A1 · Ramaswamy · 2020 [cited by examiner]
US 20200302284A1 · Garcia Garcia · 2020 [cited by examiner]
US 20200342301A1 · Miao et al. · 2020 [cited by applicant]
US 20240078417A1 · Temam · 2024 [cited by examiner]
EP 3037980 · 2016 [cited by applicant]
WO WO2017049496 · 2017 [cited by applicant]
WO WO2018053835 · 2018 [cited by applicant]
WO WO2019090325A1 · 2019 [cited by applicant]
WO WO2020046859A1 · 2020 [cited by applicant]
WO WO2020047823A1 · 2020 [cited by applicant]
WO WO2020072274A1 · 2020 [cited by applicant]
Deshpande, A beginner's guide to understanding convolutional neural networks, Jul. 20, 2016. [cited by applicant]
Alwani et al., “Fused-layer CNN accelerators.” 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2016, pp. 1-12. [cited by applicant]
Du et al., “Width Provably Matters in Optimization for Deep Linear Neural Networks”, May 27, 2019, arXiv:1901.08572v3. [cited by applicant]
Gale et al., “The State of Sparsity in Deep Neural Networks”, Feb. 25, 2019, arXiv:1902.09574v1. [cited by applicant]
Han et al., “Learning both Weights and Connections for Efficient Neural Networks”, 2015, Advances in Neural Information Processing Systems, vol. 28. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network”, Mar. 9, 2015. [cited by applicant]
Lavin et al., “Fast Algorithms for Convolutional Neural Networks”, Nov. 10, 2015. [cited by applicant]
Lecun et al., “Optimal brain damage”, Advances in neural information processing systems, 1990, pp. 598-605. [cited by applicant]
Mishra et al., “Apprentice: Using Knowledge Distillation Techniques to Improve Low-Precision Network Accuracy”, Nov. 15, 2017. [cited by applicant]
Rusu et al., “Progressive Neural Networks”, Sep. 7, 2016. [cited by applicant]
Budden et al., “Deep tensor convolution on multicores”, In Proceedings of the 34th International Conference on Machine Learning, 2017, vol. 70, pp. 615-624. [cited by applicant]
Chen, Xuhao, “Escoin: Efficient Sparse Convolutional Neural Network Inference on GPUs.” From Jul. 2017 “Conference '17”, Apr. 3, 2019 (Apr. 3, 2019) Retrieved on Jan. 17, 2020 (Jan. 17, 2020)from <https://arxiv.orq/pdf/… [cited by applicant]
Georganas et al., “Anatomy Of High-Performance Deep Learning Convolutions On SIMD Architectures.” In: SC18: International Conference for High Performance Computing, Networking, Storage and Analysis. Aug. 20, 2018 (Aug. … [cited by applicant]
Kaya et al., “Scalable sparse tensor decompositions in distributed memory systems”, SC'15: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, IEEE, 2015. (Year:… [cited by applicant]
Kim et al., “Designing Vector-Friendly Compact BLAS and LAPACK Kernels”, SC17, Nov. 12-17, 2017, Denver, CO, USA. [cited by applicant]
Lascorz et al., “Bit-Tactical: Exploiting Ineffectual Computations in Convolutional Neural Networks: Which, Why, and How”, Mar. 9, 2018. [cited by applicant]
Liu et al., “Sparse convolutional neural networks.” In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Jun. 12, 2015 (Jun. 12, 2015) Retrieved on Jan. 17, 2020. [cited by applicant]
Papyan et al., “Convolutional neural networks analyzed via convolutional sparse coding.” In: The Journal of Machine Learning Research. Jul. 17, 2017 (Jul. 17, 2017) Retrieved on Feb. 20, 2020. [cited by applicant]
Scardapane et al. “Group sparse regularization for deep neural networks.”, In: Neurocomputing. Jul. 2, 2016 (Jul. 2, 2016) Retrieved on Nov. 16, 2019 (Nov. 16, 2019). [cited by applicant]
Smith et al., “SPLATT: Efficient and parallel sparse tensor-matrix multiplication”, 2015 IEEE International Parallel and Distributed Processing Symposium, IEEE, 2015, (Year: 2015). [cited by applicant]
Wozniak et al., “GiMMiK-Generating bespoke matrix multiplication kernels for accelerators: Application to high-order Computational Fluid Dynamics”, Computer Physics Communications, vol. 202, 2016, pp. 12-22. [cited by applicant]
Zhangxiaowen Gong et al. “Sparse Train: Leveraging Dynamic Sparsity in Training DNNs on General-Purpose SIMD Processors”; 2019. [cited by applicant]
Yu, Dong, Li Deng, and Frank Seide. “The deep tensor neural network with applications to large vocabulary speech recognition”. IEEE Transactions on Audio Speech, and Language Processing 21.2 (2012): 388-396. (Year: 2012… [cited by applicant]
Kurtz, Mark, et al. “Inducing and Exploiting Activation Sparsity for Fast Neural Network Inference.” Proceedings of the International Conference on Machine Learning. 2020. [cited by applicant]
Robert Lim; “Methods for Accelerating Machine Learning in High Performance Computing”; University of Oregon—AREA-2019-01. [cited by applicant]
Zhizhou Li et al.; “A CPU-based Algorithm for Traffic Optimization Based on Sparse Convolutional Neural Networks”; 2017 IEEE 30th Canadian Conference on Electrical and Computer (CCECE). [cited by applicant]
Baoyuan Liu et al.; “Sparse Convolutional Neural Networks”; CVPR 2015—Computer Vision Foundation—IEEE. [cited by applicant]
Hesham Mostafa et al.; “Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization”; Proceedings of the 36 th International Conference on Machine Learning, Long Beach, Califo… [cited by applicant]
Israt Nisa et al.; “Sampled Dense Matrix Multiplication for High-Performance Machine Learning”; 2018 IEEE 25th International Conference on High Performance Computing (Hi PC). [cited by applicant]
Yang, Huanrui, Wei Wen, and Hai Li. “Deephoyer: Learning sparser neural network with differentiable scale-invariant sparsity measures.” arXiv preprint arXiv:1908.09979 (2019). [cited by applicant]
Yuster, Raphael, and Uri Zwick. “Fast sparse matrix multiplication.” ACM Transactions On Algorithms (TALG) 1.1 (2005): 2-13. [cited by applicant]
Paixao, Crysttian A., and Flávio Codeco Coelho. Matrix compression methods. No. e1049. PeerJ PrePrints, 2015. [cited by applicant]
Park, Jongsoo, et al. “Faster cnns with direct sparse convolutions and guided pruning.” arXiv preprint arXiv:1608.01409 (2016). [cited by applicant]
https://www.kinematicsoup.com/news/2016/9/6/data-compression-bit-packing-101, published Sep. 6, 2016. [cited by applicant]