IP Library › Granted Patent US 12,675,681
Granted Patent B2
US 12,675,681 · App. 18/359,270 · Granted Jul 7, 2026

Smart memory handling and data management for machine learning networks

Inventors: Tomer Schwartz (Even Yehuda, IL); Ehud Cohen (Kiryat Motskin, IL); Uzi Sarel (Zichron-Yaakov, IL); Amitai Armon (Tel-Aviv, IL); Yaniv Fais (Tel-Aviv, IL); Lev Faivishevsky (Kfar Saba, IL); Amit Bleiweiss (Yad Binyamin, IL); Yahav Shadmiy (Ramat Gan, IL); Jacob Subag (Kiryat Haim, IL)
Assignee: INTEL CORPORATION
G06N3/063G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,681
App. No.
18/359,270
Filed
Jul 26, 2023
Granted
Jul 7, 2026
Kind
B2
Examiner
ZHAO, DAQUAN
Art Unit
2484
USPC
706/33
Abstract

A mechanism is described for facilitating memory handling and data management in machine learning at autonomous machines. A method of embodiments, as described herein, includes detecting multiple tables associated with multiple neural networks at multiple autonomous machines, where each of the multiple tables include an index. The method may further include combining the multiple tables and multiple indexes associated with the multiple tables into a single table and a single index, respectively, where the single table is communicated to the multiple autonomous machines to allow simultaneous processing of one or more portions of the single table using one or more memory devices and one or more processors of one or more of the multiple autonomous machines.

Claims (33)

1 . An apparatus comprising:

one or more processors to process data for operation of the apparatus, the one or more processors including one or more graphical processing units (GPUs); and

a memory for storage of data, the memory including a memory address portion;

wherein the one or more processors are to perform processing with a neural network, including the one or more processors to:

generate a plurality of memory addresses for storage in the memory address portion, the plurality of memory addresses being based at least in part on a memory pattern associated with the neural network;

rearrange a plurality of feature maps for the neural network such that each feature map starts from an address of the plurality of memory addresses; and

fetch the plurality of memory addresses to load the plurality of feature maps for the one or more processors.

2 . The apparatus of claim 1 , wherein the one or more processors are further to:

convolve the plurality of feature maps with a specified kernel; and

accumulate results of each convolved feature map to generate a combined feature map.

3 . The apparatus of claim 1 , wherein the plurality of memory addresses provide memory layouts that are relatively sparse for two-dimensional (2D) data and relatively dense for three-dimensional (3D) data.

4 . The apparatus of claim 1 , wherein the rearrangement of the plurality of feature maps is performed at one or more of a memory level, a cache level, or a register level.

5 . The apparatus of claim 1 , wherein fetching the plurality of memory addresses includes the one or more processors to fetch the plurality of memory addresses in parallel.

6 . The apparatus of claim 1 , wherein the neural network is a convolutional neural network (CNN).

7 . The apparatus of claim 1 , wherein the one or more GPUs include circuitry for processing of a neural network.

8 . The apparatus of claim 7 , wherein the circuitry for processing of the neural network includes circuitry to perform memory layout and circuitry to perform feature matching.

9 . The apparatus of claim 7 , wherein the circuitry for processing of the neural network includes circuitry based on the memory pattern associated with the neural network.

10 . A method comprising:

generating a plurality of memory addresses for storage in a memory address portion of a memory, the plurality of memory addresses being based at least in part on a memory pattern associated with a neural network;

rearranging a plurality of feature maps for the neural network such that each feature map starts from an address of the plurality of memory addresses; and

fetching the plurality of memory addresses to load the plurality of feature maps for processing by one or more processors.

11 . The method of claim 10 , further comprising:

convolving the plurality of feature maps with a specified kernel; and

accumulating results of each convolved feature map to generate a combined feature map.

12 . The method of claim 10 , wherein the plurality of memory addresses provide memory layouts that are relatively sparse for two-dimensional (2D) data and relatively dense for three-dimensional (3D) data.

13 . The method of claim 10 , wherein the rearrangement of the plurality of feature maps is performed at one or more of a memory level, a cache level, or a register level.

14 . The method of claim 10 , wherein fetching the plurality of memory addresses includes fetching the plurality of memory addresses in parallel.

15 . The method of claim 10 , wherein the neural network is a convolutional neural network (CNN).

16 . At least one non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations comprising: generating a plurality of memory addresses for storage in a memory address portion of a memory, the plurality of memory addresses being based at least in part on a memory pattern associated with a neural network; rearranging a plurality of feature maps for the neural network such that each feature map starts from an address of the plurality of memory addresses; and fetching the plurality of memory addresses to load the plurality of feature maps for processing by one or more processors.

17 . The non-transitory machine-readable medium of claim 16 , wherein the operations further comprise: convolving the plurality of feature maps with a specified kernel; and accumulating results of each convolved feature map to generate a combined feature map.

18 . The non-transitory machine-readable medium of claim 16 , wherein the plurality of memory addresses provide memory layouts that are relatively sparse for two-dimensional (2D) data and relatively dense for three-dimensional (3D) data.

19 . The non-transitory machine-readable medium of claim 16 , wherein the rearrangement of the plurality of feature maps is performed at one or more of a memory level, a cache level, or a register level.

20 . The non-transitory machine-readable medium of claim 16 , wherein fetching the plurality of memory addresses includes fetching the plurality of memory addresses in parallel.

Continuity (3)
Continuation 17394671 · Aug 5, 2021
Continuation 15581045 · Apr 28, 2017
Related Publication 20240028883A1 · Jan 25, 2024
References Cited (29)
US 5131072A · Yoshizawa · 1992 [cited by examiner]
US 5808621A · Sundaresan · 1998 [cited by applicant]
US 7516129B2 · Risberg · 2009 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 20020019844A1 · Kurowski · 2002 [cited by applicant]
US 20100312735A1 · Knoblauch · 2010 [cited by examiner]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20170300059A1 · Rust · 2017 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180173571A1 · Huang · 2018 [cited by examiner]
US 20200226130A1 · Amzal · 2020 [cited by applicant]
US 20210064931A1 · Yang et al. · 2021 [cited by applicant]
DE 102018110371 · 2018 [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Stephen Junking, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Zhang, et al., “Cambricon-X: An Accelerator for Sparse Neural Networks”, Mar. 2016, IEEE, all pages (Year 2016). [cited by applicant]
Bouthaina, et al., “Shared Hardware Accelerator Architecture for Heterogeneous MPSpCs”, 2013, IEEE, all pages (Year 2013). [cited by applicant]
Li, et al., “A Multistage Dataflow Implementation of a Deep Convolutional Neural Networks Based on FPGA For High-Speed Object Recognition”, 2016, IEEE, all pages (Year 2016). [cited by applicant]
Protic, et al., “Distributed Shared Memory: Concepts and Systems”, 1996, IEEE, all pages. [cited by applicant]
Various authors, “Join (SQL)”, 2014, Wikipedia, all pages (Year:2014). [cited by applicant]
Moreno-Montiel, et al., “Parallel Classification System Based on an Ensemble of Mixture of Experts”, 2014, Proceedings of the 3rd International Conference on Pattern Recognition Applications and Methods, all pages (Year… [cited by applicant]
Abadi, et al., “TensorFlow: Large Scale Machine Learning on Heterogeneous Distributed Systems”, Mar. 16, 2023, arXiv, a;; pages (Year: 2016). [cited by applicant]
Chen, et al., “DianNao Family: Energy Efficient Hardware Accelerators for Machine Learning”, Nov. 2016, Communications of the ACM, all pages (Year: 2016). [cited by applicant]