IP Library › Granted Patent US 12,541,569
Granted Patent B2
US 12,541,569 · App. 17/554,970 · Granted Feb 3, 2026

Methods and apparatus for performing a machine learning operation using storage element pointers

Inventors: Kevin Brady (Newry, GB); Martin Power (Dublin, IE); Martin-Thomas Grymel (Leixlip, IE); Alessandro Palla (Pisa, IT); David Bernard (Kilcullen, IE); Niall Hanrahan (Galway, IE)
Assignee: Intel Corporation
G06F18/2148G06F18/213G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,569
App. No.
17/554,970
Granted
Feb 3, 2026
Kind
B2
Abstract

Methods, apparatus, systems, and articles of manufacture are disclosed for performing a machine learning operation using storage element pointers. An example computer readable medium comprises instructions that when executed, cause at least one processor to select, in response to a determination that a machine learning operation is to be performed, create first and second storage element pointers based on a type of machine learning operation to be performed, remap input tensor data of the input tensor based on the first storage element pointer without movement of the input tensor data in memory, cause execution of the machine learning operation with the remapped input tensor data to create intermediate tensor data, remap the intermediate tensor data based on the second storage element pointer without movement of the intermediate tensor data in memory, and provide the remapped intermediate tensor data as an output tensor.

Claims (42)

1 . An apparatus, comprising:

a computer processor; and a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations, the operations comprising:

generating a first pointer table for a machine learning operation of a first type, the first pointer table comprising storage element pointers pointing to storage elements of a memory, wherein a storage element is to store a particular portion of an input tensor of the machine learning operation of the first type;

generating a second pointer table by rearranging one or more storage element pointers in the first pointer table;

forming a first tensor based on the second pointer table, the first tensor comprising a data element of the input tensor, wherein a position of the data element in the first tensor is different form a position of the data element in the input tensor;

causing execution of a machine learning operation of a second type with the first tensor to generate a second tensor;

generating a third pointer group, the third pointer group comprising storage element pointers pointing to storage elements storing data elements of the second tensor; and

forming an output tensor of the machine learning operation of the first type from the second tensor based on the third pointer group.

2 . The apparatus of claim 1 , wherein the machine learning operation of the first type is a convolution operation.

3 . The apparatus of claim 2 , wherein the convolution operation of the second type is a group convolution operation.

4 . The apparatus of claim 2 , wherein the convolution operation of the second type is a dilated convolution operation.

5 . The apparatus of claim 1 , wherein the first tensor further comprises an additional data element of the input tensor, wherein the data element and the additional data element are adjacent in the first tensor, wherein the data element and the additional data element are separated by one or more other data elements in the input tensor.

6 . The apparatus of claim 1 , wherein the second pointer table comprises a subset of the storage element pointers of the first pointer table.

7 . The apparatus of claim 1 , wherein the particular portion of the input tensor comprises activations in a plurality of input channels of the machine learning operation of the first type.

8 . At least one non-transitory computer readable medium storing instructions that, when executed, cause at least one processor to perform operations, the operations comprising:

generating a first pointer table for a machine learning operation of a first type, the first pointer table comprising storage element pointers pointing to storage elements of a memory, wherein a storage element is to store a particular portion of an input tensor of the machine learning operation of the first type;

generating a second pointer table by rearranging one or more storage element pointers in the first pointer table;

forming a first tensor based on the second pointer table, the first tensor comprising a data element of the input tensor, wherein a position of the data element in the first tensor is different form a position of the data element in the input tensor;

causing execution of a machine learning operation of a second type with the first tensor to generate a second tensor;

generating a third pointer group, the third pointer group comprising storage element pointers pointing to storage elements storing data elements of the second tensor; and

forming an output tensor of the machine learning operation of the first type from the second tensor based on the third pointer group.

9 . The at least one non-transitory computer readable medium of claim 8 , wherein the machine learning operation of the first type is a convolution operation.

10 . The at least one non-transitory computer readable medium of claim 9 , wherein the convolution operation of the second type is a group convolution operation.

11 . The at least one non-transitory computer readable medium of claim 9 , wherein the convolution operation of the second type is a dilated convolution operation.

12 . The at least one non-transitory computer readable medium of claim 8 , wherein the generating of the first pointer table, the second pointer table, or the third pointer table is performed before the execution of the machine learning operation of the second type.

13 . The at least one non-transitory computer readable medium of claim 8 , wherein the first tensor further comprises an additional data element of the input tensor, wherein the data element and the additional data element are adjacent in the first tensor, wherein the data element and the additional data element are separated by one or more other data elements in the input tensor.

14 . The at least one non-transitory computer readable medium of claim 8 , wherein the second pointer table comprises a subset of the storage element pointers of the first pointer table.

15 . The at least one non-transitory computer readable medium of claim 8 , wherein the particular portion of the input tensor comprises activations in a plurality of input channels of the machine learning operation of the first type.

16 . A method, comprising:

generating a first pointer table for a machine learning operation of a first type, the first pointer table comprising storage element pointers pointing to storage elements of a memory, wherein a storage element is to store a particular portion of an input tensor of the machine learning operation of the first type;

generating a second pointer table by rearranging one or more storage element pointers in the first pointer table;

forming a first tensor based on the second pointer table, the first tensor comprising a data element of the input tensor, wherein a position of the data element in the first tensor is different form a position of the data element in the input tensor;

causing execution of a machine learning operation of a second type on the first tensor to generate a second tensor;

generating a third pointer group, the third pointer group comprising storage element pointers pointing to storage elements storing data elements of the second tensor; and

forming an output tensor of the machine learning operation of the first type from the second tensor based on the third pointer group.

17 . The method of claim 16 , wherein the machine learning operation of the first type is a convolution operation.

18 . The method of claim 17 , wherein the convolution operation of the second type is a group convolution operation.

19 . The method of claim 17 , wherein the convolution operation of the second type is a dilated convolution operation.

20 . The method of claim 16 , wherein the generating of the first pointer table, the second pointer table, or the third pointer table is performed before the execution of the machine learning operation of the second type.

21 . The method of claim 16 , wherein the first tensor further comprises an additional data element of the input tensor, wherein the data element and the additional data element are adjacent in the first tensor, wherein the data element and the additional data element are separated by one or more other data elements in the input tensor.

22 . The method of claim 16 , wherein the second pointer table comprises a subset of the storage element pointers of the first pointer table.

23 . The method of claim 16 , wherein the particular portion of the input tensor comprises activations in a plurality of input channels of the machine learning operation of the first type.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2022
From: HANRAHAN, NIALL; BERNARD, DAVID
To: INTEL CORPORATION
Reel/Frame 059008/0416 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2022
From: POWER, MARTIN; GRYMEL, MARTIN-THOMAS; PALLA, ALESSANDRO; BRADY, KEVIN
To: INTEL CORPORATION
Reel/Frame 058706/0283 →
Continuity (1)
Related Publication 20220108135A1 · Apr 7, 2022
References Cited (33)
US 10275306B1 · MacLaren · 2019 [cited by examiner]
US 10303543B1 · MacLaren · 2019 [cited by examiner]
US 11526469B1 · Mathews · 2022 [cited by examiner]
US 20080168253A1 · Garrison · 2008 [cited by applicant]
US 20160343455A1 · Lesartre et al. · 2016 [cited by applicant]
US 20190220704A1 · Schulz-Trieglaff · 2019 [cited by examiner]
US 20200019859A1 · Kyriazopoulou Panagiotopoulou · 2020 [cited by examiner]
US 20200042220A1 · Armangau · 2020 [cited by examiner]
US 20200285949A1 · Baum et al. · 2020 [cited by applicant]
US 20210064987A1 · Springer · 2021 [cited by examiner]
US 20210142154A1 · Pang · 2021 [cited by applicant]
US 20210181951A1 · Jin · 2021 [cited by applicant]
US 20210303155A1 · Meister · 2021 [cited by examiner]
US 20220108135A1 · Brady et al. · 2022 [cited by applicant]
WO 2021035397A1 · 2021 [cited by applicant]
Yu et al., “A Data-Center FPGA Acceleration Platform for Convolutional Neural Networks,” IEEE, 2019, 8 pgs. (Year: 2019). [cited by examiner]
Aimar et al., “NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps, ” IEEE Transactions on Neural Networks and Learning Systems, vol. 30, No. 1, Mar. 2019, pp. 64… [cited by examiner]
Lin et al., “MERIT: Tensor Transform for Memory-Efficient Vision Processing on Parallel Architectures,” Nov. 7, 2019, 14 pgs. (Year: 2019). [cited by examiner]
Majumder et al., “A Flexible FPGA Accelerator for Convolutional Neural Networks,” Dec. 21, 2019, 13 pgs. (Year: 2019). [cited by examiner]
Carreras et al., “Optimizing Temporal Convolutional Network Inference on FPGA-Based Accelerators,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. 10, No. 3, Sep. 2019, pp. 348-361. (Year: 201… [cited by examiner]
Srivastava et al., “Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense tensor Computations,” 2020 IEEE International Symposium on High Performance Computer Architecture, pp. 689-702. (Year: 2020). [cited by examiner]
Abts et al., “Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads,” 2020 ACM/IEEE 47th International Symposium on Computer Architecture, pp. 145-158. (Year: 2020). [cited by examiner]
Dave et al., “Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights,” IEEE, Jul. 22, 2021, 44 pgs. (Year: 2021). [cited by examiner]
International Searching Authority, “Written Opinion of the International Searching Authority,” issued in connection with International Patent Application No. PCT/US2022/050301, mailed on Mar. 27, 2023, 4 pages. [cited by applicant]
International Searching Authority, “International Search Report,” issued in connection with International Patent Application No. PCT/US2022/050301, mailed on Mar. 27, 2023, 4 pages. [cited by applicant]
Zhang et al., “ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices,” arXiv:1707.01083v2, submitted Jul. 4, 2017, revised Dec. 7, 2017, 9 pages. [cited by applicant]
Chen et al., “DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs,” arXiv:1606.00915v2, submitted Jun. 2, 2016, revised May 12, 2017, 14 pages. [cited by applicant]
Zhao et al., “ICNet for Real-Time Semantic segmentation on High-Resolution Images,” arXiv:1704.08545v2, submitted Apr. 27, 2017, revised Aug. 20, 2018, 16 pages. [cited by applicant]
Kie et al., “Aggregated Residual Transformations for Deep Neural Networks,” arXiv:1611.05431v2, submitted Nov. 16, 2016, revised Apr. 11, 2017, 10 pages. [cited by applicant]
Gao et al., “Res2Net: A New Multi-scale Backbone Architecture,” arXiv:1904.01169v3, submited Apr. 2, 2019, revised Jan. 27, 2021, 11 pages. [cited by applicant]
Wolterink et al., “Dilated Convolutional Neural Networks for Cardiovascular MR Segmentation in Congenital Heart Disease,” arXiv:1704.03669v1, submitted Apr. 12, 2017, 9 pages. [cited by applicant]
Van Den Oord et al., “WaveNet: A Generative Model for Raw Audio,” arXiv:1609.03499v2, submitted Sep. 12, 2016, revised Sep. 19, 2016, 15 pages. [cited by applicant]
Lavasani et al., An FPGA-based In-line Acclertor for Memcached, IEEE Computer Architecture Letters, vol. 13, No. 2, Jul.-Dec. 2014, 4 pages. [cited by applicant]