IP Library Granted Patent US 11,720,800
Granted Patent B2
US 11,720,800 · App. 17/455,863 · Granted Aug 8, 2023

Efficient data layouts for convolutional neural networks

Inventors: Ashkan Aliabadi (Santa Clara, CA); Gregory David Roberts (Pasadena, CA)
Assignee: Magic Leap, Inc.
G06N3/084G06F16/172G06F18/2113G06F18/2137G06F18/231G06N3/04G06N3/045G06N3/063G06V10/82G06V40/193G06V10/449
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,720,800
App. No.
17/455,863
Granted
Aug 8, 2023
Kind
B2
Abstract

Systems and methods for efficient implementation of a convolutional layer of a convolutional neural network are disclosed. In one aspect, weight values of kernels in a kernel stack of a convolutional layer can be reordered into a tile layout with tiles of runnels. Pixel values of input activation maps of the convolutional layer can be reordered into an interleaved layout comprising a plurality of clusters of input activation map pixels. The output activation maps can be determined using the clusters of the input activation map pixels and kernels tile by tile.

Claims (34)

1. A method implemented by a system of one more processors, the method comprising:

receiving a convolutional layer of a convolutional neural network comprising kernels in a kernel stack, wherein the kernels of the kernel stack are in a tile kernel layout comprising a plurality of kernel tiles of kernel runnels;

receiving input activation maps of the convolutional layer, wherein the input activation maps are in a basic input activation map layout;

reordering pixel values of the input activation maps from the basic input activation map layout into an interleaved input activation map layout comprising a plurality of clusters of input activation map pixels by striding; and

determining output activation maps of the convolutional layer from the plurality of kernel tiles and a plurality of input activation map tiles.

2. The method of claim 1 , wherein reordering the pixel values of the input activation maps from the basic input activation map layout into the interleaved input activation map layout comprises reordering pixel values of the input activation maps from the basic input activation map layout into the interleaved input activation map layout comprising the plurality of clusters of input activation map pixels by striding with a stride size of a multiple of a number of the input activation maps.

3. The method of claim 2 , wherein the multiple of the number of the input activation maps is one.

4. The method of claim 2 , wherein the multiple of the number of the input activation maps is a multiple of a register width, the register width corresponding to a size associated with a kernel runnel.

5. The method of claim 1 , wherein a dimension of a kernel is one.

6. The method of claim 1 , wherein the pixel values are contiguous in memory of the system.

7. The method of claim 1 , wherein the output activation maps are in a transposed, interleaved output activation map layout comprising a plurality of clusters of output activation maps.

8. The method of claim 1 , wherein determining the output activation maps comprises:

performing fused-multiply-add operations tile by tile on the plurality of kernel tiles and the plurality of clusters of input activation map pixels.

9. A system comprising one or more processors and non-transitory computer storage media storing instructions that when executed by the one or more processors, cause the one or more processors to:

receive a convolutional layer of a convolutional neural network comprising kernels in a kernel stack, wherein the kernels of the kernel stack are in a tile kernel layout comprising a plurality of kernel tiles of kernel runnels;

receive input activation maps of the convolutional layer, wherein the input activation maps are in a basic input activation map layout;

reorder pixel values of the input activation maps from the basic input activation map layout into an interleaved input activation map layout comprising a plurality of clusters of input activation map pixels by striding; and

determine output activation maps of the convolutional layer from the plurality of kernel tiles and a plurality of input activation map tiles.

10. The system of claim 9 , wherein reordering the pixel values of the input activation maps from the basic input activation map layout into the interleaved input activation map layout comprises reordering pixel values of the input activation maps from the basic input activation map layout into the interleaved input activation map layout comprising the plurality of clusters of input activation map pixels by striding with a stride size of a multiple of a number of the input activation maps.

11. The system of claim 10 , wherein the multiple of the number of the input activation maps is one.

12. The system of claim 10 , wherein the multiple of the number of the input activation maps is a multiple of a register width, the register width corresponding to a size associated with a kernel runnel.

13. The system of claim 9 , wherein a dimension of a kernel is one.

14. The system of claim 9 , wherein the pixel values are contiguous in memory of the system.

15. The system of claim 9 , wherein the output activation maps are in a transposed, interleaved output activation map layout comprising a plurality of clusters of output activation maps.

16. The system of claim 9 , wherein determining the output activation maps comprises:

performing fused-multiply-add operations tile by tile on the plurality of kernel tiles and the plurality of clusters of input activation map pixels.

17. Non-transitory computer storage media storing instructions that when executed by a system of one or more processors, cause the one or more processors to:

receive a convolutional layer of a convolutional neural network comprising kernels in a kernel stack, wherein the kernels of the kernel stack are in a tile kernel layout comprising a plurality of kernel tiles of kernel runnels;

receive input activation maps of the convolutional layer, wherein the input activation maps are in a basic input activation map layout;

reorder pixel values of the input activation maps from the basic input activation map layout into an interleaved input activation map layout comprising a plurality of clusters of input activation map pixels by striding; and

determine output activation maps of the convolutional layer from the plurality of kernel tiles and a plurality of input activation map tiles.

18. The computer storage media of claim 17 , wherein reordering the pixel values of the input activation maps from the basic input activation map layout into the interleaved input activation map layout comprises reordering pixel values of the input activation maps from the basic input activation map layout into the interleaved input activation map layout comprising the plurality of clusters of input activation map pixels by striding with a stride size of a multiple of a number of the input activation maps.

19. The computer storage media of claim 17 , wherein reordering the pixel values of the input activation maps from the basic input activation map layout into the interleaved input activation map layout comprises reordering pixel values of the input activation maps from the basic input activation map layout into the interleaved input activation map layout comprising the plurality of clusters of input activation map pixels by striding with a stride size of a multiple of a number of the input activation maps.

20. The computer storage media of claim 17 , wherein the multiple of the number of the input activation maps is one or wherein the multiple of the number of the input activation maps is a multiple of a register width, the register width corresponding to a size associated with a kernel runnel.

Assignments (3)
SECURITY INTEREST Recorded Oct 20, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073008/0696 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2023
From: ALIABADI, ASHKAN; ROBERTS, GREGORY DAVID
To: MAGIC LEAP, INC.
Reel/Frame 063075/0040 →
SECURITY INTEREST Recorded May 24, 2022
From: MOLECULAR IMPRINTS, INC.; MENTOR ACQUISITION ONE, LLC; MAGIC LEAP, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 060338/0665 →
Continuity (4)
Continuation 16680335 · Nov 11, 2019
Continuation 15724142 · Oct 3, 2017
Provisional Application 62403930 · Oct 4, 2016
Related Publication 20220076056A1 · Mar 10, 2022