IP Library › Granted Patent US 12,682,604
Granted Patent B2
US 12,682,604 · App. 17/435,657 · Granted Jul 14, 2026

Generic modular sparse three-dimensional (3D) convolution design utilizing sparse 3D group convolution

Inventors: Anbang Yao (Beijing, CN); Jiahui Zhang (Beijing, CN); Dawei Sun (Beijing, CN); Dian Gu (Shanghai, CN); Yurong Chen (Beijing, CN)
Assignee: INTEL CORPORATION
G06V10/764G06N3/04G06V10/765G06V10/82G06V10/95G06V10/955G06V10/513
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,604
App. No.
17/435,657
Filed
Sep 1, 2021
Granted
Jul 14, 2026
Kind
B2
Art Unit
2143
USPC
706/27
Abstract

Embodiments are generally directed to sparse 3D convolution acceleration in a convolutional layer of an artificial neural network model. An embodiment of an apparatus includes one or more processors including a graphics processor to process data; and a memory for storage of data, including feature maps. The one or more processors are to provide for sparse 3D convolution acceleration by applying a shared 3D convolutional kernel/filter to an input feature map to produce an output feature map, including increasing sparsity of the input feature map by partitioning it into multiple disjoint input groups; generation of multiple disjoint output groups corresponding to the input groups by performing a convolution calculation represented by the shared 3D convolutional kernel/filter on all feature values associated with active/valid voxels of each input group to produce corresponding feature values within corresponding output groups; and outputting the output feature map by sequentially stacking the output groups.

Claims (46)

1 . An apparatus for sparse three-dimensional (3D) convolution acceleration comprising:

one or more processors including a graphics processor to process data; and

a memory for storage of data, including feature maps;

wherein the one or more processors are to provide for sparse 3D convolution acceleration in a first convolutional layer of a plurality of convolutional layers of a neural network model to facilitate training of the neural network model to perform analysis or processing of image data by applying a shared 3D convolutional kernel/filter to an input feature map derived from the image data to produce an output feature map, including:

linearly increasing sparsity of the input feature map by partitioning the input feature map into a plurality (G) of disjoint input groups having dimensions of magnitudes equal to the input feature map, wherein said partitioning includes programmatically selecting active/valid voxels of the input feature map for inclusion in respective input groups of the plurality of disjoint input groups;

generation of a plurality of disjoint output groups corresponding to the plurality of disjoint input groups by, for each input group of the plurality of disjoint input groups, performing, a convolution operation using the shared 3D convolutional kernel/filter on all input feature values associated with all or a subset of active/valid voxels of the input group to produce corresponding output feature values within a corresponding output group of the plurality of disjoint output groups by sliding the shared 3D convolutional kernel/filter according to a stride, left-to-right and top-to-bottom, within a spatial dimension of the input group; and

creating the output feature map with enhanced feature representation for use by a subsequent convolutional layer of the plurality of convolutional layers, wherein the output feature map merges all output feature values associated with active/valid voxels of each of the plurality of disjoint output groups into the output feature map by sequentially stacking the plurality of disjoint output groups.

2 . The apparatus of claim 1 , wherein a sparsity of the input feature map is s % and a sparsity of an average input group representative of the plurality of disjoint input groups is G*s %.

3 . The apparatus of claim 1 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups is performed along the spatial dimension of the input feature map and applies to a plurality of input channels.

4 . The apparatus of claim 3 , wherein the active/valid voxels of the input group are efficiently identified with reference to a hash table containing a plurality of entries, wherein each entry of the plurality entries corresponds to a particular active/valid voxel and specifies (i) the spatial dimension of the particular active/valid voxel in terms of w, h and d coordinates and (ii) a group identifier of a particular input group of the plurality of disjoint input groups to which the particular active/valid voxel has been partitioned.

5 . The apparatus of claim 1 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups is performed independently for each of a plurality of input channel groups.

6 . The apparatus of claim 1 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups includes dividing the input feature map based on a randomly generated group identifier specifying a group of the plurality of disjoint input groups for each of the active/valid voxels.

7 . The apparatus of claim 1 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups includes generating a group identifier (GID) specifying a group of the plurality of disjoint input groups to which each of the active/valid voxels is to be partitioned based on a mapping function of a form:

GID=mod( aw+bh+cd;G )

where:

mod is a modulus operation;

(w, h, d) represents a position of a particular active/valid voxel in the input feature map; and

a, b, and c are configurable parameters to control the partitioning.

8 . A method for sparse three-dimensional (3D) convolution acceleration as part of training a neural network model implemented on a data processing system to perform analysis or processing of image data, the method comprising:

receiving, by a first convolutional layer of a plurality of convolutional layers of the neural network model an input feature map derived from the image data;

linearly increasing sparsity of the input feature map, by partitioning, by the first convolutional layer, the input feature map into a plurality (G) of disjoint input groups having dimensions of magnitudes equal to the input feature map;

generating, by the first convolutional layer, a plurality of disjoint output groups corresponding to the plurality of disjoint input groups by, for each input group of the plurality of disjoint input groups, performing, a convolution operation using a shared 3D convolutional kernel/filter associated with the first convolutional layer on all input feature values of all or a subset of active/valid voxels of the input group to produce corresponding output feature values within a corresponding output group of the plurality of disjoint output groups by sliding the convolutional kernel/filter according to a stride, left-to-right and top-to-bottom, within a spatial dimension of the input group; and

creating, by the first convolutional layer, an output feature map with enhanced feature representation for use by a subsequent convolutional layer of the plurality of convolutional layers, wherein the output feature map merges all output feature values associated with active/valid voxels of each of the plurality of disjoint output groups into the output feature map by sequentially stacking the plurality of disjoint output groups.

9 . The method of claim 8 , wherein a sparsity of the input feature map is s % and a sparsity of an average input group representative of the plurality of disjoint input groups is G*s %.

10 . The method of claim 8 , wherein the active/valid voxels of the input group are efficiently identified with reference to a hash table containing a plurality of entries, wherein each entry of the plurality entries corresponds to a particular active/valid voxel and specifies (i) a spatial dimension of the particular active/valid voxel in terms of w, h and d coordinates and (ii) a group identifier of a particular input group of the plurality of disjoint input groups to which the particular active/valid voxel has been partitioned.

11 . The method of claim 8 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups comprise independently performing said partitioning for each of a plurality of input channel groups.

12 . The method of claim 8 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups comprises dividing an equal number of voxels of the input feature map among each of the plurality of disjoint input groups based on a randomly generated group identifier specifying a group of the plurality of disjoint input groups for each of the active/valid voxels.

13 . The method of claim 8 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups comprises generating a group identifier (GID) specifying a group of the plurality of disjoint input groups to which each of the active/valid voxels is to be partitioned based on a mapping function of a form:

GID=mod( aw+bh+cd;G )

where:

mod is a modulus operation;

(w, h, d) represents a position of a particular active/valid voxel in the input feature map; and

a, b, and c are configurable parameters to control the partitioning.

14 . A non-transitory computer-readable storage medium embodying a set of instructions, which when executed by one or more processors, including a graphics processor, causes the one or more processors to, as part of training of a neural network model to perform analysis or processing of image data:

receive, by a first convolutional layer of a plurality of convolutional layers of the neural network model, an input feature map derived from the image data;

linearly increase sparsity of the input feature map, by partitioning, by the first convolutional layer, the input feature map into a plurality (G) of disjoint input groups having dimensions of magnitudes equal to the input feature map;

generate, by the first convolutional layer, a plurality of disjoint output groups corresponding to the plurality of disjoint input groups by, for each input group of the plurality of disjoint input groups, performing, a convolution operation using a shared 3D convolutional kernel/filter associated with the convolutional layer on all input feature values of all or a subset of active/valid voxels of the input group to produce corresponding output feature values within a corresponding output group of the plurality of disjoint output groups by sliding the shared 3D convolutional kernel/filter according to a stride, left-to-right and top-to-bottom, within a spatial dimension of the input group; and

create, by the first convolutional layer, an output feature map with enhanced feature representation for use by a subsequent convolutional layer of the plurality of convolutional layers, wherein the output feature map merges all output feature values associated with active/valid voxels of each of the plurality of disjoint output groups into the output feature map by sequentially stacking the plurality of disjoint output groups.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein a sparsity of the input feature map is s % and a sparsity of an average input group representative of the plurality of disjoint input groups is G*s %.

16 . The non-transitory computer-readable storage medium of claim 14 , wherein the active/valid voxels of the input group are efficiently identified with reference to a hash table containing a plurality of entries, wherein each entry of the plurality entries corresponds to a particular active/valid voxel and specifies (i) a spatial dimension of the particular active/valid voxel in terms of w, h and d coordinates and (ii) a group identifier of a particular input group of the plurality of disjoint input groups to which the particular active/valid voxel has been partitioned.

17 . The non-transitory computer-readable storage medium of claim 14 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups comprises:

independently performing said partitioning for each of a plurality of input channel groups; or

dividing an equal number of voxels of the input feature map among each of the plurality of disjoint input groups based on a randomly generated group identifier specifying a group of the plurality of disjoint input groups for each of the active/valid voxels.

18 . The apparatus of claim 1 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups is performed in accordance with a first partitioning strategy for a first input channel group of a plurality of input channel groups and in accordance with a second partitioning strategy for a second input channel group of the plurality of input channel groups.

19 . The method of claim 8 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups is performed in accordance with a first partitioning strategy for a first input channel group of a plurality of input channel groups and in accordance with a second partitioning strategy for a second input channel group of the plurality of input channel groups.

20 . The non-transitory computer-readable storage medium of claim 14 , wherein said partitioning the input feature map into a plurality (G) of disjoint input groups is performed in accordance with a first partitioning strategy for a first input channel group of a plurality of input channel groups and in accordance with a second partitioning strategy for a second input channel group of the plurality of input channel groups.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2021
From: YAO, ANBANG; ZHANG, JIAHUI; SUN, DAWEI; GU, DIAN; CHEN, YURONG
To: INTEL CORPORATION
Reel/Frame 057672/0210 →
Continuity (1)
Related Publication 20220147791A1 · May 12, 2022
References Cited (24)
US 20160196672A1 · Chertok · 2016 [cited by examiner]
US 20180181857A1 · Mathew et al. · 2018 [cited by applicant]
US 20180181864A1 · Mathew · 2018 [cited by examiner]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20190156206A1 · Graham · 2019 [cited by examiner]
US 20200050555A1 · Kim · 2020 [cited by examiner]
US 20210166464A1 · Moloney · 2021 [cited by examiner]
US 20210200993A1 · Chen · 2021 [cited by examiner]
CN 109726671A · 2019 [cited by applicant]
CN 109829399A · 2019 [cited by applicant]
EP 3987433A1 · 2022 [cited by applicant]
WO 2018084942A1 · 2018 [cited by applicant]
WO 2020252762A1 · 2020 [cited by applicant]
Xue, Jian, et al. “Efficient volume rendering methods for out-of-Core datasets by semi-adaptive partitioning.” Information Sciences 370 (2016): 463-475. (Year: 2016). [cited by examiner]
Zhang, Ting, et al. “Interleaved group convolutions.” Proceedings of the IEEE international conference on computer vision. 2017. (Year: 2017). [cited by examiner]
Li, Hongyang, Wanli Ouyang, and Xiaogang Wang. “Multi-bias non-linear activation in deep neural networks.” International conference on machine learning. PMLR, 2016. (Year: 2016). [cited by examiner]
Riegler, Gernot, Ali Osman Ulusoy, and Andreas Geiger. “Octnet: Learning deep 3d representations at high resolutions.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017. (Year: 2017). [cited by examiner]
Dumoulin, Vincent, and Francesco Visin. “A guide to convolution arithmetic for deep learning.” arXiv preprint arXiv:1603.07285 (2016). (Year: 2016). [cited by examiner]
Zhang Jiahui et al: “Efficient Semantic Scene Completion Network with Spatial Group Convolution”, Oct. 6, 2018 (Oct. 6, 2018), SAT 2015 18th International Conference, Austin, TX, USA, Sep. 24-27, 2015; [Lecture Notes in… [cited by applicant]
Extended European Search Report for EP19934327, Feb. 20, 2023, 9 pages. [cited by applicant]
Intention to Grant for EP19934327, Sep. 10, 2024, 5 pages. [cited by applicant]
Notification of Publication for CN Application No. 201980096478.X, Jan. 4, 2022, 3 pages. [cited by applicant]
Indian Hearing Notice for IN Application No. 202147041874 mailed Feb. 13, 2025, 4 pages. [cited by applicant]
Zhang et al., “Efficient Semantic Scene Completion Network with Spatial Group Convolution”, Tsinghua University, 17 pages, 2018. [cited by applicant]