IP Library Granted Patent US 12,705,683
Granted Patent B2
US 12,705,683 · App. 18/591,993 · Granted Aug 11, 2026

Methods and systems for performing a sparse submanifold deconvolution on a GPU

Inventors: Gunduz Vehbi Demirci (Hertfordshire, GB); Cagatay Dikici (Hertfordshire, GB); Grant Michael Stevens (Hertfordshire, GB); Le Yang (Hertfordshire, GB)
Assignee: Imagination Technologies Limited
G06T1/20G06F17/16G06V10/82G06V10/94
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,683
App. No.
18/591,993
Filed
Feb 29, 2024
Granted
Aug 11, 2026
Kind
B2
Art Unit
2617
USPC
345/501
Abstract

Methods of implementing a sparse submanifold deconvolution on a graphics processing unit, the sparse submanifold deconvolution being representable as a direct convolution between an input tensor to the sparse submanifold deconvolution and each of a plurality of a sub-filters, each sub-filter of the plurality of sub-filters comprising a subset of weights of a filter of the sparse submanifold deconvolution. The methods include: receiving, at the graphics processing unit, the input tensor in a dense format; receiving, at the graphics processing unit, information identifying target positions of an output tensor of the sparse submanifold deconvolution; performing, at the graphics processing unit, an indexed unfold operation on the input tensor based on the identified target positions of the output tensor to generate an input matrix comprising elements of the input tensor in each sub-window of the input tensor relevant to at least one of the identified target positions of the output tensor; and performing, at the graphics processing unit, a matrix multiplication between a weight matrix and the input matrix to generate an output matrix that comprises elements of the output tensor at the identified target positions.

Claims (24)

1 . A method of implementing a sparse submanifold deconvolution on a graphics processing unit, the sparse submanifold deconvolution being representable as a direct convolution between an input tensor to the sparse submanifold deconvolution and each of a plurality of a sub-filters, each sub-filter of the plurality of sub-filters comprising a subset of weights of a filter of the sparse submanifold deconvolution, the method comprising:

receiving, at the graphics processing unit, the input tensor in a dense format;

receiving, at the graphics processing unit, information identifying target positions of an output tensor of the sparse submanifold deconvolution;

performing, at the graphics processing unit, an indexed unfold operation on the input tensor based on the identified target positions of the output tensor to generate an input matrix comprising elements of the input tensor in each sub-window of the input tensor relevant to at least one of the identified target positions of the output tensor; and

performing, at the graphics processing unit, a matrix multiplication between a weight matrix and the input matrix to generate an output matrix that comprises elements of the output tensor at the identified target positions.

2 . The method of claim 1 , wherein the output tensor has at least a height dimension, a width dimension and a channel dimension and a target position of the output tensor is a height and width position of the output tensor.

3 . The method of claim 2 , wherein the information identifying the target positions of the output tensor comprises a target position list that comprises height and width co-ordinates of each target position of the output tensor.

4 . The method of claim 1 , wherein a sub-window of the input tensor is a window of the input tensor used to compute at least one element of an output tensor of one of the direct convolutions.

5 . The method of claim 1 , wherein performing the indexed unfold operation on the input tensor comprises identifying, from the identified target positions of the output tensor and one or more parameters of the sparse submanifold deconvolution, each sub-window of the input tensor relevant to at least one of the identified target positions of the output tensor.

6 . The method of claim 5 , wherein a sub-window of the input tensor is relevant to a target position if that sub-window is used to generate an element of the output tensor of the sparse submanifold deconvolution at that target position.

7 . The method of claim 5 , wherein the elements of a channel of the output tensor of the sparse submanifold deconvolution are divisible into a plurality of blocks wherein each element in a block is generated by a same sub-window of the input tensor and a different sub-filter of a filter, and identifying the sub-window of the input tensor relevant to an identified target position comprises identifying the block of the output tensor that the identified target position forms part of, and mapping the identified block of the output tensor to the sub-window of the input tensor used to generate that block.

8 . The method of claim 7 , wherein an identified block of the output tensor is mapped to a sub-window of the input tensor using a position in the output tensor of a predetermined element of the block and the one or more parameters of the sparse submanifold deconvolution.

9 . The method of claim 1 , wherein performing the indexed unfold operation on the input tensor comprises identifying the elements of each relevant sub-window from one or more parameters of the sparse submanifold deconvolution.

10 . The method of claim 9 , wherein identifying the elements of a relevant sub-window comprises identifying a position in the input tensor of a predetermined element in the sub-window and implementing a series of nested loops to move through the elements in the sub-window from the identified position, the series of nested loops comprising a loop for each dimension of the sub-window.

11 . The method of claim 1 , wherein performing the indexed unfold operation on the input tensor comprises storing the elements of each relevant sub-window in the input matrix.

12 . The method of claim 11 , further comprising receiving a zeroed input matrix, and the elements of the relevant sub-windows of the input tensor are stored in the received input matrix.

13 . The method of claim 1 , wherein performing the indexed unfold operation on the input tensor comprises identifying, from one or more parameters of the sparse submanifold deconvolution, which sub-filter of the plurality of sub-filters is relevant to each of the identified target positions of the output tensor.

14 . The method of claim 13 , wherein the elements of a channel of the output tensor are divisible into a plurality of blocks wherein each element in a block is generated by a same sub-window of the input tensor and a different sub-filter of a filter, and identifying which sub-filter of the plurality of sub-filters is relevant to an identified target position of the output tensor comprises identifying the block of the output tensor that the identified target position forms part of and a location of the identified target position within that block.

15 . The method of claim 1 , wherein the input matrix comprises a column for each relevant sub-window of the input tensor and each column of the input matrix comprises the elements of the input tensor in the corresponding relevant sub-window.

16 . The method of claim 1 , wherein the weight matrix comprises a row for each sub-filter relevant to at least one identified target position of the output tensor, and each row of the weight matrix comprises all weights forming the corresponding sub-filter.

17 . The method of claim 1 , further comprising performing, at the graphics processing unit, an indexed fold operation on the output matrix based on the identified target positions of the output tensor to generate the output tensor in a dense format.

18 . The method of claim 17 , wherein performing the indexed fold operation on the output matrix comprises identifying, based on the identified target positions of the output tensor and one or more parameters of the sparse submanifold deconvolution, elements in the output matrix that correspond to the identified target positions of the output tensor, and storing each element of the output matrix that corresponds to an identified target position at that target position of a channel of the output tensor.

19 . A graphics processing unit configured to perform the method as set forth in claim 1 .

20 . A non-transitory computer readable storage medium having stored thereon computer readable code configured to cause a graphics processing unit to perform the method as set forth in claim 1 when the code is run.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2026
From: DEMIRCI, GUNDUZ VEHBI; STEVENS, GRANT MICHAEL; YANG, LE
To: IMAGINATION TECHNOLOGIES LIMITED
Reel/Frame 075061/0359 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2026
From: DEMIRCI, GUNDUZ VEHBI; STEVENS, GRANT MICHAEL; YANG, LE
To: IMAGINATION TECHNOLOGIES LIMITED
Reel/Frame 075061/0759 →
SECURITY INTEREST Recorded Jul 31, 2024
From: IMAGINATION TECHNOLOGIES LIMITED
To: FORTRESS INVESTMENT GROUP (UK) LTD
Reel/Frame 068221/0001 →
Priority Claims (5)
GB 2303116 · Mar 2, 2023 · national
GB 2303117 · Mar 2, 2023 · national
GB 2303118 · Mar 2, 2023 · national
GB 2303119 · Mar 2, 2023 · national
GB 2303120 · Mar 2, 2023 · national
Continuity (1)
Related Publication 20240320779A1 · Sep 26, 2024
References Cited (18)
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20190156206A1 · Graham · 2019 [cited by examiner]
US 20200301994A1 · Dikici · 2020 [cited by examiner]
US 20210110187A1 · Pillai et al. · 2021 [cited by applicant]
US 20220058450A1 · Lin et al. · 2022 [cited by applicant]
CN 110349146A · 2019 [cited by applicant]
CN 111401554A · 2020 [cited by applicant]
CN 114119377A · 2022 [cited by applicant]
CN 115600662A · 2023 [cited by applicant]
GB 2582352A · 2020 [cited by applicant]
KR 20220114435A · 2022 [cited by applicant]
Graham et al; “3D Semantic Segmentation with Submanifold Sparse Convolutional Networks”; 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; Jun. 18, 2018; pp. 9224-9232. [cited by applicant]
Zhou; “How does sparse convolution work?”; Retrieved from the Internet: URL: https :/ /towardsdatascience.com/how-does-sparse-convolution⋅-work-3257a0a8f d1; Dec. 27, 2020; 11 pages. [cited by applicant]
Graham et al; “Sparse 3D convolutional neural networks”; URL:https://arxiv.org/abs/1505.02890; Aug. 25, 2015; pp. 1-11. [cited by applicant]
Wang et al; “An Efficient FPGA Accelerator for Point Cloud”; IEEE 35th International System-On-Chip Conference; Sep. 5, 2022; pp. 1-6. [cited by applicant]
Sharma, “Sparse Submanifold Convolutions,” Aug. 27, 2021, https://medium.com/geekculture/3d-sparse-sabmanifold-convolutions-eaa427b3a196. [cited by applicant]
Liu et al, 2020, “Deep Adaptive Inference Networks for Single Image Super-Resolution”, [online], Available from: https://arxiv.org/abs/2004. 03915. [cited by applicant]
Park et al, 2017, “Faster CNNs with Direct Sparse Convolutions and Guided Pruning”, International Conference on Learning Representations, [online], Available from: https://arxiv.org/pdf/1608.01409.pdf. [cited by applicant]