IP Library › Granted Patent US 12,450,466
Granted Patent B2
US 12,450,466 · App. 17/060,420 · Granted Oct 21, 2025

Superpixel methods for convolutional neural networks

Inventors: Reginald Clifford Young (Palo Alto, CA); Jonathan Ross (Mountain View, CA)
Assignee: Google LLC
G06N3/04G06F17/16G06N3/02G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,466
App. No.
17/060,420
Granted
Oct 21, 2025
Kind
B2
Abstract

Methods, systems, and apparatus for efficiently performing a computation of a convolutional neural network layer. One of the methods includes transforming a X by Y by Z input tensor into a X′ by Y′ by Z′ input tensor, wherein X′ is smaller than or equal to X, Y′ is smaller than or equal to Y, and Z′ is larger than or equal to Z; obtaining one or more modified weight matrices, wherein the modified weight matrices operate on the X′ by Y′ by Z′ input tensor to generate a U′ by V′ by W′ output tensor, and the U′ by V′ by W′ output tensor is a transformed U by V by W output tensor; and processing the X′ by Y′ by Z′ input tensor using the modified weight matrices to generate the U′ by V′ by W′ output tensor.

Claims (58)

1. A method performed using a convolutional neural network implemented on a hardware integrated circuit, the method comprising:

receiving an input tensor for a layer of the convolutional neural network having a stride greater than one, the input tensor having multiple dimensions and a plurality of inputs;

generating a modified input tensor from the input tensor based on the stride greater than one, the modified input tensor having multiple dimensions and a respective plurality of inputs;

processing the modified input tensor using a modified weight matrix for a superpixel layer of the convolutional neural network having a stride equal to one, comprising:

applying a convolution to inputs of the modified input tensor using the modified weight matrix; and

in response to processing the modified input tensor, generating a transformed layer output of the superpixel layer, the transformed layer output comprising outputs that correspond mathematically to a neural network output generated by processing the input tensor using an unmodified version of the modified weight matrix;

wherein each of the input tensor and the modified input tensor are multi-dimensional tensors with a respective depth dimension; and

the respective depth dimension of the modified input tensor is greater than the respective depth dimension of the input tensor.

2. The method of claim 1 , wherein applying the convolution comprises:

convolving inputs of the modified input tensor with a corresponding weight of the modified weight matrix based on a kernel stride that controls shifting of a convolutional filter corresponding to the modified weight matrix as the convolutional filter is processed against inputs of the modified input tensor.

3. The method of claim 1 , comprising:

grouping multiple inputs of the input tensor at least by trading spatial extent or indexing with regard to X and Y dimensions of the input tensor for depth extent or indexing with regard to a Z dimension of the input tensor.

4. The method of claim 3 , comprising:

based on the grouping of multiple inputs, transforming the layer of the convolutional neural network having the stride greater than one into the superpixel layer having the stride equal to one;

wherein the superpixel layer comprises a reduced number of kernel elements relative to an untransformed layer of the convolutional neural network.

5. The method of claim 4 , wherein transforming the layer of the convolutional neural network into the superpixel layer comprises:

applying a superpixel transformation to convolutional neural network layer inputs of the input tensor and an unmodified convolutional neural network layer weight matrix to generate a differently shaped, but mathematically equivalent, superpixel convolutional neural network layer.

6. The method of claim 4 , wherein:

processing the modified input tensor using the modified weight matrix for the superpixel layer requires fewer matrix multiplications than processing the input tensor using an modified weight matrix for the untransformed layer of the convolutional neural network.

7. The method of claim 1 , wherein generating the transformed layer output of the superpixel layer comprises:

generating a score or classification output corresponding to inputs of the input tensor that are extracted features of an image.

8. A system for performing neural network computations using a convolutional neural network implemented on a hardware integrated circuit, the system comprising:

a processor and a non-transitory machine-readable storage device storing instructions that are executable by the processor to perform operations comprising:

receiving an input tensor for a layer of the convolutional neural network having a stride greater than one, the input tensor having multiple dimensions and a plurality of inputs;

generating a modified input tensor from the input tensor based on the stride greater than one, the modified input tensor having multiple dimensions and a respective plurality of inputs;

processing the modified input tensor using a modified weight matrix for a superpixel layer of the convolutional neural network having a stride equal to one, comprising:

applying a convolution to inputs of the modified input tensor using the modified weight matrix; and

in response to processing the modified input tensor, generating a transformed layer output of the superpixel layer, the transformed layer output comprising outputs that correspond mathematically to a neural network output generated by processing the input tensor using an unmodified version of the modified weight matrix;

wherein each of the input tensor and the modified input tensor are multi-dimensional tensors with a respective depth dimension; and

the respective depth dimension of the modified input tensor is greater than the respective depth dimension of the input tensor.

9. The system of claim 8 , wherein applying the convolution comprises:

convolving inputs of the modified input tensor with a corresponding weight of the modified weight matrix based on a kernel stride that controls shifting of a convolutional filter corresponding to the modified weight matrix as the convolutional filter is processed against inputs of the modified input tensor.

10. The system of claim 8 , wherein the operations comprise:

grouping multiple inputs of the input tensor at least by trading spatial extent or indexing with regard to X and Y dimensions of the input tensor for depth extent or indexing with regard to a Z dimension of the input tensor.

11. The system of claim 10 , wherein the operations comprise:

based on the grouping of multiple inputs, transforming the layer of the convolutional neural network having the stride greater than one into the superpixel layer having the stride equal to one;

wherein the superpixel layer comprises a reduced number of kernel elements relative to an untransformed layer of the convolutional neural network.

12. The system of claim 11 , wherein transforming the layer of the convolutional neural network into the superpixel layer comprises:

applying a superpixel transformation to convolutional neural network layer inputs of the input tensor and an unmodified convolutional neural network layer weight matrix to generate a differently shaped, but mathematically equivalent, superpixel convolutional neural network layer.

13. The system of claim 11 , wherein:

processing the modified input tensor using the modified weight matrix for the superpixel layer requires fewer matrix multiplications than processing the input tensor using an modified weight matrix for the untransformed layer of the convolutional neural network.

14. The system of claim 8 , wherein generating the transformed layer output of the superpixel layer comprises:

generating a score or classification output corresponding to inputs of the input tensor that are extracted features of an image.

15. A non-transitory machine-readable storage device storing instructions for performing neural network computations using a convolutional neural network implemented on a hardware integrated circuit, the instructions being executable by a processor to perform operations comprising:

receiving an input tensor for a layer of the convolutional neural network having a stride greater than one, the input tensor having multiple dimensions and a plurality of inputs;

generating a modified input tensor from the input tensor based on the stride greater than one, the modified input tensor having multiple dimensions and a respective plurality of inputs;

processing the modified input tensor using a modified weight matrix for a superpixel layer of the convolutional neural network having a stride equal to one, comprising:

applying a convolution to inputs of the modified input tensor using the modified weight matrix; and

in response to processing the modified input tensor, generating a transformed layer output of the superpixel layer, the transformed layer output comprising outputs that correspond mathematically to a neural network output generated by processing the input tensor using an unmodified version of the modified weight matrix;

wherein each of the input tensor and the modified input tensor are multi-dimensional tensors with a respective depth dimension; and

the respective depth dimension of the modified input tensor is greater than the respective depth dimension of the input tensor.

16. The non-transitory machine-readable storage device of claim 15 , wherein the operations comprise:

grouping multiple inputs of the input tensor at least by trading spatial extent or indexing with regard to X and Y dimensions of the input tensor for depth extent or indexing with regard to a Z dimension of the input tensor.

17. The non-transitory machine-readable storage device of claim 16 , wherein the operations comprise:

based on the grouping of multiple inputs, transforming the layer of the convolutional neural network having the stride greater than one into the superpixel layer having the stride equal to one;

wherein the superpixel layer comprises a reduced number of kernel elements relative to an untransformed layer of the convolutional neural network.

18. The non-transitory machine-readable storage device of claim 17 , wherein transforming the layer of the convolutional neural network into the superpixel layer comprises:

applying a superpixel transformation to convolutional neural network layer inputs of the input tensor and an unmodified convolutional neural network layer weight matrix to generate a differently shaped, but mathematically equivalent, superpixel convolutional neural network layer.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME FROM GOOGLE LLC TO GOOGLE INC. PREVIOUSLY RECORDED AT REEL: 053946 FRAME: 0780. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 15, 2020
From: YOUNG, REGINALD CLIFFORD; ROSS, JONATHAN
To: GOOGLE INC.
Reel/Frame 054085/0298 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2020
From: YOUNG, REGINALD CLIFFORD; ROSS, JONATHAN
To: GOOGLE INC.
Reel/Frame 053946/0780 →
ENTITY CONVERSION Recorded Oct 1, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 053963/0926 →
Continuity (3)
Continuation 16717341 · Dec 17, 2019
Continuation 15209658 · Jul 13, 2016
Related Publication 20210125029A1 · Apr 29, 2021
References Cited (80)
US 9940573B2 · Young · 2018 [cited by examiner]
US 10223788B2 · Bozorgtabar · 2019 [cited by examiner]
US 20130030699A1 · Barnes · 2013 [cited by examiner]
US 20140067735A1 · Yu · 2014 [cited by applicant]
US 20140180989A1 · Krizhevsky et al. · 2014 [cited by applicant]
US 20140207368A1 · Barnes · 2014 [cited by examiner]
US 20150065803A1 · Douglas · 2015 [cited by examiner]
US 20150265251A1 · Cho · 2015 [cited by examiner]
US 20160035078A1 · Lin · 2016 [cited by applicant]
US 20160171707A1 · Schwartz · 2016 [cited by examiner]
US 20160217369A1 · Annapureddy et al. · 2016 [cited by applicant]
US 20170132496A1 · Shoaib · 2017 [cited by applicant]
US 20180018554A1 · Young · 2018 [cited by examiner]
US 20180018556A1 · Young · 2018 [cited by examiner]
US 20180285683A1 · Chen · 2018 [cited by examiner]
US 20180336462A1 · Brothers · 2018 [cited by applicant]
US 20190012768A1 · Tafazoli Bilandi · 2019 [cited by examiner]
US 20190057532A1 · Marzban · 2019 [cited by examiner]
US 20200125922A1 · Young et al. · 2020 [cited by applicant]
CN 104915322 · 2015 [cited by applicant]
GB 2553900A · 2018 [cited by examiner]
JP 2017068608A · 2017 [cited by applicant]
JP 2017091479A · 2017 [cited by applicant]
JP 2017520864A · 2017 [cited by applicant]
JP 2018506785A · 2018 [cited by applicant]
KR 20150125921A · 2015 [cited by applicant]
KR 102344473B1 · 2021 [cited by applicant]
WO WO2014058651 · 2014 [cited by applicant]
WO WO2014105865 · 2014 [cited by applicant]
WO WO2014205231 · 2014 [cited by applicant]
WO WO2015157526A1 · 2015 [cited by applicant]
WO WO2018119807A1 · 2018 [cited by examiner]
Liu et al, “Learning Depth from Single Monocular Images Using Deep Convolutional Neural Feilds”, IEEE, 2015, pp. 1015 (Year: 2015). [cited by examiner]
Alberto Garcia, “3D Object Recognition with Convolutional Neural Networks”, University of Alicante, Jun. 2016, pp. 1-110. (Year: 2016). [cited by examiner]
DE Office Action in German Appln. No. 10 2017 115 519.8, dated Nov. 3, 2022, 21 pages (with English Translation). [cited by applicant]
Ioannou et al., “Training CNNs with Low-Rank Filters for Efficient Image Classification,” arXiv, Feb. 7, 2016, 17 pages. [cited by applicant]
JP Office Action in Japanese Appln. No. 2021-156865, mailed Nov. 1, 2022, 4 pages (with English Translation). [cited by applicant]
AU Office Action in Australian Application No. 2017295714, dated Sep. 27, 2019, 5 pages. [cited by applicant]
AU Office Action in Australian Application No. 2017295714, dated May 20, 2020, 3 pages. [cited by applicant]
Britz et al, “Understanding Convolutional Neural Networks for NLP” WildML, Nov. 2015, 17 pages. [cited by applicant]
Canadian Office Action in Canadian Application No. 3,030,428, dated Dec. 6, 2019, 7 pages. [cited by applicant]
CN Office Action in Chinese Application No. 201710570292, dated Jan. 13, 2020, 15 pages (with English translation). [cited by applicant]
CN Office Action in Chinese Application No. 201710570292.6, dated Aug. 3, 2020, 7 page (with English translation). [cited by applicant]
Cohen et al., “Convolutional Rectifier Networks as Generalized Tensor Decompositions” arXiv, May 2016, 22 pages. [cited by applicant]
Collobert et al., “Implementing Neural Networks Efficiently,” The Handbook of brain theory and neural networks 3361.10 (1995), 20 pages. [cited by applicant]
Cong et al., “Minimizing Computation in Convolutional Neural Networks” ICANN 2014, LNCS 8681, pp. 281-290, 2014. [cited by applicant]
Gribov, “Efficient Kernel Convolution for Smooth Surfaces without Edge Effects,” arXiv:1601.04065 [stat.CO], Jan. 13, 2016, 16 pages. [cited by applicant]
He et al., “SuperCNN: A Superpixelwise Convolutional Neural Network for Salient Object Detections”, Springer, Apr. 8, 2015, pp. 1-15. [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US2017/041930, mailed on Oct. 25, 2017, 15 pages. [cited by applicant]
JP Office Action in Japanese Application No. 2019-501488, dated May 22, 2020, 14 pages (with English translation). [cited by applicant]
Karpathy et al., “Large-scale Video Classification with Convolutional Neural Networks,” Proceedings of the International Computer Vision and Pattern Recognition (CVPR 2014). [cited by applicant]
Lavin et al., “Fast Algorithms for Convolutional Neural Networks,” arXiv:1509.09308 [cs.NE], Nov. 10, 2015, pp. 1-9. [cited by applicant]
Lavin, “maxDNN: An Efficient Convolution Kernel for Deep Learning with Maxwell GPUs,” arXiv:1501.06633 [cs.NE], Jan. 30, 2015, pp. 1-7. [cited by applicant]
Liu et al. “Sparse Convolutional Neural Networks,” 2015 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 7, 2015, 9 pages. [cited by applicant]
Liu et al., “Learning Depth from Single Monocular Images Using Deep Convolutional Neural Fields”, IEEE Transactions on Pattern Analysis and Machine Learning Intelligence, pp. 1-14. [cited by applicant]
Liu, “Sparse Convolutional Neural Networks” IEEE, Dec. 2015, 9 pages. [cited by applicant]
Novikov et al. “Tensorizing Neural Networks,” arXiv 1509.06569, Sep. 22, 2015, 9 pages. [cited by applicant]
O'Shea et al..,, An Introduction to Convolutional Neural Networks, Dec. 2015, pp. 1-11. [cited by applicant]
Ovtcharov et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware,” Microsoft Research, Feb. 22, 2015, pp. 1-4. [cited by applicant]
SG Written Opinion in Singaporean Application No. 11201900240W, dated May 6, 2020, 4 pages. [cited by applicant]
Shengfeng et al. “SuperCNN: A Superpixelwise Convolutional Neural Network for Salient Object Detection,” International Journal of Computer Vision, Kluwer Academic Publishers, vol. 115(3), Apr. 8, 2015, 15 pages. [cited by applicant]
Tschopp, “Efficient Convolutional Neural Networks for Pixelwise Classification on Heterogeneous Hardware Systems,” arXiv:1509.03371v1 [cs.CV], Sep. 11, 2015, 92 pages. [cited by applicant]
SG Written Opinion in Singaporean Application No. 11201900240E, dated Sep. 22, 2021, 5 pages. [cited by applicant]
EP Office Action in European Appln. No. 17749244.4, dated Oct. 25, 2022, 5 pages. [cited by applicant]
Abdelfattah et al., “High-performance tensor contractions for GPUs,” Procedia Computer Science, 2016, 80:108-118. [cited by applicant]
Office Action in Australian Appln. No. 2022200600, dated Jan. 25, 2023, 4 pages. [cited by applicant]
JP Office Action in Japanese Application No. 2019-501488, dated Sep. 29, 2020, 7 pages (with English translation). [cited by applicant]
IN Office Action in Indian Application No. 201947000978, dated Nov. 24, 2021, 6 pages (with English translation). [cited by applicant]
Notice of Allowance in Australian Appln. No. 2022200600, mailed on Nov. 1, 2023, 3 pages. [cited by applicant]
AU Office Action in Australian Application No. 2020220126, dated Feb. 23, 2021, 4 pages. [cited by applicant]
EP Office Action in European Application No. 17749244.4, dated Feb. 8, 2021, 7 pages. [cited by applicant]
KR Office Action in Korean Application No. 10-2019-7004190, dated Feb. 16, 2021, 9 pages (with English translation). [cited by applicant]
Extended European Search Report in European Appln. No. 23219583.4, mailed on Jun. 21, 2024, 10 pages. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2017/041930, mailed on Jan. 24, 2019, 8 pages. [cited by applicant]
Long et al., “Fully convolutional networks for semantic segmentation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 1, 2015, 3431-3440. [cited by applicant]
Notice of Allowance in Korean Appln. No. 10-2021-7042308, dated Jan. 25, 2024, 4 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 202110225055.2, mailed on Jul. 23, 2024, 12 pages (with English translation). [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2023-036865, mailed on Sep. 24, 2024, 7 pages (with machine translation). [cited by applicant]
Office Action in Japanese Appln. No. 2024-186619, mailed on Aug. 12, 2025, 4 pages (with machine translation). [cited by applicant]
Tanomoto et al., “Convolutional Neural Network Processing Using a Memory Network-Based Accelerator,” IEICE Technical Report, Nov. 2014, 114(330):57-62. [cited by applicant]