IP Library Granted Patent US 12,361,571
Granted Patent B2
US 12,361,571 · App. 18/663,438 · Granted Jul 15, 2025

Method and apparatus with convolution neural network processing using shared operand

Inventor: Sehwan Lee (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06T7/33G06N3/04G06N3/063G06N3/08G06T1/20G06T2207/20081G06T2207/20084G06T2210/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,571
App. No.
18/663,438
Granted
Jul 15, 2025
Kind
B2
Abstract

A neural network apparatus including one or more processors including a controller configured to determine a shared operand to be shared in parallelized operations as being either one of a pixel value among pixel values of an input feature map and a weight value among weight values of a kernel, based on either one or both of a feature of the input feature map and a feature of the kernel, and one or more processing units configured to perform the parallelized operations based on the determined shared operand.

Claims (34)

1. A neural network apparatus, comprising:

one or more processors configured to execute instructions; and

a memory storing the instructions, wherein execution of the instructions configures the processors to:

obtain an input feature map and a kernel;

determine a shared operand to be shared in parallelized convolution operations as being either one of a pixel value among pixel values of an input feature map and a weight value among weight values of a kernel based on a shape of the input feature map and dimensions of the kernel; and

perform the parallelized convolution operations using the shared operand, wherein the shared operand is determined as:

being the pixel value of the input feature map in response to a two-dimensional area size of the input feature map being less than or equal to a first threshold value; and

being the weight value of the kernel in response to a two-dimensional area size of the input feature map greater than the first threshold value,

wherein the first threshold value is a reference value used to determine whether the shared operand is the pixel value of the input feature map or the weight value of the kernel.

2. The apparatus of claim 1 , wherein the shared operand, the pixel value of the input feature map, and the weight value of the kernel are of a first layer of a neural network, and

wherein the one or more processors are further configured to:

determine, for a second layer of the neural network, a second shared operand of the second layer to be either one of a pixel value of among pixel values of an input feature map of the second layer and a weight value among weight values of a kernel of the second layer, based on a second shape of the input feature map of the second layer and second dimensions of the kernel of the second layer.

3. The apparatus of claim 1 , wherein the one or more processors are further configured to determine the shared operand further based on either one or both of a percentage of pixels having a zero value within the input feature map and a percentage of weights having a zero value within the kernel.

4. The apparatus of claim 3 , wherein the one or more processors are further configured to determine the shared operand to be the pixel value of the input feature map in response to the percentage of the pixels having a zero value is greater than a third threshold value, and the shared operand to be the weight value of the kernel in response to the percentage of the weights having a zero value is greater than a fourth threshold value.

5. The apparatus of claim 1 , wherein the one or more processors are configured to:

determine the shared operand to be the weight value of the kernel in response to determining that a number of channels of the input feature map is smaller than a second threshold value and the two-dimensional area size of the input feature map is greater than the first threshold value; and

determine the shared operand to be the pixel value of the input feature map in response to determining that the number of the channels of the input feature map is greater than or equal to the second threshold value and the two-dimensional area size of the input feature map is less than or equal to the first threshold value.

6. A processor-implemented neural network method, the method comprising:

obtaining an input feature map and a kernel;

determining a shared operand to be shared in parallelized convolution operations as being either one of a pixel value among pixel values of an input feature map and a weight value among weight values of a kernel based on a shape of the input feature map and dimensions of the kernel; and

performing the parallelized convolution operations using the shared operand,

wherein the shared operand is determined as:

being the pixel value of the input feature map in response to a two-dimensional area size of the input feature map being less than or equal to a first threshold value; and

being the weight value of the kernel in response to a two-dimensional area size of the input feature map greater than the first threshold value,

wherein the first threshold value is a reference value used to determine whether the shared operand is the pixel value of the input feature map or the weight value of the kernel.

7. The method of claim 6 , wherein the shared operand, the pixel value of the input feature map, and the weight value of the kernel are of a first layer of a neural network, and

wherein the method further comprises:

determining, for a second layer of the neural network, a second shared operand of the second layer to be either one of a pixel value of among pixel values of an input feature map of the second layer and a weight value among weight values of a kernel of the second layer, based on a second shape of the input feature map of the second layer and second dimensions of the kernel of the second layer.

8. The method of claim 6 , wherein the determining the shared operand comprises determining the shared operand further based on either one or both of a percentage of pixels having a zero value within the input feature map and a percentage of weights having a zero value within the kernel.

9. The method of claim 8 , wherein the determining the shared operand comprises determining the shared operand to be the pixel value of the input feature map in response to the percentage of the pixels having a zero value is greater than a third threshold value, and the shared operand to be the weight value of the kernel in response to the percentage of the weights having a zero value is greater than a fourth threshold value.

10. The method of claim 6 , wherein the determining the shared operand comprises:

determining the shared operand to be the weight value of the kernel in response to determining that a number of channels of the input feature map is smaller than a second threshold value and the two-dimensional area size of the input feature map is greater than the first threshold value; and

determining the shared operand to be the pixel value of the input feature map in response to determining that the number of the channels of the input feature map is greater than or equal to the second threshold value and the two-dimensional area size of the input feature map is less than or equal to the first threshold value.

11. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 6 .

Priority Claims (1)
KR 10-2019-0038606 · Apr 2, 2019 · national
Continuity (3)
Continuation 16564215 · Sep 9, 2019
Provisional Application 62799190 · Jan 31, 2019
Related Publication 20240303837A1 · Sep 12, 2024
References Cited (38)
US 10572225B1 · Ghasemi et al. · 2020 [cited by applicant]
US 10936891B2 · Kim et al. · 2021 [cited by applicant]
US 11200092B2 · Zhang et al. · 2021 [cited by applicant]
US 11507817B2 · Abdelaziz et al. · 2022 [cited by applicant]
US 12014505B2 · Lee · 2024 [cited by examiner]
US 20160358069A1 · Brothers et al. · 2016 [cited by applicant]
US 20170024632A1 · Johnson et al. · 2017 [cited by applicant]
US 20180089562A1 · Jin et al. · 2018 [cited by applicant]
US 20180129893A1 · Son et al. · 2018 [cited by applicant]
US 20180181829A1 · Morozov et al. · 2018 [cited by applicant]
US 20180181865A1 · Adachi · 2018 [cited by applicant]
US 20180189643A1 · Kim et al. · 2018 [cited by applicant]
US 20180253636A1 · Lee et al. · 2018 [cited by applicant]
US 20180267898A1 · Henry et al. · 2018 [cited by applicant]
US 20190042923A1 · Janedula et al. · 2019 [cited by applicant]
US 20190180499A1 · Caulfield et al. · 2019 [cited by applicant]
US 20190311252A1 · Chen et al. · 2019 [cited by applicant]
US 20190340510A1 · Li et al. · 2019 [cited by applicant]
US 20200151019A1 · Yu et al. · 2020 [cited by applicant]
US 20200210175A1 · Alexander et al. · 2020 [cited by applicant]
US 20210004701A1 · Shibata · 2021 [cited by applicant]
US 20210011970A1 · Han · 2021 [cited by applicant]
US 20210192248A1 · Kim et al. · 2021 [cited by applicant]
US 20210365791A1 · Mishali · 2021 [cited by applicant]
US 20220269950A1 · Lee · 2022 [cited by applicant]
CN 108090565A · 2018 [cited by applicant]
CN 108205519A · 2018 [cited by applicant]
CN 108268931A · 2018 [cited by applicant]
KR 1020180101978A · 2018 [cited by applicant]
WO WO2018107476A1 · 2018 [cited by applicant]
Korean Office Action issued on May 17, 2024, in counterpart Korean Patent Application No. 10-2019-0038606 (5 pages in English, 7 pages in Korean). [cited by applicant]
Kim, D. et al., “ZeNA: Zero-Aware Neural Network Accelerator”, [cited by applicant]
Sombatsiri, Salita, et al. “Parallelism-flexible Convolution Core for Sparse Convolutional Neural Networks on FPGA.” [cited by applicant]
Extended European Search Report issued on Jun. 25, 2020 in counterpart European Patent Application No. 20152834.6 (10 pages in English). [cited by applicant]
Japanese Office Action issued on Feb. 27, 2024, in counterpart Japanese Patent Application No. 2020-015497 (6 pages in English, 6 pages in Japanese). [cited by applicant]
Chinese Office Action issued on Dec. 3, 2024 in corresponding Chinese Patent Application No. 201911321489.1. (7 pages in English and 8 pages in Chinese). [cited by applicant]
Wang Can, “Binary Acceleration Based on Hashing”, DOI: 10.16589/j.cnki.cn11-3571/tn.2018.20.017 (1 Page in English, 3 Pages in Chinese). [cited by applicant]
Chinese Office Action Issued on Apr. 16, 2025, in Counterpart Chinese Patent Application No. 201911321489.1 (5 Pages in English, 5 Pages in Chinese). [cited by applicant]