IP Library › Granted Patent US 12,731,369
Granted Patent B2
US 12,731,369 · App. 18/072,050 · Granted Sep 8, 2026

Image processing apparatus using convolutional neural networks and operating method thereof

Inventors: Iljun Ahn (Suwon-si, KR); Soomin Kang (Suwon-si, KR); Jaeyeon Park (Suwon-si, KR); Youngchan Song (Suwon-si, KR); Hanul Shin (Suwon-si, KR); Tammy Lee (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06V10/454G06V10/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,369
App. No.
18/072,050
Granted
Sep 8, 2026
Kind
B2
Abstract

An image processing apparatus for processing an image by using one or more convolutional neural networks includes a memory storing one or more instructions, and at least one processor configured to execute the one or more instructions stored in the memory to obtain first feature data by performing a convolution operation between input data obtained from a first image and a first kernel, divide a plurality of channels included in the first feature data into first groups, obtain second feature data by performing a convolution operation between the first feature data respectively corresponding to the first groups and second kernels respectively corresponding to the first groups, obtain shuffling data by shuffling the second feature data, obtain output data by performing a convolution operation between data obtained by summing channels included in the shuffling data and a third kernel, and generate a second image based on the output data.

Claims (75)

1 . An image processing apparatus for processing an image by using one or more convolutional neural networks, the image processing apparatus comprising:

a memory storing one or more instructions; and

at least one processor configured to execute the one or more instructions stored in the memory to:

divide a plurality of channels included in input information obtained from a first image into first groups;

obtain output data respectively corresponding to the first groups, based on input data respectively corresponding to the first groups by:

obtaining first feature data based on a first convolution operation being performed between input data and a first kernel;

dividing a plurality of channels included in the first feature data into second groups;

obtaining second feature data based on a second convolution operation being performed between the first feature data respectively corresponding to the second groups and second kernels respectively corresponding to the second groups;

obtaining shuffling data by shuffling the second feature data; and

obtaining the output data by performing a convolution operation between data obtained by concatenating channels included in the shuffling data and a third kernel,

obtain output information corresponding to the input information, by summing channels included in the output data respectively corresponding to the first groups;

obtain third feature data based on a third convolution operation being performed between the output information and a fourth kernel;

divide a plurality of channels included in the third feature data into the first groups;

obtain fourth feature data based on a fourth convolution operation being performed between the third feature data respectively corresponding to the first groups and fifth kernels respectively corresponding to the first groups;

divide a plurality of channels included in the fourth feature data into the second groups and obtain second shuffling data by shuffling the fourth feature data respectively corresponding to the second groups;

obtain fifth feature data based on a fifth convolution operation being performed between the second shuffling data and sixth kernels respectively corresponding to the second groups;

obtain sixth feature data respectively corresponding to the first groups by summing channels included in the fifth feature data;

generate an attention map including weight information corresponding to each of a plurality of pixels included in the first image, based on the sixth feature data;

generate a spatially variable kernel corresponding to each of the plurality of pixels, based on the attention map and a spatial kernel including weight information according to a position relationship between each of the plurality of pixels and at least one neighboring pixel of each of the plurality of pixels; and

generate a second image by applying the spatially variable kernel to the first image.

2 . The image processing apparatus of claim 1 , wherein the at least one processor is further configured to execute the one or more instructions to:

determine a number of channels included in each of the second kernels based on the number of channels of the first feature data respectively corresponding to the first groups.

3 . The image processing apparatus of claim 1 , wherein, in the spatial kernel, a pixel located in a center of the spatial kernel has a greatest value, and a pixel value decreases away from the center.

4 . The image processing apparatus of claim 1 , wherein

a size of the spatial kernel is K×K, and a number of channels of the attention map is K 2 ,

the at least one processor is further configured to execute the one or more instructions stored in the memory to:

convert pixel values included in the spatial kernel into a weight vector with a size of 1×1×K 2 by arranging the pixel values in a channel direction, and

generate the spatially variable kernel based on a multiplication operation being performed between each of one-dimensional vectors with the size of 1×1×K 2 included in the attention map and the weight vector, and

wherein K denotes a natural number.

5 . The image processing apparatus of claim 1 , wherein the spatially variable kernel includes a same number of kernels as a number of pixels included in the first image.

6 . An operating method of an image processing apparatus for processing an image by using one or more convolutional neural networks, the operating method comprising:

dividing a plurality of channels included in input information obtained from a first image into first groups;

obtaining output data respectively corresponding to the first groups, based on input data respectively corresponding to the first groups by:

obtaining first feature data based on a first convolution operation being performed between input data and a first kernel;

dividing a plurality of channels included in the first feature data into second groups;

obtaining second feature data based on a second convolution operation being performed between the first feature data respectively corresponding to the second groups and second kernels respectively corresponding to the second groups;

obtaining shuffling data by shuffling the second feature data; and

obtaining the output data by performing a convolution operation between data obtained by concatenating channels included in the shuffling data and a third kernel;

obtaining output information corresponding to the input information, by summing channels included in the output data respectively corresponding to the first groups;

obtaining third feature data based on a third convolution operation being performed between the output information and a fourth kernel;

dividing a plurality of channels included in the third feature data into the first groups;

obtaining fourth feature data based on a fourth convolution operation being performed between the third feature data respectively corresponding to the first groups and fifth kernels respectively corresponding to the first groups;

dividing a plurality of channels included in the fourth feature data into the second groups and obtain second shuffling data by shuffling the fourth feature data respectively corresponding to the second groups;

obtaining fifth feature data based on a fifth convolution operation being performed between the second shuffling data and sixth kernels respectively corresponding to the second groups;

obtaining sixth feature data respectively corresponding to the first groups by summing channels included in the fifth feature data;

generating an attention map including weight information corresponding to each of a plurality of pixels included in the first image, based on the sixth feature data;

generating a spatially variable kernel corresponding to each of the plurality of pixels, based on the attention map and a spatial kernel including weight information according to a position relationship between each of the plurality of pixels and at least one neighboring pixel of each of the plurality of pixels; and

generating a second image by applying the spatially variable kernel to the first image.

7 . The operating method of claim 6 , wherein a number of channels included in each of the second kernels is determined based on the number of channels of the first feature data respectively corresponding to the first groups.

8 . The operating method of claim 6 , wherein, in the spatial kernel, a pixel located in a center of the spatial kernel has a greatest value, and a pixel value decreases away from the center.

9 . The operating method of claim 6 , wherein

a size of the spatial kernel is K×K, and a number of channels of the attention map is K 2 ,

the generating of the spatially variable kernel comprises:

converting pixel values included in the spatial kernel into a weight vector with a size of 1×1×K 2 by arranging the pixel values in a channel direction, and

generating the spatially variable kernel based on a multiplication operation being performed between each of one-dimensional vectors with the size of 1×1×K 2 included in the attention map and the weight vector, and

wherein K denotes a natural number.

10 . The operating method of claim 6 , wherein the spatially variable kernel includes a same number of kernels as a number of pixels included in the first image.

11 . A non-transitory computer-readable recording medium having recorded thereon a program for performing an image processing method, the image processing method comprising:

dividing a plurality of channels included in input information obtained from a first image into first groups;

obtaining output data respectively corresponding to the first groups, based on input data respectively corresponding to the first groups by:

obtaining first feature data based on a first convolution operation being performed between input data and a first kernel;

dividing a plurality of channels included in the first feature data into second groups;

obtaining second feature data based on a second convolution operation being performed between the first feature data respectively corresponding to the second groups and second kernels respectively corresponding to the second groups;

obtaining shuffling data by shuffling the second feature data; and

obtaining the output data by performing a convolution operation between data obtained by concatenating channels included in the shuffling data and a third kernel;

obtaining output information corresponding to the input information, by summing channels included in the output data respectively corresponding to the first groups;

obtaining third feature data based on a third convolution operation being performed between the output information and a fourth kernel;

dividing a plurality of channels included in the third feature data into the first groups;

obtaining fourth feature data based on a fourth convolution operation being performed between the third feature data respectively corresponding to the first groups and fifth kernels respectively corresponding to the first groups;

dividing a plurality of channels included in the fourth feature data into the second groups and obtain second shuffling data by shuffling the fourth feature data respectively corresponding to the second groups;

obtaining fifth feature data based on a fifth convolution operation being performed between the second shuffling data and sixth kernels respectively corresponding to the second groups;

obtaining sixth feature data respectively corresponding to the first groups by summing channels included in the fifth feature data;

generating an attention map including weight information corresponding to each of a plurality of pixels included in the first image, based on the sixth feature data;

generating a spatially variable kernel corresponding to each of the plurality of pixels, based on the attention map and a spatial kernel including weight information according to a position relationship between each of the plurality of pixels and at least one neighboring pixel of each of the plurality of pixels; and

generating a second image by applying the spatially variable kernel to the first image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2022
From: AHN, ILJUN; KANG, SOOMIN; PARK, JAEYEON; SONG, YOUNGCHAN; SHIN, HANUL; LEE, TAMMY
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061924/0547 →
Priority Claims (2)
KR 10-2021-0169338 · Nov 30, 2021 · national
KR 10-2022-0095694 · Aug 1, 2022 · national
Continuity (2)
Continuation PCTKR2022018204 · Nov 17, 2022
Related Publication 20230169748A1 · Jun 1, 2023
References Cited (36)
US 9716830B2 · Baek · 2017 [cited by applicant]
US 11481574B2 · Yang et al. · 2022 [cited by applicant]
US 11714921B2 · Fang et al. · 2023 [cited by applicant]
US 12230008B2 · Park · 2025 [cited by examiner]
US 20190147319A1 · Kim et al. · 2019 [cited by applicant]
US 20200364486A1 · Park et al. · 2020 [cited by applicant]
US 20210019633A1 · Venkatesh · 2021 [cited by applicant]
US 20210093310A1 · Baldwin · 2021 [cited by applicant]
US 20210103793A1 · Huang · 2021 [cited by examiner]
US 20220284555A1 · Ahn et al. · 2022 [cited by applicant]
CN 105894013A · 2016 [cited by applicant]
CN 110309876A · 2019 [cited by applicant]
CN 112927174A · 2021 [cited by applicant]
CN 113052189A · 2021 [cited by applicant]
JP 2021103441A · 2021 [cited by applicant]
KR 1020190054770A · 2019 [cited by applicant]
KR 1020200132304A · 2020 [cited by applicant]
KR 1020210019537A · 2021 [cited by applicant]
KR 102305470B1 · 2021 [cited by applicant]
KR 1020210140757A · 2021 [cited by applicant]
Zhang, T., Qi, G., Xiao, B., & Wang, J. (2017). Interleaved Group Convolutions for Deep Neural Networks. ArXiv, abs/1707.02725. (Year: 2017). [cited by examiner]
R. Gomes, p. Rozario and N. Adhikari, “Deep Learning optimization in remote sensing image segmentation using dilated convolutions and ShuffleNet,” 2021 IEEE International Conference on Electro Information Technology (EI… [cited by examiner]
Q. -L. Zhang and Y. -B. Yang, “SA-Net: Shuffle Attention for Deep Convolutional Neural Networks,” ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, … [cited by examiner]
Su, H., Jampani, V., Sun, D., et al. (2019). Pixel-Adaptive Convolutional Neural Networks. ArXiv, abs/1904.05373. (Year: 2019). [cited by examiner]
Wu, J., Li, D., Yang, Y., Bajaj, C., Ji, X. (2019). Dynamic Filtering with Large Sampling Field for ConvNets. Arvix, abs/1803.07624. ( Year: 2019). [cited by examiner]
Esquivel, J., Vargas, A., Meyer, P., Tickoo, O. (2019). Adaptive Convolutional Kernels. 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW). doi: 10.1109/ICCVW.2019.00249. (Year: 2019). [cited by examiner]
Yanfang Zhang, Weihong Li, Zhenghao Li, Taigong Ning, “Dual attention per-pixel filter network for spatially varying image deblurring,” Digital Signal Processing, vol. 113, 2021, 103008, ISSN 1051-2004, https://doi.org/… [cited by examiner]
Roy, Swalpa & Manna, Suvojit & Song, Tiecheng & Bruzzone, Lorenzo. (2020). Attention-Based Adaptive Spectral-Spatial Kernel ResNet for Hyperspectral Image Classification. IEEE Transactions on Geoscience and Remote Sensi… [cited by examiner]
Extended European Search Report dated Dec. 18, 2024, issued by the European Patent Office in European Application No. 22901634.0. [cited by applicant]
Zhang et al., “Interleaved Group Convolutions for Deep Neural Networks”, XP080775495, 2017 (11 pages total). [cited by applicant]
Brabandere et al., “Dynamic Filter Networks”, XP093230006, 2016 (14 pages total). [cited by applicant]
Zhang et al., “Dynet: Dynamic Convolution for Accelerating Convolutional Neural Networks”, XP093006338, 2020, pp. 1-16 (16 pages total). [cited by applicant]
Zamora-Esquivel et al., “Adaptive Convolutional Kernels”, XP033732392, 2019, pp. 1998-2005 (8 pages total). [cited by applicant]
Zhang, Xiangyu et al., “ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices”, arXiv:1707.01083v1 [cs.CV], pp. 1-10, Jul. 4, 2017. [cited by applicant]
International Search Report and Written Opinion issued Feb. 27, 2023 by the International Searching Authority in counterpart International Patent Application No. PCT/KR2022/018204. (PCTISA/220, PCT/ISA/210 and PCT/ISA/2… [cited by applicant]
Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” arXiv: 1704.04861v1 [cs.CV], Apr. 17, 2017, Total 9 pages. [cited by applicant]