IP Library › Granted Patent US 12,230,008
Granted Patent B2
US 12,230,008 · App. 17/894,685 · Granted Feb 18, 2025

Image processing apparatus and operating method thereof

Inventors: Jaeyeon Park (Suwon-si, KR); Iijun Ahn (Suwon-si, KR); Soomin Kang (Suwon-si, KR); Hanul Shin (Suwon-si, KR); Tammy Lee (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06V10/761G06T3/40G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,230,008
App. No.
17/894,685
Granted
Feb 18, 2025
Kind
B2
Abstract

An image processing apparatus, including a processor configured to execute instructions stored in a memory to: obtain characteristic information of a first image, divide the characteristic information into a plurality of groups, input each group into a respective layer of a plurality of layers included in a convolutional neural network and perform a convolution operation using one or more kernels to obtain a plurality of pieces of output information, generate an attention map including weight information corresponding to each pixel included in the first image, based on the plurality of pieces of output information, generate a spatially variant kernel including a kernel corresponding to the each pixel, based on the attention map and a spatial kernel including weight information according to a position relationship between the each pixel and a neighboring pixel, and generate a second image by applying the spatially variant kernel to the first image.

Claims (61)

1. An image processing apparatus comprising:

a memory configured to store one or more instructions; and

a processor configured to execute the one or more instructions stored in the memory to:

obtain characteristic information of a first image,

divide the characteristic information into a plurality of groups,

input each group of the plurality of groups into a respective layer of a plurality of layers included in a convolutional neural network and perform a convolution operation using one or more kernels to obtain a plurality of pieces of output information corresponding to the plurality of groups from the plurality of layers,

generate an attention map including weight information corresponding to each pixel of pixels included in the first image, based on the plurality of pieces of output information,

generate a spatially variant kernel including a kernel corresponding to the each pixel, based on the attention map and a spatial kernel including weight information according to a position relationship between the each pixel and a neighboring pixel, and

generate a second image by applying the spatially variant kernel to the first image.

2. The image processing apparatus of claim 1 , wherein the characteristic information of the first image includes similarity information representing a similarity between the each pixel and the neighboring pixel.

3. The image processing apparatus of claim 2 , wherein the processor is further configured to execute the one or more instructions to:

obtain first similarity information based on a difference between the each pixel and a first neighboring pixel having a first relative position with respect to the each pixel, and

obtain second similarity information based on a difference between the each pixel and a second neighboring pixel having a second relative position with respect to the each pixel.

4. The image processing apparatus of claim 1 , wherein the processor is further configured to execute the one or more instructions to divide the characteristic information according to channels to obtain the plurality of groups based on the characteristic information including a plurality of pieces of channel information.

5. The image processing apparatus of claim 4 , wherein the processor is further configured to execute the one or more instructions to:

divide the plurality of pieces of channel information included in the characteristic information into a first group and a second group,

input pieces of first channel information included in the first group into a first layer of the convolutional neural network, and

input pieces of second channel information included in the second group into a second layer located after the first layer in the convolutional neural network.

6. The image processing apparatus of claim 5 , wherein the processor is further configured to execute the one or more instructions to input the pieces of the second channel information and pieces of information output from the first layer into the second layer.

7. The image processing apparatus of claim 5 , wherein the processor is further configured to execute the one or more instructions to:

obtain first output information corresponding to the first group from a third layer of the convolutional neural network, and

obtain second output information corresponding to the second group from a fourth layer of the convolutional neural network, the fourth layer being located after the third layer in the convolutional neural network.

8. The image processing apparatus of claim 1 , wherein the processor is further configured to execute the one or more instructions to downscale the characteristic information of the first image and divide the downscaled characteristic information into the plurality of groups.

9. The image processing apparatus of claim 1 , wherein the processor is further configured to execute the one or more instructions to:

obtain quality information for the each pixel,

obtain a plurality of output values corresponding to a plurality of pieces of preset quality information with respect to a group,

obtain weights for the each pixel with respect to the plurality of output values based on the quality information for the each pixel, and

apply the weights for the each pixel and sum the plurality of output values to obtain output information corresponding to the group.

10. The image processing apparatus of claim 1 , wherein within the spatial kernel, a pixel located at a center of the spatial kernel has a greatest value, and pixel values decrease away from the center.

11. The image processing apparatus of claim 1 , wherein a size of the spatial kernel is K×K and a number of channels of the attention map is K 2 , and

wherein the processor is further configured to execute the one or more instructions to:

arrange pixel values included in the spatial kernel in a channel direction to convert the spatial kernel into a weight vector having a size of 1×1×K 2 , and

generate the spatially variant kernel by performing a multiplication operation between the weight vector and each vector of a plurality of one-dimensional vectors having the size of 1×1×K 2 included in the attention map.

12. The image processing apparatus of claim 1 , wherein a number of kernels included in the spatially variant kernel is same as a number of pixels included in the first image.

13. The image processing apparatus of claim 12 , wherein the processor is further configured to execute the one or more instructions to generate the second image by:

performing filtering by applying a first kernel included in the spatially variant kernel to a first region centered on a first pixel included in the first image, and

performing filtering by applying a second kernel included in the spatially variant kernel to a second region centered on a second pixel included in the first image.

14. An operating method of an image processing apparatus, the operating method comprising:

obtaining characteristic information of a first image;

dividing the characteristic information into a plurality of groups;

inputting each group of the plurality of groups into a respective layer of a plurality of layers included in a convolutional neural network and performing a convolution operation using one or more kernels to obtain a plurality of pieces of output information corresponding to the plurality of groups from the plurality of layers;

generating an attention map including weight information corresponding to each pixel of pixels included in the first image, based on the plurality of pieces of output information;

generating a spatially variant kernel including a kernel corresponding to the each pixel, based on the attention map and a spatial kernel including weight information according to a position relationship between the each pixel and a neighboring pixel; and

generating a second image by applying the spatially variant kernel to the first image.

15. The operating method of claim 14 , wherein the obtaining of the characteristic information of the first image comprises obtaining similarity information representing a similarity between the each pixel included in the first image and the neighboring pixel.

16. The operating method of claim 15 , wherein the obtaining of the similarity information comprises:

obtaining first similarity information based on a difference between the each pixel included in the first image and a first neighboring pixel having a first relative position with respect to the each pixel; and

obtaining second similarity information based on a difference between the each pixel included in the first image and a second neighboring pixel having a second relative position with respect to the each pixel.

17. The operating method of claim 14 , wherein the dividing of the characteristic information into the plurality of groups comprises dividing the characteristic information according to channels to obtain the plurality of groups based on the characteristic information including a plurality of pieces of channel information.

18. The operating method of claim 17 , wherein the dividing of the characteristic information into the plurality of groups further comprises dividing the plurality of pieces of channel information included in the characteristic information into a first group and a second group, and

wherein the obtaining of the plurality of pieces of output information comprises:

inputting pieces of first channel information included in the first group into a first layer of the convolutional neural network; and

inputting pieces of second channel information included in the second group into a second layer located after the first layer in the convolutional neural network.

19. The operating method of claim 18 , wherein the obtaining of the plurality of pieces of output information further comprises inputting the pieces of the second channel information and pieces of information output from the first layer into the second layer.

20. A non-transitory computer-readable recording medium configured to store instructions which, when executed by at least one processor, cause the at least one processor to:

obtain characteristic information of a first image;

divide the characteristic information into a plurality of groups;

input each group of the plurality of groups into a respective layer of a plurality of layers included in a convolutional neural network and performing a convolution operation using one or more kernels to obtain a plurality of pieces of output information corresponding to the plurality of groups from the plurality of layers;

generate an attention map including weight information corresponding to each pixel of pixels included in the first image, based on the plurality of pieces of output information;

generate a spatially variant kernel including a kernel corresponding to the each pixel, based on the attention map and a spatial kernel including weight information according to a position relationship between the each pixel and a neighboring pixel; and

generate a second image by applying the spatially variant kernel to the first image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2022
From: PARK, JAEYEON; AHN, ILJUN; KANG, SOOMIN; SHIN, HANUL; LEE, TAMMY
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 060890/0539 →
Priority Claims (1)
KR 10-2021-0160711 · Nov 19, 2021 · national
Continuity (2)
Continuation PCTKR2022010264 · Jul 14, 2022
Related Publication 20230169752A1 · Jun 1, 2023
References Cited (26)
US 8953896B2 · Lee et al. · 2015 [cited by applicant]
US 9330442B2 · Kang · 2016 [cited by applicant]
US 9741107B2 · Xu et al. · 2017 [cited by applicant]
US 10671886B2 · Price et al. · 2020 [cited by applicant]
US 20190188586A1 · Rajabizadeh et al. · 2019 [cited by applicant]
US 20200342328A1 · Revaud et al. · 2020 [cited by applicant]
US 20200364486A1 · Park · 2020 [cited by examiner]
US 20210104021A1 · Sohn et al. · 2021 [cited by applicant]
US 20220019844A1 · Park et al. · 2022 [cited by applicant]
US 20220284555A1 · Ahn et al. · 2022 [cited by applicant]
CN 109948699A · 2019 [cited by applicant]
KR 101797673B1 · 2017 [cited by applicant]
KR 101967089B1 · 2019 [cited by applicant]
KR 1020200067631A · 2020 [cited by applicant]
KR 102144994A · 2020 [cited by applicant]
KR 1020200132304A · 2020 [cited by applicant]
KR 1020200125468A · 2020 [cited by applicant]
KR 1020220125124A · 2022 [cited by applicant]
WO 2021097728A1 · 2021 [cited by applicant]
Howard, et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications”, 2017, Google Inc., arXiv:1704.04861v1 [cs.CV], 9 pages total. [cited by applicant]
International Search Report (PCT/ISA/210) and Written Opinion (PCT/ISA/237) dated Oct. 19, 2022 issued by the International Searching Authority in International Application No. PCT/KR2022/010264. [cited by applicant]
Zhang, Y., et al., “Dual attention per-pixel filter network for spatially varying image deblurring” Digital Signal Processing, vol. 113, 2021, pp. 1-17 (18 pages). [cited by applicant]
B. Zhang, S. Jin, Y. Xia, Y. Huang and Z. Xiong, “Attention Mechanism Enhanced Kernel Prediction Networks for Denoising of Burst Images,” ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Pr… [cited by applicant]
Ronneberger, O., Fischer, P., Brox, T., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, In: Navab, N., Hornegger, J., Wells, W., Frangi, A. (eds) Medical Image Computing and Computer-Assisted Interven… [cited by applicant]
S. Niklaus, L. Mai and F. Liu, “Video Frame Interpolation via Adaptive Separable Convolution,” 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 2017, pp. 261-270, doi: 10.1109/ICCV.2017.37. [cited by applicant]
European Extended Search Report issued Oct. 31, 2024 by the European Patent Office for EP Patent Application No. 22895779.1. [cited by applicant]
Cited By (1)
US 12,731,369