Image processing apparatus using convolutional neural networks and operating method thereof
An image processing apparatus for processing an image by using one or more convolutional neural networks includes a memory storing one or more instructions, and at least one processor configured to execute the one or more instructions stored in the memory to obtain first feature data by performing a convolution operation between input data obtained from a first image and a first kernel, divide a plurality of channels included in the first feature data into first groups, obtain second feature data by performing a convolution operation between the first feature data respectively corresponding to the first groups and second kernels respectively corresponding to the first groups, obtain shuffling data by shuffling the second feature data, obtain output data by performing a convolution operation between data obtained by summing channels included in the shuffling data and a third kernel, and generate a second image based on the output data.
1 . An image processing apparatus for processing an image by using one or more convolutional neural networks, the image processing apparatus comprising:
a memory storing one or more instructions; and
at least one processor configured to execute the one or more instructions stored in the memory to:
divide a plurality of channels included in input information obtained from a first image into first groups;
obtain output data respectively corresponding to the first groups, based on input data respectively corresponding to the first groups by:
obtaining first feature data based on a first convolution operation being performed between input data and a first kernel;
dividing a plurality of channels included in the first feature data into second groups;
obtaining second feature data based on a second convolution operation being performed between the first feature data respectively corresponding to the second groups and second kernels respectively corresponding to the second groups;
obtaining shuffling data by shuffling the second feature data; and
obtaining the output data by performing a convolution operation between data obtained by concatenating channels included in the shuffling data and a third kernel,
obtain output information corresponding to the input information, by summing channels included in the output data respectively corresponding to the first groups;
obtain third feature data based on a third convolution operation being performed between the output information and a fourth kernel;
divide a plurality of channels included in the third feature data into the first groups;
obtain fourth feature data based on a fourth convolution operation being performed between the third feature data respectively corresponding to the first groups and fifth kernels respectively corresponding to the first groups;
divide a plurality of channels included in the fourth feature data into the second groups and obtain second shuffling data by shuffling the fourth feature data respectively corresponding to the second groups;
obtain fifth feature data based on a fifth convolution operation being performed between the second shuffling data and sixth kernels respectively corresponding to the second groups;
obtain sixth feature data respectively corresponding to the first groups by summing channels included in the fifth feature data;
generate an attention map including weight information corresponding to each of a plurality of pixels included in the first image, based on the sixth feature data;
generate a spatially variable kernel corresponding to each of the plurality of pixels, based on the attention map and a spatial kernel including weight information according to a position relationship between each of the plurality of pixels and at least one neighboring pixel of each of the plurality of pixels; and
generate a second image by applying the spatially variable kernel to the first image.
2 . The image processing apparatus of claim 1 , wherein the at least one processor is further configured to execute the one or more instructions to:
determine a number of channels included in each of the second kernels based on the number of channels of the first feature data respectively corresponding to the first groups.
3 . The image processing apparatus of claim 1 , wherein, in the spatial kernel, a pixel located in a center of the spatial kernel has a greatest value, and a pixel value decreases away from the center.
4 . The image processing apparatus of claim 1 , wherein
a size of the spatial kernel is K×K, and a number of channels of the attention map is K 2 ,
the at least one processor is further configured to execute the one or more instructions stored in the memory to:
convert pixel values included in the spatial kernel into a weight vector with a size of 1×1×K 2 by arranging the pixel values in a channel direction, and
generate the spatially variable kernel based on a multiplication operation being performed between each of one-dimensional vectors with the size of 1×1×K 2 included in the attention map and the weight vector, and
wherein K denotes a natural number.
5 . The image processing apparatus of claim 1 , wherein the spatially variable kernel includes a same number of kernels as a number of pixels included in the first image.
6 . An operating method of an image processing apparatus for processing an image by using one or more convolutional neural networks, the operating method comprising:
dividing a plurality of channels included in input information obtained from a first image into first groups;
obtaining output data respectively corresponding to the first groups, based on input data respectively corresponding to the first groups by:
obtaining first feature data based on a first convolution operation being performed between input data and a first kernel;
dividing a plurality of channels included in the first feature data into second groups;
obtaining second feature data based on a second convolution operation being performed between the first feature data respectively corresponding to the second groups and second kernels respectively corresponding to the second groups;
obtaining shuffling data by shuffling the second feature data; and
obtaining the output data by performing a convolution operation between data obtained by concatenating channels included in the shuffling data and a third kernel;
obtaining output information corresponding to the input information, by summing channels included in the output data respectively corresponding to the first groups;
obtaining third feature data based on a third convolution operation being performed between the output information and a fourth kernel;
dividing a plurality of channels included in the third feature data into the first groups;
obtaining fourth feature data based on a fourth convolution operation being performed between the third feature data respectively corresponding to the first groups and fifth kernels respectively corresponding to the first groups;
dividing a plurality of channels included in the fourth feature data into the second groups and obtain second shuffling data by shuffling the fourth feature data respectively corresponding to the second groups;
obtaining fifth feature data based on a fifth convolution operation being performed between the second shuffling data and sixth kernels respectively corresponding to the second groups;
obtaining sixth feature data respectively corresponding to the first groups by summing channels included in the fifth feature data;
generating an attention map including weight information corresponding to each of a plurality of pixels included in the first image, based on the sixth feature data;
generating a spatially variable kernel corresponding to each of the plurality of pixels, based on the attention map and a spatial kernel including weight information according to a position relationship between each of the plurality of pixels and at least one neighboring pixel of each of the plurality of pixels; and
generating a second image by applying the spatially variable kernel to the first image.
7 . The operating method of claim 6 , wherein a number of channels included in each of the second kernels is determined based on the number of channels of the first feature data respectively corresponding to the first groups.
8 . The operating method of claim 6 , wherein, in the spatial kernel, a pixel located in a center of the spatial kernel has a greatest value, and a pixel value decreases away from the center.
9 . The operating method of claim 6 , wherein
a size of the spatial kernel is K×K, and a number of channels of the attention map is K 2 ,
the generating of the spatially variable kernel comprises:
converting pixel values included in the spatial kernel into a weight vector with a size of 1×1×K 2 by arranging the pixel values in a channel direction, and
generating the spatially variable kernel based on a multiplication operation being performed between each of one-dimensional vectors with the size of 1×1×K 2 included in the attention map and the weight vector, and
wherein K denotes a natural number.
10 . The operating method of claim 6 , wherein the spatially variable kernel includes a same number of kernels as a number of pixels included in the first image.
11 . A non-transitory computer-readable recording medium having recorded thereon a program for performing an image processing method, the image processing method comprising:
dividing a plurality of channels included in input information obtained from a first image into first groups;
obtaining output data respectively corresponding to the first groups, based on input data respectively corresponding to the first groups by:
obtaining first feature data based on a first convolution operation being performed between input data and a first kernel;
dividing a plurality of channels included in the first feature data into second groups;
obtaining second feature data based on a second convolution operation being performed between the first feature data respectively corresponding to the second groups and second kernels respectively corresponding to the second groups;
obtaining shuffling data by shuffling the second feature data; and
obtaining the output data by performing a convolution operation between data obtained by concatenating channels included in the shuffling data and a third kernel;
obtaining output information corresponding to the input information, by summing channels included in the output data respectively corresponding to the first groups;
obtaining third feature data based on a third convolution operation being performed between the output information and a fourth kernel;
dividing a plurality of channels included in the third feature data into the first groups;
obtaining fourth feature data based on a fourth convolution operation being performed between the third feature data respectively corresponding to the first groups and fifth kernels respectively corresponding to the first groups;
dividing a plurality of channels included in the fourth feature data into the second groups and obtain second shuffling data by shuffling the fourth feature data respectively corresponding to the second groups;
obtaining fifth feature data based on a fifth convolution operation being performed between the second shuffling data and sixth kernels respectively corresponding to the second groups;
obtaining sixth feature data respectively corresponding to the first groups by summing channels included in the fifth feature data;
generating an attention map including weight information corresponding to each of a plurality of pixels included in the first image, based on the sixth feature data;
generating a spatially variable kernel corresponding to each of the plurality of pixels, based on the attention map and a spatial kernel including weight information according to a position relationship between each of the plurality of pixels and at least one neighboring pixel of each of the plurality of pixels; and
generating a second image by applying the spatially variable kernel to the first image.