Image processing apparatus, method, and program for smoothing boundary of segmented region
A processor is configured to convert a size of a target image to derive a size-converted image, segment the size-converted image into regions of at least one class by using a segmentation model constructed by machine-learning a neural network to derive a plurality of class images in which a pixel value of each pixel represents class-likeness for the at least one class, convert a size of at least one class image into the size of the target image to derive at least one converted class image, and segment the target image based on a pixel value in each pixel of the at least one converted class image.
1 . An image processing apparatus comprising:
at least one processor, configured to:
convert a size of a target image to derive a size-converted image;
segment the size-converted image into regions of at least one class by using a segmentation model constructed by machine-learning a neural network to derive a plurality of class images in which a pixel value of each pixel represents class-likeness for the at least one class;
convert a size of at least one class image into the size of the target image to derive at least one converted class image; and
segment the target image based on a pixel value in each pixel of the at least one converted class image.
2 . The image processing apparatus according to claim 1 ,
wherein the size conversion is enlargement, reduction, or normalization in at least one direction in which pixels are arranged in the target image.
3 . The image processing apparatus according to claim 1 ,
wherein the pixel value of the class image is a score, which is derived by the neural network and represents a probability of being in the at least one class.
4 . The image processing apparatus according to claim 1 ,
wherein the at least one processor is configured to convert the size of the class image into the size of the target image through an interpolation calculation.
5 . The image processing apparatus according to claim 1 ,
wherein the at least one processor is configured to derive argmax of a pixel value of a corresponding pixel in the at least one converted class image to segment the target image.
6 . The image processing apparatus according to claim 1 ,
wherein the at least one processor is configured to sequentially perform derivation of the converted class image and segmentation of the target image for each class.
7 . An image processing method implemented by an image processing apparatus having at least one processor, the method comprising:
converting, by the at least one processor, a size of a target image to derive a size-converted image;
segmenting, by the at least one processor, the size-converted image into regions of at least one class by using a segmentation model constructed by machine-learning a neural network to derive a plurality of class images in which a pixel value of each pixel represents class-likeness for the at least one class;
converting, by the at least one processor, a size of at least one class image into the size of the target image to derive at least one converted class image; and
segmenting, by the at least one processor, the target image based on a pixel value in each pixel of the at least one converted class image.
8 . A non-transitory computer-readable storage medium that stores an image processing program causing a computer having at least one processor to execute:
a procedure of converting, by the at least one processor, a size of a target image to derive a size-converted image;
a procedure of segmenting, by the at least one processor, the size-converted image into regions of at least one class by using a segmentation model constructed by machine-learning a neural network to derive a plurality of class images in which a pixel value of each pixel represents class-likeness for the at least one class;
a procedure of converting, by the at least one processor, a size of at least one class image into the size of the target image to derive at least one converted class image; and
a procedure of segmenting, by the at least one processor, the target image based on a pixel value in each pixel of the at least one converted class image.