Method for image segmentation and system therefor
Provided are a method for image segmentation and a system therefor. The method according to some embodiments may include acquiring a deep learning model trained through an image segmentation task, extracting motion information associated with a current frame of a given image, and performing image segmentation for the current frame by reflecting the extracted motion information into class-specific feature maps of the deep learning model, the class-specific feature maps being generated by the deep learning model based on the current frame.
1 . A method for image segmentation performed by at least one processor, the method comprising:
acquiring a deep learning model trained through an image segmentation task;
extracting motion information associated with a current frame of a given image; and
performing image segmentation for the current frame by reflecting the extracted motion information into class-specific feature maps, the class-specific feature maps being generated by the deep learning model based on the current frame,
wherein the performing the image segmentation comprises reflecting the extracted motion information into the class-specific feature maps based on a weight.
2 . The method of claim 1 , wherein the extracted motion information is not used in training of the deep learning model.
3 . The method of claim 1 , wherein the performing the image segmentation comprises performing the image segmentation for the current frame based on an amount of motion associated with the current frame being equal or greater than a threshold value, and
wherein a result of image segmentation for a previous frame is used in performing the image segmentation for the current frame based on the amount of motion associated with the current frame being less than the threshold value.
4 . The method of claim 1 , wherein the extracting the motion information comprises:
determining a reference frame from among a plurality of frames included in the given image; and
extracting the motion information associated with the current frame based on a difference between the current frame and the reference frame.
5 . The method of claim 4 , wherein a difference in a frame number between the current frame and the reference frame is determined to be greater based on a higher frame rate of a device that has captured the given image.
6 . The method of claim 4 , wherein a difference in a frame number between the current frame and the reference frame is determined to be smaller based on a higher resolution of a display that outputs the given image.
7 . The method of claim 4 , wherein a difference in a frame number between the current frame and the reference frame is determined to be smaller based on a higher importance of an object within the current frame.
8 . The method of claim 7 , wherein
the object corresponds to a user that participates in a video conference, and
an importance of the object is determined based on at least one of an amount of an utterance or a role of the user during the video conference.
9 . The method of claim 1 , wherein
the extracted motion information includes motion information of a first object and motion information of a second object, the first object and the second object being within the current frame, and
the performing the image segmentation comprises:
reflecting the motion information of the first object into a feature map of a first class corresponding to the first object; and
reflecting the motion information of the second object into a feature map of a second class corresponding to the second object.
10 . The method of claim 9 , wherein
the extracted motion information is two-dimensional (2D) data, and
the reflecting the motion information of the first object comprises:
determining an activated region within the feature map of the first class based on feature values exceeding a threshold value;
detecting an object motion region within the 2D data that spatially corresponds to the activated region; and
reflecting values of the object motion region into the feature map of the first class.
11 . The method of claim 9 , wherein
the extracted motion information is 2D data, and
the reflecting the motion information of the first object comprises:
detecting a motion region of the first object from the 2D data using attribute information of the first object; and
reflecting values of the motion region into the feature map of the first class.
12 . The method of claim 1 , wherein
the weight is determined to be greater based on a lower performance of the deep learning model.
13 . The method of claim 1 , wherein
the given image is an image of a user who participates in a video conference, and
the method further comprises:
applying a virtual background set by the user to the current frame using a result of the image segmentation for the current frame.
14 . A system for image segmentation comprising:
at least one processor; and
a memory configured to store at least one instruction,
wherein the at least one processor is configured to, by executing the at least one instruction stored in the memory, perform:
acquiring a deep learning model trained through an image segmentation task;
extracting motion information associated with a current frame of a given image; and
performing image segmentation for the current frame by reflecting the extracted motion information into class-specific feature maps, the class-specific feature maps being generated by the deep learning model based on the current frame,
wherein the performing the image segmentation comprises reflecting the extracted motion information into the class-specific feature maps based on a weight.
15 . A non-transitory computer-readable recording medium storing computer program executable by at least one processor to perform:
acquiring a deep learning model trained through an image segmentation task;
extracting motion information associated with a current frame of a given image; and
performing image segmentation for the current frame by reflecting the extracted motion information into class-specific feature maps, the class-specific feature maps being generated by the deep learning model based on the current frame,
wherein the performing the image segmentation comprises reflecting the extracted motion information into the class-specific feature maps based on a weight.