IP Library › Granted Patent US 12,737,896
Granted Patent B2
US 12,737,896 · App. 18/535,369 · Granted Sep 15, 2026

Method for image segmentation and system therefor

Inventors: Hee Sung Yang (Seoul, KR); Jun Ho Kang (Seoul, KR); Do Hyung Im (Seoul, KR)
Assignee: SAMSUNG SDS CO., LTD.
G06T7/215G06T7/12G06V10/25G06V10/764G06V10/771H04N5/272H04N7/157
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,896
App. No.
18/535,369
Filed
Dec 11, 2023
Granted
Sep 15, 2026
Kind
B2
Art Unit
2669
USPC
382/173
Abstract

Provided are a method for image segmentation and a system therefor. The method according to some embodiments may include acquiring a deep learning model trained through an image segmentation task, extracting motion information associated with a current frame of a given image, and performing image segmentation for the current frame by reflecting the extracted motion information into class-specific feature maps of the deep learning model, the class-specific feature maps being generated by the deep learning model based on the current frame.

Claims (52)

1 . A method for image segmentation performed by at least one processor, the method comprising:

acquiring a deep learning model trained through an image segmentation task;

extracting motion information associated with a current frame of a given image; and

performing image segmentation for the current frame by reflecting the extracted motion information into class-specific feature maps, the class-specific feature maps being generated by the deep learning model based on the current frame,

wherein the performing the image segmentation comprises reflecting the extracted motion information into the class-specific feature maps based on a weight.

2 . The method of claim 1 , wherein the extracted motion information is not used in training of the deep learning model.

3 . The method of claim 1 , wherein the performing the image segmentation comprises performing the image segmentation for the current frame based on an amount of motion associated with the current frame being equal or greater than a threshold value, and

wherein a result of image segmentation for a previous frame is used in performing the image segmentation for the current frame based on the amount of motion associated with the current frame being less than the threshold value.

4 . The method of claim 1 , wherein the extracting the motion information comprises:

determining a reference frame from among a plurality of frames included in the given image; and

extracting the motion information associated with the current frame based on a difference between the current frame and the reference frame.

5 . The method of claim 4 , wherein a difference in a frame number between the current frame and the reference frame is determined to be greater based on a higher frame rate of a device that has captured the given image.

6 . The method of claim 4 , wherein a difference in a frame number between the current frame and the reference frame is determined to be smaller based on a higher resolution of a display that outputs the given image.

7 . The method of claim 4 , wherein a difference in a frame number between the current frame and the reference frame is determined to be smaller based on a higher importance of an object within the current frame.

8 . The method of claim 7 , wherein

the object corresponds to a user that participates in a video conference, and

an importance of the object is determined based on at least one of an amount of an utterance or a role of the user during the video conference.

9 . The method of claim 1 , wherein

the extracted motion information includes motion information of a first object and motion information of a second object, the first object and the second object being within the current frame, and

the performing the image segmentation comprises:

reflecting the motion information of the first object into a feature map of a first class corresponding to the first object; and

reflecting the motion information of the second object into a feature map of a second class corresponding to the second object.

10 . The method of claim 9 , wherein

the extracted motion information is two-dimensional (2D) data, and

the reflecting the motion information of the first object comprises:

determining an activated region within the feature map of the first class based on feature values exceeding a threshold value;

detecting an object motion region within the 2D data that spatially corresponds to the activated region; and

reflecting values of the object motion region into the feature map of the first class.

11 . The method of claim 9 , wherein

the extracted motion information is 2D data, and

the reflecting the motion information of the first object comprises:

detecting a motion region of the first object from the 2D data using attribute information of the first object; and

reflecting values of the motion region into the feature map of the first class.

12 . The method of claim 1 , wherein

the weight is determined to be greater based on a lower performance of the deep learning model.

13 . The method of claim 1 , wherein

the given image is an image of a user who participates in a video conference, and

the method further comprises:

applying a virtual background set by the user to the current frame using a result of the image segmentation for the current frame.

14 . A system for image segmentation comprising:

at least one processor; and

a memory configured to store at least one instruction,

wherein the at least one processor is configured to, by executing the at least one instruction stored in the memory, perform:

acquiring a deep learning model trained through an image segmentation task;

extracting motion information associated with a current frame of a given image; and

performing image segmentation for the current frame by reflecting the extracted motion information into class-specific feature maps, the class-specific feature maps being generated by the deep learning model based on the current frame,

wherein the performing the image segmentation comprises reflecting the extracted motion information into the class-specific feature maps based on a weight.

15 . A non-transitory computer-readable recording medium storing computer program executable by at least one processor to perform:

acquiring a deep learning model trained through an image segmentation task;

extracting motion information associated with a current frame of a given image; and

performing image segmentation for the current frame by reflecting the extracted motion information into class-specific feature maps, the class-specific feature maps being generated by the deep learning model based on the current frame,

wherein the performing the image segmentation comprises reflecting the extracted motion information into the class-specific feature maps based on a weight.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2023
From: YANG, HEE SUNG; KANG, JUN HO; IM, DO HYUNG
To: SAMSUNG SDS CO., LTD.
Reel/Frame 065856/0213 →
Priority Claims (1)
KR 10-2023-0051755 · Apr 20, 2023 · national
Continuity (1)
Related Publication 20240354963A1 · Oct 24, 2024
References Cited (18)
US 10311578B1 · Kim et al. · 2019 [cited by applicant]
US 12067733B2 · Cui · 2024 [cited by examiner]
US 20190355128A1 · Grauman · 2019 [cited by examiner]
US 20200027247A1 · Minnen · 2020 [cited by examiner]
US 20220109838A1 · Guruva reddiar · 2022 [cited by examiner]
US 20220148284A1 · Kim · 2022 [cited by examiner]
US 20230145028A1 · Oh · 2023 [cited by examiner]
US 20230319413A1 · Manzari · 2023 [cited by examiner]
US 20230419522A1 · Yang · 2023 [cited by examiner]
US 20230421719A1 · Zhu · 2023 [cited by examiner]
US 20240056322A1 · Kare · 2024 [cited by examiner]
US 20250104423A1 · Wen · 2025 [cited by examiner]
CN 111950444A · 2020 [cited by applicant]
CN 110321761B · 2022 [cited by applicant]
CN 112862828B · 2022 [cited by applicant]
CN 115393491A · 2022 [cited by applicant]
KR 1020190067680A · 2019 [cited by applicant]
Ding et al., “Language-Bridged Spatial-Temporal Interaction for Referring Video Object Segmentation”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, 10 total pages, doi:10.1109/CVPR526… [cited by applicant]