IP Library Granted Patent US 9,633,446
Granted Patent B2
US 9,633,446 · App. 14/585,839 · Granted Apr 25, 2017

Method, apparatus and computer program product for segmentation of objects in media content

Inventor: Tinghuai Wang (Tampere, FI)
Assignee: Nokia Technologies Oy
G06T7/2006G06F17/30858G06T7/0081G06T7/0087G06T7/0097G06T7/2033G06T7/2066H04N5/147G06T2207/10016G06T2207/10024G06T2207/20072G06T2207/20076G06T2207/20081G06T2207/30241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,633,446
App. No.
14/585,839
Granted
Apr 25, 2017
Kind
B2
Abstract

In an example embodiment a method, apparatus and computer program product are provided. The method includes extracting a first set of target-object regions and at least one set of non-target object regions from a plurality of regions of a video content based at least on likelihood information. A plurality of unlabeled regions of the video content are classified based on the first set of target-object regions and the at least one set of non-target object regions to generate a second set of target object regions that is denser than the first set of target object regions. A model is learned for modelling at least one of the target object and non-target objects of the video content based at least on the second set of target object regions. The target object in the video content is segmented based on the model and the second set of target object regions.

Claims (56)

1. A method comprising:

extracting a first set of target object regions and at least one set of non-target object regions from a plurality of regions of a video content based at least on a likelihood information, the likelihood information being indicative of a likelihood of the plurality of regions to be associated with a target object of the video content;

classifying a plurality of unlabeled regions of the video content based on the first set of target object regions and the at least one set of non-target object regions to generate a second set of target object regions, the second set of target object regions being denser than the first set of target object regions;

learning a model for modelling at least one of the target object and non-target objects of the video content based at least on the second set of target object regions; and

segmenting the target object in the video content based on the model and the second set of target object regions.

2. The method as claimed in claim 1 , wherein extracting the first set of target object regions and the at least one set of non-target object regions further comprises:

determining likelihood scores for a plurality of frame-regions associated with the video content based on the likelihood information, the plurality of frame-regions being generated based on a partitioning of a plurality of frames of the video content;

selecting the plurality of regions from the plurality of frame-regions having the likelihood scores greater than or equal to a threshold value of the likelihood score;

clustering the plurality of regions into a plurality of clusters based on a similarity information associated with respective region pairs of the plurality of regions; and

ranking the plurality of clusters based on the likelihood score of corresponding regions forming a respective cluster of the plurality of clusters, wherein the respective cluster having a rank greater than or equal to a predetermined rank is determined to be associated with the first set of target object regions, and wherein the respective cluster having the rank less than the predetermined rank is determined to be associated with the at least one set of non-target object regions.

3. The method as claimed in claim 1 , further comprising determining the likelihood information associated with respective frame-regions of the plurality of frame-regions based on an appearance information, a motion information, and a spatial location information associated with the respective frame-regions of the plurality of frame-regions.

4. The method as claimed in claim 1 , further comprising training a classifier for labeling the plurality of unlabeled regions of the video content based on the plurality of clusters associated with the first set of target object regions and the at least one set of non-target object regions.

5. The method as claimed in claim 4 , wherein the plurality of unlabeled regions of the video content comprises regions of the video content not associated with the first set of target object regions and the at least one set of non-target object regions.

6. The method as claimed in claim 1 , wherein the second set of target object regions being denser than the first set of target object regions comprises the second set of target object regions being spatially and temporally denser than the second set of target object regions.

7. The method as claimed in claim 1 , wherein classifying the plurality of unlabeled regions comprises:

applying the classifier to the plurality of unlabeled regions of the video content;

assigning weights (Y) to the plurality of unlabeled regions based on the classifier; and

determining optimal labels (X) for one or more unlabeled regions of the plurality of unlabeled regions based on a minimization of an energy function associated with the one or more unlabeled regions.

8. An apparatus comprising:

at least one processor; and

at least one memory comprising computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to at least perform:

extract a first set of target object regions and at least one set of non-target object regions from a plurality of regions of a video content based at least on a likelihood information, the likelihood information being indicative of a likelihood of the plurality of regions to be associated with a target object of the video content;

classify a plurality of unlabeled regions of the video content based on the first set of target object regions and the at least one set of non-target object regions to generate a second set of target object regions, the second set of target object regions being denser than the first set of target object regions;

learn a model for modelling at least one of the target object and non-target objects of the video content based at least on the second set of target object regions; and

segment the target object in the video content based on the model and the second set of target object regions.

9. The apparatus as claimed in claim 8 , wherein the apparatus is further caused at least in part to determine the likelihood information associated with respective frame-regions of the plurality of frame-regions based on an appearance information, a motion information, and a spatial location information associated with the respective frame-regions of the plurality of frame-regions.

10. The apparatus as claimed in claim 8 , wherein the apparatus is further caused at least in part to train a classifier for labeling the plurality of unlabeled regions of the video content based on the plurality of clusters associated with the first set of target object regions and the at least one set of non-target object regions.

11. The apparatus as claimed in claim 10 , wherein the plurality of unlabeled regions of the video content comprises regions of the video content not associated with the first set of target object regions and the at least one set of non-target object regions.

12. The apparatus as claimed in claim 8 , wherein the second set of target object regions being denser than the first set of target object regions comprises the second set of target object regions being spatially and temporally denser than the second set of target object regions.

13. The apparatus as claimed in claim 8 , wherein classifying the plurality of unlabeled regions comprises:

applying the classifier to the plurality of unlabeled regions of the video content;

assigning weights (Y) to the plurality of unlabeled regions based on the classifier; and

determining optimal labels (X) for one or more unlabeled regions of the plurality of unlabeled regions based on a minimization of an energy function associated with the one or more unlabeled regions.

14. The apparatus as claimed in claim 8 , wherein for extracting the first set of target object regions and the at least one set of non-target object regions, the apparatus is further caused, at least in part to:

determine likelihood scores for a plurality of frame-regions associated with the video content based on the likelihood information, the plurality of frame-regions being generated based on a partitioning of a plurality of frames of the video content;

select the plurality of regions from the plurality of frame-regions having the likelihood scores greater than or equal to a threshold value of the likelihood score;

cluster the plurality of regions into a plurality of clusters based on a similarity information associated with respective region pairs of the plurality of regions; and

rank the plurality of clusters based on the likelihood score of corresponding regions forming a respective cluster of the plurality of clusters, wherein the respective cluster having a rank greater than or equal to a predetermined rank is determined to be associated with the first set of target object regions, and wherein the respective cluster having the rank less than the predetermined rank is determined to be associated with the at least one set of non-target object regions.

15. A computer program product comprising at least one non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium comprising a set of instructions, which, when executed by one or more processors, cause an apparatus to at least perform:

extract a first set of target object regions and at least one set of non-target object regions from a plurality of regions of a video content based at least on a likelihood information, the likelihood information being indicative of a likelihood of the plurality of regions to be associated with a target object of the video content;

classify a plurality of unlabeled regions of the video content based on the first set of target object regions and the at least one set of non-target object regions to generate a second set of target object regions, the second set of target object regions being denser than the first set of target object regions;

learn a model for modelling at least one of the target object and non-target objects of the video content based at least on the second set of target object regions; and

segment the target object in the video content based on the model and the second set of target object regions.

16. The computer program product as claimed in claim 15 , wherein for extracting the first set of target object regions and the at least one set of non-target object regions, the apparatus is further caused, at least in part to:

determine likelihood scores for a plurality of frame-regions associated with the video content based on the likelihood information, the plurality of frame-regions being generated based on a partitioning of a plurality of frames of the video content;

select the plurality of regions from the plurality of frame-regions having the likelihood scores greater than or equal to a threshold value of the likelihood score;

cluster the plurality of regions into a plurality of clusters based on a similarity information associated with respective region pairs of the plurality of regions; and

rank the plurality of clusters based on the likelihood score of corresponding regions forming a respective cluster of the plurality of clusters, wherein the respective cluster having a rank greater than or equal to a predetermined rank is determined to be associated with the first set of target object regions, and wherein the respective cluster having the rank less than the predetermined rank is determined to be associated with the at least one set of non-target object regions.

17. The computer program product as claimed in claim 15 , wherein the apparatus is further caused at least in part to determine the likelihood information associated with respective frame-regions of the plurality of frame-regions based on an appearance information, a motion information, and a spatial location information associated with the respective frame-regions of the plurality of frame-regions.

18. The computer program product as claimed in claim 15 , wherein the apparatus is further caused at least in part to train a classifier for labeling the plurality of unlabeled regions of the video content based on the plurality of clusters associated with the first set of target object regions and the at least one set of non-target object regions.

19. The computer program product as claimed in claim 18 , wherein the plurality of unlabeled regions of the video content comprises regions of the video content not associated with the first set of target object regions and the at least one set of non-target object regions.

20. The computer program product as claimed in claim 15 , wherein the second set of target object regions being denser than the first set of target object regions comprises the second set of target object regions being spatially and temporally denser than the second set of target object regions.

21. The computer program product as claimed in claim 15 , wherein classifying the plurality of unlabeled regions comprises:

applying the classifier to the plurality of unlabeled regions of the video content;

assigning weights (Y) to the plurality of unlabeled regions based on the classifier; and

determining optimal labels (X) for one or more unlabeled regions of the plurality of unlabeled regions based on a minimization of an energy function associated with the one or more unlabeled regions.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2015
From: NOKIA CORPORATION
To: NOKIA TECHNOLOGIES OY
Reel/Frame 034781/0200 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2015
From: WANG, TINGHUAI
To: NOKIA CORPORATION
Reel/Frame 034730/0729 →
Priority Claims (1)
GB 1402977.1 · Feb 20, 2014 · national
Continuity (1)
Related Publication 20150235377A1 · Aug 20, 2015