IP Library Granted Patent US 11,748,888
Granted Patent B2
US 11,748,888 · App. 17/096,778 · Granted Sep 5, 2023

End-to-end merge for video object segmentation (VOS)

Inventors: Abdalla Ahmed (Toronto, CA); Irina Kezele (Toronto, CA); Parham Aarabi (Richmond Hill, CA); Brendan Duke (Toronto, CA)
Assignee: L'Oreal
G06T7/11G06N3/08G06T7/174G06T7/20G06T9/002G06V10/273G06V10/771G06V10/82G06V20/40G06V40/107G06V40/12G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/20132G06V10/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,748,888
App. No.
17/096,778
Granted
Sep 5, 2023
Kind
B2
Abstract

There are provided methods and computing devices using semi-supervised learning to perform end-to-end video object segmentation, tracking respective object(s) from a single-frame annotation of a reference frame through a video sequence of frames. A known deep learning model may be used to annotate the reference frame to provide ground truth locations and masks for each respective object. A current frame is processed to determine current frame object locations, defining object scoremaps as a normalized cross-correlation between encoded object features of the current frame and encoded object features of a previous frame. Scoremaps for each of more than one previous frame may be defined. An Intersection over Union (IoU) function, responsive to the scoremaps, ranks candidate object proposals defined from the reference frame annotation to associate the respective objects to respective locations in the current frame. Pixel-wise overlap may be removed using a merge function responsive to the scoremaps.

Claims (22)

1. A method of semi-supervised video object segmentation to track and segment one or more objects throughout a video clip comprising a sequence of frames including respective previous frames and a current frame, each of the respective previous frames defining a respective target and the current frame a source, the method comprising,

for each respective object:

encoding features of the respective object in the source;

for each of the respective targets, defining respective attention maps between the features encoded from the source and features of the respective object encoded from the respective target;

associating the respective objects to respective locations in the current frame using an Intersection over Union (IoU) function responsive to the respective attention maps to rank candidate object proposals for the respective locations, where each candidate object proposal is tracked from a single-frame reference annotation of a reference frame of the video clip providing ground truth locations of the respective objects in the reference frame; and

defining a video segmentation mask for the respective object in the current frame in accordance with the associating.

2. The method of claim 1 , wherein, for each candidate object proposal, the Intersection over Union (IoU) function provides an IoU value that is responsive to each attention map with which to rank the candidate object proposals, the ranking selecting a highest IoU value for the respective object to identify the respective object in the current frame thereby tracking the respective object from the reference frame through the previous frame to the current frame.

3. The method of claim 2 , wherein the IoU value is a fusion of respective IoU values for the respective object from each of the respective attention maps.

4. The method of claim 1 , wherein defining a video segmentation mask comprises removing pixel-wise overlap from the respective objects in the current frame.

5. The method of claim 1 comprising using the video segmentation mask to apply an effect to the respective object for display.

6. The method of claim 5 , wherein the effect is associated with a product and/or service and the method comprises providing an interface to purchase one or both of the product and/or service.

7. A computing device comprising a non-transient storage device and a processor coupled thereto, the storage device storing instructions, which when executed by the processor, configure the computing device to perform semi-supervised video object segmentation to track and segment one or more objects throughout a video clip comprising a sequence of frames including respective previous frames and a current frame, each of the respective previous frames defining a respective target and the current frame a source, the computing device operating to:

for each respective object:

encode features of the respective object in the source; and

for each of the respective targets, define respective attention maps between the features encoded from the source and features of the respective object encoded from the respective target;

associate the respective objects to respective locations in the current frame using an Intersection over Union (IoU) function responsive to the respective attention maps to rank candidate object proposals for the respective locations, where each candidate object proposal is tracked from a single-frame reference annotation of a reference frame of the video clip providing ground truth locations of the respective objects in the reference frame; and

define a video segmentation mask for the respective object in the current frame in accordance with associating the respective objects to respective locations in the current frame using the Intersection over Union (IoU) function.

8. The computing device of claim 7 , wherein, for each candidate object proposal, the Intersection over Union (IoU) function provides an IoU value that is responsive to each attention map with which to rank the candidate object proposals, the ranking selecting a highest IoU value for the respective object to identify the respective object in the current frame thereby tracking the respective object from the reference frame through the previous frame to the current frame.

9. The computing device of claim 8 , wherein the IoU value is a fusion of respective IoU values for the respective object from each of the respective attention maps.

10. The computing device of claim 7 , wherein to define a video segmentation mask comprises removing pixel-wise overlap from the respective objects in the current frame.

11. The computing device of claim 7 comprising using the video segmentation mask to apply an effect to the respective object for display.

12. The computing device claim of 11 , wherein the effect is associated with a product and/or service and the computing device is configured to provide an interface to purchase one or both of the product and/or service.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2022
From: MODIFACE INC.
To: L'OREAL
Reel/Frame 062203/0481 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2022
From: AHMED, ABDALLA; KEZELE, IRINA; AARABI, PARHAM; DUKE, BRENDAN
To: MODIFACE INC.
Reel/Frame 059276/0600 →
Continuity (2)
Provisional Application 62935851 · Nov 15, 2019
Related Publication 20210150728A1 · May 20, 2021
Cited By (1)
US 12,327,337