IP Library Granted Patent US 12,205,299
Granted Patent B2
US 12,205,299 · App. 17/396,055 · Granted Jan 21, 2025

Video matting

Inventors: Linjie Yang (Los Angeles, CA); Peter Lin (Los Angeles, CA); Imran Saleemi (Los Angeles, CA)
Assignee: Lemon Inc.
G06T7/194G06T3/40G06T7/11G06V20/46G06T2207/10016G06T2207/20081G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,299
App. No.
17/396,055
Granted
Jan 21, 2025
Kind
B2
Abstract

The present disclosure describes techniques of improving video matting. The techniques comprise extracting features from each frame of a video by an encoder of a model, wherein the video comprises a plurality of frames; incorporating, by a decoder of the model, into any particular frame temporal information extracted from one or more frames previous to the particular frame, wherein the particular frame and the one or more previous frames are among the plurality of frames of the video, and the decoder is a recurrent decoder; and generating a representation of a foreground object included in the particular frame by the model, wherein the model is trained using segmentation dataset and matting dataset.

Claims (32)

1. A method of improving video matting, comprising:

extracting features from each frame of a video by an encoder of a model, wherein the video comprises a plurality of frames;

incorporating, by a decoder of the model, into any particular frame temporal information extracted from one or more frames previous to the particular frame, wherein the particular frame and the one or more previous frames are among the plurality of frames of the video, and the decoder is a recurrent decoder; and

generating a representation of a foreground object included in the particular frame by the model, wherein the model is trained using segmentation dataset and matting dataset, and wherein an application of the segmentation dataset to the model and an application of the matting dataset to the model are interleaved in a process of training the model.

2. The method of claim 1 , further comprising:

generating a low-resolution representative of each frame by downsampling each frame.

3. The method of claim 2 , further comprising:

encoding the low-resolution representative using the encoder, wherein the encoder comprises a plurality of convolution and pooling layers.

4. The method of claim 1 , wherein the decoder comprises a plurality of convolutional gated recurrent unit (ConvGRU) for incorporating the temporal information.

5. The method of claim 1 , wherein the model further comprises a Deep Guided Filter (DGF) for processing a high-resolution video.

6. The method of claim 1 , wherein the model is further trained by applying at least one loss function, and wherein the at least one loss function comprises a Least Absolute Deviations (L1) loss, a Laplacian loss, and a temporal coherence loss, or a foreground prediction loss.

7. The method of claim 1 , wherein the foreground object included in the particular frame is a human being.

8. A system of improving video matting, comprising:

at least one processor; and

at least one memory communicatively coupled to the at least one processor and storing instructions that upon execution by the at least one processor cause the system to perform operations, the operations comprising:

extracting features from each frame of a video by an encoder of a model, wherein the video comprises a plurality of frames;

incorporating, by a decoder of the model, into any particular frame temporal information extracted from one or more frames previous to the particular frame, wherein the particular frame and the one or more previous frames are among the plurality of frames of the video, and the decoder is a recurrent decoder; and

generating a representation of a foreground object included in the particular frame by the model, wherein the model is trained using segmentation dataset and matting dataset, and wherein an application of the segmentation dataset to the model and an application of the matting dataset to the model are interleaved in a process of training the model.

9. The system of claim 8 , the operations further comprising:

generating a low-resolution representative of each frame by downsampling each frame.

10. The system of claim 9 , further comprising:

encoding the low-resolution representative using the encoder, wherein the encoder comprises a plurality of convolution and pooling layers.

11. The system of claim 8 , wherein the decoder comprises a plurality of convolutional gated recurrent unit (ConvGRU) for incorporating the temporal information.

12. The system of claim 8 , wherein the model further comprises a Deep Guided Filter (DGF) for processing a high-resolution video.

13. The system of claim 8 , wherein the model is further trained by applying at least one loss function, and wherein the at least one loss function comprises a Least Absolute Deviations (L1) loss, a Laplacian loss, and a temporal coherence loss, or a foreground prediction loss.

14. The system of claim 8 , wherein the foreground object included in the particular frame is a human being.

15. A non-transitory computer-readable storage medium, storing computer-readable instructions that upon execution by a processor cause the processor to implement operations, the operation comprising:

extracting features from each frame of a video by an encoder of a model, wherein the video comprises a plurality of frames;

incorporating, by a decoder of the model, into any particular frame temporal information extracted from one or more frames previous to the particular frame, wherein the particular frame and the one or more previous frames are among the plurality of frames of the video, and the decoder is a recurrent decoder; and

generating a representation of a foreground object included in the particular frame by the model, wherein the model is trained using segmentation dataset and matting dataset, and wherein an application of the segmentation dataset to the model and an application of the matting dataset to the model are interleaved in a process of training the model.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the model further comprises a Deep Guided Filter (DGF) for processing a high-resolution video.

17. The non-transitory computer-readable storage medium of claim 15 , wherein the decoder comprises a plurality of convolutional gated recurrent unit (ConvGRU) for incorporating the temporal information.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2022
From: YANG, LINJIE; LIN, PETER; SALEEMI, IMRAN
To: BYTEDANCE INC.
Reel/Frame 059886/0075 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2022
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 059886/0115 →
Continuity (1)
Related Publication 20230044969A1 · Feb 9, 2023
References Cited (12)
US 9430715B1 · Wang · 2016 [cited by examiner]
US 20120020554A1 · Sun · 2012 [cited by examiner]
US 20120075331A1 · Mallick · 2012 [cited by examiner]
US 20140002746A1 · Bai · 2014 [cited by examiner]
US 20140003719A1 · Bai · 2014 [cited by examiner]
US 20170116481A1 · Chen · 2017 [cited by examiner]
US 20190037150A1 · Srikanth · 2019 [cited by examiner]
US 20190206066A1 · Saleemi · 2019 [cited by examiner]
US 20200340798A1 · Izatt · 2020 [cited by examiner]
US 20210383242A1 · Ostyakov · 2021 [cited by examiner]
US 20230360177A1 · Lee · 2023 [cited by examiner]
US 20240153255A1 · Kaneko · 2024 [cited by examiner]