IP Library › Granted Patent US 12,541,829
Granted Patent B2
US 12,541,829 · App. 18/142,304 · Granted Feb 3, 2026

Motion-based pixel propagation for video inpainting

Inventors: Jenhao Hsiao (Palo Alto, CA); Zheng Zhang (Palo Alto, CA)
Assignee: INNOPEAK TECHNOLOGY, INC.
G06T5/77G06T5/50G06T7/246G06T7/33G06T7/73G06V10/776G06V20/46G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,829
App. No.
18/142,304
Granted
Feb 3, 2026
Kind
B2
Abstract

Techniques for pixel propagation and video inpainting are described. A video processing application selects, from a sequence of video frames, a reference video frame that includes replacement pixels corresponding to a mask of an undesired object. The video processing application aligns the reference video frame with the target video frame. The video processing application identifies, in the reference video frame, pixels corresponding to the mask. The video processing application uses the pixels from the reference video frame to replace pixels in the target video frame that correspond to the undesired object.

Claims (75)

1 . A method implemented by a computer system, the method comprising: receiving a sequence of video frames and a set of masks, wherein the sequence of video frames comprises an initial video frame including an undesired object and the set of masks comprises an initial mask corresponding to the undesired object; identifying from the sequence of video frames, a subset of video frames; creating, using the set of masks, a combined mask corresponding to the undesired object;

selecting, from the sequence of video frames, a reference video frame that comprises replacement pixels corresponding to pixels in a target video frame that are associated with the combined mask when:

an overlap between the combined mask in the reference video frame and the combined mask in the target video frame is less than a first threshold; and

an optical flow metric of the combined mask in the reference video frame compared to the combined mask in the target video frame is less than a second threshold;

aligning the reference video frame with the target video frame; and replacing, in the target video frame, pixels corresponding to the combined mask with pixels corresponding to the combined mask in the reference video frame and removing the undesired object from the target video frame.

2 . The method of claim 1 , wherein receiving the set of masks comprises generating the set of masks by identifying, in each video frame of the sequence of video frames, a set of pixels that correlates with a set of pixels of the initial mask.

3 . The method of claim 1 , wherein creating the combined mask comprises:

determining a subset of the set of masks for which a stability score is greater than a third threshold; and

forming, from the subset of the set of masks, the combined mask as a union of pixels corresponding to the undesired object in the subset of the set of masks.

4 . The method of claim 1 , wherein determining the overlap comprises:

identifying, in the reference video frame, a first set of pixels that correspond to pixels in the combined mask;

locating, in the target video frame, a second set of pixels that corresponds to the first set of pixels; and

computing a stability score using the first set of pixels and the second set of pixels.

5 . The method of claim 1 , further comprising outputting the target video frame to a display device.

6 . The method of claim 1 , further comprising determining the optical flow metric by:

determining a difference between pixels that correspond to the combined mask in the reference video frame and pixels that correspond to the combined mask in the target video frame; and

computing the optical flow metric based on the difference.

7 . The method of claim 1 , wherein aligning the reference video frame with the target video frame comprises:

a) determining a plurality of feature points in the target video frame and the reference video frame;

b) sampling a subset of the plurality of feature points to provide a validation feature set;

c) performing feature matching for the validation feature set;

d) calculating an overall matching score for the validation feature set;

iteratively performing b) through d) a predetermined number of times;

selecting the validation feature set having a highest overall matching score; and

computing a homography for the selected validation feature set.

8 . The method of claim 7 , wherein sampling the subset of the plurality of feature points comprises performing a random sampling.

9 . A computer system, including: one or more processors; and

one or more non-transitory computer-storage media storing instructions that, upon execution on by the one or more processors, cause the computer system to perform operations including:

receiving a sequence of video frames and a set of masks, wherein the sequence of video frames comprises an initial video frame including an undesired object and the set of masks comprises an initial mask corresponding to the undesired object;

identifying from the sequence of video frames, a subset of video frames;

creating, using the set of masks, a combined mask corresponding to the undesired object;

selecting, from the sequence of video frames, a reference video frame that comprises replacement pixels corresponding to pixels in a target video frame that are associated with the combined mask when:

an overlap between the combined mask in the reference video frame and the combined mask in the target video frame is less than a first threshold; and

an optical flow metric of the combined mask in the reference video frame compared to the combined mask in the target video frame is less than a second threshold;

aligning the reference video frame with the target video frame; and

replacing, in the target video frame, pixels corresponding to the combined mask with pixels corresponding to the combined mask in the reference video frame and removing the undesired object from the target video frame.

10 . The computer system of claim 9 , wherein receiving the set of masks comprises generating the set of masks by identifying, in each video frame of the sequence of video frames, a set of pixels that correlates with a set of pixels of the initial mask.

11 . The computer system of claim 9 , wherein creating the combined mask comprises:

determining a subset of the set of masks for which a stability score is greater than a third threshold; and

forming, from the subset of the set of masks, the combined mask as a union of pixels corresponding to the undesired object in the subset of the set of masks.

12 . The computer system of claim 9 , wherein the operations further comprise determining an optical flow metric by:

determining a difference between pixels that correspond to the combined mask in the reference video frame and pixels that correspond to the combined mask in the target video frame; and

computing the optical flow metric based on the difference.

13 . The computer system of claim 9 , wherein aligning the reference video frame with the target video frame comprises:

a) determining a plurality of feature points in the target video frame and the reference video frame;

b) sampling a subset of the plurality of feature points to provide a validation feature set;

c) performing feature matching for the validation feature set; d) calculating an overall matching score for the validation feature set;

iteratively performing b) through d) a predetermined number of times;

selecting the validation feature set having a highest overall matching score; and

computing a homography for the selected validation feature set.

14 . The computer system of claim 9 , wherein sampling the subset of the plurality of feature points comprises performing a random sampling.

15 . One or more non-transitory computer-storage media storing instructions that, upon execution on a computer system, cause the computer system to perform operations including:

receiving a sequence of video frames and a set of masks, wherein the sequence of video frames comprises an initial video frame including an undesired object and the set of masks comprises an initial mask corresponding to the undesired object;

identifying from the sequence of video frames, a subset of video frames; creating, using the set of masks, a combined mask corresponding to the undesired object;

selecting, from the sequence of video frames, a reference video frame that comprises replacement pixels corresponding to pixels in a target video frame that are associated with the combined mask when:

an overlap between the combined mask in the reference video frame and the combined mask in the target video frame is less than a first threshold; and

an optical flow metric of the combined mask in the reference video frame compared to the combined mask in the target video frame is less than a second threshold;

aligning the reference video frame with the target video frame; and

replacing, in the target video frame, pixels corresponding to the combined mask with pixels corresponding to the combined mask in the reference video frame and removing the undesired object from the target video frame.

16 . The non-transitory computer-storage media of claim 15 , wherein receiving the set of masks comprises generating the set of masks by identifying, in each video frame of the sequence of video frames, a set of pixels that correlates with a set of pixels of the initial mask.

17 . The non-transitory computer-storage media of claim 15 , wherein creating the combined mask comprises:

determining a subset of the set of masks for which a stability score is greater than a third threshold; and

forming, from the subset of the set of masks, the combined mask as a union of pixels corresponding to the undesired object in the subset of the set of masks.

18 . The non-transitory computer-storage media of claim 15 , wherein the operations further comprise determining the optical flow metric by:

determining a difference between pixels that correspond to the combined mask in the reference video frame and pixels that correspond to the combined mask in the target video frame; and

computing the optical flow metric based on the difference.

19 . The non-transitory computer-storage media of claim 15 , wherein aligning the reference video frame with the target video frame comprises:

a) determining a plurality of feature points in the target video frame and the reference video frame;

b) sampling a subset of the plurality of feature points to provide a validation feature set;

c) performing feature matching for the validation feature set;

d) calculating an overall matching score for the validation feature set;

iteratively performing b) through d) a predetermined number of times;

selecting the validation feature set having a highest overall matching score; and

computing a homography for the selected validation feature set.

20 . The non-transitory computer-storage media of claim 19 , wherein sampling the subset of the plurality of feature points comprises performing a random sampling.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2023
From: HSIAO, JENHAO; ZHANG, ZHENG
To: INNOPEAK TECHNOLOGY, INC.
Reel/Frame 063509/0454 →
Continuity (2)
Continuation PCTUS2020058552 · Nov 2, 2020
Related Publication 20230274399A1 · Aug 31, 2023
References Cited (16)
US 8411966B2 · Zhang · 2013 [cited by examiner]
US 10645364B2 · Kumar · 2020 [cited by examiner]
US 20090196475A1 · Demirli et al. · 2009 [cited by applicant]
US 20160350936A1 · Korchev et al. · 2016 [cited by applicant]
US 20180068451A1 · Leung · 2018 [cited by examiner]
US 20200118594A1 · Oxholm · 2020 [cited by examiner]
US 20200175654A1 · Tagra · 2020 [cited by examiner]
US 20210287007A1 · Wang · 2021 [cited by examiner]
US 20210342983A1 · Lin · 2021 [cited by examiner]
US 20220417590A1 · Jiao · 2022 [cited by examiner]
Granados, M. et all., (2012). Background Inpainting for Videos with Dynamic Objects and a Free-Moving Camera. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds) Computer Vision â ECCV 2012. ECCV 20… [cited by examiner]
Woo, S. et al., (2019). Align-and-Attend Network for Globally and Locally Coherent Video Inpainting. ArXiv, abs/1905.13066. (Year: 2019). [cited by examiner]
International Search Report and Written Opinion dated Feb. 2, 2021 in International Application No. PCT/US2020/058552. [cited by applicant]
Ma et al. “Image Matching from Handcrafted to Deep Features: A Survey”. international Journal of Computer Vision[online] published Aug. 4, 2020 retrieved Dec. 31, 2020]. Retrieved from the Internet <URL: https://link.sp… [cited by applicant]
Sungho Lee et al. “Copy-and-paste networks for deep video inpainting” In:ICCV. pp. 4413-4421, 2019. [cited by applicant]
Rui Xu et al.“Deep flow-guided video inpainting”In IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2019. [cited by applicant]