IP Library › Granted Patent US 12,400,303
Granted Patent B2
US 12,400,303 · App. 17/641,700 · Granted Aug 26, 2025

Image replacement inpainting

Inventors: Nathan James Frey (Los Angeles, CA); Vinay Kotikalapudi Sriram (Boyds, MD)
Assignee: Google LLC
G06T5/77G06T5/50G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,303
App. No.
17/641,700
Granted
Aug 26, 2025
Kind
B2
Abstract

A method for replacing an object in an image. The method may include identifying a first object at a position within a first image, masking, based on the first image and the position of the first object, a target area to produce a masked image, generating, based on the masked image and an inpainting machine learning model, a second image different from the first image, the inpainting machine learning model being trained using a difference between the target area of training images and content of generated images at location corresponding to the target area of the training images, generating, based on the masked image and the second image, a third image, and adding, to the third image, a new object different from the first object.

Claims (40)

1. A computer-implemented method for replacing an object in an image comprising:

identifying a first object at a position within a first image;

masking, based on the first image and the position of the first object, a target area to produce a masked image;

generating, based on the masked image and an inpainting machine learning model, a second image different from the first image, the inpainting machine learning model being trained using a difference between the target area of training images and content of generated images at a location corresponding to the target area of the training images;

generating, based on the masked image and the second image, a third image, wherein generating the third image comprises:

masking, based on the second image and the location corresponding to the target area of the first image, an inverse target area to produce an inverse masked image; and

generating, based on the masked image and the inverse masked image, the third image; and

adding, to the third image, a new object different from the first object.

2. The computer-implemented method of claim 1 , wherein the first image is a frame of a video.

3. The computer-implemented method of claim 1 , wherein the inpainting machine learning model is trained using a loss function that represents a difference between the target area of training images and content of generated images at a location corresponding to the target area of the training images.

4. The computer-implemented method of claim 1 , wherein the inverse target area comprises an area of the second image that is outside of the location corresponding to the target area of the first image.

5. The computer-implemented method of claim 1 , wherein masking, based on the second image and the location corresponding to the target area of the first image, an inverse target area to produce an inverse masked image comprises generating, from the second image, an inverse masked image that includes at least some content of the second image that is inside the target area and that does not include at least some content of the second image that is outside the target area.

6. The computer-implemented method of claim 1 , wherein generating the third image comprises compositing the inverse masked image with the masked image.

7. The computer-implemented method of claim 1 , further comprising extrapolating, based on the third image, a fourth image, wherein each of the first image, second image, third image, and fourth image is a frame of a video.

8. A system comprising:

one or more processors; and

one or more memory elements including instructions that, when executed, cause the one or more processors to perform operations including:

identifying a first object at a position within a first image;

masking, based on the first image and the position of the first object, a target area to produce a masked image;

generating, based on the masked image and an inpainting machine learning model, a second image different from the first image, the inpainting machine learning model being trained using a difference between the target area of training images and content of generated images at a location corresponding to the target area of the training images;

generating, based on the masked image and the second image, a third image, wherein generating the third image comprises:

masking, based on the second image and the location corresponding to the target area of the first image, an inverse target area to produce an inverse masked image; and

generating, based on the masked image and the inverse masked image, the third image; and

adding, to the third image, a new object different from the first object.

9. The system of claim 8 , wherein the inpainting machine learning model is trained using a loss function that represents a difference between the target area of training images and content of generated images at a location corresponding to the target area of the training images.

10. The system of claim 8 , wherein the first image is a frame of a video.

11. The system of claim 8 , wherein the inverse target area comprises an area of the second image that is outside of the location corresponding to the target area of the first image.

12. The system of claim 8 , wherein masking, based on the second image and the location corresponding to the target area of the first image, an inverse target area to produce an inverse masked image comprises generating, from the second image, an inverse masked image that includes at least some content of the second image that is inside the target area and that does not include at least some content of the second image that is outside the target area.

13. The system of claim 8 , wherein generating the third image comprises compositing the inverse masked image with the masked image.

14. The system of claim 8 , the operations further comprising extrapolating, based on the third image, a fourth image, wherein each of the first image, second image, third image, and fourth image is a frame of a video.

15. A non-transitory computer storage medium encoded with instructions that when executed by a distributed computing system cause the distributed computing system to perform operations comprising:

identifying a first object at a position within a first image;

masking, based on the first image and the position of the first object, a target area to produce a masked image;

generating, based on the masked image and an inpainting machine learning model, a second image different from the first image, the inpainting machine learning model being trained using a difference between the target area of training images and content of generated images at a location corresponding to the target area of the training images;

generating, based on the masked image and the second image, a third image, wherein generating the third image comprises:

masking, based on the second image and the location corresponding to the target area of the first image, an inverse target area to produce an inverse masked image; and

generating, based on the masked image and the inverse masked image, the third image; and

adding, to the third image, a new object different from the first object.

16. The non-transitory computer storage medium of claim 15 , wherein the inpainting machine learning model is trained using a loss function that represents a difference between the target area of training images and content of generated images at a location corresponding to the target area of the training images.

17. The non-transitory computer storage medium of claim 15 , wherein the first image is a frame of a video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2022
From: FREY, NATHAN JAMES; SRIRAM, VINAY KOTIKALAPUDI
To: GOOGLE LLC
Reel/Frame 059405/0995 →
Continuity (1)
Related Publication 20220301118A1 · Sep 22, 2022
References Cited (46)
US 9519984B2 · Masuko · 2016 [cited by applicant]
US 10503998B2 · Harron et al. · 2019 [cited by applicant]
US 10540757B1 · Bouhnik · 2020 [cited by examiner]
US 10755391B2 · Lin · 2020 [cited by examiner]
US 11776184B2 · Zhang · 2023 [cited by examiner]
US 11854203B1 · Gafni · 2023 [cited by examiner]
US 12223672B2 · Assouline · 2025 [cited by examiner]
US 20130094780A1 · Tang · 2013 [cited by examiner]
US 20190200050A1 · Zhang et al. · 2019 [cited by applicant]
US 20190236759A1 · Lai et al. · 2019 [cited by applicant]
US 20190311505A1 · Helm · 2019 [cited by examiner]
US 20190355102A1 · Lin · 2019 [cited by examiner]
US 20200327675A1 · Lin · 2020 [cited by examiner]
US 20210073961A1 · Safdarnejad · 2021 [cited by examiner]
US 20210133936A1 · Chandra · 2021 [cited by examiner]
US 20210334935A1 · Grigoriev · 2021 [cited by examiner]
US 20210342983A1 · Lin · 2021 [cited by examiner]
US 20210342984A1 · Lin · 2021 [cited by examiner]
US 20220366544A1 · Kudelski · 2022 [cited by examiner]
CN 108830827 · 2018 [cited by applicant]
CN 109472260 · 2019 [cited by applicant]
CN 109564575 · 2019 [cited by applicant]
EP 2492843 · 2012 [cited by applicant]
JP 2003333424 · 2003 [cited by applicant]
KR 1020130140904 · 2013 [cited by applicant]
KR 1020190021967 · 2019 [cited by applicant]
Dahun Kim [online], “Deep Video Inpainting—VINet (CVPR 2019)” Presented at the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 21, 2019, retrieved on Oct. 11, 2022, <https://www.youtube.com/watch?v=RtTh… [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2020/032677, dated Feb. 4, 2021, 14 pages. [cited by applicant]
Isola et al., “Image-to-image translation with conditional adversarial networks.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017, 1125-1134. [cited by applicant]
Jin et al., “Video logo removal detection based on sparse representation.” Multimedia Tools and Applications 77.22, Nov. 2018, 29303-29322. [cited by applicant]
Mosleh et al., “Automatic inpainting scheme for video text detection and removal.” IEEE Transactions on Image processing 22.11 Jul. 17, 2013, 4460-4472. [cited by applicant]
Wang et al., “Automatic TV logo detection, tracking and removal in broadcast video.” International Conference on Multimedia Modeling. Springer, Berlin, Heidelberg, Jan. 9, 2007, 63-72. [cited by applicant]
Office Action in Chinese Appln. No. 202080070503.X, mailed on Aug. 7, 2024, 17 pages (with English translation). [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2020/032677, mailed on Nov. 24, 2022, 7 pages. [cited by applicant]
Notice of Allowance in Korean Appln. No. 10-2022-7011445, mailed on Nov. 16, 2023, 3 pages (with English translation). [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2022-521020, mailed on Feb. 19, 2024, 5 pages (with English translation). [cited by applicant]
Office Action in India Appln. No. 202227012575, mailed on Jun. 10, 2024, 3 pages (with English translation). [cited by applicant]
Office Action in India Appln. No. 202227012575, mailed on Jun. 25, 2024, 3 pages (with English translation). [cited by applicant]
Office Action in Japanese Appln. No. 2022-521020, mailed on Sep. 11, 2023, 7 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 202080070503.X, mailed on Nov. 19, 2024, 14 pages (with English translation). [cited by applicant]
Din et al., “A novel GAN-based network for unmasking of masked face.” IEEE Access 8, Mar. 2020, 44276-44287. [cited by applicant]
Office Action in Japanese Appln. No. 2022-521020, dated Mar. 27, 2023, 9 pages (with English translation). [cited by applicant]
Office Action in Korean Appln. No. 10-2022-7011445, dated May 9, 2023, 11 pages (with English translation). [cited by applicant]
Office Action in India Appln. No. 202227012575, dated Jan. 5, 2023, 6 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 202080070503.X, mailed on Mar. 28, 2025, 17 pages (with English translation). [cited by applicant]
Office Action in European Appln. No. 20729539.5, mailed on Jan. 2, 2025, 6 pages. [cited by applicant]
Cited By (1)
US 12,725,498