IP Library › Granted Patent US 12,217,472
Granted Patent B2
US 12,217,472 · App. 17/968,634 · Granted Feb 4, 2025

Segmenting and removing objects from visual media items

Inventors: Orly Liba (Mountain View, CA); Nikhil Karnad (Mountain View, CA); Nori Kanazawa (Mountain View, CA); Yael Pritch Knaan (Mountain View, CA); Huizhong Chen (Mountain View, CA); Longqi Cai (Mountain View, CA)
Assignee: Google LLC
G06V10/273G06T5/20G06T5/77G06T5/94G06T11/00G06V10/764G06V10/774G06V20/20G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,472
App. No.
17/968,634
Granted
Feb 4, 2025
Kind
B2
Abstract

A media application generates training data that includes a first set of visual media items and a second set of visual media items, where the first set of visual media items correspond to the second set of visual items and include distracting objects that are manually segmented. The media application trains a segmentation machine-learning model based on the training data to receive a visual media item with one or more distracting objects and to output a segmentation mask for one or more segmented objects that correspond to the one or more distracting objects.

Claims (45)

1. A computer-implemented method comprising:

generating training data that includes a first set of media items and a second set of media items, wherein the first set of media items include distracting objects and the second set of media items include manual segmentations of the distracting objects;

identifying one or more original media items in the first set of media items that include one or more broken powerlines;

generating one or more corrected media items that correct the one or more broken powerlines;

generating one or more augmented media items for the training data by blending portions of the one or more corrected media items with portions of respective one or more original media items to increase a randomness of augmentation; and

training a segmentation machine-learning model based on the training data to receive a media item with one or more distracting objects and to output a segmentation mask for one or more segmented objects that correspond to the one or more distracting objects.

2. The method of claim 1 , wherein generating the one or more augmented media items includes blending the one or more corrected media items and the respective one or more original media items with a checkerboard mask.

3. The method of claim 1 , wherein generating the one or more corrected media items to correct the one or more broken powerlines includes:

modifying a local contrast in the one or more original media items to generate corresponding one or more enhanced media items.

4. The method of claim 3 , wherein the local contrast is modified using a gain curve that adds two bias curves together.

5. The method of claim 1 , wherein generating training data includes augmenting one or more of the first set of media items by applying a dilation to a segmentation mask of the one or more distracting objects.

6. The method of claim 1 , wherein the one or more distracting objects are organized into categories, the categories including at least one selected from a group of powerlines, power poles, towers, and combinations thereof.

7. The method of claim 1 , wherein training the segmentation machine-learning model comprises:

generating a first machine-learning model based on the training data; and

distilling the first machine-learning model to a trained segmentation machine-learning model by running inference on the training data that is segmented by the first machine-learning model.

8. The method of claim 1 , wherein the training data further includes synthesized images with the distracting objects added in front of outdoor environment objects.

9. A computer-implemented method to remove a distracting object from a media item, the method comprising:

receiving a media item from a user;

identifying one or more distracting objects in the media item;

providing the media item to a trained segmentation machine-learning model;

outputting, with the trained segmentation machine-learning model, a segmentation mask for the one or more distracting objects in the media item; and

inpainting a portion of the media item that matches the segmentation mask to obtain an output media item, wherein the one or more distracting objects are absent from the output media item;

wherein the trained segmentation machine-learning model is trained by generating training data by:

identifying one or more original media items in a first set of media items that include one or more broken powerlines;

generating one or more corrected media items that correct the one or more broken powerlines; and

generating one or more augmented media items for the training data by blending portions of the one or more corrected media items with portions of respective one or more original media items to increase a randomness of augmentation.

10. The method of claim 9 , wherein the one or more distracting objects are organized into categories, the categories including at least one selected from a group of powerlines, power poles, towers, and combinations thereof.

11. The method of claim 9 , further comprising providing a suggestion to a user to remove the one or more distracting objects from the media item.

12. The method of claim 9 , wherein the trained segmentation machine-learning model is trained using training data that includes the first set of media items and a second set of media items, wherein the first set of media items include distracting objects and the second set of media items include manual segmentations of the distracting objects.

13. A non-transitory computer-readable medium with instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:

generating training data that includes a first set of media items and a second set of media items, wherein the first set of media items include distracting objects and the second set of media items include manual segmentations of the distracting objects;

identifying one or more original media items in the first set of media items that include one or more broken powerlines;

generating one or more corrected media items that correct the one or more broken powerlines;

generating one or more augmented media items for the training data by blending portions of the one or more corrected media items with portions of respective one or more original media items to increase a randomness of augmentation; and

training a segmentation machine-learning model based on the training data to receive a media item with one or more distracting objects and to output a segmentation mask for one or more segmented objects that correspond to the one or more distracting objects.

14. The computer-readable medium of claim 13 , wherein generating the one or more augmented media items includes blending the one or more corrected media items and the respective one or more original media items with a checkerboard mask.

15. The computer-readable medium of claim 13 , wherein generating the one or more corrected media items to correct the one or more broken powerlines includes:

modifying a local contrast in the one or more original media items to generate corresponding one or more enhanced media items.

16. The computer-readable medium of claim 15 , wherein the local contrast is modified using a gain curve that adds two bias curves together.

17. The computer-readable medium of claim 13 , wherein generating training data includes augmenting one or more of the first set of media items by applying a dilation to a segmentation mask of the one or more distracting objects.

18. The computer-readable medium of claim 13 , wherein the one or more distracting objects are organized into categories, the categories including at least one selected from a group of powerlines, power poles, towers, and combinations thereof.

19. The computer-readable medium of claim 13 , wherein training the segmentation machine-learning model comprises:

generating a first machine-learning model based on the training data; and

distilling the first machine-learning model to a trained segmentation machine-learning model by running inference on the training data that is segmented by the first machine-learning model.

20. The computer-readable medium of claim 13 , wherein the training data further includes synthesized images with the distracting objects added in front of outdoor environment objects.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: LIBA, ORLY; KARNAD, NIKHIL; KANAZAWA, NORI; KNAAN, YAEL PRITCH; CHEN, HUIZHONG; CAI, LONGQI
To: GOOGLE LLC
Reel/Frame 061650/0979 →
Continuity (2)
Provisional Application 63257114 · Oct 18, 2021
Related Publication 20230118460A1 · Apr 20, 2023
References Cited (27)
US 20120209287A1 · Zhao · 2012 [cited by examiner]
US 20150269717A1 · Umit · 2015 [cited by examiner]
US 20180005063A1 · Chan et al. · 2018 [cited by applicant]
US 20200125927A1 · Kim · 2020 [cited by applicant]
US 20200364913A1 · Bradski · 2020 [cited by applicant]
US 20210279595A1 · Sridhar et al. · 2021 [cited by applicant]
US 20220129670A1 · Lin · 2022 [cited by examiner]
US 20220375100A1 · Qi · 2022 [cited by examiner]
JP 2014110624 · 2014 [cited by applicant]
JP 2015118596 · 2015 [cited by applicant]
JP 2019053732 · 2019 [cited by applicant]
JP 2021521982 · 2021 [cited by applicant]
Li et al.; “Transmission line detection in aerial images: An instance segmentation approach based on multitask neural networks;” Signal Processing: Image Communication 96 (2021) 116278; Elsevier B.V.; 9 pages (Year: 202… [cited by examiner]
Abdelfattah et al.; “TTPLA: An Aerial-Image Dataset for Detection and Segmentation of Transmission Towers and Power Lines;” Computer Vision Foundation; ACCV 2020; 17 pages (Year: 2020). [cited by examiner]
Zhang et al.; “Detecting Power Lines in UAV Images with Convolutional Features and Structured Constraints;” Remote Sensing; MDPI, Basel, Switzerland; 18 pages (Year: 2019). [cited by examiner]
Madaan et al.; “Wire Detection using Synthetic Data and Dilated Convolutional Networks for Unmanned Aerial Vehicles;” 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); Sep. 24-28, 2017, Va… [cited by examiner]
EPO, International Search Report for International Patent Application No. PCT/US2022/047033, Jan. 26, 2023, 3 pages. [cited by applicant]
EPO, Written Opinion for International Patent Application No. PCT/US2022/047033, Jan. 26, 2023, 8 pages. [cited by applicant]
Saurav, et al., “Power Line Segmentation in Aerial Images Using Convolutional Neural Networks”, 16th European Conference—Computer Vision—ECCV 2020, Cornell University Library, Nov. 25, 2019, pp. 623-632. [cited by applicant]
Yin, et al., “A Co-Random Walks Segmentation Method for Aerial Insulator Video Images”, 2019 12th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), IEEE, 2019, pp… [cited by applicant]
Abdelfattah, et al., “TTPLA: an aerial-image dataset for detection and segmentation of transmission towers and power lines”, Proceedings of the Asian Conference on Computer Vision, 2020, 17 pages. [cited by applicant]
Hong, et al., “Weakly Supervised Learning with Deep Convolutional Neural Networks for Semantic Segmentation: Understanding Semantic Layout of Images with Minimum Human Supervision”, IEEE Signal Processing Magazine 34.6,… [cited by applicant]
JPO, Office Action for Japanese Patent Application No. 2023-561387, Sep. 17, 2024, 8 pages. [cited by applicant]
Li, et al., “Transmission line detection in aerial images: an instance segmentation approach based on multitask neural networks”, Signal Processing: Image Communication 96: 116278, 2021. [cited by applicant]
Madaan, et al., “Wire detection using synthetic data and dilated convolutional networks for unmanned aerial vehicles”, 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2017, 8 pages. [cited by applicant]
Miyayama, et al., “An image coding method based on a foreground and a background”, IEICE Technical Report; IEICE Tech. Rep. 115.335, 2015, pp. 5-8. [cited by applicant]
Niwa, et al., “FPGA Implementation of Semantic Segmentation on LWIR Images for Autonomous Robot”, IEICE Technical Report; IEICE Tech. Rep. 120.339, 2021, pp. 101-106. [cited by applicant]