IP Library Granted Patent US 12,322,167
Granted Patent B2
US 12,322,167 · App. 17/563,146 · Granted Jun 3, 2025

Computerized system and method for image creation using generative adversarial networks

Inventors: Yueh-Ning Ku (Sunnyvale, CA); Mikhail Kuznetsov (New York, NY); Shaunak Mishra (New York, NY); Paloma de Juan (New York, NY)
Assignee: YAHOO AD TECH LLC
G06V10/82G06N3/02G06T5/50G06T5/77G06T7/11G06V10/26G06V10/462G06V10/88
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,167
App. No.
17/563,146
Filed
Dec 28, 2021
Granted
Jun 3, 2025
Kind
B2
Art Unit
2669
USPC
382/153
Abstract

Disclosed frameworks for generating an image including a salient object and a staged background include extracting a salient object from a source image and applying a generative model to the salient object to generate the image. According to some embodiments, extracting a salient object from a source image involves using salient object detection method to identify the relevant portions of the source image corresponding to the salient object. In some embodiments, the generative model is a generative adversarial network trained using a domain relevant dataset.

Claims (52)

1. A method comprising:

retrieving a first image corresponding to a domain, the first image including a first salient object;

retrieving a second image corresponding to the domain, the second image including a second salient object;

generating a segmented background by removing the second salient object from the second image, the segmented background having removed pixels;

generating a target background by inpainting the removed pixels of the segmented background; and

generating a third image including the first salient object and the target background by aligning a shape and a center mass of a mask of the first salient object to a mask of the second salient object.

2. The method of claim 1 , wherein the step of generating a segmented background comprises detecting and segmenting the second salient object using a salient object detection model.

3. The method of claim 1 , further comprising extracting, by a salient object detection model, the first salient object from the first image.

4. The method of claim 1 , wherein the step of generating a target background by inpainting is performed by an inpainting process using at least one neural network having a loss function comprising a weighted boundary loss.

5. The method of claim 1 , further comprising animating the third image, the step of animating comprising:

generating a salient object animation frame by:

segmenting the first salient object from the third image;

translating the first salient object from a first position to a second position; and,

inpainting blank pixels in the salient object animation frame using an inpainting process;

generating an N number of salient object animation frames by repeating the steps of segmenting, translating, and inpainting; and,

generating a salient object animation comprising at least some of the salient object animation frames.

6. The method of claim 5 , wherein the inpainting process uses at least one neural network having a loss function comprising a weighted boundary loss.

7. A method comprising:

retrieving a first image corresponding to a domain, the first image including a first salient object;

retrieving a domain relevant dataset corresponding to the domain, the domain relevant dataset including a plurality of second images, the second images having second salient objects;

ranking the plurality of second images based on a similarity measure of each second image to the first image;

selecting a subset of second images from the plurality of second images based on the ranking;

generating a set of segmented backgrounds by removing the second salient objects from each second image from the subset of second images, each of the segmented background of the set of segmented backgrounds having removed pixels;

generating a set of target backgrounds by inpainting the removed pixels of each segmented background of the set of segmented backgrounds; and,

generating a set of third images each image of the set of third images including the first salient object and a target background from the set of target backgrounds by aligning a shape and a center mass of a mask of the first salient object to a mask of each of the second salient objects of the second images of the subset of second images.

8. The method of claim 7 , wherein the step of generating a set of segmented backgrounds comprises detecting and segmenting the second salient objects using a salient object detection model.

9. The method of claim 7 , further comprising extracting, by a salient object detection model, the first salient object from the first image.

10. The method of claim 7 , wherein the step of generating a set of target backgrounds by inpainting is performed by an inpainting process using at least one neural network having a loss function comprising a weighted boundary loss.

11. The method of claim 7 , further comprising animating each of the set of third images, the step of animating comprising:

generating a salient object animation frame by:

segmenting the first salient object from a third image from the set of third images;

translating the first salient object from a first position to a second position; and,

inpainting blank pixels in the salient object animation frame using an inpainting process;

generating an N number of salient object animation frames by repeating the steps of segmenting, translating, and inpainting; and,

generating a salient object animation comprising at least some of the salient object animation frames.

12. The method of claim 11 , wherein the inpainting process uses at least one neural network having a loss function comprising a weighted boundary loss.

13. A non-transitory computer-readable storage medium for tangibly storing computer program instructions capable of being executed by a computer processor, the computer program instructions defining steps of:

retrieving a first image corresponding to a domain, the first image including a first salient object;

retrieving a second image corresponding to the domain, the second image including a second salient object;

generating a segmented background by removing the second salient object from the second image, the segmented background having removed pixels;

generating a target background by inpainting the removed pixels of the segmented background; and

generating a third image including the first salient object and the target background by aligning a shape and a center mass of a mask of the first salient object to a mask of the second salient object.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the step of generating a segmented background comprises detecting and segmenting the second salient object using a salient object detection model.

15. The non-transitory computer-readable storage medium of claim 13 , further comprising extracting, by a salient object detection model, the first salient object from the first image.

16. The non-transitory computer-readable storage medium of claim 13 , wherein the step of generating a target background by inpainting is performed by an inpainting process using at least one neural network having a loss function comprising a weighted boundary loss.

17. The non-transitory computer-readable storage medium of claim 13 , further comprising:

retrieving a domain relevant dataset corresponding to the domain, the domain relevant dataset including a plurality of second images, the second images having second salient objects;

ranking the plurality of second images based on a similarity measure of each second image to the first image;

selecting a subset of second images from the plurality of second images based on the ranking;

generating a set of segmented backgrounds by removing the second salient objects from each second image from the subset of second images, each of the segmented background of the set of segmented backgrounds having removed pixels;

generating a set of target backgrounds by inpainting the removed pixels of each segmented background of the set of segmented backgrounds; and,

generating a set of fourth images each image of the set of fourth images including the first salient object and a target background from the set of target backgrounds.

Assignments (2)
CHANGE OF NAME Recorded Mar 22, 2022
From: VERIZON MEDIA INC.
To: YAHOO AD TECH LLC
Reel/Frame 059472/0328 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: KU, YUEH-NING; DE JUAN, PALOMA; MISHRA, SHAUNAK; KUZNETSOV, MIKHAIL
To: VERIZON MEDIA INC.
Reel/Frame 058486/0969 →
Continuity (1)
Related Publication 20230206614A1 · Jun 29, 2023
References Cited (20)
US 20160364419A1 · Stanton · 2016 [cited by examiner]
US 20230252072A1 · Kislyuk · 2023 [cited by examiner]
KR 102482262B1 · 2022 [cited by examiner]
Xuebin Qin et al., “U2-Net: Going deeper with nested U-structure for salient object detection”, 2020, Pattern Recognition, vol. 106, p. 1-16 (Year: 2020). [cited by examiner]
Qin et al. “U2-Net: Going Deeper with Nested U-Structure for Sallent Object Detection,” 15 pages (2020). [cited by applicant]
Wang et al., “Image Inpainting with External-internal Learning and Monochromic Bottleneck,” 15 pages (2021). [cited by applicant]
Nazeri et al. “EdgeConnect: Generative Image Inpainting with Adversarial Edge Learning,” 17 pages (2019). [cited by applicant]
Mishra et al. “Learning to Create Better Ads: Generation and Ranking Approaches for Ad Creative Refinement,” 9 bages (2020). [cited by applicant]
Zhou et al., “Understanding Consumer Journey using Attention based Recurrent Neural Networks,” Applied Data Science Track Paper, 10 pages (2019). [cited by applicant]
Mishra et al. “Guiding Creative Design in Online Advertising,” 5 pages (2019). [cited by applicant]
Yu et al., “Free-Form Image Inpainting with Gated Convolution,” 17 pages (2019). [cited by applicant]
Zhou et al., “Recommending Themes for Ad Creative Design via Visual-Linguistic Representations,” 7 pages (2020). [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets,” 9 pages (2014). [cited by applicant]
Mirza et al. “Conditional Generative Adversarial Nets,” 7 pages (2014). [cited by applicant]
Ronneberger et al., “U-Net Convolutional Networks for Biomedical Image Segmentation,” 8 pages (2015). [cited by applicant]
Szegedy et al., “Rethinking the Inception architecture for Computer Vision,” 10 pages (2015). [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks,” 17 pages (2018). [cited by applicant]
Heusel et al., “GANs Trained by a Two Time-Scale Update Rule Coverage to a Local Nash Equilibrium,” 38 pages (2018). [cited by applicant]
Hussain et al.,, “Automatic Understanding of Image and Video Adverstisements,” 11 pages (2017). [cited by applicant]
Wang et al., “High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs,” 14 pages (2018). [cited by applicant]