IP Library › Granted Patent US 11,631,156
Granted Patent B2
US 11,631,156 · App. 17/088,120 · Granted Apr 18, 2023

Image generation and editing with latent transformation detection

Inventors: Mayank Singh (Noida, IN); Parth Patel (Vadodara, IN); Nupur Kumari (Noida, IN); Balaji Krishnamurthy (Noida, IN)
Assignee: Adobe Inc.
G06T3/0075G06K9/6256G06N3/04G06N3/08G06T3/4007G06T7/73G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,631,156
App. No.
17/088,120
Granted
Apr 18, 2023
Kind
B2
Abstract

This disclosure includes technologies for image processing, particularly for image generation and editing in a configurable semantic direction. A generative adversarial network is trained with an auxiliary network with an auxiliary task that is designed to disentangle the latent space of the generative adversarial network. Resultantly, a new type of GAN is created to improve image generation or editing in both conditional and unconditional settings.

Claims (35)

1. A computer-implemented method, comprising:

generating two GAN-induced transformations by adding a perturbation code to two latent codes in a latent space of a generative adversarial network (GAN);

determining, at an auxiliary network (AN) coupled to the GAN, a difference between respective feature representations of the two GAN-induced transformations; and

training the GAN based on the difference between the respective feature representations of the two GAN-induced transformations, wherein the training comprises disentangling the latent space of the GAN based on a prediction of the two GAN-induced transformations being caused by the perturbation code.

2. The method of claim 1 , wherein the training comprises training the GAN to generate a same type of image transformation on a plurality of training images when the perturbation code is added to different latent codes in the latent space of the GAN.

3. The method of claim 1 , wherein the training comprises training the AN and the GAN together with a self-supervised task, wherein the self-supervised task comprises adding different latent space perturbations to induce different GAN-induced latent transformations.

4. The method of claim 1 , wherein the training comprises training the AN and the GAN together with an adversarial loss and an auxiliary loss, wherein the auxiliary loss is configured to cause a disentanglement of semantics encoded in the latent space of the GAN.

5. The method of claim 4 , wherein the auxiliary loss is determined based on a prediction of whether a pair of GAN-induced latent transformations are caused by a same perturbation in the latent space.

6. The method of claim 1 , further comprising:

sampling a distribution to obtain a plurality of latent codes and a plurality of perturbation codes;

generating a pair of baseline images from the GAN based on a pair of latent codes of the plurality of latent codes; and

generating a pair of transformed images from the GAN based on additions of respective perturbation codes to the pair of latent codes.

7. The method of claim 6 , further comprising:

computing an auxiliary loss of the AN based on respective differences between the pair of baseline images and the pair of transformed images.

8. The method of claim 7 , further comprising:

updating first learnable parameters of the AN based on the auxiliary loss of the AN; and

updating second learnable parameters of the GAN based on the auxiliary loss of the AN and an adversarial loss of the GAN.

9. The method of claim 1 , further comprising:

in response to a request for a semantic edit, generating an image transformation corresponding to the semantic edit.

10. A computing system, comprising:

a generative adversarial network (GAN) including one or more generators connected with one or more discriminators;

an auxiliary network (AN), connected to the one or more discriminators of the GAN; and

a processor, operationally connected to the GAN and the AN, being configured to execute instructions to train the AN to cause a disentanglement of a latent space of the one or more generators based on a prediction of whether two GAN-induced transformations are caused by a common perturbation in the latent space.

11. The computing system of claim 10 , wherein the AN is further trained to promote the one or more generators to generate images such that the two GAN-induced transformations are distinguishable at a feature representation level.

12. The computing system of claim 10 , wherein the AN is configured to determine a difference between respective feature representations of two GAN-induced transformations.

13. The computing system of claim 12 , wherein the GAN and the AN are trained with an adversarial loss and an auxiliary loss, wherein the auxiliary loss is based on the difference between the respective feature representations of the two GAN-induced transformations.

14. The computing system of claim 10 , wherein the common perturbation comprises a perturbation code being added to a latent code in the latent space.

15. The computing system of claim 14 , wherein the GAN is trained to generate a same type of image transformation on a plurality of training images when the perturbation code is added to different latent codes in the latent space of the GAN.

16. The computing system of claim 10 , wherein the one or more generators are trained to generate, via the processor, image variations based on a plurality of source images.

17. The computing system of claim 10 , wherein the one or more generators are trained to generate, via the processor, an image transformation at a semantic direction in response to a semantic edit request.

18. The computing system of claim 17 , wherein the semantic edit request comprises altering an image of a person in a semantic direction of age, expression, or pose.

19. A computer system, comprising:

means for generating two image transformations;

means for determining a difference between respective feature representations of the two image transformations; and

means for training a generative adversarial network (GAN) and an auxiliary network (AN) together with an auxiliary loss computed based on the difference between the respective feature representations of the two image transformations, wherein the auxiliary loss is configured to cause a disentanglement of semantics encoded in a latent space of the GAN.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2020
From: SINGH, MAYANK; PATEL, PARTH; KUMARI, NUPUR; KRISHNAMURTHY, BALAJI
To: ADOBE INC.
Reel/Frame 054359/0911 →
Continuity (1)
Related Publication 20220138897A1 · May 5, 2022
Cited By (3)
US 12,367,546 US 12,488,509 US 12,677,069