IP Library Granted Patent US 12,008,739
Granted Patent B2
US 12,008,739 · App. 17/452,529 · Granted Jun 11, 2024

Automatic photo editing via linguistic request

Inventors: Ning Xu (Milpitas, CA); Zhe Lin (Clyde Hill, WA); Franck Dernoncourt (San Jose, CA)
Assignee: ADOBE INC.
G06T5/77G06N3/08G06T5/50G06T5/90G06T11/60G10L15/22G06T2207/20081G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,008,739
App. No.
17/452,529
Granted
Jun 11, 2024
Kind
B2
Abstract

The present disclosure relates to systems and methods for automatically processing images based on a user request. In some examples, a request is divided into a retouching command (e.g., a global edit) and an inpainting command (e.g., a local edit). A retouching mask and an inpainting mask are generated to indicate areas where the edits will be applied. A photo-request attention and a multi-modal modulation process are applied to features representing the image, and a modified image that incorporates the user's request is generated using the modified features.

Claims (70)

1. A method of image processing, comprising:

identifying an image and an edit command for the image;

encoding the edit command to obtain an inpainting vector and a retouching vector, wherein the inpainting vector indicates a local edit of the image and the retouching vector indicates a global edit of the image;

generating an inpainting mask based on the inpainting vector and a retouching mask based on the retouching vector;

generating an image feature representation based on the image and the inpainting mask;

generating a modified image feature representation based on the image feature representation, the retouching mask, and an attention matrix calculated using the retouching vector; and

generating a modified image based on the modified image feature representation, wherein the modified image represents an application of the edit command to the image.

2. The method of claim 1 , further comprising:

receiving the edit command from a user via an audio input device; and

transcribing the edit command, wherein the encoding is based on the transcribed edit command.

3. The method of claim 1 , further comprising:

generating a word embedding for each of a plurality of words of the edit command;

generating a probability vector that includes an inpainting probability and a retouching probability for each of the plurality of words;

applying the inpainting probability to a corresponding word of the plurality of words to obtain an inpainting weighted word representation, wherein the inpainting vector includes the inpainting weighted word representation; and

applying the retouching probability to the corresponding word of the plurality of words to obtain a retouching weighted word representation, wherein the retouching vector includes the retouching weighted word representation.

4. The method of claim 1 , further comprising:

encoding the image to generate image features;

identifying an inpainting object based on the edit command;

generating attention weights based on the edit command in the inpainting object; and

generating the inpainting mask based on the image features and the attention weights.

5. The method of claim 1 , further comprising:

determining whether the retouching vector indicates the global edit; and

generating the retouching mask based on the determination.

6. The method of claim 1 , further comprising:

encoding the image to generate image features;

applying the inpainting mask to the image features to obtain an inpainting image features; and

performing a convolution operation on the inpainting image features to obtain the image feature representation.

7. The method of claim 1 , further comprising:

calculating the attention matrix based on the image feature representation and the retouching vector;

expanding the retouching vector based on the retouching mask to obtain an expanded retouching vector; and

weighting the expanded retouching vector based on the attention matrix to obtain a weighted retouching vector.

8. The method of claim 7 , further comprising:

generating a scaling parameter and a shifting parameter based on the weighted retouching vector; and

adding the shifting parameter to a product of the scaling parameter and the image feature representation to obtain the modified image feature representation.

9. The method of claim 1 , further comprising:

performing a convolution operation on the modified image feature representation to obtain the modified image.

10. The method of claim 1 , wherein:

the inpainting vector indicates a local edit of the image.

11. The method of claim 1 , wherein:

the retouching vector indicates a global edit of the image.

12. A method of training a neural network for image processing, comprising:

receiving training data comprising a training image, an edit command, and a ground truth image representing an application of the edit command to the training image;

encoding the edit command to obtain an inpainting vector and a retouching vector wherein the inpainting vector indicates a local edit of the training image and the retouching vector indicates a global edit of the training image;

generating an inpainting mask based on the inpainting vector and a retouching mask based on the retouching vector;

generating an image feature representation based on the training image and the inpainting mask;

generating a modified image feature representation based on the image feature representation, the retouching mask, and an attention matrix calculated using the retouching vector; and

generating a modified image based on the modified image feature representation, wherein the modified image represents the application of the edit command to the training image;

calculating a loss function based on the modified image and the ground truth image; and

training the neural network based on the loss function.

13. The method of claim 12 , further comprising:

computing an unconditional adversarial loss using a discriminator network based on the modified image and the ground truth image; and

training the neural network based on the unconditional adversarial loss.

14. The method of claim 12 , wherein:

the loss function comprises an L1 loss.

15. An apparatus for image processing, comprising:

an encoder configured to encode an edit command for an image to obtain an inpainting vector and a retouching vector, wherein the inpainting vector indicates a local edit of the image and the retouching vector indicates a global edit of the image;

a phrase conditioned grounding (PCG) network configured to generate an inpainting mask based on the inpainting vector and a retouching mask based on the retouching vector;

a convolutional network configured to generate an image feature representation based on the image and the inpainting mask;

a multi-modal modulation network configured to generate a modified image feature representation based on the image feature representation, the retouching mask, and an attention matrix calculated using the retouching vector; and

a decoder configured to generate a modified image based on the modified image feature representation, wherein the modified image represents an application of the edit command to the image.

16. The apparatus of claim 15 , further comprising:

a photo request attention (PRA) network configured to calculate the attention matrix based on the image feature representation and the retouching vector.

17. The apparatus of claim 16 , wherein:

the PRA network comprises an embedding network configured to embed the image feature representation and the retouching vector into a common embedding space.

18. The apparatus of claim 15 , further comprising:

an audio input device configured to receive the edit command from a user, and to transcribe the edit command, wherein the encoding is based on the transcribed edit command.

19. The apparatus of claim 15 , wherein:

the encoder comprises a Bi-LSTM.

20. The apparatus of claim 15 , wherein:

the convolutional network comprises a gated convolution layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2021
From: XU, NING; LIN, ZHE; DERNONCOURT, FRANCK
To: ADOBE INC.
Reel/Frame 057937/0166 →
Continuity (1)
Related Publication 20230126177A1 · Apr 27, 2023
Cited By (1)
US 12,699,727