IP Library › Granted Patent US 11,670,023
Granted Patent B2
US 11,670,023 · App. 17/007,693 · Granted Jun 6, 2023

Artificial intelligence techniques for performing image editing operations inferred from natural language requests

Inventors: Ning Xu (Milpitas, CA); Trung Bui (San Jose, CA); Jing Shi (Rochester, NY); Franck Dernoncourt (Sunnyvale, CA)
Assignee: Adobe Inc.
G06T11/60G10L15/16G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,670,023
App. No.
17/007,693
Granted
Jun 6, 2023
Kind
B2
Abstract

This disclosure involves executing artificial intelligence models that infer image editing operations from natural language requests spoken by a user. Further, this disclosure performs the inferred image editing operations using inferred parameters for the image editing operations. Systems and methods may be provided that infer one or more image editing operations from a natural language request associated with a source image, locate areas of the source that are relevant to the one or more image editing operations to generate image masks, and performing the one or more image editing operations to generate a modified source image.

Claims (60)

1. A system comprising:

one or more processors; and

a non-transitory computer-readable medium communicatively coupled to the one or more processors and storing program code executable by the one or more processors, the program code implementing a natural language-based image editor comprising:

an operation classifier model configured to infer an image editing operation from a natural language request associated with a source image;

a grounding model configured to (a) locate an object or region of the source image that is inferred to correspond to the image editing operation and (b) generate an image mask for the object or region; and

an operation modular network configured to generate a modified source image by using a submodule to perform the image editing operation, the submodule configured to infer one or more parameters used for performing the image editing operation.

2. The system of claim 1 , wherein the operation classifier model comprises:

a first neural network configured to encode the source image;

a second neural network configured to encode the natural language request; and

an operation classifier configured to infer the image editing operation in response to receiving the encoded source image and the encoded natural language request as an input.

3. The system of claim 1 , wherein the grounding model further comprises:

an attention layer configured to classify the image editing operation as a local operator or a global operator, wherein the local operator applies to a local area of the source image, and wherein the global operator applies to an entirety of the source image; and

a language attention network configured to ground the image editing operation to the object or region within the source image when the image editing operation is classified as a local operator.

4. The system of claim 3 , wherein the language attention network further comprises:

a subject module configured to generate a subject attention weight indicating a relevance between the object or region of the source image and a subject depicted in the source image;

a location module configured to generate a location attention weight indicating a relevance between the object or region of the source image and a location of the subject depicted in the source image;

a relationship module configured to generate a relationship attention weight indicating a relevance between the object or region of the source image and another object depicted within the source image; and

an operation attention module configured to locate the object or region within the source image by modifying the subject attention weight, the location attention weight, and the relationship attention weight using an operation attention weight.

5. The system of claim 1 , wherein the submodule for the image editing operation includes a differentiable filter.

6. The system of claim 5 , wherein the differentiable filter modifies the source image based on the source image, the natural language request, and the generated image mask, and wherein the modified source image is modified according to the image editing operation.

7. The system of claim 1 , wherein the natural language request is generated based on audio data representing a natural language expression spoken by a user.

8. A computer-implemented method comprising:

retrieving a source image and a natural language request;

inferring, using an operation classifier model, an image editing operation from the natural language request;

generating, using a grounding model, an image mask for an object or region of the source image that is inferred to correspond to the image editing operation;

performing, using a submodule of an operation modular network, the image editing operation on the source image, wherein the submodule is configured to infer one or more parameters used for performing the image editing operation; and

outputting a modified source image.

9. The computer-implemented method of claim 8 , wherein inferring the image editing operation from the natural language request further comprises:

encoding the source image using a first neural network;

encoding the natural language request using a second neural network; and

inferring, using an operation classifier, the image editing operation from the natural language request in response to receiving the encoded source image and the encoded natural language request as an input.

10. The computer-implemented method of claim 8 , wherein generating the image mask further comprises:

classifying, using an operation region model, the image editing operation as a local operator or a global operator, wherein the local operator applies to a local area of the source image, and wherein the global operator applies to an entirety of the source image; and

grounding, using a language attention network, the image editing operation to the object or region within the source image when the image editing operation is classified as a local operator.

11. The computer-implemented method of claim 10 , further comprising:

generating a subject attention weight indicating a relevance between the object or region of the source image and a subject depicted in the source image;

generating a location attention weight indicating a relevance between the object or region of the source image and a location of the subject depicted in the source image;

generating a relationship attention weight indicating a relevance between the object or region of the source image and another object depicted within the source image; and

locating the object or region within the source image by modifying the subject attention weight, the location attention weight, and the relationship attention weight using an operation attention weight.

12. The computer-implemented method of claim 8 , further comprising generating the modified source image, wherein the modified source image is modified according to the image editing operation.

13. The computer-implemented method of claim 12 , further comprising:

generating the modified source image in response to receiving the source image, the natural language request, and the generated image mask.

14. The computer-implemented method of claim 8 , wherein the natural language request is generated based on audio data representing a natural language expression spoken by a user.

15. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause a processing apparatus to perform operations including:

retrieving a source image and a natural language request;

inferring an image editing operation from the natural language request in response to receiving an encoded source image and an encoded natural language request as an input;

classifying an image editing operation as a local operator or a global operator, wherein the local operator applies to a local area of the source image, and wherein the global operator applies to an entirety of the source image;

grounding the image editing operation to an object or region within the source image to generate an image mask when the image editing operation is classified as a local operator;

editing the source image based on the image editing operation as inferred from the natural language request to produce a modified source image; and

generating the modified source image in response to receiving the source image, the natural language request, and the generated image mask, and wherein the modified source image is modified according to the image editing operation.

16. The non-transitory machine-readable storage medium of claim 15 , wherein editing the source image further comprises:

encoding the source image using a first neural network; and

encoding the natural language request using a second neural network.

17. The non-transitory machine-readable storage medium of claim 15 , wherein editing the source image further comprises:

generating a subject attention weight indicating a relevance between the object or region of the source image and a subject depicted in the source image; and

generating a location attention weight indicating a relevance between the object or region of the source image and a location of the subject depicted in the source image.

18. The non-transitory machine-readable storage medium of claim 17 , wherein editing the source image further comprises:

generating a relationship attention weight indicating a relevance between the object or region of the source image and another object depicted within the source image; and

locating the object or region within the source image by modifying the subject attention weight, the location attention weight, and the relationship attention weight using an operation attention weight.

19. The non-transitory machine-readable storage medium of claim 15 , wherein the operations further comprise performing the image editing operation by inferring one or more parameters used for performing the image editing operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2020
From: XU, NING; BUI, TRUNG; SHI, JING; DERNONCOURT, FRANCK
To: ADOBE INC.
Reel/Frame 053645/0514 →
Continuity (1)
Related Publication 20220067992A1 · Mar 3, 2022
Cited By (6)
US 12,210,835 US 12,405,700 US 12,542,862 US 12,602,154 US 12,619,303 US 12,737,107