IP Library Granted Patent US 12669914
Granted Patent B2
US 12669914 · App. 18/190,513 · Granted Jun 30, 2026

Utilizing a generative machine learning model and graphical user interface for creating modified digital images from an infill semantic map

Inventors: Qing Liu (Santa Clara, CA); Jianming Zhang (Campbell, CA); Krishna Kumar Singh (San Jose, CA); Scott Cohen (Sunnyvale, CA); Zhe Lin (Fremont, CA)
Assignee: Adobe Inc.
G06T5/77G06T7/11G06T7/40G06V10/25G06V10/764G06V10/82G06T2207/20104
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12669914
App. No.
18/190,513
Granted
Jun 30, 2026
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.

Claims (72)

1 . A computer-implemented method comprising:

providing, for display via a user interface of a client device, an input digital image and a semantic map of the input digital image;

receiving, via the user interface, a selection of a selectable expansion option that indicates an infill modification comprising a region corresponding to the input digital image to fill;

in response to receiving, via the user interface, a selection of a semantic completion option, providing, for display via the user interface, an infill semantic map comprising semantic classifications of pixels within the region of the input digital image to fill in accordance with the selection of the selectable expansion option; and

in response to an image completion interaction via the user interface, generating a modified digital image comprising infilled pixels for the region according to the infill semantic map.

2 . The computer-implemented method of claim 1 , further comprising:

receiving the selectable expansion option of the infill modification by receiving a user selection of an expanded region beyond the input digital image; and

providing the infill semantic map by providing, for display via the user interface, the infill semantic map of the input digital image within an expanded digital image frame corresponding to the expanded region indicated by the selectable expansion option.

3 . The computer-implemented method of claim 2 , further comprising generating, in response to the selection of the semantic completion option indicating an infill completion via the user interface, the infill semantic map, wherein the semantic classifications of the pixels extends into the expanded digital image frame corresponding to the expanded region indicated by the selectable expansion option.

4 . The computer-implemented method of claim 1 , wherein:

providing the infill semantic map comprises providing, for display via the user interface, a plurality of infill semantic maps; and

in response to user selection of the infill semantic map from the plurality of infill semantic maps, generating the modified digital image.

5 . The computer-implemented method of claim 4 , wherein providing the plurality of infill semantic maps comprises:

providing, for display via the user interface, the infill semantic map comprising the semantic classifications within a semantic boundary; and

providing, for display via the user interface, an additional infill semantic map comprising additional semantic classifications within an additional semantic boundary.

6 . The computer-implemented method of claim 1 , wherein generating the modified digital image comprises:

generating a plurality of modified digital images according to the infill semantic map;

providing, for display via the user interface, the plurality of modified digital images; and

receiving a user selection of the modified digital image from the plurality of modified digital images generated according to the infill semantic map.

7 . The computer-implemented method of claim 1 further comprising:

receiving, via the user interface, a semantic editing input for a region of the semantic map;

generating the infill semantic map guided by the semantic editing input; and

providing, for display via the user interface, the infill semantic map guided by the semantic editing input.

8 . The computer-implemented method of claim 7 , further comprising:

providing, for display via the user interface, a semantic editing tool; and

determining the semantic editing input comprises receiving an input semantic classification or an input semantic boundary based on user interaction with the semantic editing tool.

9 . The computer-implemented method of claim 1 , further comprising:

providing, via the user interface, a segmentation option to apply a segmentation model to the input digital image;

in response to receiving a selection of the segmentation option, generating a segmented digital image; and

providing, via the user interface, the segmented digital image, wherein the segmented digital image indicates the region of the input digital image to fill.

10 . The computer-implemented method of claim 1 , further comprising:

providing, for display via the user interface, a selectable texture option;

in response to receiving a user interaction with the selectable texture option, identifying an input texture; and

generating the modified digital image by filling the region corresponding to the indicated infill modification guided by the selected texture option.

11 . The computer-implemented method of claim 1 , further comprising:

providing, for display via the user interface, a diffusion iteration option; and

in response to user interaction with the diffusion iteration option:

determining a number of diffusion iterations; and

generating the modified digital image by utilizing a diffusion neural network comprising the number of diffusion iterations.

12 . A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

providing, for display via a user interface of a client device, an input digital image and a semantic map of the input digital image;

receiving, via a semantic editing tool in the user interface, an indication of a semantic editing input for a region of the semantic map;

in response to receiving a selection of a semantic completion option, generating, utilizing a generative semantic machine learning model, an infill semantic map from the semantic map and the indication of the semantic editing input;

providing, for display via the user interface, the infill semantic map comprising the semantic editing input; and

in response to user interaction with an image completion element, generate a modified digital image comprising infilled pixels for the region.

13 . The non-transitory computer-readable medium of claim 12 , wherein receiving the indication of the semantic editing input further comprises receiving an input semantic classification or an input semantic boundary.

14 . The non-transitory computer-readable medium of claim 12 , wherein receiving the indication of the semantic editing input further comprises:

providing, for display via the user interface, a semantic editing tool comprising a drawing tool to modify the infill semantic map; and

receiving, via the semantic editing tool, an indication of an input semantic classification or an input semantic boundary received via the drawing tool.

15 . The non-transitory computer-readable medium of claim 12 , further comprising:

providing, for display via the user interface, a diffusion iteration option; and

in response to user interaction with the diffusion iteration option:

determining a number of diffusion iterations; and

generating the infill semantic map by utilizing a diffusion neural network comprising the number of diffusion iterations guided by the indication of the semantic editing input.

16 . The non-transitory computer-readable medium of claim 12 , further comprises:

providing, via the user interface, a segmentation option to apply a segmentation model to the input digital image;

in response to receiving a selection of the segmentation option, generating a segmented digital image; and

generating the modified digital image from the segmented digital image indicating the region of the input digital image to fill and the semantic editing input.

17 . A system comprising:

one or more memory devices comprising an input digital image and a semantic map; and

one or more processors configured to cause the system to:

provide, for display via a user interface of a client device, the input digital image and the semantic map of the input digital image;

receive, via the user interface, a selection of a selectable expansion option that indicates an infill modification comprising a region corresponding to the input digital image to fill;

in response to receiving, via the user interface, a selection of a semantic completion option, provide, for display via the user interface, an infill semantic map comprising semantic classifications of pixels within the region of the input digital image to fill in accordance with the selection of the selectable expansion option; and

in response to an image completion interaction via the user interface, generate a modified digital image comprising infilled pixels for the region according to the infill semantic map.

18 . The system of claim 17 , wherein the one or more processors are configured to cause the system to:

receive the selectable expansion option of the infill modification by receiving a user selection of an expanded region beyond the input digital image; and

provide the infill semantic map by providing, for display via the user interface, the infill semantic map of the input digital image within an expanded digital image frame corresponding to the expanded region indicated by the selectable expansion option.

19 . The system of claim 18 , wherein the one or more processors are configured to cause the system to generate, in response to the selection of the semantic completion option indicating an infill completion via the user interface, the infill semantic map, wherein the semantic classifications of the pixels extends into the expanded digital image frame corresponding to the expanded region indicated by the selectable expansion option.

20 . The system of claim 17 , wherein the one or more processors are configured to cause the system to:

provide, for display via the user interface, a plurality of infill semantic maps; and

in response to user selection of the infill semantic map from the plurality of infill semantic maps, generate the modified digital image.