IP Library Granted Patent US 12,669,914
Granted Patent B2
US 12,669,914 · App. 18/190,513 · Granted Jun 30, 2026

Utilizing a generative machine learning model and graphical user interface for creating modified digital images from an infill semantic map

Inventors: Qing Liu (Santa Clara, CA); Jianming Zhang (Campbell, CA); Krishna Kumar Singh (San Jose, CA); Scott Cohen (Sunnyvale, CA); Zhe Lin (Fremont, CA)
Assignee: Adobe Inc.
G06T5/77G06T7/11G06T7/40G06V10/25G06V10/764G06V10/82G06T2207/20104
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,669,914
App. No.
18/190,513
Granted
Jun 30, 2026
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.

Claims (72)

1 . A computer-implemented method comprising:

providing, for display via a user interface of a client device, an input digital image and a semantic map of the input digital image;

receiving, via the user interface, a selection of a selectable expansion option that indicates an infill modification comprising a region corresponding to the input digital image to fill;

in response to receiving, via the user interface, a selection of a semantic completion option, providing, for display via the user interface, an infill semantic map comprising semantic classifications of pixels within the region of the input digital image to fill in accordance with the selection of the selectable expansion option; and

in response to an image completion interaction via the user interface, generating a modified digital image comprising infilled pixels for the region according to the infill semantic map.

2 . The computer-implemented method of claim 1 , further comprising:

receiving the selectable expansion option of the infill modification by receiving a user selection of an expanded region beyond the input digital image; and

providing the infill semantic map by providing, for display via the user interface, the infill semantic map of the input digital image within an expanded digital image frame corresponding to the expanded region indicated by the selectable expansion option.

3 . The computer-implemented method of claim 2 , further comprising generating, in response to the selection of the semantic completion option indicating an infill completion via the user interface, the infill semantic map, wherein the semantic classifications of the pixels extends into the expanded digital image frame corresponding to the expanded region indicated by the selectable expansion option.

4 . The computer-implemented method of claim 1 , wherein:

providing the infill semantic map comprises providing, for display via the user interface, a plurality of infill semantic maps; and

in response to user selection of the infill semantic map from the plurality of infill semantic maps, generating the modified digital image.

5 . The computer-implemented method of claim 4 , wherein providing the plurality of infill semantic maps comprises:

providing, for display via the user interface, the infill semantic map comprising the semantic classifications within a semantic boundary; and

providing, for display via the user interface, an additional infill semantic map comprising additional semantic classifications within an additional semantic boundary.

6 . The computer-implemented method of claim 1 , wherein generating the modified digital image comprises:

generating a plurality of modified digital images according to the infill semantic map;

providing, for display via the user interface, the plurality of modified digital images; and

receiving a user selection of the modified digital image from the plurality of modified digital images generated according to the infill semantic map.

7 . The computer-implemented method of claim 1 further comprising:

receiving, via the user interface, a semantic editing input for a region of the semantic map;

generating the infill semantic map guided by the semantic editing input; and

providing, for display via the user interface, the infill semantic map guided by the semantic editing input.

8 . The computer-implemented method of claim 7 , further comprising:

providing, for display via the user interface, a semantic editing tool; and

determining the semantic editing input comprises receiving an input semantic classification or an input semantic boundary based on user interaction with the semantic editing tool.

9 . The computer-implemented method of claim 1 , further comprising:

providing, via the user interface, a segmentation option to apply a segmentation model to the input digital image;

in response to receiving a selection of the segmentation option, generating a segmented digital image; and

providing, via the user interface, the segmented digital image, wherein the segmented digital image indicates the region of the input digital image to fill.

10 . The computer-implemented method of claim 1 , further comprising:

providing, for display via the user interface, a selectable texture option;

in response to receiving a user interaction with the selectable texture option, identifying an input texture; and

generating the modified digital image by filling the region corresponding to the indicated infill modification guided by the selected texture option.

11 . The computer-implemented method of claim 1 , further comprising:

providing, for display via the user interface, a diffusion iteration option; and

in response to user interaction with the diffusion iteration option:

determining a number of diffusion iterations; and

generating the modified digital image by utilizing a diffusion neural network comprising the number of diffusion iterations.

12 . A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

providing, for display via a user interface of a client device, an input digital image and a semantic map of the input digital image;

receiving, via a semantic editing tool in the user interface, an indication of a semantic editing input for a region of the semantic map;

in response to receiving a selection of a semantic completion option, generating, utilizing a generative semantic machine learning model, an infill semantic map from the semantic map and the indication of the semantic editing input;

providing, for display via the user interface, the infill semantic map comprising the semantic editing input; and

in response to user interaction with an image completion element, generate a modified digital image comprising infilled pixels for the region.

13 . The non-transitory computer-readable medium of claim 12 , wherein receiving the indication of the semantic editing input further comprises receiving an input semantic classification or an input semantic boundary.

14 . The non-transitory computer-readable medium of claim 12 , wherein receiving the indication of the semantic editing input further comprises:

providing, for display via the user interface, a semantic editing tool comprising a drawing tool to modify the infill semantic map; and

receiving, via the semantic editing tool, an indication of an input semantic classification or an input semantic boundary received via the drawing tool.

15 . The non-transitory computer-readable medium of claim 12 , further comprising:

providing, for display via the user interface, a diffusion iteration option; and

in response to user interaction with the diffusion iteration option:

determining a number of diffusion iterations; and

generating the infill semantic map by utilizing a diffusion neural network comprising the number of diffusion iterations guided by the indication of the semantic editing input.

16 . The non-transitory computer-readable medium of claim 12 , further comprises:

providing, via the user interface, a segmentation option to apply a segmentation model to the input digital image;

in response to receiving a selection of the segmentation option, generating a segmented digital image; and

generating the modified digital image from the segmented digital image indicating the region of the input digital image to fill and the semantic editing input.

17 . A system comprising:

one or more memory devices comprising an input digital image and a semantic map; and

one or more processors configured to cause the system to:

provide, for display via a user interface of a client device, the input digital image and the semantic map of the input digital image;

receive, via the user interface, a selection of a selectable expansion option that indicates an infill modification comprising a region corresponding to the input digital image to fill;

in response to receiving, via the user interface, a selection of a semantic completion option, provide, for display via the user interface, an infill semantic map comprising semantic classifications of pixels within the region of the input digital image to fill in accordance with the selection of the selectable expansion option; and

in response to an image completion interaction via the user interface, generate a modified digital image comprising infilled pixels for the region according to the infill semantic map.

18 . The system of claim 17 , wherein the one or more processors are configured to cause the system to:

receive the selectable expansion option of the infill modification by receiving a user selection of an expanded region beyond the input digital image; and

provide the infill semantic map by providing, for display via the user interface, the infill semantic map of the input digital image within an expanded digital image frame corresponding to the expanded region indicated by the selectable expansion option.

19 . The system of claim 18 , wherein the one or more processors are configured to cause the system to generate, in response to the selection of the semantic completion option indicating an infill completion via the user interface, the infill semantic map, wherein the semantic classifications of the pixels extends into the expanded digital image frame corresponding to the expanded region indicated by the selectable expansion option.

20 . The system of claim 17 , wherein the one or more processors are configured to cause the system to:

provide, for display via the user interface, a plurality of infill semantic maps; and

in response to user selection of the infill semantic map from the plurality of infill semantic maps, generate the modified digital image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: LIU, QING; ZHANG, JIANMING; SINGH, KRISHNA KUMAR; COHEN, SCOTT; LIN, ZHE
To: ADOBE INC.
Reel/Frame 063124/0628 →
Continuity (8)
Continuation In Part 18058622 · Nov 23, 2022
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18058630 · Nov 23, 2022
Continuation In Part 18058575 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Provisional Application 63378616 · Oct 6, 2022
Related Publication 20240135510A1 · Apr 25, 2024
References Cited (68)
US 11462040B2 · Lin et al. · 2022 [cited by applicant]
US 11972599B2 · Xu · 2024 [cited by examiner]
US 12026845B2 · Pardeshi · 2024 [cited by applicant]
US 12260530B2 · Kumar Singh · 2025 [cited by examiner]
US 12333691B2 · Liu · 2025 [cited by examiner]
US 12347080B2 · Singh · 2025 [cited by examiner]
US 20170358059A1 · Zhang · 2017 [cited by examiner]
US 20190114748A1 · Lin et al. · 2019 [cited by applicant]
US 20190228508A1 · Price · 2019 [cited by examiner]
US 20190236394A1 · Price · 2019 [cited by examiner]
US 20190340462A1 · Pao · 2019 [cited by examiner]
US 20190355102A1 · Lin · 2019 [cited by examiner]
US 20200074674A1 · Guo · 2020 [cited by examiner]
US 20200234480A1 · Volkov et al. · 2020 [cited by applicant]
US 20200342574A1 · Meinke · 2020 [cited by examiner]
US 20200394828A1 · Shukla et al. · 2020 [cited by applicant]
US 20210056348A1 · Berlin et al. · 2021 [cited by applicant]
US 20210236936A1 · Tureaud · 2021 [cited by examiner]
US 20210264207A1 · Smith et al. · 2021 [cited by applicant]
US 20210357684A1 · Amirghodsi · 2021 [cited by examiner]
US 20220068037A1 · Pardeshi · 2022 [cited by applicant]
US 20220101532A1 · Shi · 2022 [cited by examiner]
US 20220129682A1 · Tang · 2022 [cited by examiner]
US 20220207262A1 · Jeong et al. · 2022 [cited by applicant]
US 20220392133A1 · Volkov et al. · 2022 [cited by applicant]
US 20230037339A1 · Villegas et al. · 2023 [cited by applicant]
US 20230110206A1 · Karras et al. · 2023 [cited by applicant]
US 20230123820A1 · Wang et al. · 2023 [cited by applicant]
US 20230319223A1 · Naruiec et al. · 2023 [cited by applicant]
US 20230410447A1 · Cheng et al. · 2023 [cited by applicant]
US 20240135511A1 · Singh et al. · 2024 [cited by applicant]
US 20240135513A1 · Singh et al. · 2024 [cited by applicant]
US 20240135572A1 · Singh · 2024 [cited by examiner]
US 20240153047A1 · Smith et al. · 2024 [cited by applicant]
US 20240169624A1 · Brandt et al. · 2024 [cited by applicant]
US 20240169701A1 · Kulal et al. · 2024 [cited by applicant]
US 20240171848A1 · Figueroa et al. · 2024 [cited by applicant]
US 20240249459A1 · Bradley et al. · 2024 [cited by applicant]
US 20240281918A1 · Shin · 2024 [cited by examiner]
US 20240331322A1 · Smith · 2024 [cited by applicant]
CN 113240613B · 2021 [cited by applicant]
CN 114862697A · 2022 [cited by applicant]
CN 114943656A · 2022 [cited by applicant]
GB 2606253A · 2022 [cited by applicant]
WO 2022083504A1 · 2022 [cited by applicant]
Michail Christos Doukas, Stefanos Zafeiriou, Viktorija Sharmanska HeadGAN: One-Shot Neural Head Synthesis and Editing. [cited by applicant]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer—High-Resolution Image Synthesis with Latent Diffusion Models arXiv:2112.10752 Wed, Apr. 13, 2022. [cited by applicant]
Badour AlBahar, Jingwan Lu, Jimei Yang, Zhixin Shu, Eli Shechtman, Jia-Bin Huang—Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN Badour et al., SIGGRAPH ASIA 2021. [cited by applicant]
Artur Grigorev, Artem Sevastopolsky, Alexander Vakhiov, and Victor Lempitsky. Coordinate-based texture inpainting for pose-guided image generation. arXiv preprint arXiv:1811.11459, 2018. [cited by applicant]
Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Kripasindhu Sarkar, Vladislav Golyanik, Lingjie Liu, and Christian Theobalt. Style and pose control for image synthesis of humans from a single monocular view, 2021. [cited by applicant]
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [cited by applicant]
Combined Search and Examination Report as received in GB2318853.5 dated Jun. 12, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319660.3 dated Jun. 14, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319084.6 dated Jun. 25, 2024. [cited by applicant]
Wiles, 0., Koepke, A and Zisserman, A, 2018. “X2face: A network for controlling face generation using images, audio, and pose codes.” in Proceedings of the European conference on computer vision (ECCV) (pp. 690-706). [cited by applicant]
U.S. Appl. No. 18/190,544, Nov. 20, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,556, Dec. 13, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, Jan. 27, 2025, Office Action. [cited by applicant]
Qiao, Fengchun, et al. “Geometry-contrastive gan for facial expression transfer.” arXiv preprint arXiv: 1802.01822 (2018). (Year: 2018). [cited by applicant]
Chen, Yajing, et al. “Self-supervised learning of detailed 3d face reconstruction.” IEEE Transactions on Image Processing 29 (2020): 8696-8705. (Year: 2020). [cited by applicant]
Screen captures from YouTube video clip entitled “How to Use xpression camera—For Video Chat, Vlogging, Live Streaming, Content Creation, Gaming,” 4 pages, uploaded on Feb. 8, 2023 by user “EmbodyMe”. Retrieved from Int… [cited by applicant]
U.S. Appl. No. 18/190,500, Feb. 26, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,500, Apr. 15, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,556, Mar. 12, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,673, Mar. 13, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, Mar. 12, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, May 7, 2025, Office Action. [cited by applicant]