IP Library Granted Patent US 12,333,691
Granted Patent B2
US 12,333,691 · App. 18/190,500 · Granted Jun 17, 2025

Utilizing a generative machine learning model to create modified digital images from an infill semantic map

Inventors: Qing Liu (Santa Clara, CA); Jianming Zhang (Campbell, CA); Krishna Kumar Singh (San Jose, CA); Scott Cohen (Sunnyvale, CA); Zhe Lin (Fremont, CA)
Assignee: Adobe Inc.
G06T5/77G06T5/70G06T7/11G06T11/60G06V10/764G06V10/82G06V20/70G06T2200/24G06T2207/20021G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,691
App. No.
18/190,500
Granted
Jun 17, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.

Claims (57)

1. A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

generating a semantic map corresponding to a digital image, wherein the semantic map comprises semantic classifications of pixels within the digital image;

generating, utilizing a semantic map diffusion neural network, an infill semantic map from the semantic map, wherein the infill semantic map comprises semantic classifications for an infill modification that indicates a region to fill for the digital image;

generating a diffusion representation of the digital image by utilizing diffusion layers of a digital image diffusion neural network; and

generating, utilizing denoising layers of the digital image diffusion neural network, a modified digital image from the diffusion representation and the infill semantic map.

2. The non-transitory computer-readable medium of claim 1 , wherein generating the modified digital image further comprises generating the modified digital image from the digital image by utilizing the digital image diffusion neural network conditioned on the infill semantic map.

3. The non-transitory computer-readable medium of claim 1 , wherein generating the infill semantic map comprises:

generating a diffusion representation of the semantic map utilizing diffusion layers of the semantic map diffusion neural network; and

generating, utilizing denoising layers of the semantic map diffusion neural network, the infill semantic map from the diffusion representation of the semantic map.

4. The non-transitory computer-readable medium of claim 3 , wherein the operations further comprise:

receiving a semantic editing input corresponding to the semantic map; and

generating the infill semantic map from the semantic map by conditioning the denoising layers of the semantic map diffusion neural network with the semantic editing input.

5. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise:

determine an input texture for modifying the digital image; and

generating the modified digital image from the infill semantic map by conditioning the denoising layers of the digital image diffusion neural network with the input texture.

6. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise:

generating semantic classifications for the infill semantic map; and

generating the modified digital image by conditioning the denoising layers of the digital image diffusion neural network with the semantic classifications.

7. A system comprising:

one or more memory devices comprising an input digital image, a semantic map model, a generative semantic machine learning model, and a generative image machine learning model; and

one or more processors configured to cause the system to:

generate, utilizing the semantic map model, a semantic map from the input digital image, wherein the semantic map comprises semantic classifications of pixels within the input digital image;

determine an infill modification indicating a region to fill for the input digital image;

generate, utilizing the generative semantic machine learning model, an infill semantic map from the infill modification and the semantic map, wherein the infill semantic map comprises semantic classifications for the infill modification that indicates the region to fill for the input digital image; and

generate, utilizing the generative image machine learning model, a modified digital image from the infill semantic map and the input digital image.

8. The system of claim 7 , wherein the one or more processors are configured to cause the system to:

determine an input texture for the region; and

generate the modified digital image from the input digital image by utilizing a digital image diffusion neural network conditioned on the infill semantic map and the input texture.

9. The system of claim 7 , wherein the one or more processors are configured to cause the system to:

segment, utilizing a segmentation neural network, an object to remove from the input digital image; and

generate the infill semantic map from the input digital image and the object to remove from the input digital image.

10. The system of claim 9 , wherein the one or more processors are configured to cause the system to generate the infill semantic map by generating, utilizing the generative semantic machine learning model, semantic classifications of pixels within the region to be filled.

11. The system of claim 8 , wherein the one or more processors are configured to cause the system to:

generate, utilizing a semantic map diffusion neural network, a diffusion representation of the semantic map; and

generate the infill semantic map from the diffusion representation.

12. The system of claim 8 , wherein the one or more processors are configured to cause the system to generate the modified digital image by:

generating a diffusion representation of the input digital image by utilizing diffusion layers of a digital image diffusion neural network; and

generating, utilizing denoising layers of the digital image diffusion neural network, the modified digital image from the diffusion representation and the infill semantic map.

13. A computer-implemented method comprising:

determining an infill modification indicating a region to fill for a digital image;

generating, utilizing a semantic map model, a semantic map from the digital image, wherein the semantic map comprises semantic classifications of pixels within the digital image;

generating, utilizing a generative semantic machine learning model, an infill semantic map from the digital image and the infill modification, wherein the infill semantic map comprises semantic classifications for the infill modification that indicates the region to fill for the digital image; and

generating, utilizing a generative image machine learning model, a modified digital image from the infill semantic map and the digital image.

14. The computer-implemented method of claim 13 , wherein generating the infill semantic map further comprises:

generating, utilizing the generative semantic machine learning model, the infill semantic map from the semantic map and the infill modification.

15. The computer-implemented method of claim 13 , further comprising determining the region to fill by generating, utilizing a segmentation neural network, a segmented digital image from the digital image.

16. The computer-implemented method of claim 13 , wherein determining the infill modification further comprises segmenting, utilizing a segmentation neural network, an object to remove from the digital image.

17. The computer-implemented method of claim 16 , further comprising generating the infill semantic map from the digital image to remove the object from the digital image by generating, utilizing the generative semantic machine learning model, semantic classifications for pixels within the region to be filled.

18. The computer-implemented method of claim 13 , wherein generating the infill semantic map comprises:

generating, utilizing diffusion layers of a semantic map diffusion neural network, a diffusion representation of the semantic map; and

generating, utilizing denoising layers of the semantic map diffusion neural network, the infill semantic map from the diffusion representation of the semantic map.

19. The computer-implemented method of claim 13 , wherein generating the infill semantic map comprises:

receiving a semantic editing input corresponding to the semantic map; and

generating the infill semantic map from the semantic map utilizing a semantic map diffusion neural network conditioned on the semantic editing input.

20. The computer-implemented method of claim 13 , wherein generating the modified digital image comprises:

generating a diffusion representation of the digital image by utilizing diffusion layers of a digital image diffusion neural network; and

generating, utilizing denoising layers of the digital image diffusion neural network, the modified digital image from the diffusion representation and the infill semantic map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: LIU, QING; ZHANG, JIANMING; SINGH, KRISHNA KUMAR; COHEN, SCOTT; LIN, ZHE
To: ADOBE INC.
Reel/Frame 063129/0666 →
Continuity (9)
Continuation In Part 18058538 · Nov 23, 2022
Continuation In Part 18058554 · Nov 23, 2022
Continuation In Part 18058575 · Nov 23, 2022
Continuation In Part 18058601 · Nov 23, 2022
Continuation In Part 18025622 · Mar 9, 2023
Continuation In Part 18058622 · Nov 23, 2022
Continuation In Part 18058630 · Nov 23, 2022
Provisional Application 63378616 · Oct 6, 2022
Related Publication 20240135509A1 · Apr 25, 2024
References Cited (38)
US 11462040B2 · Lin · 2022 [cited by examiner]
US 12026845B2 · Pardeshi · 2024 [cited by applicant]
US 20190114748A1 · Lin et al. · 2019 [cited by applicant]
US 20200234480A1 · Volkov et al. · 2020 [cited by applicant]
US 20200394828A1 · Shukla et al. · 2020 [cited by applicant]
US 20210056348A1 · Berlin et al. · 2021 [cited by applicant]
US 20220068037A1 · Pardeshi · 2022 [cited by examiner]
US 20220207262A1 · Jeong et al. · 2022 [cited by applicant]
US 20220392133A1 · Volkov et al. · 2022 [cited by applicant]
US 20230037339A1 · Villegas et al. · 2023 [cited by applicant]
US 20230123820A1 · Wang et al. · 2023 [cited by applicant]
US 20240135511A1 · Singh · 2024 [cited by examiner]
US 20240135513A1 · Singh · 2024 [cited by examiner]
US 20240153047A1 · Smith · 2024 [cited by examiner]
US 20240169624A1 · Brandt et al. · 2024 [cited by applicant]
US 20240169701A1 · Kulal et al. · 2024 [cited by applicant]
US 20240171848A1 · Figueroa et al. · 2024 [cited by applicant]
US 20240331322A1 · Smith · 2024 [cited by applicant]
CN 113240613B · 2021 [cited by applicant]
CN 114862697A · 2022 [cited by applicant]
CN 114943656A · 2022 [cited by applicant]
GB 2606253A · 2022 [cited by applicant]
WO 2022083504A1 · 2022 [cited by applicant]
U.S. Appl. No. 18/190,544, filed Nov. 20, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,556, filed Dec. 13, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, filed Jan. 27, 2025, Office Action. [cited by applicant]
Combined Search and Examination Report as received in GB2318853.5 dated Jun. 12, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319660.3 dated Jun. 14, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319084.6 dated Jun. 25, 2024. [cited by applicant]
Wiles, 0., Koepke, A and Zisserman, A, 2018. “X2face: A network for controlling face generation using images, audio, and pose codes.” in Proceedings of the European conference on computer vision (ECCV) (pp. 690-706). [cited by applicant]
Michail Christos Doukas, Stefanos Zafeiriou, Viktorija Sharmanska HeadGAN: One-Shot Neural Head Synthesis and Editing. [cited by applicant]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer—High-Resolution Image Synthesis with Latent Diffusion Models arXiv:2112.10752 Wed, Apr. 13, 2022. [cited by applicant]
Badour AlBahar, Jingwan Lu, Jimei Yang, Zhixin Shu, Eli Shechtman, Jia-Bin Huang—Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN Badour et al., SIGGRAPH Asia 2021. [cited by applicant]
Artur Grigorev, Artem Sevastopolsky, Alexander Vakhiov, and Victor Lempitsky. Coordinate-based texture inpainting for pose-guided image generation. arXiv preprint arXiv:1811.11459, 2018. [cited by applicant]
Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Kripasindhu Sarkar, Vladislav Golyanik, Lingjie Liu, and Christian Theobalt. Style and pose control for image synthesis of humans from a single monocular view, 2021. [cited by applicant]
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [cited by applicant]
U.S. Appl. No. 18/190,556, filed Mar. 12, 2025, Notice of Allowance. [cited by applicant]
Cited By (1)
US 12,669,914