IP Library Granted Patent US 12,430,829
Granted Patent B2
US 12,430,829 · App. 18/057,851 · Granted Sep 30, 2025

Multi-modal image editing

Inventors: Shaoan Xie (Pittsburgh, PA); Zhifei Zhang (San Jose, CA); Zhe Lin (Clyde Hill, WA); Tobias Hinz (Campbell, CA)
Assignee: ADOBE INC.
G06T11/60G06T7/11G06T11/001G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,829
App. No.
18/057,851
Filed
Nov 22, 2022
Granted
Sep 30, 2025
Kind
B2
Art Unit
2612
USPC
345/629
Abstract

Systems and methods for multi-modal image editing are provided. In one aspect, a system and method for multi-modal image editing includes identifying an image, a prompt identifying an element to be added to the image, and a mask indicating a first region of the image for depicting the element. The system then generates a partially noisy image map that includes noise in the first region and image features from the image in a second region outside the first region. A diffusion model generates a composite image map based on the partially noisy image map and the prompt. In some cases, the composite image map includes the target element in the first region that corresponds to the mask.

Claims (55)

1. A method for image editing, comprising:

identifying, by a processor, an image, a prompt identifying an element to be added to the image, a preliminary mask, and a mask precision indicator;

expanding, by the processor, the preliminary mask based on the mask precision indicator to obtain a mask indicating a first region of the image for depicting the element;

generating, by the processor, a partially noisy image map that includes noise in the first region and image features from the image in a second region outside the first region; and

generating, by the processor, a composite image map using a diffusion model based on the partially noisy image map and the prompt, wherein the composite image map includes the element in the first region that corresponds to the mask.

2. The method of claim 1 , further comprising:

receiving a brush tool input from a user; and

generating the preliminary mask based on the brush tool input.

3. The method of claim 1 , further comprising:

encoding the image to obtain an image feature map, wherein the partially noisy image map is based on the image feature map; and

decoding the composite image map to obtain a composite image.

4. The method of claim 1 , further comprising:

generating a predicted mask for the element based on the composite image map; and

combining the image and the composite image map based on the predicted mask.

5. The method of claim 1 , further comprising:

providing the mask as an input to the diffusion model.

6. The method of claim 1 , further comprising:

generating intermediate denoising data using the diffusion model; and

combining the intermediate denoising data with the partially noisy image map to obtain an intermediate composite image map.

7. A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device configured to perform operations comprising:

identifying an image, a prompt identifying an element to be added to the image, a preliminary mask, and a mask precision indicator;

expanding the preliminary mask based on the mask precision indicator to obtain a mask indicating a first region of the image for depicting the element;

generating a partially noisy image map that includes noise in the first region and image features from the image in a second region outside the first region; and

generating a composite image map using a diffusion model based on the partially noisy image map and the prompt, wherein the composite image map includes the element in the first region that corresponds to the mask.

8. The system of claim 7 , further comprising:

a user interface that includes a brush tool for generating the mask and a precision control element for obtaining a mask precision indicator.

9. The system of claim 7 , wherein:

the diffusion model is further configured to generate image features based on the image, and decode the composite image map to obtain a composite image.

10. The system of claim 7 , wherein:

the diffusion model comprises a U-Net architecture.

11. The system of claim 7 , wherein:

the diffusion model comprises a channel corresponding to the mask.

12. The system of claim 7 , the processing device being further configured to perform operations comprising:

generating a predicted mask based on the composite image map using a mask network.

13. A non-transitory computer readable medium storing code for data processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

identifying an image, a prompt identifying an element to be added to the image, a preliminary mask, and a mask precision indicator;

expanding the preliminary mask based on the mask precision indicator to obtain a mask indicating a first region of the image for depicting the element;

generating a partially noisy image map that includes noise in the first region and image features from the image in a second region outside the first region; and

generating a composite image map using a diffusion model based on the partially noisy image map and the prompt, wherein the composite image map includes the element in the first region that corresponds to the mask.

14. The non-transitory computer readable medium of claim 13 , the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

receiving a brush tool input from a user; and

generating the preliminary mask based on the brush tool input.

15. The non-transitory computer readable medium of claim 13 , the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

encoding the image to obtain an image feature map, wherein the partially noisy image map is based on the image feature map; and

decoding the composite image map to obtain a composite image.

16. The non-transitory computer readable medium of claim 13 , the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

generating a predicted mask for the element based on the composite image map; and

combining the image and the composite image map based on the predicted mask.

17. The non-transitory computer readable medium of claim 13 , the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

providing the mask as an input to the diffusion model.

18. The non-transitory computer readable medium of claim 13 , the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

generating intermediate denoising data using the diffusion model; and

combining the intermediate denoising data with the partially noisy image map to obtain an intermediate composite image map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2022
From: XIE, SHAOAN; ZHANG, ZHIFEI; LIN, ZHE; HINZ, TOBIAS
To: ADOBE INC.
Reel/Frame 061851/0543 →
Continuity (1)
Related Publication 20240169622A1 · May 23, 2024
References Cited (18)
US 11830159B1 · Mann · 2023 [cited by examiner]
US 20230103638A1 · Saharia · 2023 [cited by examiner]
US 20240161468A1 · Li · 2024 [cited by examiner]
CN 116152087A · 2023 [cited by applicant]
WO WO2023250088A1 · 2023 [cited by examiner]
WO WO2024081778A1 · 2024 [cited by examiner]
Avrahami et al., “Blended Diffusion for Text-driven Editing of Natural Images,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 18-24, 2022, New Orleans, LA, USA. (Year: 2022). [cited by examiner]
1Goodfellow, et al., “Generative Adversarial Networks”, Communications of the ACM, vol. 63, No. 11, pp. 139-144, Nov. 2020, 6 pages. [cited by applicant]
2Avrahami, et al., “Blended Diffusion for Text-driven Editing of Natural Images”, arXiv preprint arXiv:2111.14818 [cs. CV] Mar. 28, 2022, 32 pages. [cited by applicant]
3Ho, et al., “Denoising Diffusion Probabilistic Models”, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Dec. 2020, 12 pages. [cited by applicant]
4Radford, et al., “Learning Transferable Visual Models from Natural Language Supervision”, Proceedings of the 38th International Conference on Machine Learning, PMLR 139, Jul. 2021, 16 pages. [cited by applicant]
5Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, arXiv preprint arXiv:2112.10752v2 [cs.CV] Apr. 13, 2022, 45 pages. [cited by applicant]
6Nichol, et al., “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models”, arXiv preprint arXiv:2112.10741v3 [cs.CV] Mar. 8, 2022, 20 pages. [cited by applicant]
Combined Search and Examination Report dated Mar. 6, 2024 in corresponding United Kingdom Patent Application No. 2314576.6, 7 pages. [cited by applicant]
Avrahami, et al., “Blended Diffusion for Text-driven Editing of Natural Images” , 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition CVPR) Jun. 18, 2022-Jun. 24, 2022 New Orleans, LA, USA Jun. 18, 2022,… [cited by applicant]
Lugmayr, et al., “RePaint: Inpainting using Denoising Diffusion Probabilistic Models”, arXiv preprint arXiv:2201.09865v4 [cs.CV] Aug. 31, 2022, 25 page. [cited by applicant]
Wang, et al., “Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting ”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Vancouver, BC, Canada Jun. 17, 2023, pp. 18… [cited by applicant]
Kim, et al., DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation, arXiv preprint arXiv:2110.02711v6 [cs.CV] Aug. 11, 2022, 28 pages. [cited by applicant]
Cited By (2)
US 12,579,712 US 12,711,678