IP Library › Granted Patent US 12,361,616
Granted Patent B2
US 12,361,616 · App. 18/169,444 · Granted Jul 15, 2025

Image generation using a diffusion model

Inventors: Nicholas Isaac Kolkin (San Francisco, CA); Elya Shechtman (Seattle, WA)
Assignee: ADOBE INC.
G06T11/60G06T5/70G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,616
App. No.
18/169,444
Granted
Jul 15, 2025
Kind
B2
Abstract

Systems and methods for image generation are provided. An aspect of the systems and methods for image generation includes obtaining an original image depicting an element and a target prompt describing a modification to the element. The system may then compute a first output and a second output using a diffusion model. The first output is based on a description of the element and the second output is based on the target prompt. The system then computes a difference between the first output and the second output, and generates a modified image including the modification to the element of the original image based on the difference.

Claims (52)

1. A method for image generation, comprising:

obtaining an original image depicting an element and a target prompt describing a modification to the element;

computing a first output and a second output using a diffusion model, wherein the first output is based on a description of the element and the second output is based on the target prompt;

computing a difference between the first output and the second output by subtracting the first output from the second output; and

generating a modified image including the modification to the element of the original image based on the difference.

2. The method of claim 1 , further comprising:

obtaining an anchor prompt that includes a first modifier of the element, wherein the description is based on the anchor prompt and the target prompt includes a second modifier that describes the modification to the element.

3. The method of claim 1 , further comprising:

generating the description of the element based on the original image and the target prompt.

4. The method of claim 1 , further comprising:

adding first noise to the original image to obtain a noise image; and

removing second noise from the noise image based on the difference to obtain the modified image.

5. The method of claim 4 , further comprising:

computing a weighted sum of the first output and the difference, wherein the second noise is based on the weighted sum.

6. The method of claim 4 , further comprising:

determining a number of noise addition steps for adding the first noise; and

computing the difference at each of a plurality of noise removal steps for removing the second noise.

7. The method of claim 1 , further comprising:

obtaining a mask indicating a region corresponding to the element of the original image, wherein the modified image is generated based on the mask.

8. The method of claim 7 , further comprising:

generating a noise map based on the original image and the mask, wherein the modified image is generated based on the noise map.

9. The method of claim 1 , further comprising:

encoding an anchor prompt and the target prompt using a text encoder to obtain an encoded anchor prompt and an encoded target prompt, wherein the first output and the second output are based on the encoded anchor prompt and the encoded target prompt, respectively.

10. A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to:

obtain an original image depicting an element, an anchor prompt describing the element, and a target prompt describing a modification to the element;

compute a first output based on the anchor prompt and a second output based on the target prompt using a diffusion model;

compute a difference between the first output and the second output by subtracting the first output from the second output; and

generate a modified image including the modification to the element of the original image based on the difference.

11. The non-transitory computer readable medium of claim 10 , wherein the instructions further cause the processor to:

add first noise to the original image to obtain a noise image; and

remove second noise from the noise image based on the difference to obtain the modified image.

12. The non-transitory computer readable medium of claim 11 , wherein the instructions further cause the processor to:

compute a weighted sum of the first output and the difference, wherein the second noise is based on the weighted sum.

13. The non-transitory computer readable medium of claim 11 , wherein the instructions further cause the processor to:

determine a number of noise addition steps for adding the first noise; and

compute the difference at each of a plurality of noise removal steps for removing the second noise.

14. The non-transitory computer readable medium of claim 10 , wherein the instructions further cause the processor to:

obtain a mask indicating a region corresponding to the element of the original image, wherein the modified image is generated based on the mask.

15. The non-transitory computer readable medium of claim 10 , wherein:

the anchor prompt includes a first modifier of the element and the target prompt includes a second modifier of the element, wherein the second modifier describes the modification to the element.

16. A system for image generation, comprising:

one or more processors;

one or more memory components coupled with the one or more processors; and

a diffusion model configured to compute a first output based on an anchor prompt and a second output based on a target prompt, compute a difference between the first output and the second output by subtracting the first output from the second output, and generate a modified image including a modification to an element of an original image based on the difference.

17. The system of claim 16 , further comprising:

a mask generation model configured to generate a mask indicating a location of the element, wherein the modified image is generated based on the mask.

18. The system of claim 16 , further comprising:

a noise component configured to add first noise to the original image to obtain a noise image, wherein the diffusion model is further configured to remove second noise from the noise image based on the difference to obtain the modified image.

19. The system of claim 16 , further comprising:

a user interface configured to obtain the original image and the target prompt.

20. The system of claim 16 , further comprising:

a text encoder configured to encode the anchor prompt and the target prompt to obtain an encoded anchor prompt and an encoded target prompt, wherein the first output and the second output are based on the encoded anchor prompt and the encoded target prompt, respectively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2023
From: KOLKIN, NICHOLAS ISAAC; SHECHTMAN, ELYA
To: ADOBE INC.
Reel/Frame 062708/0011 →
Continuity (2)
Provisional Application 63379808 · Oct 17, 2022
Related Publication 20240135610A1 · Apr 25, 2024
References Cited (22)
US 12118976B1 · Chen · 2024 [cited by examiner]
US 20230095092A1 · Xiao · 2023 [cited by examiner]
US 20230118966A1 · Liu · 2023 [cited by examiner]
US 20230377690A1 · Anand · 2023 [cited by examiner]
US 20240338859A1 · Marri · 2024 [cited by examiner]
US 20240355018A1 · Aggarwal · 2024 [cited by examiner]
US 20240355022A1 · Shi · 2024 [cited by examiner]
Brooks T, Holynski A, Efros AA. Instructpix2pix: Learning to follow image editing instructions. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2023 (pp. 18392-18402). [cited by examiner]
Song, J., Meng, C. and Ermon, S., 2020. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502. [cited by examiner]
Gu J, Wang S, Zhao H, Lu T, Zhang X, Wu Z, Xu S, Zhang W, Jiang YG, Xu H. Reuse and diffuse: Iterative denoising for text-to-video generation. arXiv preprint arXiv:2309.03549. Sep. 7, 2023. [cited by examiner]
Dhariwal P, Nichol A. Diffusion models beat gans on image synthesis. Advances in neural information processing systems. Dec. 6, 2021;34:8780-94. [cited by examiner]
Croitoru FA, Hondru V, Ionescu RT, Shah M. Diffusion models in vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence. Mar. 27, 2023;45(9):10850-69. [cited by examiner]
Avrahami O, Lischinski D, Fried O. Blended Diffusion for Text-driven Editing of Natural Images. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Jun. 18, 2022 (pp. 18187-18197). IEEE. [cited by examiner]
Ho, J., Jain, A. and Abbeel, P., 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33, pp. 6840-6851. [cited by examiner]
Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”; arXiv preprint: arXiv:2204.06125v1 [cs.CV] Apr. 13, 2022, 27 pages. [cited by applicant]
Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv preprint: arXiv:2205.11487v1 [cs.CV] May 23, 2022, 46 pages. [cited by applicant]
Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, arXiv preprint: arXiv:2112.10752v2 [cs.CV] Apr. 13, 2022, 45 pages. [cited by applicant]
Meng, et al., “SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations”, arXiv preprint: arXiv:2108.01073v2 [cs.CV] Jan. 5, 2022, 33 pages. [cited by applicant]
Ho, et al., “Classifier-Free Diffusion Guidance”, arXiv preprint: arXiv:2207.12598v1 [cs.LG] Jul. 26, 2022, 14 pages. [cited by applicant]
Karras, et al., “Elucidating the Design Space of Diffusion-Based Generative Models”, arXiv preprint: arXiv:2206.00364v2 [cs.CV] Oct. 11, 2022, 47 pages. [cited by applicant]
Hertz, et al., “Prompt-to-Prompt Image Editing with Cross Attention Control”, arXiv preprint: arXiv:2208.01626v1 [cs.CV] Aug. 2, 2022, 19 pages. [cited by applicant]
Liu, et al., “Compositional Visual Generation with Composable Diffusion Models”, arXiv preprint: arXiv:2206.01714v4 [cs.CV] Jul. 27, 2022, 28 pages. [cited by applicant]