IP Library › Granted Patent US 12,462,348
Granted Patent B2
US 12,462,348 · App. 18/165,141 · Granted Nov 4, 2025

Multimodal diffusion models

Inventors: Cusuh Ham (Marietta, GA); Tobias Hinz (Campbell, CA); Jingwan Lu (Sunnyvale, CA); Krishna Kumar Singh (San Jose, CA); Zhifei Zhang (San Jose, CA)
Assignee: ADOBE INC.
G06T5/70G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,348
App. No.
18/165,141
Granted
Nov 4, 2025
Kind
B2
Abstract

Systems and methods for image processing are described. Embodiments of the present disclosure obtain a noise image and guidance information for generating an image. A diffusion model generates an intermediate noise prediction for the image based on the noise image. A conditioning network generates noise modulation parameters. The intermediate noise prediction and the noise modulation parameters are combined to obtain a modified intermediate noise prediction. The diffusion model generates the image based on the modified intermediate noise prediction, wherein the image depicts a scene based on the guidance information.

Claims (58)

1 . A method comprising:

obtaining a noise input and guidance information indicating a target image control;

generating, using an image generation model, a noise prediction based on the noise input, wherein the noise prediction comprises a prediction of noise to remove from the noise input;

generating noise modulation parameters based on the guidance information using a conditioning network, wherein the noise modulation parameters represent the target image control;

combining the noise prediction and the noise modulation parameters to obtain a modified noise prediction, wherein the modified noise prediction comprises a modified prediction of noise to remove from the noise input based on the target image control; and

generating, using the image generation model, a synthetic image based on the modified noise prediction, wherein the synthetic image depicts a scene based on the target image control.

2 . The method of claim 1 , wherein:

the guidance information comprises non-textual guidance.

3 . The method of claim 1 , wherein:

the guidance information comprises an additional modality other than a training modality used for training the image generation model.

4 . The method of claim 1 , wherein:

the guidance information comprises multiple modalities.

5 . The method of claim 1 , wherein the combining comprises:

performing element-wise multiplication of the noise prediction and a first portion of the noise modulation parameters; and

performing element-wise addition of the noise prediction and a second portion of the noise modulation parameters.

6 . The method of claim 1 , further comprising:

iteratively generating a plurality of intermediate noise predictions corresponding to a plurality of diffusion steps, respectively; and

iteratively generating a plurality of noise modulation parameters corresponding to the plurality of intermediate noise predictions, respectively, wherein the synthetic image is generated based on the plurality of noise modulation parameters.

7 . The method of claim 1 , further comprising:

generating an intermediate image prediction, wherein the noise modulation parameters are generated based on the intermediate image prediction.

8 . The method of claim 7 , further comprising:

generating a subsequent image prediction based on the intermediate image prediction and the modified noise prediction, wherein the synthetic image is generated based on the subsequent image prediction.

9 . An apparatus comprising:

a memory component;

a processing device coupled to the memory component, the processing device configured to perform operations comprising:

obtaining a noise input and guidance information indicating a target image control;

generating, using an image generation model, a noise prediction based on the noise input, wherein the noise prediction comprises a prediction of noise to remove from the noise input;

generating noise modulation parameters based on the guidance information using a conditioning network, wherein the noise modulation parameters represent the target image control;

combining the noise prediction and the noise modulation parameters to obtain a modified noise prediction, wherein the modified noise prediction comprises a modified prediction of noise to remove from the noise input based on the target image control; and

generating, using the image generation model, a synthetic image based on the modified noise prediction, wherein the synthetic image depicts a scene based on the target image control.

10 . The apparatus of claim 9 , wherein:

the image generation model comprises a U-Net architecture, wherein the noise modulation parameters are applied to an output of the U-Net architecture of the image generation model.

11 . The apparatus of claim 9 , wherein:

the image generation model comprises a text-guided diffusion model, and the image generation model takes a text prompt as an input.

12 . The apparatus of claim 9 , wherein:

the conditioning network comprises a U-Net architecture, wherein the noise modulation parameters comprise an output of the U-Net architecture of the conditioning network.

13 . A non-transitory computer readable medium storing code for video processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

obtaining a noise input and guidance information indicating a target image control;

generating, using an image generation model, a noise prediction based on the noise input, wherein the noise prediction comprises a prediction of noise to remove from the noise input;

generating noise modulation parameters based on the guidance information using a conditioning network, wherein the noise modulation parameters represent the target image control;

combining the noise prediction and the noise modulation parameters to obtain a modified noise prediction, wherein the modified noise prediction comprises a modified prediction of noise to remove from the noise input based on the target image control; and

generating, using the image generation model, a synthetic image based on the modified noise prediction, wherein the synthetic image depicts a scene based on the target image control.

14 . The non-transitory computer readable medium of claim 13 , wherein:

the guidance information comprises non-textual guidance.

15 . The non-transitory computer readable medium of claim 13 , wherein:

the guidance information comprises an additional modality other than a training modality used for training the image generation model.

16 . The non-transitory computer readable medium of claim 13 , wherein:

the guidance information comprises multiple modalities.

17 . The non-transitory computer readable medium of claim 13 , the code further comprising instructions executable by the at least one processor to perform operations comprising:

performing element-wise multiplication of the noise prediction and a first portion of the noise modulation parameters; and

performing element-wise addition of the noise prediction and a second portion of the noise modulation parameters.

18 . The non-transitory computer readable medium of claim 13 , the code further comprising instructions executable by the at least one processor to perform operations comprising:

iteratively generating a plurality of intermediate noise predictions corresponding to a plurality of diffusion steps, respectively; and

iteratively generating a plurality of noise modulation parameters corresponding to the plurality of intermediate noise predictions, respectively, wherein the synthetic image is generated based on the plurality of noise modulation parameters.

19 . The non-transitory computer readable medium of claim 13 , the code further comprising instructions executable by the at least one processor to perform operations comprising:

generating an intermediate image prediction, wherein the noise modulation parameters are generated based on the intermediate image prediction.

20 . The non-transitory computer readable medium of claim 19 , the code further comprising instructions executable by the at least one processor to perform operations comprising:

generating a subsequent image prediction based on the intermediate image prediction and the modified noise prediction, wherein the synthetic image is generated based on the subsequent image prediction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2023
From: HAM, CUSUH; HINZ, TOBIAS; LU, JINGWAN; SINGH, KRISHNA KUMAR; ZHANG, ZHIFEI
To: ADOBE INC.
Reel/Frame 062605/0567 →
Continuity (1)
Related Publication 20240265505A1 · Aug 8, 2024
References Cited (31)
US 11393166B2 · Wei · 2022 [cited by examiner]
US 11908180B1 · Ho · 2024 [cited by examiner]
US 12020403B2 · Kulkarni · 2024 [cited by examiner]
US 12067659B2 · Wang · 2024 [cited by examiner]
US 12148119B2 · Zhang · 2024 [cited by examiner]
US 12169850B1 · Luguev · 2024 [cited by examiner]
US 20190251721A1 · Hua · 2019 [cited by examiner]
US 20220067281A1 · Hu · 2022 [cited by examiner]
US 20220335285A1 · Sundaresan · 2022 [cited by examiner]
US 20230109379A1 · Kreis · 2023 [cited by examiner]
US 20230162413A1 · Batra · 2023 [cited by examiner]
US 20230377214A1 · Kansy · 2023 [cited by examiner]
US 20240037948A1 · Chen · 2024 [cited by examiner]
US 20240346629A1 · Harikumar · 2024 [cited by examiner]
US 20250005809A1 · Kim · 2025 [cited by examiner]
US 20250005829A1 · Kim · 2025 [cited by examiner]
Mao, Jiafeng, Xueting Wang, and Kiyoharu Aizawa. “Guided image synthesis via initial image editing in diffusion model.” Proceedings of the 31st ACM International Conference on Multimedia. 2023. (Year: 2023). [cited by examiner]
Rombach, Robin, et al. “High-resolution image synthesis with latent diffusion models.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022. (Year: 2022). [cited by examiner]
Hertz, Amir, et al. “Prompt-to-prompt image editing with cross attention control.” arXiv preprint arXiv:2208.01626 (2022). (Year: 2022). [cited by examiner]
Voynov et al, Sketch-Guided Text-to-Image Diffusion Models, Sketch-Guided Text-to-Image Diffusion Models, https://arxiv.org/abs/2211.13752 Nov. 2022 (Year: 2022). [cited by examiner]
Valevski et al, UniTune: Text-Driven Image Editing by Fine Tuning an Image Generation Model on a Single Image, https://arxiv.org/abs/2210.09477v3, 2022 (Year: 2022). [cited by examiner]
1Gafni, et al., “Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors”, arXiv preprint arXiv:2203.13131v1 [cs.CV] Mar. 24, 2022, 17 pages. [cited by applicant]
2Gatys, et al., “A Neural Algorithm of Artistic Style”, arXiv preprint arXiv:1508.06576v2 [cs.CV] Sep. 2, 2015, 16 pages. [cited by applicant]
3Huang, et al., “Multimodal Conditional Image Synthesis with Product-of-Experts GANs”, arXiv preprint arXiv:2112.05130v1 [cs.CV] Dec. 9, 2021, 18 pages. [cited by applicant]
4Liu, et al., “More Control for Free! Image Synthesis with Semantic Diffusion Guidance”, arXiv preprint arXiv:2112.05744v2 [cs.CV] Dec. 14, 2021, 16 pages. [cited by applicant]
5Nichol, et al., “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models”, arXiv preprint arXiv:2112.10741v3 [cs.CV] Mar. 8, 2022, 20 pages. [cited by applicant]
6Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv preprint arXiv:2204.06125v1 [cs.CV] Apr. 13, 2022, 27 pages. [cited by applicant]
7Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, arXiv preprint arXiv:2112.10752v2 [cs.CV] Apr. 13, 2022, 45 pages. [cited by applicant]
8Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv preprint arXiv:2205.11487v1 [cs.CV] May 23, 2022, 46 pages. [cited by applicant]
9Song, et al., “Denoising Diffusion Implicit Models”, arXiv preprint arXiv:2010.02502v4 [cs.LG] Oct. 5, 2022, 22 pages. [cited by applicant]
10Zhang, et al. “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”, In CVPR, arXiv preprint arXiv:1801.03924v1 [cs.CV] Jan. 11, 2018, 14 pages. [cited by applicant]