IP Library › Granted Patent US 12,749,293
Granted Patent B2
US 12,749,293 · App. 18/817,692 · Granted Sep 29, 2026

Modality specific learnable attention for multi-conditioned diffusion models

Inventors: Hareesh Ravi (San Jose, CA); Aashish Kumar Misraa (Santa Clara, CA); Ajinkya Gorakhnath Kale (San Jose, CA)
Assignee: ADOBE INC.
G06T11/00G06V10/771
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,749,293
App. No.
18/817,692
Granted
Sep 29, 2026
Kind
B2
Abstract

A method, apparatus, non-transitory computer readable medium, and system for image generation include encoding a text prompt to obtain a text embedding. An image prompt is encoded to obtain an image embedding. Cross-attention is performed on the text embedding and then on the image embedding to obtain a text attention output and an image attention output, respectively. A synthesized image is generated based on the text attention output and the image attention output.

Claims (60)

1 . A method comprising:

obtaining a text prompt, an image prompt, and a noise input;

encoding the text prompt to obtain a text embedding;

encoding the image prompt to obtain an image embedding;

generating, using an image generation model, an intermediate feature map based on the noise input;

performing, using a text attention layer of the image generation model, cross-attention on the text embedding and the intermediate feature map to obtain a text attention output;

performing, using an image attention layer of the image generation model, cross-attention on the image embedding and the same intermediate feature map provided to the text attention layer to obtain an image attention output; and

generating, using a generator network of the image generation model, a synthesized image based on the text attention output and the image attention output.

2 . The method of claim 1 , wherein encoding the image prompt comprises:

encoding, using an image encoder, the image prompt to obtain a preliminary image encoding; and

projecting, using an image projector, the preliminary image encoding to obtain the image embedding.

3 . The method of claim 1 , wherein:

the text embedding comprises a first plurality of tokens in a text embedding space and the image embedding comprises a second plurality of tokens in the text embedding space.

4 . The method of claim 1 , wherein:

the text embedding comprises a same number of tokens as the image embedding.

5 . The method of claim 1 , further comprising:

combining the text attention output and the image attention output to obtain a combined attention output, wherein the synthesized image is generated based on the combined attention output.

6 . The method of claim 1 , wherein generating the synthesized image comprises:

performing a diffusion process on the noise input.

7 . The method of claim 1 , wherein:

the text attention output and the image attention output are located in a common embedding space.

8 . A method of training a machine learning model, the method comprising:

obtaining a training set including a training text prompt, a training image prompt, and a noise input; and

training, using the training set, an image generation model to generate a synthesized image, the training comprising:

encoding the training text prompt to obtain a text embedding;

encoding the training image prompt to obtain an image embedding;

generating, using the image generation model, an intermediate feature map based on the noise input;

training a text attention layer of the image generation model to perform cross-attention on the text embedding and the intermediate feature map to obtain a text attention output; and

training an image attention layer of the image generation model to perform cross-attention on the embedding and the same intermediate feature map provided to the text attention layer to obtain an image attention output.

9 . The method of claim 8 , wherein the training the image generation model comprises:

computing a diffusion loss; and

updating parameters of the image generation model based on the diffusion loss.

10 . The method of claim 8 , wherein obtaining the training set comprises:

generating the training text prompt based on the training image prompt.

11 . The method of claim 8 ,

wherein the image generation model is trained to generate the synthesized image based on the text embedding and the image embedding.

12 . The method of claim 11 , further comprising:

projecting, using an image projector, a preliminary image encoding to obtain the image embedding, wherein the image generation model is trained to generate the synthesized image based on the image embedding.

13 . The method of claim 12 , wherein:

the image projector is jointly trained with the image generation model.

14 . An apparatus comprising:

at least one processor; and

at least one memory including instructions executable by the at least one processor to perform operations comprising:

obtaining a text prompt, an image prompt, and a noise input;

encoding the text prompt to obtain a text embedding;

encoding the image prompt to obtain an image embedding;

generating, using an image generation model, an intermediate feature map based on the noise input;

performing, using a text attention layer of the image generation model, cross-attention on the text embedding and the intermediate feature map to obtain a text attention output;

performing, using an image attention layer of the image generation model, cross-attention on the image embedding and the same intermediate feature map provided to the text attention layer to obtain an image attention output; and

generating, using a generator network of the image generation model, a synthesized image based on the text attention output and the image attention output.

15 . The apparatus of claim 14 , further comprising:

a text encoder configured to encode the text prompt to obtain the text embedding.

16 . The apparatus of claim 15 , wherein:

the text encoder includes a transformer architecture.

17 . The apparatus of claim 14 , further comprising:

an image encoder configured to encode the image prompt to obtain the image embedding.

18 . The apparatus of claim 17 , further comprising:

an image projector configured to project a preliminary image encoding to obtain the image embedding.

19 . The apparatus of claim 14 , wherein:

the image generation model comprises a diffusion model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2024
From: RAVI, HAREESH; MISRAA, AASHISH KUMAR; KALE, AJINKYA GORAKHNATH
To: ADOBE INC.
Reel/Frame 068425/0608 →
Continuity (2)
Provisional Application 63588403 · Oct 6, 2023
Related Publication 20250117972A1 · Apr 10, 2025
References Cited (35)
US 11922550B1 · Ramesh · 2024 [cited by examiner]
US 12198224B2 · Yuan · 2025 [cited by examiner]
US 12333636B2 · Li · 2025 [cited by examiner]
US 12430829B2 · Xie · 2025 [cited by examiner]
US 12518358B2 · Li · 2026 [cited by examiner]
US 12626431B2 · Li · 2026 [cited by examiner]
US 20230260164A1 · Yuan · 2023 [cited by examiner]
US 20230377214A1 · Kansy et al. · 2023 [cited by applicant]
US 20240144544A1 · Liew · 2024 [cited by examiner]
US 20240161258A1 · Maschmeyer et al. · 2024 [cited by applicant]
US 20240169622A1 · Xie · 2024 [cited by examiner]
US 20240169623A1 · Zeng · 2024 [cited by examiner]
US 20240249446A1 · Atzmon et al. · 2024 [cited by applicant]
US 20240256793A1 · Maschmeyer · 2024 [cited by applicant]
US 20240281924A1 · Park · 2024 [cited by examiner]
US 20240282016A1 · Liu et al. · 2024 [cited by applicant]
US 20240296607A1 · Li · 2024 [cited by examiner]
US 20240311693A1 · Smith et al. · 2024 [cited by applicant]
US 20240320867A1 · Bean · 2024 [cited by applicant]
US 20240330381A1 · Sadr · 2024 [cited by applicant]
US 20240331236A1 · Li · 2024 [cited by examiner]
US 20240338799A1 · Li · 2024 [cited by examiner]
US 20240386249A1 · Harvey et al. · 2024 [cited by applicant]
US 20240420407A1 · Ntavelis et al. · 2024 [cited by applicant]
US 20250022258A1 · Li et al. · 2025 [cited by applicant]
US 20250037323A1 · Srinivasan et al. · 2025 [cited by applicant]
US 20250054322A1 · Ye et al. · 2025 [cited by applicant]
US 20250078392A1 · Shi · 2025 [cited by examiner]
US 20250094484A1 · Zheng · 2025 [cited by examiner]
US 20250111655A1 · Povalyaev et al. · 2025 [cited by applicant]
US 20260094327A1 · Cohen · 2026 [cited by examiner]
Ronneberger, Olaf, Philipp Fischer, and Thomas Brox, “U-net: Convolutional networks for biomedical image segmentation”, International Conference on Medical image computing and computer-assisted intervention. Cham: Sprin… [cited by applicant]
Office Action dated Nov. 7, 2025 in related U.S. Appl. No. 18/637,654. [cited by applicant]
Balaji, et al., eDiff-i: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers, arXiv preprint arXiv:2211.01324v5 [cs.CV] Mar. 14, 2023, 24 pages. [cited by applicant]
Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv preprint arXiv:2204.06125v1 [cs.CV] Apr. 13, 2022, 27 pages. [cited by applicant]