Modality specific learnable attention for multi-conditioned diffusion models
A method, apparatus, non-transitory computer readable medium, and system for image generation include encoding a text prompt to obtain a text embedding. An image prompt is encoded to obtain an image embedding. Cross-attention is performed on the text embedding and then on the image embedding to obtain a text attention output and an image attention output, respectively. A synthesized image is generated based on the text attention output and the image attention output.
1 . A method comprising:
obtaining a text prompt, an image prompt, and a noise input;
encoding the text prompt to obtain a text embedding;
encoding the image prompt to obtain an image embedding;
generating, using an image generation model, an intermediate feature map based on the noise input;
performing, using a text attention layer of the image generation model, cross-attention on the text embedding and the intermediate feature map to obtain a text attention output;
performing, using an image attention layer of the image generation model, cross-attention on the image embedding and the same intermediate feature map provided to the text attention layer to obtain an image attention output; and
generating, using a generator network of the image generation model, a synthesized image based on the text attention output and the image attention output.
2 . The method of claim 1 , wherein encoding the image prompt comprises:
encoding, using an image encoder, the image prompt to obtain a preliminary image encoding; and
projecting, using an image projector, the preliminary image encoding to obtain the image embedding.
3 . The method of claim 1 , wherein:
the text embedding comprises a first plurality of tokens in a text embedding space and the image embedding comprises a second plurality of tokens in the text embedding space.
4 . The method of claim 1 , wherein:
the text embedding comprises a same number of tokens as the image embedding.
5 . The method of claim 1 , further comprising:
combining the text attention output and the image attention output to obtain a combined attention output, wherein the synthesized image is generated based on the combined attention output.
6 . The method of claim 1 , wherein generating the synthesized image comprises:
performing a diffusion process on the noise input.
7 . The method of claim 1 , wherein:
the text attention output and the image attention output are located in a common embedding space.
8 . A method of training a machine learning model, the method comprising:
obtaining a training set including a training text prompt, a training image prompt, and a noise input; and
training, using the training set, an image generation model to generate a synthesized image, the training comprising:
encoding the training text prompt to obtain a text embedding;
encoding the training image prompt to obtain an image embedding;
generating, using the image generation model, an intermediate feature map based on the noise input;
training a text attention layer of the image generation model to perform cross-attention on the text embedding and the intermediate feature map to obtain a text attention output; and
training an image attention layer of the image generation model to perform cross-attention on the embedding and the same intermediate feature map provided to the text attention layer to obtain an image attention output.
9 . The method of claim 8 , wherein the training the image generation model comprises:
computing a diffusion loss; and
updating parameters of the image generation model based on the diffusion loss.
10 . The method of claim 8 , wherein obtaining the training set comprises:
generating the training text prompt based on the training image prompt.
11 . The method of claim 8 ,
wherein the image generation model is trained to generate the synthesized image based on the text embedding and the image embedding.
12 . The method of claim 11 , further comprising:
projecting, using an image projector, a preliminary image encoding to obtain the image embedding, wherein the image generation model is trained to generate the synthesized image based on the image embedding.
13 . The method of claim 12 , wherein:
the image projector is jointly trained with the image generation model.
14 . An apparatus comprising:
at least one processor; and
at least one memory including instructions executable by the at least one processor to perform operations comprising:
obtaining a text prompt, an image prompt, and a noise input;
encoding the text prompt to obtain a text embedding;
encoding the image prompt to obtain an image embedding;
generating, using an image generation model, an intermediate feature map based on the noise input;
performing, using a text attention layer of the image generation model, cross-attention on the text embedding and the intermediate feature map to obtain a text attention output;
performing, using an image attention layer of the image generation model, cross-attention on the image embedding and the same intermediate feature map provided to the text attention layer to obtain an image attention output; and
generating, using a generator network of the image generation model, a synthesized image based on the text attention output and the image attention output.
15 . The apparatus of claim 14 , further comprising:
a text encoder configured to encode the text prompt to obtain the text embedding.
16 . The apparatus of claim 15 , wherein:
the text encoder includes a transformer architecture.
17 . The apparatus of claim 14 , further comprising:
an image encoder configured to encode the image prompt to obtain the image embedding.
18 . The apparatus of claim 17 , further comprising:
an image projector configured to project a preliminary image encoding to obtain the image embedding.
19 . The apparatus of claim 14 , wherein:
the image generation model comprises a diffusion model.