Text-to-image synthesis utilizing diffusion models with test-time attention segregation and retention optimization
The present disclosure relates to systems, methods, and non-transitory computer-readable media that utilizes attention segregation loss and/or attention retention loss at inference time of a diffusion neural network to generate a text-conditioned image. In particular, in some embodiments, the disclosed systems utilize the attention segregation loss to reduce overlap between concepts by comparing attention maps for multiple concepts of a text query corresponding to a denoising step. Further, in some embodiments, the disclosed systems utilize the attention retention loss to improve information retention for concepts across denoising steps by comparing attention maps between different denoising steps. Accordingly, in some embodiments, by utilizing the attention segregation loss and the attention retention loss, the disclosed systems accurately maintain multiple concepts from a text query when generating a text-conditioned image.
1 . A computer-implemented method comprising:
generating, from a text query and a first noise representation from a first denoising step of a diffusion neural network, a second noise representation utilizing a second denoising step of the diffusion neural network, wherein the text query comprises a first concept and a second concept;
determining an attention segregation loss between a first attention map and a second attention map corresponding to the second denoising step by comparing a first text embedding of the first concept with the second noise representation to generate the first attention map and comparing a second text embedding of the second concept with the second noise representation to generate the second attention map;
determining an attention retention loss between the first denoising step and the second denoising step based on comparing the first attention map and the second attention map corresponding to the second denoising step with additional attention maps corresponding to the first denoising step; and
generating a text-conditioned image based on a modified noise representation generated from the second noise representation based on the attention segregation loss and the attention retention loss.
2 . The computer-implemented method of claim 1 , further comprising:
generating, utilizing a text encoder, the first text embedding of the first concept from the text query; and
generating, utilizing the text encoder, the second text embedding of the second concept from the text query.
3 . The computer-implemented method of claim 2 , further comprising:
conditioning the second denoising step with the first text embedding and the second text embedding to generate the second noise representation; and
conditioning an additional denoising step with the first text embedding and the second text embedding to generate an additional noise representation,
wherein conditioning the second denoising step and the additional denoising step comprises providing the first text embedding and the second text embedding as context data to cause the diffusion neural network to focus on specific portions of input data.
4 . The computer-implemented method of claim 1 , wherein determining the attention retention loss comprises:
generating the first attention map of the first concept of the text query corresponding to the second denoising step and the second attention map of the second concept of the text query corresponding to the second denoising step;
generating the additional attention maps comprising a third attention map of the first concept of the text query corresponding to the first denoising step and a fourth attention map of the second concept of the text query corresponding to the first denoising step; and
determining the attention retention loss between the first denoising step and the second denoising step by comparing the first attention map corresponding to the second denoising step with the third attention map corresponding to the first denoising step and comparing the second attention map corresponding to the second denoising step with the fourth attention map corresponding to the first denoising step.
5 . The computer-implemented method of claim 4 , wherein comparing the first attention map corresponding to the second denoising step and the third attention map corresponding to the first denoising step comprises:
determining a threshold activation region for the third attention map corresponding to the first denoising step; and
generating a binary mask for the third attention map corresponding to the first denoising step based on the threshold activation region.
6 . The computer-implemented method of claim 5 , wherein comparing the first attention map corresponding to the second denoising step and the third attention map corresponding to the first denoising step comprises comparing the binary mask with the first attention map corresponding to the second denoising step.
7 . The computer-implemented method of claim 1 , wherein generating the text-conditioned image further comprises:
generating a combined loss from the attention segregation loss and the attention retention loss;
generating the modified noise representation from the combined loss; and
utilizing additional steps of the diffusion neural network to generate the text-conditioned image from the modified noise representation.
8 . The computer-implemented method of claim 1 , wherein generating the second noise representation comprises:
generating, utilizing a text encoder, a text vector representation from the text query; and
conditioning the second denoising step utilizing the text vector representation.
9 . A system comprising:
one or more memory devices comprising a diffusion neural network, a text query comprising a first text concept and a second text concept, and a noise vector; and
one or more processors configured to cause the system to:
generate a noise representation from the noise vector and the text query utilizing a denoising step of the diffusion neural network;
generate a first attention map by comparing a first text embedding of the first text concept with the noise representation from the noise vector;
generate a second attention map by comparing a second text embedding of the second text concept with the noise representation from the noise vector;
determine an attention segregation loss between the first attention map of the first text concept and the second attention map of the second text concept by comparing the first attention map and the second attention map;
generate a modified noise representation from the noise representation utilizing the attention segregation loss; and
generate, utilizing additional steps of the diffusion neural network, a text-conditioned image from the modified noise representation.
10 . The system of claim 9 , wherein the one or more processors are configured to cause the system to determine an attention retention loss based on the denoising step and a previous denoising step.
11 . The system of claim 10 , wherein the one or more processors are configured to cause the system to determine the attention retention loss by:
generating a previous noise representation utilizing the previous denoising step of the diffusion neural network; and
generating a previous attention map corresponding to the previous denoising step.
12 . The system of claim 11 , wherein the one or more processors are configured to cause the system to compare an attention map corresponding to the denoising step and the previous attention map corresponding to the previous denoising step to determine the attention retention loss.
13 . The system of claim 10 , wherein the one or more processors are configured to cause the system to generate the modified noise representation from the noise representation utilizing the attention segregation loss and the attention retention loss.
14 . The system of claim 9 , wherein the one or more processors are configured to cause the system to generate an additional noise representation corresponding to an additional denoising step from the modified noise representation.
15 . The system of claim 14 , wherein the one or more processors are configured to cause the system to:
generate, for the additional denoising step, a third attention map corresponding to the first text concept and a fourth attention map corresponding to the second text concept;
determine an additional attention segregation loss by comparing the third attention map and the fourth attention map from the additional denoising step; and
generate an additional modified noise representation from the additional noise representation utilizing the additional attention segregation loss.
16 . The system of claim 15 , wherein the one or more processors are configured to cause the system to generate, utilizing a text encoder, a text vector representation from the text query to condition the additional noise representation utilizing the text vector representation.
17 . A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:
generating, from a text query and a first noise representation from a first denoising step of a diffusion neural network, a second noise representation utilizing a second denoising step of the diffusion neural network;
determining a first attention map and a second attention map for the first denoising step, wherein the first attention map corresponds to a first concept in the text query and the second attention map corresponds to a second concept in the text query;
determining a third attention map and a fourth attention map for the second denoising step, wherein the third attention map corresponds to the first concept in the text query and the fourth attention map corresponds to the second concept in the text query;
determining an first attention retention loss for the first concept by comparing the first attention map for the first denoising step and the third attention map for the second denoising step, wherein the first attention retention loss indicates a retention of information for the first concept in the text query from the first denoising step to the second denoising step;
determining a second attention retention loss for the second concept by comparing the second attention map for the first denoising step and the fourth attention map for the second denoising step;
generating a modified noise representation from the second noise representation utilizing the first attention retention loss for the first concept and the second attention retention loss for the second concept; and
generating a text-conditioned image from the modified noise representation.
18 . The non-transitory computer-readable medium of claim 17 , wherein generating the second noise representation further comprises generating, utilizing a text encoder, a text vector representation from the text query to condition the second denoising step utilizing the text vector representation.
19 . The non-transitory computer-readable medium of claim 17 , wherein determining the first attention retention loss further comprises:
determining a threshold activation region for the first attention map corresponding to the first denoising step;
generating a binary mask for the first attention map corresponding to the first denoising step based on the threshold activation region; and
comparing the binary mask with the third attention map corresponding to the second denoising step to determine the first attention retention loss.
20 . The non-transitory computer-readable medium of claim 17 , wherein the operations further comprise determining an attention segregation loss from the text query comprising a first concept and a second concept by:
generating, for the first denoising step, a first attention map corresponding to the first concept and a second attention map corresponding to the second concept; and
comparing the first attention map corresponding to the first concept to the second attention map corresponding to the second concept to determine the attention segregation loss.