IP Library › Granted Patent US 12,633,000
Granted Patent B2
US 12,633,000 · App. 18/337,634 · Granted May 19, 2026

Text-to-image synthesis utilizing diffusion models with test-time attention segregation and retention optimization

Inventors: Aishwarya Agarwal (Bengaluru, IN); Srikrishna Karanam (Bangalore, IN); Joseph Koonthanam Jose (Kottayam, IN); Apoorv Umang Saxena (Bengaluru, IN); Koustava Goswami (Bangalore, IN); Balaji Vasan Srinivasan (Bangalore, IN)
Assignee: Adobe Inc.
G06T11/00G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,000
App. No.
18/337,634
Filed
Jun 20, 2023
Granted
May 19, 2026
Kind
B2
Examiner
HE, WEIMING
Art Unit
2611
USPC
345/441
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that utilizes attention segregation loss and/or attention retention loss at inference time of a diffusion neural network to generate a text-conditioned image. In particular, in some embodiments, the disclosed systems utilize the attention segregation loss to reduce overlap between concepts by comparing attention maps for multiple concepts of a text query corresponding to a denoising step. Further, in some embodiments, the disclosed systems utilize the attention retention loss to improve information retention for concepts across denoising steps by comparing attention maps between different denoising steps. Accordingly, in some embodiments, by utilizing the attention segregation loss and the attention retention loss, the disclosed systems accurately maintain multiple concepts from a text query when generating a text-conditioned image.

Claims (64)

1 . A computer-implemented method comprising:

generating, from a text query and a first noise representation from a first denoising step of a diffusion neural network, a second noise representation utilizing a second denoising step of the diffusion neural network, wherein the text query comprises a first concept and a second concept;

determining an attention segregation loss between a first attention map and a second attention map corresponding to the second denoising step by comparing a first text embedding of the first concept with the second noise representation to generate the first attention map and comparing a second text embedding of the second concept with the second noise representation to generate the second attention map;

determining an attention retention loss between the first denoising step and the second denoising step based on comparing the first attention map and the second attention map corresponding to the second denoising step with additional attention maps corresponding to the first denoising step; and

generating a text-conditioned image based on a modified noise representation generated from the second noise representation based on the attention segregation loss and the attention retention loss.

2 . The computer-implemented method of claim 1 , further comprising:

generating, utilizing a text encoder, the first text embedding of the first concept from the text query; and

generating, utilizing the text encoder, the second text embedding of the second concept from the text query.

3 . The computer-implemented method of claim 2 , further comprising:

conditioning the second denoising step with the first text embedding and the second text embedding to generate the second noise representation; and

conditioning an additional denoising step with the first text embedding and the second text embedding to generate an additional noise representation,

wherein conditioning the second denoising step and the additional denoising step comprises providing the first text embedding and the second text embedding as context data to cause the diffusion neural network to focus on specific portions of input data.

4 . The computer-implemented method of claim 1 , wherein determining the attention retention loss comprises:

generating the first attention map of the first concept of the text query corresponding to the second denoising step and the second attention map of the second concept of the text query corresponding to the second denoising step;

generating the additional attention maps comprising a third attention map of the first concept of the text query corresponding to the first denoising step and a fourth attention map of the second concept of the text query corresponding to the first denoising step; and

determining the attention retention loss between the first denoising step and the second denoising step by comparing the first attention map corresponding to the second denoising step with the third attention map corresponding to the first denoising step and comparing the second attention map corresponding to the second denoising step with the fourth attention map corresponding to the first denoising step.

5 . The computer-implemented method of claim 4 , wherein comparing the first attention map corresponding to the second denoising step and the third attention map corresponding to the first denoising step comprises:

determining a threshold activation region for the third attention map corresponding to the first denoising step; and

generating a binary mask for the third attention map corresponding to the first denoising step based on the threshold activation region.

6 . The computer-implemented method of claim 5 , wherein comparing the first attention map corresponding to the second denoising step and the third attention map corresponding to the first denoising step comprises comparing the binary mask with the first attention map corresponding to the second denoising step.

7 . The computer-implemented method of claim 1 , wherein generating the text-conditioned image further comprises:

generating a combined loss from the attention segregation loss and the attention retention loss;

generating the modified noise representation from the combined loss; and

utilizing additional steps of the diffusion neural network to generate the text-conditioned image from the modified noise representation.

8 . The computer-implemented method of claim 1 , wherein generating the second noise representation comprises:

generating, utilizing a text encoder, a text vector representation from the text query; and

conditioning the second denoising step utilizing the text vector representation.

9 . A system comprising:

one or more memory devices comprising a diffusion neural network, a text query comprising a first text concept and a second text concept, and a noise vector; and

one or more processors configured to cause the system to:

generate a noise representation from the noise vector and the text query utilizing a denoising step of the diffusion neural network;

generate a first attention map by comparing a first text embedding of the first text concept with the noise representation from the noise vector;

generate a second attention map by comparing a second text embedding of the second text concept with the noise representation from the noise vector;

determine an attention segregation loss between the first attention map of the first text concept and the second attention map of the second text concept by comparing the first attention map and the second attention map;

generate a modified noise representation from the noise representation utilizing the attention segregation loss; and

generate, utilizing additional steps of the diffusion neural network, a text-conditioned image from the modified noise representation.

10 . The system of claim 9 , wherein the one or more processors are configured to cause the system to determine an attention retention loss based on the denoising step and a previous denoising step.

11 . The system of claim 10 , wherein the one or more processors are configured to cause the system to determine the attention retention loss by:

generating a previous noise representation utilizing the previous denoising step of the diffusion neural network; and

generating a previous attention map corresponding to the previous denoising step.

12 . The system of claim 11 , wherein the one or more processors are configured to cause the system to compare an attention map corresponding to the denoising step and the previous attention map corresponding to the previous denoising step to determine the attention retention loss.

13 . The system of claim 10 , wherein the one or more processors are configured to cause the system to generate the modified noise representation from the noise representation utilizing the attention segregation loss and the attention retention loss.

14 . The system of claim 9 , wherein the one or more processors are configured to cause the system to generate an additional noise representation corresponding to an additional denoising step from the modified noise representation.

15 . The system of claim 14 , wherein the one or more processors are configured to cause the system to:

generate, for the additional denoising step, a third attention map corresponding to the first text concept and a fourth attention map corresponding to the second text concept;

determine an additional attention segregation loss by comparing the third attention map and the fourth attention map from the additional denoising step; and

generate an additional modified noise representation from the additional noise representation utilizing the additional attention segregation loss.

16 . The system of claim 15 , wherein the one or more processors are configured to cause the system to generate, utilizing a text encoder, a text vector representation from the text query to condition the additional noise representation utilizing the text vector representation.

17 . A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

generating, from a text query and a first noise representation from a first denoising step of a diffusion neural network, a second noise representation utilizing a second denoising step of the diffusion neural network;

determining a first attention map and a second attention map for the first denoising step, wherein the first attention map corresponds to a first concept in the text query and the second attention map corresponds to a second concept in the text query;

determining a third attention map and a fourth attention map for the second denoising step, wherein the third attention map corresponds to the first concept in the text query and the fourth attention map corresponds to the second concept in the text query;

determining an first attention retention loss for the first concept by comparing the first attention map for the first denoising step and the third attention map for the second denoising step, wherein the first attention retention loss indicates a retention of information for the first concept in the text query from the first denoising step to the second denoising step;

determining a second attention retention loss for the second concept by comparing the second attention map for the first denoising step and the fourth attention map for the second denoising step;

generating a modified noise representation from the second noise representation utilizing the first attention retention loss for the first concept and the second attention retention loss for the second concept; and

generating a text-conditioned image from the modified noise representation.

18 . The non-transitory computer-readable medium of claim 17 , wherein generating the second noise representation further comprises generating, utilizing a text encoder, a text vector representation from the text query to condition the second denoising step utilizing the text vector representation.

19 . The non-transitory computer-readable medium of claim 17 , wherein determining the first attention retention loss further comprises:

determining a threshold activation region for the first attention map corresponding to the first denoising step;

generating a binary mask for the first attention map corresponding to the first denoising step based on the threshold activation region; and

comparing the binary mask with the third attention map corresponding to the second denoising step to determine the first attention retention loss.

20 . The non-transitory computer-readable medium of claim 17 , wherein the operations further comprise determining an attention segregation loss from the text query comprising a first concept and a second concept by:

generating, for the first denoising step, a first attention map corresponding to the first concept and a second attention map corresponding to the second concept; and

comparing the first attention map corresponding to the first concept to the second attention map corresponding to the second concept to determine the attention segregation loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: AGARWAL, AISHWARYA; KARANAM, SRIKRISHNA; JOSE, JOSEPH KOONTHANAM; SAXENA, APOORV UMANG; GOSWAMI, KOUSTAVA; SRINIVASAN, BALAJI VASAN
To: ADOBE INC.
Reel/Frame 063994/0812 →
Continuity (1)
Related Publication 20240428468A1 · Dec 26, 2024
References Cited (42)
US 11809523B2 · Kastaniotis · 2023 [cited by examiner]
US 20250225700A1 · Li · 2025 [cited by examiner]
CN 113140023A · 2021 [cited by applicant]
WO WO2024243527A1 · 2024 [cited by examiner]
Omri Avrahami, Dani Lischinski, Ohad Fried, Blended Diffusion for Text-driven Editing of Natural Images, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (Year: 2022). [cited by examiner]
Aaron Van Den Oord, et al. Neural discrete representation learning. Advances in neural information processing systems, 31st Conference on Neural Information Processing Systems, Nov. 2, 2017. [cited by applicant]
Aditya Ramesh, et al. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, Apr. 13, 2022. [cited by applicant]
Alec Radford, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748-8763. PMLR, Feb. 26, 2021. [cited by applicant]
Alex Nichol, et al. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, Dec. 20, 2021. [cited by applicant]
Amir Hertz, et al. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, Aug. 2, 2022. [cited by applicant]
Ashish Vaswani, et al. Attention is all you need. Advances in neural information processing systems, 31st Conference on Neural Information Processing Systems, Jun. 12, 2017. [cited by applicant]
Chitwan Saharia, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, May 23, 2022. [cited by applicant]
Colin Raffel, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485-5551, 2020. [cited by applicant]
Diederik P Kingma, et al. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, Dec. 20, 2013. [cited by applicant]
Han Zhang, et al. Cross-modal contrastive learning for text-to-image generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 833-842, Jan. 12, 2021. [cited by applicant]
Hila Chefer, et al. Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. arXiv preprint arXiv:2301.13826, SIGGRAPH 2023, Jan. 31, 2023. [cited by applicant]
Huaibo Huang, et al. Introvae: Introspective variational autoencoders for photographic image synthesis. Advances in neural information processing systems, 31, Jul. 17, 2018. [cited by applicant]
Hui Ye, et al. Improving text-to-image synthesis using contrastive learning. arXiv preprint arXiv:2107.02423, Jul. 6, 2021. [cited by applicant]
Jonathan Ho, et al. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, Jul. 26, 2022. [cited by applicant]
Jun-Yan Zhu, et al. Toward multimodal image-to-image translation. Advances in neural information processing systems, 31st Conference on Neural Information Processing Systems, Nov. 30, 2017. [cited by applicant]
Jun-Yan Zhu, et al. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision, pp. 2223-2232, Mar. 30, 2017. [cited by applicant]
Junnan Li, et al. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pp. 12888-12900. PMLR, 2022. [cited by applicant]
Minfeng Zhu, et al. Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5802-5810, Apr. 2, 2019. [cited by applicant]
Ming Tao, et al. Df-gan: A simple and effective baseline for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16515-16525, 2022. [cited by applicant]
Nan Liu, et al. Compositional visual generation with composable diffusion models. In Computer Vision-ECCV 2022: 17th European Conference, Tel Aviv, Israel, Oct. 23-27, 2022, Proceedings, Part XVII, pp. 423-439. Springer… [cited by applicant]
Olaf Ronneberger, et al. U-950 net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015: 18th International Conference, Munich, Germany, Oc… [cited by applicant]
Phillip Isola, et al. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1125-1134, 2017. [cited by applicant]
Robin Rombach, et al. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684-10695, 2022. [cited by applicant]
Ron Mokady, et al. Null-text inversion for editing real images using guided diffusion models. arXiv preprint arXiv:2211.09794, Nov. 17, 2022. [cited by applicant]
Sam Witteveen, et al. Investigating prompt engineering in diffusion models. arXiv preprint arXiv:2211.15462, Nov. 21, 2022. [cited by applicant]
Taesung Park, et al. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2337-2346, Mar. 18, 2019. [cited by applicant]
Tao Xu, et al. Attngan: Fine- grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1316-1324, 2018. [cited by applicant]
Tero Karras, et al. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4401-4410, 2019. [cited by applicant]
Tim Brooks, et al. In- structpix2pix: Learning to follow image editing instructions. arXiv preprint arXiv:2211.09800, CVPR 2023, pp. 1-15. Nov. 17, 2022. [cited by applicant]
Vivian Liu, et al. Design guidelines for prompt engineering text-to-image generative models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pp. 1-23, 2022. [cited by applicant]
Weixi Feng, et al. Training-free structured diffusion guidance for compositional text-to-image synthesis. arXiv preprint arXiv:2212.05032, Dec. 9, 2022. [cited by applicant]
Yaru Hao, et al. Optimizing prompts for text-to-image generation. arXiv preprint arXiv:2212.09611, Dec. 19, 2022. [cited by applicant]
Yu Zeng, et al. Scenecomposer: Any-level semantic image synthesis. arXiv preprint arXiv:2211.11742, Nov. 21, 2022. [cited by applicant]
Yuheng Li, et al. Gligen: Open-set grounded text-to-image generation. arXiv preprint arXiv:2301.07093, Jan. 17, 2023. [cited by applicant]
Zhengyuan Yang, et al. Reco: Region-controlled text-to-image generation. arXiv preprint arXiv:2211.15518, Nov. 23, 2022. [cited by applicant]
Zijie J Wang, et al. Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models. arXiv preprint arXiv:2210.14896, Oct. 26, 2022. [cited by applicant]
Combined Search and Examination Report as received in United Kingdom application No. 2405389.4 dated Sep. 10, 2024. [cited by applicant]