IP Library › Granted Patent US 12,626,420
Granted Patent B2
US 12,626,420 · App. 18/511,692 · Granted May 12, 2026

Segmentation free guidance in diffusion models

Inventors: Kambiz Azarian Yazdi (San Diego, CA); Fatih Murat Porikli (San Diego, CA); Qiqi Hou (San Diego, CA); Debasmit Das (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06T11/00G06F40/284G06T5/70G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,420
App. No.
18/511,692
Granted
May 12, 2026
Kind
B2
Abstract

Certain aspects of the present disclosure provide techniques for generating an output image based on a text prompt. A method may include receiving the text prompt; providing a user interface comprising one or more input elements associated with one or more words of the text prompt; receiving input corresponding to at least one of the one or more input elements, the input indicating a semantic importance for each of at least one of the one or more words associated with the at least one of the one or more input elements; and generating the output image based on the text prompt and the input.

Claims (78)

1 . An apparatus configured to generate an output image based on a text prompt, comprising:

one or more memories configured to store a latent image representation; and

one or more processors, coupled to the one or more memories, configured to:

obtain the text prompt;

encode the text prompt into a plurality of conditioning tokens;

for each of a plurality of patches of the latent image representation:

calculate a respective plurality of cross-attention weights corresponding to the plurality of conditioning tokens based on the patch as a query and the plurality of conditioning tokens as a key; and

modify a maximum value cross-attention weight among the respective plurality of cross-attention weights to generate a modified respective plurality of cross-attention weights;

perform an iteration of denoising using the modified respective plurality of cross-attention weights for each of the plurality of patches to obtain a modified latent image representation; and

generate the output image based on the modified latent image representation.

2 . The apparatus of claim 1 , wherein to modify the maximum value cross-attention weight among the respective plurality of cross-attention weights comprises to reduce the maximum value cross-attention weight.

3 . The apparatus of claim 1 , wherein to modify the maximum value cross-attention weight among the respective plurality of cross-attention weights comprises to multiply the maximum value cross-attention weight by a negative scalar value.

4 . The apparatus of claim 3 , wherein the one or more processors are configured to receive the negative scalar value as a user-specified parameter.

5 . The apparatus of claim 1 , wherein to modify the maximum value cross-attention weight among the respective plurality of cross-attention weights comprises to set the maximum value cross-attention weight to zero.

6 . The apparatus of claim 1 , wherein the one or more processors are configured to perform one or more initial denoising iterations using classifier-free guidance prior to performing the iteration of denoising using the modified respective plurality of cross-attention weights for each of the plurality of patches.

7 . The apparatus of claim 1 , further comprising a display, coupled to the one or more processors, configured to display a user interface configured to receive input indicative of an emphasis strength associated with one or more words of the text prompt, wherein the emphasis strength controls an amount to modify the maximum value cross-attention weight among the respective plurality of cross-attention weights.

8 . The apparatus of claim 7 , wherein the user interface includes one or more interface elements configured to receive the input, the one or more interface elements including at least one of a slider, a numerical input, or a keyword highlight.

9 . The apparatus of claim 1 , wherein the one or more processors are configured to:

for each of a plurality of patches of the modified latent image representation:

calculate a respective second plurality of cross-attention weights corresponding to the plurality of conditioning tokens based on the patch as another query and the plurality of conditioning tokens as another key; and

modify another maximum value cross-attention weight among the respective second plurality of cross-attention weights to generate another modified respective plurality of second cross-attention weights; and

perform a second iteration of denoising using the modified respective plurality of second cross-attention weights for each of the plurality of patches of the modified latent image representation to obtain a second modified latent image representation, wherein to generate the output image based on the modified latent image representation comprises to decode the second modified latent image representation using a decoder to generate the output image.

10 . The apparatus of claim 1 , further comprising a display, coupled to the one or more processors, configured to display the output image.

11 . The apparatus of claim 1 , further comprising:

a modem, coupled to the one or more processors, configured to modulate one or more carrier wave signals with data indicative of the output image.

12 . The apparatus of claim 11 , further comprising:

one or more antennas, coupled to the modem, configured to transmit the one or more carrier wave signals to a device.

13 . The apparatus of claim 1 , wherein to perform the iteration of denoising, the one or more processors are configured to:

generate a first output based on the latent image representation and the respective plurality of cross-attention weights for each of the plurality of patches of the latent image representation;

generate a second output based on the latent image representation and the modified respective plurality of cross-attention weights for each of the plurality of patches of the latent image representation; and

subtract the second output from the first output to obtain the modified latent image representation.

14 . The apparatus of claim 1 , wherein to perform the iteration of denoising, the one or more processors are configured to:

generate the modified latent image representation based on the latent image representation and the modified respective plurality of cross-attention weights for each of the plurality of patches of the latent image representation.

15 . An apparatus configured to generate an output image based on a text prompt, comprising:

one or more memories configured to store the output image; and

one or more processors, coupled to the one or more memories, configured to:

receive the text prompt;

provide a user interface comprising one or more input elements associated with one or more words of the text prompt;

receive input corresponding to at least one of the one or more input elements, the input indicating a semantic importance for each of at least one of the one or more words associated with the at least one of the one or more input elements; and

generate the output image based on the text prompt and the input.

16 . The apparatus of claim 15 , further comprising a display, coupled to the one or more processors, configured to:

display the user interface; and

display the output image.

17 . The apparatus of claim 15 , wherein to generate the output image, the one or more processors are configured to generate the output image using the text prompt and the input as inputs to a generative artificial intelligence (AI) model.

18 . The apparatus of claim 15 , wherein a first input element of the one or more input elements is a slider element configured to increase or decrease the importance of a first word of the one or more words.

19 . The apparatus of claim 15 , wherein a first input element of the one or more input elements is a dial element configured to increase or decrease the importance of a first word of the one or more words.

20 . The apparatus of claim 15 , wherein the one or more processors are configured to modify an appearance of the at least one of the one or more words based on the indicated semantic importance.

21 . The apparatus of claim 20 , wherein to modify the appearance of the at least one of the one or more words comprises to highlight the at least one of the one or more words.

22 . The apparatus of claim 15 , wherein the output image emphasizes one or more objects associated with the at least one of the one or more words indicated as having higher semantic importance as compared to other words of the one or more words.

23 . The apparatus of claim 15 , further comprising a display, coupled to the one or more processors, configured to:

prior to the one or more processors receiving the input, display a first image associated with the text prompt; and

after the one or more processors receiving the input, display the output image.

24 . The apparatus of claim 23 , wherein the display is configured to:

display a transition from the first image to the output image.

25 . The apparatus of claim 23 , wherein the output image emphasizes one or more objects associated with the at least one of the one or more words as compared to the one or more objects in the first image.

26 . The apparatus of claim 23 , wherein:

the one or more processors are configured to generate the first image using the text prompt as input to a generative artificial intelligence (AI) model; and

to generate the output image, the one or more processors are configured to generate the output image using the text prompt and the input as inputs to the generative AI model.

27 . The apparatus of claim 15 , wherein the one or more processors are configured to:

encode the text prompt into a plurality of conditioning tokens;

for each of one or more patches of a latent image representation:

calculate a respective plurality of cross-attention weights corresponding to the plurality of conditioning tokens based on the patch as a query and the plurality of conditioning tokens as a key; and

modify a maximum value cross-attention weight among the respective plurality of cross-attention weights to generate a modified respective plurality of cross-attention weights; and

perform an iteration of denoising using the modified respective plurality of cross-attention weights for each of the one or more patches to generate the output image.

28 . The apparatus of claim 27 , wherein at least one cross-attention weight among the respective plurality of cross-attention weights is modified based on the indicated semantic importance.

29 . A method comprising:

obtaining a text prompt;

encoding the text prompt into a plurality of conditioning tokens;

for each of a plurality of patches of a latent image representation:

calculating a respective plurality of cross-attention weights corresponding to the plurality of conditioning tokens based on the patch as a query and the plurality of conditioning tokens as a key; and

modifying a maximum value cross-attention weight among the respective plurality of cross-attention weights to generate a modified respective plurality of cross-attention weights;

performing an iteration of denoising using the modified respective plurality of cross-attention weights for each of the plurality of patches to obtain a modified latent image representation; and

generating an output image based on the modified latent representation.

30 . A method comprising:

receiving a text prompt;

providing a user interface comprising one or more input elements associated with one or more words of the text prompt;

receiving input corresponding to at least one of the one or more input elements, the input indicating a semantic importance for each of at least one of the one or more words associated with the at least one of the one or more input elements; and

generating an output image based on the text prompt and the input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2024
From: AZARIAN YAZDI, KAMBIZ; PORIKLI, FATIH MURAT; HOU, QIQI; DAS, DEBASMIT
To: QUALCOMM INCORPORATED
Reel/Frame 066039/0749 →
Continuity (1)
Related Publication 20250166236A1 · May 22, 2025
References Cited (32)
US 20240232580A1 · Jaegle · 2024 [cited by examiner]
US 20240331236A1 · Li · 2024 [cited by examiner]
US 20240355022A1 · Shi · 2024 [cited by examiner]
US 20240404144A1 · Aggarwal · 2024 [cited by examiner]
US 20250117967A1 · Tambi · 2025 [cited by examiner]
US 20250117973A1 · Chen · 2025 [cited by examiner]
CN 117056540A · 2023 [cited by applicant]
Karn et al.; “Image Synthesis Using GANs and Diffusion Models,” 2023 IEEE International Conference on Contemporary Computing and Communications (InC4), Bangalore, India, 2023, pp. 1-6; Date of Conference: Apr. 21-22, 20… [cited by examiner]
Chefer H., et al., “Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models”, arxig.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, ACM Transactions… [cited by applicant]
International Search Report and Written Opinion—PCT/US2024/052908—ISA/EPO—Feb. 4, 2025. [cited by applicant]
Rombach R., et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, arXiv: 2112.10752v2 [cs.Cv], XP091196342, … [cited by applicant]
Bansal A., et al., “Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise”, arXiv:2208.09392v1 [cs.CV], Aug. 19, 2022, pp. 1-22. [cited by applicant]
Bansal A., et al., “Universal Guidance for Diffusion Models”, arXiv:2302.07121v1 [cs.CV], Feb. 14, 2023, pp. 1-15. [cited by applicant]
Chung H., et al., “Diffusion Posterior Sampling for General Noisy Inverse Problems”, arXiv:2209.14687v3 [stat.ML], Feb. 27, 2023, pp. 1-30. [cited by applicant]
Chung H., et al., “Improving Diffusion Models for Inverse Problems using Manifold Constraints”, arXiv:2206.00941v2 [cs.LG], Oct. 3, 2022, pp. 1-28. [cited by applicant]
Dhariwal P., et al., “Diffusion Models Beat GANs on Image Synthesis”, arXiv:2105.05233v4 [cs.LG], Jun. 1, 2021, pp. 1-44. [cited by applicant]
Dieleman S., “Diffusion Models are Autoencoders”, sander.ai, Jan. 31, 2022, pp. 1-11. [cited by applicant]
Dieleman S., “Guidance—A Cheat Code for Diffusion Models”, sander.ai, May 26, 2022, pp. 1-13. [cited by applicant]
Graikos A., et al., “Diffusion Models as Plug and Play Priors”, arXiv:2206.09012v3 [cs.LG], Jan. 8, 2023, pp. 1-22. [cited by applicant]
Ho J., et al., “Classifier-Free Diffusion Guidance”, arXiv:2207.12598v1 [cs.LG], Jul. 26, 2022, pp. 1-14. [cited by applicant]
Ho J., et al., “Denoising Diffusion Probabilistic Models”, 34th Conference on Neural Information Processing Systems, 2020, pp. 1-12. [cited by applicant]
Kawar B., et al., “Denoising Diffusion Restoration Models”, arXiv:2201.11793v3 [eess.IV], Oct. 12, 2022, pp. 1-32. [cited by applicant]
Lugmayr A., et al., “RePaint: Inpainting using Denoising Diffusion Probabilistic Models”, arXiv:2201.09865v4 [cs.CV], Aug. 31, 2022, pp. 1-25. [cited by applicant]
Mongaras G., et al., “Diffusion Models—DDPMs, DDIMs, and Classifier Free Guidance”, Better Programming, Accessed on Nov. 16, 2023, pp. 1-45. [cited by applicant]
Nichol A., et al., “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models”, arXiv:2112.10741v3 [cs.CV], Mar. 8, 2022, 20 Pages. [cited by applicant]
Radford A., et al., “Learning Transferable Visual Models From Natural Language Supervision”, arXiv:2103.00020v1 [cs.CV], Feb. 26, 2021, pp. 1-48. [cited by applicant]
Rombach R., et al., “High-Resolution Image Synthesis with Latent Diffusion Models Robin”, IEEE Explore, CVF, Dec. 20, 2021, pp. 10684-10695. [cited by applicant]
Sohl-Dickstein J., et al., “Deep Unsupervised Learning using Nonequilibrium Thermodynamics”, arXiv:1503.03585v8 [cs.LG], Nov. 18, 2015, pp. 1-18. [cited by applicant]
Song Y., et al., “Generative Modeling by Estimating Gradients of the Data Distribution”, arXiv:1907.05600v3 [cs.LG], Oct. 10, 2020, pp. 1-23. [cited by applicant]
Wang W., et al., “Semantic Image Synthesis via Diffusion Models”, arXiv:2207.00050v2 [cs.CV], Nov. 22, 2022, 16 Pages. [cited by applicant]
Wang Y., et al., “Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model”, arXiv:2212.00490v2 [cs.CV], Dec. 7, 2022, pp. 1-31. [cited by applicant]
Whang J., et al., “Deblurring via Stochastic Refinement”, arXiv:2112.02475v2 [cs.CV], Dec. 28, 2021, pp. 1-28. [cited by applicant]