IP Library › Granted Patent US 12,493,937
Granted Patent B2
US 12,493,937 · App. 18/301,671 · Granted Dec 9, 2025

Prior guided latent diffusion

Inventors: Midhun Harikumar (Sunnyvale, CA); Venkata Naveen Kumar Yadav Marri (Newark, CA); Ajinkya Gorakhnath Kale (San Jose, CA); Pranav Vineet Aggarwal (Santa Clara, CA); Vinh Ngoc Khuc (Campbell, CA)
Assignee: ADOBE INC.
G06T5/73G06F40/279G06T5/50
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,937
App. No.
18/301,671
Granted
Dec 9, 2025
Kind
B2
Abstract

Systems and methods for image processing are described. Embodiments of the present disclosure obtain a text prompt for text guided image generation. A multi-modal encoder of an image processing apparatus encodes the text prompt to obtain a text embedding. A diffusion prior model of the image processing apparatus converts the text embedding to an image embedding. A latent diffusion model of the image processing apparatus generates an image based on the image embedding, wherein the image includes an element described by the text prompt.

Claims (60)

1 . A method comprising:

obtaining a text prompt;

encoding the text prompt to obtain a text embedding;

converting the text embedding to an image embedding space using a diffusion prior model by performing diffusion based denoising to predict a denoised image embedding; and

generating an image based on the denoised image embedding using a latent diffusion model (LDM) different from the diffusion prior model, wherein the LDM receives input noise and the denoised image embedding as input and provides the image as output, and wherein the image includes an element described by the text prompt.

2 . The method of claim 1 , further comprising:

generating a plurality of image embeddings based on the text embedding, wherein the denoised image embedding is selected from the plurality of image embeddings.

3 . The method of claim 2 , further comprising:

generating a plurality of images based on the plurality of image embeddings, respectively, wherein each of the plurality of images includes the element described by the text prompt.

4 . The method of claim 2 , further comprising:

comparing each of the plurality of image embeddings to the text embedding; and

computing a similarity score based on the comparison, wherein the denoised image embedding is selected based on the similarity score.

5 . The method of claim 1 , further comprising:

generating a plurality of images based on the denoised image embedding, wherein each of the plurality of images includes the element described by the text prompt.

6 . The method of claim 5 , further comprising:

generating a plurality of noise maps, wherein each of the plurality of images is generated based on one of the plurality of noise maps.

7 . The method of claim 5 , wherein:

the plurality of images has different aspect ratios.

8 . The method of claim 1 , wherein:

the denoised image embedding is in a same embedding space as the text embedding.

9 . The method of claim 1 , further comprising:

generating a modified image embedding using the LDM, wherein the image is generated based on the modified image embedding.

10 . The method of claim 1 , further comprising:

obtaining an image prompt;

encoding the image prompt to obtain an additional image embedding; and

generating an additional image based on the additional image embedding using the LDM, wherein the additional image includes an additional element of the image prompt.

11 . A method comprising:

obtaining training data including a text embedding;

predicting a denoised image embedding based on the text embedding;

computing a loss function based on the denoised image embedding; and

training a diffusion prior model based on the loss function, wherein a latent diffusion model (LDM) receives input noise and a denoised output from the diffusion prior model as input and provides an image as output.

12 . The method of claim 11 , further comprising:

encoding a training image to obtain an image embedding using a multi-modal encoder;

initializing the latent diffusion model (LDM); and

training the LDM to generate a synthetic image based on the image embedding.

13 . The method of claim 12 , further comprising:

obtaining additional training data including an additional image and a text describing the additional image;

encoding the additional image and the text to obtain an additional image embedding and the text embedding respectively, using the multi-modal encoder; and

generating a predicted image embedding using the diffusion prior model based on the additional image embedding.

14 . The method of claim 13 , further comprising:

adding a corresponding noise from a plurality of noise to the additional image embedding to obtain a plurality of additional noise image embeddings;

generating a plurality of predicted image embeddings based on the plurality of additional noise image embeddings;

comparing the plurality of predicted image embeddings to the text embedding; and

selecting the predicted image embedding from the plurality of predicted image embeddings based on the comparison.

15 . The method of claim 13 , wherein:

the image embedding is based on the predicted image embedding.

16 . An apparatus comprising:

at least one processor; and

at least one memory including instructions executable by the at least one processor to perform operations including:

encoding, using a multi-modal encoder, a text prompt to obtain a text embedding;

converting, using a diffusion prior model, the text embedding to an image embedding space by performing diffusion based denoising to predict a denoised image embedding; and

generating, using a latent diffusion model (LDM) different from the diffusion prior model, an image based on the denoised image embedding, wherein the LDM receives input noise and the denoised image embedding as input and provides the image as output, and wherein the image includes an element described by the text prompt.

17 . The apparatus of claim 16 , wherein:

the diffusion prior model comprises a transformer architecture.

18 . The apparatus of claim 16 , wherein:

the LDM comprises a U-Net architecture.

19 . The apparatus of claim 16 , wherein:

the LDM comprises an image decoder.

20 . The apparatus of claim 16 , further comprising:

a training component configured to train the LDM.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 17, 2023
From: HARIKUMAR, MIDHUN; MARRI, VENKATA NAVEEN KUMAR YADAV; KALE, AJINKYA GORAKHNATH; AGGARWAL, PRANAV VINEET; KHUC, VINH NGOC
To: ADOBE INC.
Reel/Frame 063346/0161 →
Continuity (1)
Related Publication 20240346629A1 · Oct 17, 2024
References Cited (25)
US 10713821B1 · Surya et al. · 2020 [cited by applicant]
US 11216505B2 · Motiian et al. · 2022 [cited by applicant]
US 11922550B1 · Ramesh et al. · 2024 [cited by applicant]
US 11995803B1 · Karpman et al. · 2024 [cited by applicant]
US 20110069325A1 · Kawashima et al. · 2011 [cited by applicant]
US 20220343561A1 · Aggarwal et al. · 2022 [cited by applicant]
US 20230118966A1 · Liu et al. · 2023 [cited by applicant]
US 20240112088A1 · Yu et al. · 2024 [cited by applicant]
US 20240153152A1 · Liu et al. · 2024 [cited by applicant]
US 20240185035A1 · Yu et al. · 2024 [cited by applicant]
US 20240282131A1 · Ren · 2024 [cited by examiner]
US 20240303764A1 · Pang et al. · 2024 [cited by applicant]
US 20240331235A1 · Smock · 2024 [cited by examiner]
Aditya Ramesh et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, pub. Apr. 13, 2022 (Year: 2022). [cited by examiner]
1Li, et al., “Universal Style Transfer via Feature Transforms”, 31st Conference on Neural Information Processing Systems (NIPS 2017), 11 pages. [cited by applicant]
2Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, In International Conference on Machine Learning (pp. 8748-8763), PMLR, arXiv preprint: arXiv:2103.00020v1 [cs.CV] Feb. 26, 2021. [cited by applicant]
3Rramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv preprint: arXiv:2204.06125v1 [cs.CV] Apr. 13, 2022, 27 pages. [cited by applicant]
4Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10684-10695), arXiv preprint: arXiv:2112.10752v… [cited by applicant]
5Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv preprint: arXiv:2205.11487v1 [cs.CV] May 23, 2022, 46 pages. [cited by applicant]
7Align: Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. Tuesday, May 11, 2021. Posted by Chao Jia and Yinfei Yang, Software Engineers, Google Research. Found on the internet at… [cited by applicant]
8GitHub—FreddeFrallan/Multilingual-CLIP: OpenAI CLIP text encoders for multiple languages, 7 pages. Found on the internet, https://github.com/FreddeFrallan/Multilingual-CLIP#readme. [cited by applicant]
Office Action dated Apr. 29, 2025 in related U.S. Appl. No. 18/329,111. [cited by applicant]
Office Action dated Jun. 3, 2025 in related U.S. Appl. No. 18/296,002. [cited by applicant]
Office Action dated Aug. 14, 2025 in related U.S. Appl. No. 18/296,002. [cited by applicant]
Office Action dated Sep. 18, 2025 in related U.S. Appl. No. 18/319,111. [cited by applicant]