IP Library Granted Patent US 12,586,271
Granted Patent B2
US 12,586,271 · App. 18/329,111 · Granted Mar 24, 2026

Color conditioned diffusion prior

Inventors: Pranav Vineet Aggarwal (Santa Clara, CA); Venkata Naveen Kumar Yadav Marri (Newark, CA); Midhun Harikumar (Sunnyvale, CA); Sachin Madhav Kelkar (Santa Clara, CA); Hareesh Ravi (San Jose, CA); Ajinkya Gorakhnath Kale (San Jose, CA)
Assignee: ADOBE INC.
G06T11/60G06F40/40G06N3/045G06N3/08G06T11/001
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,271
App. No.
18/329,111
Granted
Mar 24, 2026
Kind
B2
Abstract

Systems and methods for image processing are described. Embodiments of the present disclosure, via a multi-modal encoder of an image processing apparatus, encodes a text prompt to obtain a text embedding. A color encoder of the image processing apparatus encodes a color prompt to obtain a color embedding. A diffusion prior model of the image processing apparatus generates an image embedding based on the text embedding and the color embedding. A latent diffusion model of the image processing apparatus generates an image based on the image embedding, where the image includes an element from the text prompt and a color from the color prompt.

Claims (54)

1 . A method comprising:

encoding a text prompt to obtain a text embedding;

encoding a color prompt to obtain a color embedding, wherein the color prompt comprises a different modality than the text prompt;

generating an image embedding using a diffusion prior model based on the text embedding and the color embedding, wherein the image embedding represents the text prompt and the color prompt; and

generating an image by denoising a noise input based on the image embedding that represents the text prompt and the color prompt using a latent diffusion model (LDM), wherein the image includes an element from the text prompt and a color from the color prompt.

2 . The method of claim 1 , wherein:

the text embedding and the image embedding are in a multi-modal embedding space.

3 . The method of claim 2 , wherein:

the text embedding is in a first region of the multi-modal embedding space corresponding to text and the image embedding is in a second region of the multi-modal embedding space corresponding to images.

4 . The method of claim 1 , wherein generating the image embedding comprises:

performing an attention process using shared attention weights among the text embedding and the color embedding.

5 . The method of claim 1 , wherein:

the color embedding comprises a color histogram.

6 . The method of claim 1 , further comprising:

identifying a candidate image embedding for a candidate image;

comparing the image embedding to the candidate image embedding; and

providing the candidate image as a search result based on the comparison.

7 . The method of claim 1 , further comprising:

generating a plurality of image embeddings based on the text embedding and the color embedding, wherein the image embedding is selected from the plurality of image embeddings.

8 . The method of claim 7 , further comprising:

generating a plurality of images based on the plurality of image embeddings, respectively, wherein each of the plurality of images includes the element from the text prompt and the color from the color prompt.

9 . The method of claim 1 , further comprising:

generating a plurality of images based on the image embedding, wherein each of the plurality of images includes the element from the text prompt and the color from the color prompt.

10 . The method of claim 1 , further comprising:

generating a modified image embedding using the LDM, wherein the image is generated based on the modified image embedding.

11 . A method comprising:

obtaining training data including a text embedding representing a text prompt and a color embedding representing a color prompt, wherein the color prompt comprises a different modality than the text prompt;

initializing a diffusion prior model; and

training the diffusion prior model to generate an image embedding based on the text embedding and the color embedding, wherein the image embedding represents features corresponding to the text embedding and a color corresponding to the color embedding, and wherein the image embedding represents the text prompt and the color prompt.

12 . The method of claim 11 , further comprising:

encoding the text prompt describing a ground-truth image to obtain the text embedding; and

encoding the ground-truth image to obtain the color embedding.

13 . The method of claim 11 , further comprising:

training a latent diffusion model (LDM) to generate an image by denoising a noise input based on the image embedding that represents the text prompt and the color prompt.

14 . The method of claim 11 , further comprising:

generating a predicted image embedding using the diffusion prior model; and

computing a loss function by comparing the predicted image embedding to a ground-truth image embedding, wherein the diffusion prior model is trained based on the loss function.

15 . An apparatus comprising:

at least one processor; and

at least one memory including instructions executable by the at least one processor to perform operations including:

encoding, using a multi-modal encoder, a text prompt to obtain a text embedding;

encoding, using a color encoder, a color prompt to obtain a color embedding, wherein the color prompt comprises a different modality than the text prompt;

generating, using a diffusion prior model, an image embedding based on the text embedding and the color embedding, wherein the image embedding represents the text prompt and the color prompt; and

generating, using a latent diffusion model (LDM), an image by denoising a noise input based on the image embedding that represents the text prompt and the color prompt, wherein the image includes an element from the text prompt and a color from the color prompt.

16 . The apparatus of claim 15 , wherein:

the color encoder comprises a color histogram extractor configured to extract a color histogram from the color prompt.

17 . The apparatus of claim 15 , wherein:

the diffusion prior model comprises a transformer architecture.

18 . The apparatus of claim 15 , wherein:

the LDM comprises a U-Net architecture.

19 . The apparatus of claim 15 , wherein:

the LDM comprises an image decoder.

20 . The apparatus of claim 15 , further comprising:

a training component configured to train the diffusion prior model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2023
From: AGGARWAL, PRANAV VINEET; MARRI, VENKATA NAVEEN KUMAR YADAV; HARIKUMAR, MIDHUN; KELKAR, SACHIN MADHAV; RAVI, HAREESH; KALE, AJINKYA GORAKHNATH
To: ADOBE INC.
Reel/Frame 063854/0432 →
Continuity (1)
Related Publication 20240404144A1 · Dec 5, 2024
References Cited (23)
US 10713821B1 · Surya · 2020 [cited by examiner]
US 11216505B2 · Motiian et al. · 2022 [cited by applicant]
US 11922550B1 · Ramesh · 2024 [cited by examiner]
US 11995803B1 · Karpman et al. · 2024 [cited by applicant]
US 20110069325A1 · Kawashima · 2011 [cited by examiner]
US 20220343561A1 · Aggarwal · 2022 [cited by examiner]
US 20230118966A1 · Liu et al. · 2023 [cited by applicant]
US 20240112088A1 · Yu et al. · 2024 [cited by applicant]
US 20240153152A1 · Liu et al. · 2024 [cited by applicant]
US 20240185035A1 · Yu et al. · 2024 [cited by applicant]
US 20240282131A1 · Ren et al. · 2024 [cited by applicant]
US 20240303764A1 · Pang et al. · 2024 [cited by applicant]
US 20240331235A1 · Smock · 2024 [cited by examiner]
1Li, et al., “Universal Style Transfer via Feature Transforms”, 31st Conference on Neural Information Processing Systems (NIPS 2017), 11 pages. [cited by applicant]
2Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, In International Conference on Machine Learning (pp. 8748-8763), PMLR, arXiv preprint: arXiv:2103.00020v1 [cs.CV] Feb. 26, 2021. [cited by applicant]
3Bramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv preprint: arXiv:2204.06125v1 [cs.CV] Apr. 13, 2022, 27 pages. [cited by applicant]
4Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10684-10695), arXiv preprint: arXiv:2112.10752v… [cited by applicant]
5Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv preprint: arXiv:2205.11487v1 [cs.CV] May 23, 2022, 46 pages. [cited by applicant]
7Align: Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. Tuesday, May 11, 2021. Posted by Chao Jia and Yinfei Yang, Software Engineers, Google Research. Found on the internet at… [cited by applicant]
8GitHub—FreddeFrallan/Multilingual-CLIP: OpenAI CLIP text encoders for multiple languages, 7 pages. Found on the Internet, https://github.com/FreddeFrallan/Multilingual-CLIP#readme. [cited by applicant]
Office Action dated Jun. 3, 2025 in related U.S. Appl. No. 18/296,002. [cited by applicant]
Office Action dated Aug. 14, 2025 in related U.S. Appl. No. 18/296,002. [cited by applicant]
Office Action dated May 19, 2025 in related U.S. Appl. No. 18/301,671. [cited by applicant]