IP Library › Granted Patent US 12,725,318
Granted Patent B2
US 12,725,318 · App. 18/296,002 · Granted Sep 1, 2026

Multilingual text-to-image generation

Inventors: Venkata Naveen Kumar Yadav Marri (Newark, CA); Ajinkya Gorakhnath Kale (San Jose, CA)
Assignee: ADOBE INC.
G06T11/00G06F40/58G06V10/74G06V10/774G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,318
App. No.
18/296,002
Granted
Sep 1, 2026
Kind
B2
Abstract

Systems and methods for image processing are provided. One aspect of the systems and methods includes obtaining a text prompt in a first language. Another aspect of the systems and methods includes encoding the text prompt using a multilingual encoder to obtain a multilingual text embedding. Yet another aspect of the systems and methods includes processing the multilingual text embedding using a diffusion prior model to obtain an image embedding, wherein the diffusion prior model is trained to process multilingual text embeddings from the first language and a second language based on training data from the first language and the second language. Yet another aspect of the systems and methods includes generating an image using a diffusion model based on the image embedding, wherein the image includes an element corresponding to the text prompt.

Claims (53)

1 . A method comprising:

obtaining a text prompt in a first language;

encoding the text prompt using a multilingual encoder to obtain a multilingual text embedding;

processing the multilingual text embedding using a diffusion prior model to obtain an image embedding, wherein the diffusion prior model is trained to perform denoising based on multilingual text embeddings from the first language and a second language to obtain input guidance for an image generation model different from the diffusion prior model, and wherein the diffusion prior model is trained based on training data from the first language and the second language; and

generating an image using the image generation model by performing denoising based on the image embedding, wherein the image includes an element corresponding to the text prompt.

2 . The method of claim 1 , further comprising:

obtaining an additional text prompt in the second language;

encoding the additional text prompt using the multilingual encoder to obtain an additional multilingual text embedding;

processing the additional multilingual text embedding using the diffusion prior model to obtain an additional image embedding; and

generating an additional image using the diffusion model based on the additional image embedding, wherein the additional image includes an additional element corresponding to the additional text prompt.

3 . The method of claim 1 , further comprising:

generating a plurality of intermediate image embeddings corresponding to a plurality of diffusion time steps, wherein the image is generated based on the plurality of intermediate image embeddings.

4 . The method of claim 1 , further comprising:

obtaining a causal attention mask, wherein the image embedding is generated based on the causal attention mask.

5 . The method of claim 1 , further comprising:

generating a plurality of image embeddings using the diffusion prior model;

computing a similarity score between each of the plurality of image embeddings and the multilingual text embedding; and

selecting the image embedding from the plurality of image embeddings based on the similarity score.

6 . The method of claim 1 , wherein:

the image embedding is in a same embedding space as the multilingual text embedding.

7 . A method comprising:

obtaining training data including a plurality of images, a first plurality of image captions in a first language, and a second plurality of image captions in a second language;

encoding the first plurality of image captions and the second plurality of image captions using a multilingual encoder to obtain a plurality of multilingual text embeddings;

processing the plurality of multilingual text embeddings using a diffusion prior model to obtain a plurality of predicted image embeddings corresponding to the first plurality of image captions in the first language and the second plurality of image captions in the second language to obtain input guidance for an image generation model different from the diffusion prior model, and wherein the diffusion prior model is trained based on training data from the first language and the second language; and

training the diffusion prior model to generate image embeddings based on multilingual text embeddings from the first language and the second language, wherein the diffusion prior model is trained based on the plurality of predicted image embeddings and the plurality of images.

8 . The method of claim 7 , further comprising:

identifying a plurality of ground-truth image embeddings corresponding to the plurality of images, respectively; and

comparing the plurality of predicted image embeddings to the plurality of ground-truth image embeddings, wherein the diffusion prior model is trained based on the comparison.

9 . The method of claim 7 , further comprising:

generating a plurality of predicted images based on the plurality of predicted image embeddings using a diffusion model; and

comparing the plurality of predicted images to the plurality of images, respectively, wherein the diffusion prior model is trained based on the comparison.

10 . The method of claim 9 , wherein:

the diffusion model is pretrained prior to training the diffusion prior model.

11 . The method of claim 7 , further comprising:

translating the first plurality of image captions to obtain the second plurality of image captions.

12 . The method of claim 7 , wherein:

the plurality of images includes a first subset of images corresponding to the first language and a second subset of images corresponding to the second language, the first subset of images being different from the second subset of images.

13 . The method of claim 7 , wherein:

the multilingual encoder is pretrained prior to training the diffusion prior model.

14 . A system comprising:

at least one memory component; and

at least one processing device coupled to the at least one memory component, wherein the processing device is configured to execute instructions stored in the at least one memory component to perform operations comprising:

processing, using a diffusion prior model, a multilingual text embedding from a first language to obtain an image embedding, wherein the diffusion prior model is trained to perform denoising based on multilingual text embeddings from the first language and a second language to obtain input guidance for an image generation model different from the diffusion prior model, and wherein the diffusion prior model is trained based on training data from the first language and the second language; and

generating, using the image generation model comprising a diffusion model, an image by performing denoising based on the image embedding, wherein the image includes an element corresponding to the multilingual text embedding.

15 . The system of claim 14 , wherein the at least one processing device is configured to execute instructions stored in the at least one memory component to perform operations comprising:

encoding, using a multilingual encoder, a text prompt in the first language to obtain the multilingual text embedding.

16 . The system of claim 15 , wherein the multilingual encoder comprises a multimodal encoder for text and images.

17 . The system of claim 14 , wherein the at least one processing device is configured to execute instructions stored in the at least one memory component to perform operations comprising:

training the diffusion prior model to generate image embeddings based on multilingual text embeddings from the first language and the second language, wherein the diffusion prior model is trained based on a plurality of predicted images.

18 . The system of claim 14 , wherein:

the image embedding is in a same embedding space as the multilingual text embedding.

19 . The system of claim 14 , wherein the diffusion prior model comprises a transformer architecture.

20 . The system of claim 14 , wherein the diffusion model comprises a UNet architecture.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2023
From: MARRI, VENKATA NAVEEN KUMAR YADAV; KALE, AJINKYA GORAKHNATH
To: ADOBE INC.
Reel/Frame 063229/0937 →
Continuity (1)
Related Publication 20240338859A1 · Oct 10, 2024
References Cited (23)
US 10713821B1 · Surya et al. · 2020 [cited by applicant]
US 11216505B2 · Motiian et al. · 2022 [cited by applicant]
US 11922550B1 · Ramesh et al. · 2024 [cited by applicant]
US 11995803B1 · Karpman · 2024 [cited by examiner]
US 20110069325A1 · Kawashima et al. · 2011 [cited by applicant]
US 20220343561A1 · Aggarwal et al. · 2022 [cited by applicant]
US 20230118966A1 · Liu · 2023 [cited by examiner]
US 20240112088A1 · Yu · 2024 [cited by examiner]
US 20240153152A1 · Liu · 2024 [cited by examiner]
US 20240185035A1 · Yu · 2024 [cited by examiner]
US 20240282131A1 · Ren et al. · 2024 [cited by applicant]
US 20240303764A1 · Pang · 2024 [cited by examiner]
US 20240331235A1 · Smock et al. · 2024 [cited by applicant]
Office Action dated Apr. 29, 2025 in related U.S. Appl. No. 18/329,111. [cited by applicant]
Office Action dated May 19, 2025 in related U.S. Appl. No. 18/301,671. [cited by applicant]
Li, et al., “Universal Style Transfer via Feature Transforms”, 31st Conference on Neural Information Processing Systems (NIPS 2017), 11 pages. [cited by applicant]
Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, In International Conference on Machine Learning (pp. 8748-8763), PMLR, arXiv preprint: arXiv:2103.00020v1 [cs.CV] Feb. 26, 2021. [cited by applicant]
Rramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv preprint: arXiv:2204.06125v1 [cs.CV] Apr. 13, 2022, 27 pages. [cited by applicant]
Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10684-10695), arXiv preprint: arXiv:2112.10752v2… [cited by applicant]
Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv preprint: arXiv:2205.11487v1 [cs.CV] May 23, 2022, 46 pages. [cited by applicant]
Align: Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. Tuesday, May 11, 2021. Posted by Chao Jia and Yinfei Yang, Software Engineers, Google Research. Found on the internet at:… [cited by applicant]
GitHub—FreddeFrallan/Multilingual-CLIP: OpenAI CLIP text encoders for multiple languages, 7 pages. Found on the Internet, https://github.com/FreddeFrallan/Multilingual-CLIP#readme. [cited by applicant]
Office Action dated Sep. 18, 2025 in related U.S. Appl. No. 18/319,111. [cited by applicant]