IP Library › Granted Patent US 12,462,457
Granted Patent B2
US 12,462,457 · App. 18/419,675 · Granted Nov 4, 2025

Systems and methods for hierarchical text-conditional image generation

Inventors: Aditya Ramesh (San Francisco, CA); Prafulla Dhariwal (San Francisco, CA); Alexander Nichol (San Francisco, CA); Casey Chu (San Francisco, CA); Mark Chen (Cupertino, CA)
Assignee: OpenAI Opco, LLC
G06T11/60G06F40/284G06F40/30G06N3/045G06N3/08G06T9/002G06T11/001
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,457
App. No.
18/419,675
Granted
Nov 4, 2025
Kind
B2
Abstract

Disclosed herein are methods, systems, and computer-readable media for generating an image corresponding to a text input. In an embodiment, operations may include accessing a text description and inputting the text description into a text encoder. The operations may include receiving, from the text encoder, a text embedding, and inputting at least one of the text description or the text embedding into a first sub-model configured to generate, based on at least one of the text description or the text embedding, a corresponding image embedding. The operations may include inputting at least one of the text description or the corresponding image embedding, generated by the first sub-model, into a second sub-model configured to generate, based on at least one of the text description or the corresponding image embedding, an output image. The operations may include making the output image, generated by the first second sub-model, accessible to a device.

Claims (60)

1 . A system comprising:

at least one memory storing instructions;

at least one processor configured to execute the instructions to perform operations for generating an image corresponding to a text input, the operations comprising:

generating, by a first sub-model, an image embedding based on at least one of a text description or a text embedding, wherein the text embedding is an encoding of the text description; and

generating, by a second sub-model, an output image based on at least one of the text description or the image embedding, wherein the second sub-model is different than the first sub-model;

based on a second text description, accessing a second text embedding;

generating a vector representation of the text embedding corresponding to the output image and the second text embedding;

performing an interpolation between the image embedding of the output image and the vector representation of the text embedding corresponding to the output image and the second text embedding; and

based on the interpolation, generating a modified instance of the output image.

2 . The system of claim 1 , wherein the at least one processor is further configured to execute the instructions to perform operations comprising:

jointly training an image encoder on a first data set and a text encoder on a second data set, wherein the text encoder is configured to generate one or more text embeddings, wherein the first data set comprises a set of images and the second data set comprises a set of text descriptions corresponding to the set of images;

receiving, from the text encoder, text embeddings associated with the second data set; and

receiving, from the image encoder, image embeddings associated with the first data set.

3 . The system of claim 1 , wherein the at least one processor is further configured to execute the instructions to perform operations comprising:

making the output image accessible to a device, wherein the device is associated with an image generation request.

4 . The system of claim 1 , wherein prior to generating the output image, the second sub-model further comprises: a diffusion model, the diffusion model including at least one of the text description, the text embedding, or the corresponding image embedding.

5 . The system of claim 1 , wherein the at least one processor is further configured to execute the instructions to perform operations comprising training an image generation model using the output image.

6 . The system of claim 1 , wherein the second sub-model includes at least one upsampler model configured for upsampling prior to generating the output image.

7 . The system of claim 1 , wherein the second text description provides a modification directive relative to the text description associated with the output image.

8 . The system of claim 1 , wherein the interpolation between the image embedding and the vector representation is performed in a latent embedding space.

9 . The system of claim 1 , wherein the interpolation comprises generating a weighted combination of the image embedding and the vector representation based on one or more weighting factors.

10 . A method for generating an image corresponding to a text input, comprising:

generating, by a first sub-model, an image embedding based on inputting at least one of a text description or a text embedding to the first sub-model, wherein the text embedding is an encoding of the text description;

generating, by a second sub-model, an output image based on inputting at least one of the text description or the image embedding to the second sub-model, wherein the second sub-model is different than the first sub-model;

based on a second text description, accessing a second text embedding;

generating a vector representation of the text embedding corresponding to the output image and the second text embedding;

performing an interpolation between the image embedding of the output image and the vector representation of the text embedding corresponding to the output image and the second text embedding; and

based on the interpolation, generating a modified instance of the output image.

11 . The method of claim 10 , further comprising:

training an image encoder on a first data set and a text encoder on a second data set, wherein the text encoder is configured to generate one or more text embeddings, wherein the first data set comprises a set of images and the second data set comprises a set of text descriptions corresponding to the set of images;

receiving, from the text encoder, text embeddings associated with the second data set; and

receiving, from the image encoder, image embeddings associated with the first data set.

12 . The method of claim 11 , further comprising:

encoding the output image with the image encoder;

applying a decoder to the image; and

obtaining a joint latent representation of the image.

13 . The method of claim 10 , wherein prior to generating the output image, the second sub-model further comprises: a diffusion model, the diffusion model including at least one of the text description, the text embedding, or the corresponding image embedding.

14 . The method of claim 10 , wherein prior to generating the image embedding, the first sub-model further comprises: a diffusion model including a transformer.

15 . The method of claim 10 , wherein the second sub-model includes at least one upsampler model configured for upsampling prior to generating the output image.

16 . A system comprising:

at least one memory storing instructions;

at least one processor configured to execute the instructions to perform operations, the operations comprising:

generating, with a text encoder, a text embedding of a text description;

generating, by a first sub-model, an image embedding based on at least one of the text description or the text embedding, and generating, by a second sub-model, an output image based on at least one of the text description or the image embedding, wherein the second sub-model is different than the first sub-model;

based on a second text description, accessing a second text embedding;

generating a vector representation of the text embedding corresponding to the output image and the second text embedding;

performing an interpolation between the image embedding of the output image and the vector representation of the text embedding corresponding to the output image and the second text embedding; and

based on the interpolation, generating a modified instance of the output image.

17 . The system of claim 16 , further comprising:

training an image encoder on a first data set;

training a text encoder on a second data set;

receiving, from the text encoder, text embeddings associated with the second data set; and

receiving, from the image encoder, image embeddings associated with the first data set;

wherein:

the first data set comprises a set of images; and

the second data set comprises a set of text descriptions corresponding to the set of images.

18 . The system of claim 16 , wherein prior to generating the output image, the second sub-model further comprises: a diffusion model, the diffusion model including at least one of the text description, the text embedding, or the corresponding image embedding.

19 . The system of claim 16 , further comprising:

training an image generation model using the output image.

20 . The system of claim 16 , wherein the first sub-model is configured to encode, prior to generating the corresponding image embedding, at least one of the text description or the text embedding, via a transformer, as a sequence of tokens predicted autoregressively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2024
From: RAMESH, ADITYA; DHARIWAL, PRAFULLA; NICHOL, ALEXANDER; CHU, CASEY; CHEN, MARK
To: OPENAI OPCO LLC
Reel/Frame 066209/0693 →
Continuity (2)
Continuation 18193427 · Mar 30, 2023
Related Publication 20240331237A1 · Oct 3, 2024
References Cited (25)
US 11544880B2 · Park et al. · 2023 [cited by applicant]
US 20210073272A1 · Garrett et al. · 2021 [cited by applicant]
US 20210365727A1 · Aggarwal et al. · 2021 [cited by applicant]
US 20220036127A1 · Lin et al. · 2022 [cited by applicant]
US 20220058340A1 · Aggarwal et al. · 2022 [cited by applicant]
US 20220156992A1 · Harikumar et al. · 2022 [cited by applicant]
US 20220253478A1 · Jain et al. · 2022 [cited by applicant]
US 20220399017A1 · Xu et al. · 2022 [cited by applicant]
US 20230022550A1 · Guo · 2023 [cited by examiner]
US 20230154188A1 · Li · 2023 [cited by examiner]
US 20240153194A1 · Liu · 2024 [cited by examiner]
US 20240161462A1 · Gandelsman · 2024 [cited by examiner]
US 20240169488A1 · Liu · 2024 [cited by examiner]
US 20240185588A1 · Kumari · 2024 [cited by examiner]
US 20240220722A1 · Khattak · 2024 [cited by examiner]
US 20240221235A1 · Gafni · 2024 [cited by examiner]
US 20240281924A1 · Park · 2024 [cited by examiner]
US 20240282016A1 · Liu · 2024 [cited by examiner]
US 20240282025A1 · Park · 2024 [cited by examiner]
US 20240331345A1 · Li · 2024 [cited by examiner]
Federico A. Galatolo, et al., “Generating images from caption and vice versa via CLIP-Guided Generative Latent Space Search”, arXiv:2102.01645, Proc. of the International Conf. on Image Processing and Vision Engineering… [cited by examiner]
P. Aggarwal, et al., “Controlled and Conditional Text to Image Generation with Diffusion Prior”, arXiv:2302.11710v1, at https://arxiv.org/pdf/2302.11710.pdf, Feb. 23, 2023 (Year: 2023). [cited by examiner]
Jie Shi, et al., “DiVAE: Photorealistic Images Synthesis with Denoising Diffusion Decoder”, at https://arxiv.org/pdf/2206.00386.pdf, Jun. 1, 2022 (Year: 2022). [cited by examiner]
Chitwan Saharia, et al., “Photorealistic text-to-image diffusion models with deep language understanding”, at https://arxiv.org/pdf/2205.11487.pdf, May 23, 2022 (Year: 2022). [cited by examiner]
Robin Rombach, et al. “High-resolution image synthesis with latent diffusion models”, arXiv:2112.10752, at https://arxiv.org/pdf/2112.10752.pdf, Apr. 13, 2021 (Year: 2021). [cited by examiner]