IP Library Granted Patent US 12,499,519
Granted Patent B1
US 12,499,519 · App. 18/652,139 · Granted Dec 16, 2025

Training and deployment of image generation models

Inventors: Dmitriy Karpman (San Francisco, CA); Kevin Guo (San Francisco, CA); Ryan Weber (San Francisco, CA)
Assignee: Castle Global, Inc.
G06T5/70G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,519
App. No.
18/652,139
Granted
Dec 16, 2025
Kind
B1
Abstract

In some embodiments, a method receives a text prompt. A text encoder is executed on the text prompt to generate a representation. The method generates a set of images based on the representation and a set of parameters of an image generation model. The set of images is ranked using reward values that are generated by a reward model. The reward model is trained using human input that provided feedback on a quality of generated images using the image generation model. The method outputs one or more images based on the ranking in response to the text prompt.

Claims (74)

1 . A method comprising:

receiving a text prompt;

executing a text encoder on the text prompt to generate a representation;

generating a set of images based on the representation and a set of parameters of an image generation model;

ranking the set of images using reward values that are generated by a reward model, wherein the reward model is trained using human input that provided feedback on a quality of generated images using the image generation model; and

outputting one or more images based on the ranking in response to the text prompt.

2 . The method of claim 1 , wherein executing the text encoder comprises:

interpreting text of the text prompt to encode a semantic representation represented by the text in the representation.

3 . The method of claim 1 , wherein generating the set of images comprises:

initializing a random distribution of noise; and

iteratively denoise the random distribution of noise according to the parameters of the image generation model and a semantic representation of an image description encoded in the representation.

4 . The method of claim 1 , wherein executing the text encoder comprises:

generating multiple embedding representations in different embedding spaces using multiple text encoder models.

5 . The method of claim 4 , further comprising:

inputting the multiple text embedding representations into one or more image generation models to generate the set of images; and

ranking the set of images to select a subset of images as the set of images.

6 . The method of claim 1 , further comprising:

executing a high resolution model to upsample one or more base images based on parameters of the high resolution model to generate the set of images.

7 . The method of claim 6 , wherein executing the high resolution model comprises:

iteratively denoising the base image to progressively upsample the base image to a higher resolution.

8 . The method of claim 7 , wherein executing the high resolution model comprises:

upsampling visual aspects of the one or more base images generated by the image generation model.

9 . The method of claim 1 , further comprising:

training the image generation model to train the parameters for the image generation model.

10 . The method of claim 9 , wherein training the image generation model comprises:

receiving a set of training images and text captions describing training images in the set of training images;

executing the image generation model on the set of training images to train the set of parameters, wherein training the set of parameters comprises:

generating a first set of generated images based on the text captions;

comparing the first set of generated images to corresponding training images associated with the text captions;

adjusting the set of parameters of the image generation model based on the comparing to generate an adjusted set of parameters; and

generating a final set of parameters from the adjusted set of parameters based on reward values for a second set of generated images by the image generation model using the adjusted set of parameters, wherein the reward values are generated by a reward model that is trained using human input that provided feedback on a quality of the generated images.

11 . The method of claim 10 , wherein the set of training images comprises a first set of training images, the method further comprising:

receiving a second set of training images; and

removing images from the second set of training images to form the first set of training images, wherein images that are removed do not meet a threshold based on quality.

12 . The method of claim 10 , wherein executing the model on the set of training images to train the set of parameters comprises:

generating a set of representations for the set of training images; and

using the set of representations to generate a set of base images using the image generation model.

13 . The method of claim 12 , wherein executing the image generation model on the set of training images to train the set of parameters comprises:

training the set of parameters based on learning a correlation between text captions and associated base images.

14 . The method of claim 13 , wherein using the set of representations to generate base images comprises:

initializing a random distribution of noise for a base image in the set of base images;

iteratively denoise the random distribution of noise according to the set of parameters of the image generation model and a semantic representation of an image description encoded in the representation; and

iteratively training the set of parameters based on the iteratively denoising of the random distribution of noise.

15 . The method of claim 14 , wherein executing the image generation model on the set of training images to train the set of parameters comprises:

iteratively denoising the base image to progressively upsample the base image to a higher resolution according to a set of high resolution parameters of the image generation model; and

iteratively training the set of high resolution parameters based on the iteratively denoising of the base image.

16 . The method of claim 15 , wherein generating the final set of parameters from the set of parameters based on reward values for the generated images comprises:

generating a first generated image in the second set of generated images using the set of parameters of the image generation model;

generating a second generated image in the second set of generated images using the adjusted set of parameters of the image generation model;

generating a first reward value for the first generated image and a second reward value for the second generated image using the reward model;

generating a third reward value based on a difference between the first reward value and the second reward value; and

refining the adjusted set of parameters based on the third reward value.

17 . The method of claim 1 , wherein training the reward model comprises:

generating a reward value for a generated image;

comparing the reward value to feedback from the human input; and

adjusting one or more parameters of the reward model based on the comparing.

18 . The method of claim 1 , wherein ranking the set of images using reward values comprises:

analyzing the set of images;

computing a set of reward values for the set of images; and

ranking the set of images based on respective reward values of images in the set of images.

19 . A non-transitory computer-readable storage medium having stored thereon computer executable instructions, which when executed by a computing device, cause the computing device to be operable for:

receiving a text prompt;

executing a text encoder on the text prompt to generate a representation;

generating a set of images based on the representation and a set of parameters of an image generation model;

ranking the set of images using reward values that are generated by a reward model, wherein the reward model is trained using human input that provided feedback on a quality of generated images using the image generation model; and

outputting one or more images based on the ranking in response to the text prompt.

20 . An apparatus comprising:

one or more computer processors; and

a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable for:

receiving a text prompt;

executing a text encoder on the text prompt to generate a representation;

generating a set of images based on the representation and a set of parameters of an image generation model;

ranking the set of images using reward values that are generated by a reward model, wherein the reward model is trained using human input that provided feedback on a quality of generated images using the image generation model; and

outputting one or more images based on the ranking in response to the text prompt.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2024
From: KARPMAN, DMITRIY; GUO, KEVIN; WEBER, RYAN
To: CASTLE GLOBAL, INC.
Reel/Frame 067282/0933 →
Continuity (2)
Continuation 18525628 · Nov 30, 2023
Provisional Application 63487552 · Feb 28, 2023
References Cited (31)
US 8429173B1 · Rosenberg et al. · 2013 [cited by applicant]
US 11216506B1 · Ranzinger · 2022 [cited by applicant]
US 11514337B1 · Karpman et al. · 2022 [cited by applicant]
US 20160042252A1 · Sawhney et al. · 2016 [cited by applicant]
US 20160042253A1 · Sawhney et al. · 2016 [cited by applicant]
US 20200314507A1 · Yen · 2020 [cited by applicant]
US 20210365500A1 · Gunaselara et al. · 2021 [cited by applicant]
US 20220164643A1 · Charnock et al. · 2022 [cited by applicant]
US 20220198779A1 · Saraee et al. · 2022 [cited by applicant]
US 20220398538A1 · Jakobsson et al. · 2022 [cited by applicant]
WO 2020022956A1 · 2020 [cited by applicant]
Hao, Yaru, et al. “Optimizing Prompts for Text-to-Image Generation.” arXiv preprint arXiv:2212.09611 (2022). (Year: 2022). [cited by examiner]
Lee, Kimin, et al. “Aligning text-to-image models using human feedback.” arXiv preprint arXiv:2302.12192 (2023). (Year: 2023). [cited by examiner]
Prabhudesai, Mihir, et al. “Aligning Text-to-Image Diffusion Models with Reward Backpropagation.” arXiv preprint arXiv:2310.03739 (2023). (Year: 2023). [cited by examiner]
Dong, Hanze, et al. “Raft: Reward ranked finetuning for generative foundation model alignment.” arXiv preprint arXiv:2304.06767 (2023). (Year: 2023). [cited by examiner]
Xu, Jiazheng, et al. “ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation.” arXiv preprint arXiv:2304.05977 (2023). (Year: 2023). [cited by examiner]
Sun, Jiao, et al. “Dreamsync: Aligning text-to-image generation with image understanding feedback.” arXiv preprint arXiv: 2311.17946 (2023). (Year: 2023). [cited by examiner]
Kirstain, Yuval, et al. “Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation.” arXiv preprint arXiv: 2305.01569 (2023). (Year: 2023). [cited by examiner]
U.S. Appl. No. 17/544,615, filed Dec. 7, 2022, Inventor Dmitriy Karpman, Titled: “Logo Detection and Procesing Data Model”, 54 pages. [cited by applicant]
U.S. Appl. No. 18/525,628, filed Nov. 30, 2023, Inventor Dmitriy Karpman, Titled: “Training and Deployment of Image Generation Models”, 63 pages. [cited by applicant]
U.S. Appl. No. 63/316,371, filed Mar. 3, 2022, Inventor Dmitriy Karpman. [cited by applicant]
U.S. Appl. No. 63/481,375, filed Jan. 24, 2023, Inventor Dmitriy Karpman, Titled: “Detecting and Monitoring Unauthorized Uses of Visual Content”, 52 pages. [cited by applicant]
Clark, Kevin, et al. “Directly fine-tuning diffusion models on differentiable rewards.” arXiv preprint arXiv:2309.17 400 (2023). (Year: 2023). [cited by applicant]
Gal, Rinon, et al. “An image is worth one word: Personalizing text-to-image generation using textual inversion.” arXiv preprint arXiv: 2208.01618 (2022). (Year: 2022). [cited by applicant]
Lu, Haoming, et al. “Specialist Diffusion: Plug-and-Play Sample-Efficient Fine-Tuning of Text-to-Image Diffusion Models To Learn Any Unseen Style.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R… [cited by applicant]
Rombach, Robin, et al. “High-resolution image synthesis with latent diffusion models.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022. (Year: 2022). [cited by applicant]
Sun, Gan, et al. “Create your world: Lifelong text-to-image diffusion.” arXiv preprint arXiv:2309.04430 (2023). (Year: 2023). [cited by applicant]
Wen, Song, et al. “Improving compositional text-to-image generation with large vision-language models.” arXiv preprint arXiv: 2310.06311 (2023). (Year: 2023). [cited by applicant]
Wu, Qiucheng, et al. “Harnessing the spatial-temporal attention of diffusion models for high-fidelity text-to-image synthesis.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. (Year: 2023). [cited by applicant]
Wu, Xiaoshi, et al. “Human Preference Score: Better Aligning Text-to-image Models with Human Preference.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. (Year: 2023). [cited by applicant]
Xie, Jinheng, et al. “Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. (Year: 2023). [cited by applicant]