IP Library › Granted Patent US 12,475,607
Granted Patent B2
US 12,475,607 · App. 18/050,349 · Granted Nov 18, 2025

Generating objects of mixed concepts using text-to-image diffusion models

Inventors: Jun Hao Liew (Singapore, SG); Hanshu Yan (Singapore, SG); Daquan Zhou (Los Angeles, CA); Jiashi Feng (Singapore, SG)
Assignee: Lemon Inc.
G06T11/00G06F40/40G06T5/20G06T5/70G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,607
App. No.
18/050,349
Filed
Oct 27, 2022
Granted
Nov 18, 2025
Kind
B2
Art Unit
2673
USPC
382/275
Abstract

Generating an object using a diffusion model includes obtaining a first input and a second input, and synthesizing an output object from the first input and the second input. The synthesizing of the output object includes generating a layout of the output object from the first input, injecting the second input as a content conditioner to the layout of the output object, and de-noising the layout of the output object injected with the content conditioner to generate a content of the output object.

Claims (43)

1 . A method for generating an object using a diffusion model, the method comprising:

obtaining a first input and a second input; and

synthesizing an output object from the first input and the second input by:

generating a layout of the output object from the first input by corrupting the first input by adding Gaussian noise;

adding the second input as a content conditioner to the layout of the output object; and

de-noising the layout of the output object added with the content conditioner to generate a content of the output object, wherein a number of operations for corrupting the first input is equal to a number of operations for de-noising the layout of the output object added with the content conditioner, and wherein, when an object corresponding to the first input and an object corresponding to the second input have a similarity greater than a threshold value, decreasing the number of operations for de-noising the layout of the output object added with the content conditioner.

2 . The method of claim 1 , wherein the first input is an image, the second input is a text, and the synthesizing of the output object is performed using a pre-trained text-to-image diffusion-based generative model.

3 . The method of claim 1 , further comprising:

when an object corresponding to the first input and an object corresponding to the second input have a similarity not greater than a threshold value, increasing the number of operations for de-noising the layout of the output object added with the content conditioner.

4 . The method of claim 1 , wherein the generating of the layout of the output object includes:

applying a random noise to the first input; and

de-noising the first input applied with the random noise to generate the layout of the output object.

5 . The method of claim 4 , wherein the first input is a text, and the second input is a text.

6 . The method of claim 5 , wherein de-noising the first input applied with the random noise includes de-noising the first input applied with the random noise using a pre-trained text-to-image diffusion-based generative model.

7 . The method of claim 6 , wherein a number of operations for de-noising the first input applied with the random noise, combined with a number of operations for de-noising the layout of the output object added with the content conditioner, is equal to a number of operations for the model de-noising a noise input to generate an image.

8 . The method of claim 7 , further comprising:

when an object corresponding to the first input and an object corresponding to the second input have a similarity greater than a threshold value, decreasing the number of operations for de-noising the layout of the output object added with the content conditioner.

9 . The method of claim 7 , further comprising:

when an object corresponding to the first input and an object corresponding to the second input have a similarity not greater than a threshold value, increasing the number of operations for de-noising the layout of the output object added with the content conditioner.

10 . The method of claim 1 , further comprising:

applying a weight to the content conditioner to adjust a number of elements corresponding to the second input to be added to the layout of the output object.

11 . The method of claim 1 , further comprising:

during de-noising the layout of the output object added with the content conditioner, diffusing an intermediate output object and then de-noising the diffused intermediate output object.

12 . A non-transitory computer-readable medium having computer-executable instructions stored thereon that, upon execution, cause one or more processors to perform operations comprising:

obtaining a first input and a second input; and

synthesizing an output object from the first input and the second input by:

generating a layout of the output object from the first input by corrupting the first input by adding Gaussian noise;

adding the second input as a content conditioner to the layout of the output object; and

de-noising the layout of the output object added with the content conditioner to generate a content of the output object, wherein a number of operations for corrupting the first input is equal to a number of operations for de-noising the layout of the output object added with the content conditioner, and wherein, when an object corresponding to the first input and an object corresponding to the second input have a similarity greater than a threshold value, decreasing the number of operations for de-noising the layout of the output object added with the content conditioner.

13 . The computer-readable medium of claim 12 , wherein the generating of the layout of the output object includes:

applying a random noise to the first input; and

de-noising the first input applied with the random noise to generate the layout of the output object.

14 . The computer-readable medium of claim 12 , wherein the operations further comprise:

applying a weight to the content conditioner to adjust a number of elements corresponding to the second input to be added to the layout of the output object.

15 . A generator for generating an object using a diffusion model, the generator comprising:

control logic configured to obtain a first input and a second input;

a machine learning model configured to synthesize an output object from the first input and the second input by:

generating a layout of the output object from the first input by corrupting the first input by adding Gaussian noise;

adding the second input as a content conditioner to the layout of the output object; and

de-noising the layout of the output object added with the content conditioner to generate a content of the output object, wherein a number of operations for corrupting the first input is equal to a number of operations for de-noising the layout of the output object added with the content conditioner, and wherein, when an object corresponding to the first input and an object corresponding to the second input have a similarity greater than a threshold value, decreasing the number of operations for de-noising the layout of the output object added with the content conditioner.

16 . The generator of claim 15 , further comprising:

an enhancer to apply a weight to the content conditioner to adjust a number of elements corresponding to the second input to be added to the layout of the output object.

17 . The generator of claim 15 , wherein the model is a pre-trained text-to-image diffusion-based generative model.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2025
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 072673/0495 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2025
From: TIKTOK PTE. LTD.
To: LEMON INC.
Reel/Frame 072673/0529 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2025
From: LIEW, JUN HAO; YAN, HANSHU; FENG, JIASHI
To: TIKTOK PTE. LTD.
Reel/Frame 073266/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2025
From: ZHOU, DAQUAN
To: BYTEDANCE INC.
Reel/Frame 073266/0775 →
Continuity (1)
Related Publication 20240144544A1 · May 2, 2024
References Cited (30)
US 11908180B1 · Ho · 2024 [cited by examiner]
US 12248884B2 · Xiong · 2025 [cited by examiner]
US 20140223271A1 · Racklyeft · 2014 [cited by examiner]
US 20210358086A1 · Jørgensen · 2021 [cited by examiner]
US 20220398450A1 · Jin · 2022 [cited by examiner]
US 20230290135A1 · Zhou · 2023 [cited by examiner]
US 20230377226A1 · Saharia · 2023 [cited by examiner]
US 20240153247A1 · Zhang · 2024 [cited by examiner]
US 20240169479A1 · Wang · 2024 [cited by examiner]
US 20240193412A1 · Bai · 2024 [cited by examiner]
Valevski, Dani, et al. “Unitune: Text-driven image editing by fine tuning an image generation model on a single image.” arXiv preprint arXiv:2210.09477 2.3 (2022): 5. (Year: 2022). [cited by examiner]
Buades, Antoni, Bartomeu Coll, and Jean-Michel Morel. “Image denoising methods. A new nonlocal principle.” SIAM review 52.1 (2010): 113-147. (Year: 2010). [cited by examiner]
Poole, Ben, et al. “Dreamfusion: Text-to-3d using 2d diffusion.” arXiv preprint arXiv:2209.14988 (2022). (Year: 2022). [cited by examiner]
Weinbach, Samuel, et al. “M-vader: A model for diffusion with multimodal context.” arXiv preprint arXiv:2212.02936 (2022). (Year: 2022). [cited by examiner]
Liew, Jun Hao, et al. “Magicmix: Semantic mixing with diffusion models.” arXiv preprint arXiv:2210.16056 (2022). (Year: 2022). [cited by examiner]
Gal et al., “An Image isWorth One Word: Personalizing Text-to-Image Generation using Textual Inversion”, Computer Science, Computer Vision and Pattern Recognition, Submitted on Aug. 2, 2022, https://arxiv.org/abs/2208.0… [cited by applicant]
Gatys et al., “A Neural Algorithm of Artistic Style”, Computer Science, Computer Vision and Pattern Recognition, last revised Sep. 2, 2015, https://arxiv.org/abs/1508.06576. [cited by applicant]
Goodfellow et al., “Generative Adversarial Networks”, Nov. 2020, vol. 63, No. 11, pp. 139-144, Communications of the ACM, https://dl.acm.org/doi/10.1145/3422622. [cited by applicant]
Hertz et al., “Prompt-to-Prompt Image Editing with Cross Attention Control”, Computer Science, Computer Vision and Pattern Recognition, submitted on Aug. 2, 2022, https://arxiv.org/abs/2208.01626. [cited by applicant]
Ho et al., “Denoising Diffusion Probabilistic Models”, Computer Science, Machine Learning, last revised Dec. 16, 2020, https://arxiv.org/abs/2006.11239. [cited by applicant]
Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, Computer Science, Neural and Evolutionary Computing, last revised Mar. 29, 2019, https://arxiv.org/abs/1812.04948. [cited by applicant]
Kingma et al., “Auto-Encoding Variational Bayes”, Statistics, Machine Learning, revised May 1, 2014, https://arxiv.org/abs/1312.6114v10. [cited by applicant]
Liu et at., “Compositional Visual Generation with Composable Diffusion Models”, Computer Science, Computer Vision and Pattern Recognition, Jun. 3, 2022, https://arxiv.org/abs/2206.01714v1. [cited by applicant]
Luan et al., “Deep Photo Style Transfer”, Computer Science, Computer Vision and Pattern Recognition, last revised Apr. 11, 2017, https://arxiv.org/abs/1703.07511. [cited by applicant]
Lugmayr et al., “RePaint: Inpainting using Denoising Diffusion Probabilistic Models”, Computer Vision Foundation, 2021, pp. 11461-11471, https://openaccess.thecvf.com/content/CVPR2022/papers/Lugmayr_RePaint_Inpainting_U… [cited by applicant]
Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, Computer Science, Computer Vision and Pattern Recognition, last revised Apr. 13, 2022, https://arxiv.org/abs/2112.10752. [cited by applicant]
Song et al., “Score-Based Generative Modeling through Stochastic Differential Equations”, Computer Science, Machine Learning, last revised Feb. 10, 2021, https://arxiv.org/abs/2011.13456. [cited by applicant]
Ulyanov et al., “Texture Networks: Feed-forward Synthesis of Textures and Stylized Images”, Computer Science, Computer Vision and Pattern Recognition, Submitted Mar. 10, 2016, https://arxiv.org/abs/1603.03417. [cited by applicant]
Zhu et al., “Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks”, Computer Science, Computer Vision and Pattern Recognition, last revised Aug. 24, 2020, https://arxiv.org/abs/1703.10593. [cited by applicant]
Stenbit et al., “A walk through latent space with Stable Diffusion”, Keras, Code examples/Generative Deep Learning, Last modified Sep. 28, 2022, https://keras.io/examples/generative/random_walks_with_stable_diffusion/. [cited by applicant]