IP Library › Granted Patent US 12,555,288
Granted Patent B2
US 12,555,288 · App. 18/459,526 · Granted Feb 17, 2026

Controllable diffusion model

Inventors: Wonwoong Cho (West Lafayette, IN); Hareesh Ravi (San Jose, CA); Midhun Harikumar (Sunnyvale, CA); Vinh Ngoc Khuc (Campbell, CA); Krishna Kumar Singh (San Jose, CA); Jingwan Lu (Sunnyvale, CA); Ajinkya Gorakhnath Kale (San Jose, CA)
Assignee: ADOBE INC.
G06T11/60G06T11/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,288
App. No.
18/459,526
Granted
Feb 17, 2026
Kind
B2
Abstract

A method, apparatus, and non-transitory computer readable medium for image generation are described. Embodiments of the present disclosure obtain a content input and a style input via a user interface or from a database. The content input includes a target spatial layout and the style input includes a target style. A content encoder of an image processing apparatus encodes the content input to obtain a spatial layout mask representing the target spatial layout. A style encoder of the image processing apparatus encodes the style input to obtain a style embedding representing the target style. An image generation model of the image processing apparatus generates an image based on the spatial layout mask and the style embedding, where the image includes the target spatial layout and the target style.

Claims (53)

1 . A method comprising:

obtaining a content input and a style input, wherein the content input comprises a target spatial layout and the style input comprises a target style;

encoding, by a content encoder, the content input to obtain a spatial layout mask representing the target spatial layout;

encoding, by a style encoder, the style input to obtain a style embedding representing the target style; and

generating, by an image generation model, an image by denoising noisy features based on the spatial layout mask and the style embedding, wherein the image includes the target spatial layout and the target style.

2 . The method of claim 1 , wherein:

the content input comprises a content image and the style input comprises a style image.

3 . The method of claim 1 , further comprising:

performing a spatial-wise operation based on the spatial layout mask, wherein the image is generated based on the spatial-wise operation.

4 . The method of claim 1 , further comprising:

performing a channel-wise operation based on the style embedding, wherein the image is generated based on the channel-wise operation.

5 . The method of claim 1 , further comprising:

computing a content weight based on a diffusion timestep, wherein the image is generated based on the spatial layout mask according to the content weight.

6 . The method of claim 1 , further comprising:

computing a style weight based on a diffusion timestep, wherein the image is generated based on the style embedding according to the style weight.

7 . The method of claim 1 , further comprising:

generating a noise vector, wherein the image is generated based on the noise vector using a reverse diffusion process.

8 . The method of claim 1 , wherein:

the style embedding includes global semantic information representing the target style.

9 . The method of claim 1 , wherein:

the spatial layout mask comprises a plurality of values corresponding to a plurality of locations of the content input, respectively, and wherein the style embedding comprises a tuple of values that together represent the target style.

10 . A method comprising:

initializing a content encoder, a style encoder, and an image generation model;

receiving training data including an image comprising spatial content and a style attribute;

computing an objective function based on the spatial content and the style attribute; and

jointly training the content encoder, the style encoder, and the image generation model using an end-to-end process based on the objective function.

11 . The method of claim 10 , wherein:

the content encoder is trained to generate a spatial layout mask representing a target spatial layout.

12 . The method of claim 10 , wherein:

the style encoder is trained to generate a style embedding representing a target style.

13 . The method of claim 10 , wherein:

the image generation model is trained to generate a predicted image including a target spatial layout and a target style based on an output of the content encoder and an output of the style encoder.

14 . The method of claim 10 , further comprising:

generating a latent code based on the image using an image encoder;

generating a noisy latent code based on the latent code using a forward diffusion process; and

generating a predicted image using the image generation model, wherein the objective function is computed based on the predicted image.

15 . The method of claim 14 , further comprising:

generating a predicted spatial layout mask using the content encoder; and

generating a predicted style embedding using the style encoder, wherein the predicted image is generated based on the predicted spatial layout mask and the predicted style embedding.

16 . An apparatus comprising:

at least one processor;

at least one memory including instructions executable by the at least one processor;

a content encoder comprising parameters stored in the at least one memory and trained to encode a content input to obtain a spatial layout mask representing a target spatial layout;

a style encoder comprising parameters stored in the at least one memory and trained to encode a style input to obtain a style embedding representing a target style; and

an image generation model comprising parameters stored in the at least one memory and trained to generate an image by denoising noisy features based on the spatial layout mask and the style embedding, wherein the image includes the target spatial layout and the target style.

17 . The apparatus of claim 16 , wherein:

the content encoder and the style encoder each comprise a residual neural network.

18 . The apparatus of claim 16 , wherein:

the image generation model comprises a denoising unit.

19 . The apparatus of claim 16 , further comprising:

an image encoder configured to generate a latent code based on the image.

20 . The apparatus of claim 16 , further comprising:

a timestep scheduling component configured to compute a content weight based on a diffusion timestep, wherein the image is generated based on the spatial layout mask according to the content weight, and to compute a style weight based on the diffusion timestep, wherein the image is generated based on the style embedding according to the style weight.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2023
From: CHO, WONWOONG; RAVI, HAREESH; HARIKUMAR, MIDHUN; KHUC, VINH NGOC; SINGH, KRISHNA KUMAR; LU, JINGWAN; KALE, AJINKYA GORAKHNATH
To: ADOBE INC.
Reel/Frame 064771/0419 →
Continuity (1)
Related Publication 20250078349A1 · Mar 6, 2025
References Cited (11)
US 10504267B2 · Simons · 2019 [cited by examiner]
US 20200160154A1 · Taslakian · 2020 [cited by examiner]
US 20210264236A1 · Xu · 2021 [cited by examiner]
US 20230162409A1 · Yu · 2023 [cited by examiner]
US 20230351566A1 · Jeon · 2023 [cited by examiner]
Cho, et al., “Towards Enhanced Controllability of Diffusion Models”, arXiv preprint arXiv:2302.14368v2 [cs.CV] Mar. 15, 2023, 28 pages. [cited by applicant]
Balaji, et al., “eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers”, arXiv preprint arXiv:2211.01324v5 [cs.CV] Mar. 14, 2023, 24 pages. [cited by applicant]
Ho, et al., “Denoising Diffusion Probabilistic Models”, Advances in Neural Information Processing Systems, 33:6840-6851, 2020. [cited by applicant]
Kwon, et al. “Diagonal Attention and Style-Based GAN for Content-Style Disentanglement in Image Generation and Translation”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 13980-13989, 2… [cited by applicant]
Preechakul, et al., “Diffusion Autoencoders: Toward a Meaningful and Decodable Representation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10619-10629, 2022. [cited by applicant]
Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684-10695, 2022. [cited by applicant]