Unified diffusion model for image generation and/or editing
Systems and methods are disclosed for training and executing a unified diffusion model that jointly reconstructs images and layouts under a unified objective. The system may access image-layout-prompt triplets, each comprising an image, a layout, and a guiding prompt. The system may independently sample an image noise level from a first distribution and a layout noise level from a second distribution, then corrupt each modality according to the selected noise levels. The system may encode the prompt into conditioning embeddings and provide the corrupted modalities and embeddings to the unified diffusion model. The system may cause the model to generate a reconstructed image and layout, and update model parameters based on a loss function balancing fidelity across both modalities. By independently sampling noise levels, the system may expose the model to diverse training conditions, enabling generalization across tasks such as image-to-layout prediction, layout-to-image generation, joint synthesis, and instruction-based editing.
1 . A computer-implemented method for training a unified diffusion model with a unified objective that enables independent injection of image noise and layout noise, the method comprising:
accessing a plurality of image-layout-prompt triplets, each image-layout-prompt triplet comprising an image, a layout, and a prompt;
for each image-layout-prompt triplet:
independently sampling an image noise level from a first predefined distribution and a layout noise level from a second predefined distribution;
corrupting the image and the layout based on the respective noise levels to produce corrupted modalities;
encoding the prompt into conditioning embeddings;
providing the corrupted modalities and conditioning embeddings to the unified diffusion model;
causing the unified diffusion model to generate a reconstructed image and a reconstructed layout; and
updating one or more parameters of the unified diffusion model based on a training objective that balances fidelity of the reconstructed image and fidelity of the reconstructed layout.
2 . The method of claim 1 , wherein the first predefined distribution and the second predefined distribution differ, such that the image and layout modalities are subjected to different noise schedules during training.
3 . The method of claim 1 , wherein the first predefined distribution is configured to emphasize low or moderate noise levels, and the second predefined distribution is configured to emphasize higher noise levels.
4 . The method of claim 1 , wherein corrupting the image and the layout comprises combining each modality with randomly generated noise samples in proportion to the independently selected noise levels.
5 . The method of claim 1 , wherein the independent sampling exposes the unified diffusion model to diverse training conditions including: a nearly uncorrupted image with a heavily corrupted layout, a heavily corrupted image with a nearly uncorrupted layout, and both modalities partially corrupted.
6 . The method of claim 5 , wherein under the diverse training conditions the unified diffusion model learns to reconstruct a heavily corrupted modality using contextual information from a less corrupted modality.
7 . The method of claim 1 , wherein the training objective applies a configurable weighting factor to adjust the relative emphasis placed on reconstructing the image versus reconstructing the layout.
8 . The method of claim 1 , further comprising reseeding the first predefined distribution and the second predefined distribution across training epochs such that previously used triplets are encountered under different combinations of noise levels.
9 . The method of claim 1 , wherein the independent sampling improves generalization of the unified diffusion model to inference tasks including image-to-layout conversion, layout-to-image generation, joint synthesis of images and layouts, and instruction-based editing, without requiring task-specific retraining.
10 . A system for generating an image, comprising:
a processor programmed to:
receive inference inputs comprising at least one of an image, a layout, or a prompt;
independently select an image noise level and a layout noise level from respective predefined distributions;
apply corruption to the image and the layout based on the independently selected noise levels to generate corrupted modalities;
encode the prompt into conditioning embeddings;
provide the corrupted modalities and conditioning embeddings to a unified diffusion model; and
generate a reconstructed image and a reconstructed layout that are mutually consistent and semantically aligned with the prompt.
11 . The system of claim 10 , wherein the unified diffusion model applies bidirectional cross-modal attention to refine the corrupted modalities using contextual information across both the image and layout channels.
12 . The system of claim 10 , wherein the predefined distribution for the image noise level and the predefined distribution for the layout noise level differ to expose the unified diffusion model to asymmetric corruption conditions during inference.
13 . The system of claim 10 , wherein the processor is further programmed to select an inference mode from a group comprising: image-to-layout prediction, layout-to-image generation, and joint generation of image and layout.
14 . The system of claim 13 , wherein in the image-to-layout mode the image noise level is selected to be low while the layout noise level is selected across a broader range, thereby enabling layout inference conditioned on a substantially uncorrupted image.
15 . The system of claim 13 , wherein in the layout-to-image mode the layout noise level is selected to be low while the image noise level is selected across a broader range, thereby enabling image generation conditioned on a substantially uncorrupted layout.
16 . The system of claim 13 , wherein in the joint generation mode both the image noise level and the layout noise level are selected across a range, thereby enabling simultaneous synthesis of an image and layout under guidance of the prompt.
17 . The system of claim 10 , wherein to perform instruction-based editing, the processor is further programmed to:
access an initial image, an initial layout, and an editing instruction;
encode the editing instruction into conditioning embeddings;
apply independent noise to the initial image and layout; and
generate an updated image and an updated layout that reflect the editing instruction.
18 . The system of claim 17 , wherein the unified diffusion model aligns embeddings of the editing instruction with structural elements of the layout such that edits are applied consistently across both the image and the layout modalities.
19 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:
receive inference inputs including at least one of an image, a layout, or a prompt;
independently select noise levels for an image modality and a layout modality;
corrupt the inference inputs based on the independently selected noise levels;
assemble a conditioning package comprising the corrupted modalities, prompt embeddings, and metadata identifying the noise levels;
forward the conditioning package into a unified diffusion model; and
iteratively denoise the corrupted modalities using the unified diffusion model to produce inference results comprising a generated image and/or a generated layout.
20 . The non-transitory computer-readable medium of claim 19 , wherein independent sampling of the noise levels enables the unified diffusion model to generalize across multiple inference tasks without requiring task-specific retraining.