IP Library Granted Patent US 12,437,437
Granted Patent B2
US 12,437,437 · App. 18/052,658 · Granted Oct 7, 2025

Diffusion models having continuous scaling through patch-wise image generation

Inventors: Yinbo Chen (La Jolla, CA); Michaël Gharbi (San Francisco, CA); Oliver Wang (Seattle, WA); Richard Zhang (Burlingame, CA); Elya Shechtman (Seattle, WA)
Assignee: ADOBE INC.
G06T7/70G06T3/40G06T5/73G06T7/10G06T2207/20084G06T2207/20132G06T2207/20212
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,437
App. No.
18/052,658
Granted
Oct 7, 2025
Kind
B2
Abstract

Aspects of the methods, apparatus, non-transitory computer readable medium, and systems include obtaining a noise map and a global image code encoded from an original image and representing semantic content of the original image; generating a plurality of image patches based on the noise map and the global image code using a diffusion model; and combining the plurality of image patches to produce an output image including the semantic content.

Claims (57)

1. A method comprising:

obtaining a noise map and a global image code encoded from an original image and representing semantic content of the original image;

generating a plurality of image patches based on the noise map and the global image code using a diffusion model, wherein each image patch of the plurality of image patches is generated by denoising a noisy patch of the noise map based on the global image code; and

combining the plurality of image patches to produce an output image including the semantic content.

2. The method of claim 1 , wherein:

the diffusion model is conditioned based on the global image code.

3. The method of claim 1 , further comprising:

identifying a text prompt; and

encoding the text prompt to obtain the global image code.

4. The method of claim 1 , wherein the original image is a high resolution image.

5. The method of claim 1 , wherein:

the noise map comprises a same resolution as the output image.

6. The method of claim 1 , wherein:

each of the plurality of image patches is generated based on a region of the noise map that overlaps at least one other region used to generate another of the plurality of image patches.

7. The method of claim 6 , wherein:

the plurality of image patches do not overlap each other.

8. The method of claim 1 , further comprising:

identifying a position indicator corresponding to each of the plurality of image patches, wherein each of the plurality of image patches is generated based on the corresponding position indicator and the plurality of image patches are combined based on the position indicator.

9. The method of claim 1 , further comprising:

training the diffusion model to generate the plurality of image patches based on the global image code.

10. The method of claim 1 , wherein:

the global image code includes information representing a spatial layout of the output image.

11. A method comprising:

initializing parameters of a diffusion model;

obtaining a noise map and a global image code encoded from a training image and representing semantic content of the training image;

generating a plurality of predicted image patches based on the noise map and the global image code using the diffusion model, wherein each predicted image patch of the plurality of predicted image patches is generated by denoising a noisy patch of the noise map based on the global image code;

computing a loss function based on the plurality of predicted image patches; and

training the diffusion model to generate image patches by updating the parameters based on the loss function.

12. The method of claim 11 , wherein:

identifying a high-resolution training image;

generating a high-resolution noise map and a low-resolution noise map based on the high-resolution training image;

generating a first image patch based on the high-resolution noise map and a second image patch based on the low-resolution noise map; and

computing a patch consistency loss by comparing the first image patch and the second image patch, wherein the loss function includes the patch consistency loss.

13. The method of claim 12 , further comprising:

cropping the high-resolution training image to obtain a high-resolution training patch; and

adding noise to the high-resolution training patch to obtain the high-resolution noise map.

14. The method of claim 12 , further comprising:

down-sampling the high-resolution training image to obtain a low-resolution training image;

cropping the low-resolution training image to obtain a low-resolution training patch; and

adding noise to the low-resolution training patch to obtain the low-resolution noise map.

15. The method of claim 11 , further comprising:

combining the plurality of predicted image patches to produce a predicted image; and

computing a reconstruction loss by comparing the predicted image to a ground truth image, wherein the loss function includes the reconstruction loss.

16. The method of claim 11 , further comprising:

identifying a position indicator corresponding each of the plurality of predicted image patches, wherein each of the plurality of predicted image patches is generated based on the corresponding position indicator.

17. An apparatus comprising:

one or more processors; and

one or memories including instructions executable by the one or more processors to:

obtain a noise map and a global image code encoded from an original image and representing semantic content of the original image;

generate a plurality of image patches based on the noise map and the global image code using a diffusion model, wherein each image patch of the plurality of image patches is generated by denoising a noisy patch of the noise map based on the global image code; and

combine the plurality of image patches to produce an output image including the semantic content.

18. The apparatus of claim 17 , wherein:

the diffusion model comprises a U-Net architecture.

19. The apparatus of claim 17 , wherein the instructions are further executable to:

encode an input prompt to obtain the global image code.

20. The apparatus of claim 19 , wherein:

the global image code is encoded using a multimodal encoder, and wherein the original image is a high resolution image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2022
From: CHEN, YINBO; GHARBI, MICHAEL; WANG, OLIVER; ZHANG, RICHARD; SHECHTMAN, ELYA
To: ADOBE INC.
Reel/Frame 061656/0914 →
Continuity (1)
Related Publication 20240161327A1 · May 16, 2024
References Cited (20)
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (Apr. 13, 2022). High-Resolution Image Synthesis with Latent Diffusion Models. ArXiv:2112.10752 [Cs]. https://arxiv.org/abs/2112.10752 (Year: 2022). [cited by examiner]
Özdenizci, O., & Legenstein, R. (Jul. 29, 2022). Restoring Vision in Adverse Weather Conditions with Patch-Based Denoising Diffusion Models. ArXiv.org. https://arxiv.org/abs/2207.14626 (Year: 2022). [cited by examiner]
Wandell, B. (1995). Foundations of Vision» Chapter 8: Multiresolution Image Representations. Stanford.edu. https://foundationsofvision.stanford.edu/chapter-8-multiresolution-image-representations/#Threshold_and_Recognit… [cited by examiner]
Chen, Y., Wang, O., Zhang, R., Shechtman, E., Wang, X., & Gharbi, M. (2024). Image Neural Field Diffusion Models. ArXiv.org. https://arxiv.org/abs/2406.07480 (Year: 2024). [cited by examiner]
Xia, W., Cong, W., & Wang, G. (Nov. 18, 2022). Patch-Based Denoising Diffusion Probabilistic Model for Sparse-View CT Reconstruction. ArXiv.org. https://doi.org/10.48550/arXiv.2211.10388 (Year: 2022). [cited by examiner]
Lin, C. H., Chang, C.-C., Chen, Y.-S., Juan, D.-C., Wei, W., & Chen, H.-T. (Jan. 5, 2020). COCO-GAN: Generation by Parts via Conditional Coordinating. ArXiv.org. https://doi.org/10.48550/arXiv.1904.00284 (Year: 2020). [cited by examiner]
Ntavelis, E., Shahbazi, M., Kastanis, I., Timofte, R., Danelljan, M., & Gool, V. (Apr. 5, 2022). Arbitrary-Scale Image Synthesis. ArXiv.org. https://arxiv.org/abs/2204.02273 (Year: 2022). [cited by examiner]
Luhman, T., & Luhman, E. (Jul. 9, 2022). Improving Diffusion Model Efficiency Through Patching. ArXiv.org. https://arxiv.org/abs/2207.04316 (Year: 2022). [cited by examiner]
Peebles, W., & Xie, S. (Mar. 2, 2023). Scalable Diffusion Models with Transformers. https://arxiv.org/pdf/2212.09748 (Year: 2023). [cited by examiner]
Wang, Y., Yu, J., Yu, R., & Zhang, J. (Mar. 1, 2023). Unlimited-Size Diffusion Restoration. ArXiv.org. https://arxiv.org/abs/2303.00354v1 (Year: 2023). [cited by examiner]
Wang, W., Bao, J., Zhou, W., Chen, D., Chen, D., Yuan, L., & Li, H. (Nov. 22, 2022). SinDiffusion: Learning a Diffusion Model from a Single Natural Image. ArXiv.org. https://arxiv.org/abs/2211.12445 (Year: 2022). [cited by examiner]
Ding, Z., Zhang, M., Wu, J., & Tu, Z. (Aug. 2, 2023). Patched Denoising Diffusion Models For High-Resolution Image Synthesis. ArXiv.org. https://arxiv.org/abs/2308.01316 (Year: 2023). [cited by examiner]
1Chai, et al, “Any-resolution Training for High-resolution Image Synthesis”, arXiv preprint arXiv:2204.07156v2 [cs.CV] Aug. 5, 2022, 33 pages. [cited by applicant]
2Chen, et al, “Learning Continuous Image Representation with Local Implicit Image Function”, arXiv preprint arXiv:2012.09161v2 [cs.CV] Apr. 1, 2021, 11 pages. [cited by applicant]
Bho, et al, “Cascaded Diffusion Models for High Fidelity Image Generation”, arXiv preprint arXiv:2106.15282v3 [cs.CV] Dec. 17, 2021, 33 pages. [cited by applicant]
4Ho, et al, “Denoising Diffusion Probabilistic Models”, NIPS'20: Proceedings of the 34th International Conference on Neural Information Processing Systems, arXiv preprint arXiv:2006.11239v2 [cs.LG] Dec. 16, 2020, 25 pag… [cited by applicant]
5Kingma, et al, “Auto-Encoding Variational Bayes”, arXiv preprint arXiv:1312.6114v1 [stat.ML] Dec. 20, 2013, 9 pages. [cited by applicant]
6Kingma, et al, “Auto-Encoding Variational Bayes”, arXiv preprint arXiv:1312.6114v10 [stat.ML] May 1, 2014, 14 pages. [cited by applicant]
7Rombach, et al, “High-Resolution Image Synthesis with Latent Diffusion Models”, arXiv preprint arXiv:2112.10752v2 [cs.CV] Apr. 13, 2022, 45 pages. [cited by applicant]
8Saharia, et al, “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv preprint arXiv:2205.11487v1 [cs.CV] May 23, 2022, 46 pages. [cited by applicant]