IP Library Granted Patent US 12,651,313
Granted Patent B2
US 12,651,313 · App. 18/585,957 · Granted Jun 9, 2026

High-resolution image generation using a diffusion model and a generative adversarial network

Inventors: Tobias Hinz (Campbell, CA); Taesung Park (San Francisco, CA); Jingwan Lu (Sunnyvale, CA); Elya Shechtman (Seattle, WA); Richard Zhang (Burlingame, CA); Oliver Wang (Seattle, WA)
Assignee: ADOBE INC.
G06T3/4053G06N3/0475G06T3/4046G06T11/00G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,313
App. No.
18/585,957
Granted
Jun 9, 2026
Kind
B2
Abstract

A method, non-transitory computer readable medium, apparatus, and system for image generation include obtaining an input image having a first resolution, where the input image includes random noise, and generating a low-resolution image based on the input image, where the low-resolution image has the first resolution. The method, non-transitory computer readable medium, apparatus, and system further include generating a high-resolution image based on the low-resolution image, where the high-resolution image has a second resolution that is greater than the first resolution.

Claims (47)

1 . A method for image generation, comprising:

obtaining an input image having a first resolution, wherein the input image includes random noise;

generating, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution; and

generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution.

2 . The method of claim 1 , further comprising:

generating, using a text encoder, a text embedding, wherein the high-resolution image is generated based on the text embedding.

3 . The method of claim 2 , wherein:

the diffusion model and the GAN each take the text embedding as input.

4 . The method of claim 1 , further comprising:

generating, using an image encoder, an image embedding, wherein the high-resolution image is generated based on the image embedding.

5 . The method of claim 4 , wherein:

the diffusion model and the GAN each take the image embedding as input.

6 . The method of claim 1 , wherein:

the diffusion model contains more parameters than the GAN.

7 . The method of claim 1 , wherein:

the low-resolution image is generated using multiple iterations of the diffusion model and the high-resolution image is generated using a single iteration of the GAN.

8 . The method of claim 1 , wherein:

at least one side of the low-resolution image comprises 128 pixels and at least one side of the high-resolution image comprises 1024 pixels.

9 . The method of claim 1 , wherein:

an aspect ratio of the low-resolution image is different from 1:1 and the same as an aspect ratio of the high-resolution image.

10 . The method of claim 1 , wherein:

the diffusion model and the GAN take variable resolution inputs.

11 . A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to:

obtain an input image having a first resolution, wherein the input image includes random noise;

generate, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution; and

generate, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution.

12 . The non-transitory computer readable medium of claim 11 , wherein the instructions further cause the processor to:

generate, using a text encoder, a text embedding, wherein the high-resolution image is generated based on the text embedding.

13 . The non-transitory computer readable medium of claim 12 , wherein:

the diffusion model and the GAN each take the text embedding as input.

14 . The non-transitory computer readable medium of claim 11 , wherein the instructions further cause the processor to:

generate an image embedding using an image encoder, wherein the high-resolution image is generated based on the image embedding.

15 . The non-transitory computer readable medium of claim 14 , wherein:

the diffusion model and the GAN each take the image embedding as input.

16 . The non-transitory computer readable medium of claim 11 , wherein:

the low-resolution image is generated using multiple iterations of the diffusion model and the high-resolution image is generated using a single iteration of the GAN.

17 . A system for image generation, wherein the system comprises:

one or more processors;

one or more memory components coupled with the one or more processors;

a diffusion model comprising diffusion parameters stored in the one or more memory components, the diffusion model trained to generate a low-resolution image; and

a generative adversarial network (GAN) comprising GAN parameters stored in the one or more memory components, the GAN trained to generate a high-resolution image based on the low-resolution image from the diffusion model.

18 . The system of claim 17 , the system further comprising:

a text encoder comprising text encoding parameters, the text encoder trained to generate a text embedding.

19 . The system of claim 17 , the system further comprising:

an image encoder comprising image encoding parameters, the image encoder trained to generate an image embedding.

20 . The system of claim 17 , wherein:

the low-resolution image is generated using multiple iterations of the diffusion model and the high-resolution image is generated using a single iteration of the GAN.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2024
From: HINZ, TOBIAS; PARK, TAESUNG; LU, JINGWAN; SHECHTMAN, ELYA; ZHANG, RICHARD; WANG, OLIVER
To: ADOBE INC.
Reel/Frame 066549/0978 →
Continuity (2)
Provisional Application 63491237 · Mar 20, 2023
Related Publication 20240320789A1 · Sep 26, 2024
References Cited (91)
US 10713821B1 · Surya et al. · 2020 [cited by applicant]
US 12020400B2 · Hsieh et al. · 2024 [cited by applicant]
US 20170024855A1 · Liang et al. · 2017 [cited by applicant]
US 20180240257A1 · Li et al. · 2018 [cited by applicant]
US 20190095730A1 · Fu et al. · 2019 [cited by applicant]
US 20200242823A1 · Gehlaut et al. · 2020 [cited by applicant]
US 20200242964A1 · Wu · 2020 [cited by examiner]
US 20210342984A1 · Lin et al. · 2021 [cited by applicant]
US 20220038620A1 · Demers · 2022 [cited by applicant]
US 20220108417A1 · Liu et al. · 2022 [cited by applicant]
US 20220114698A1 · Liu · 2022 [cited by applicant]
US 20220130499A1 · Zhou et al. · 2022 [cited by applicant]
US 20220335572A1 · Seresht et al. · 2022 [cited by applicant]
US 20230081171A1 · Zhang et al. · 2023 [cited by applicant]
US 20230082567A1 · Suresha et al. · 2023 [cited by applicant]
US 20230108422A1 · Brauer · 2023 [cited by examiner]
US 20230154161A1 · Pham et al. · 2023 [cited by applicant]
US 20230177810A1 · Xu et al. · 2023 [cited by applicant]
US 20230230198A1 · Zhang et al. · 2023 [cited by applicant]
US 20240037732A1 · Gong et al. · 2024 [cited by applicant]
US 20240037822A1 · Aberman et al. · 2024 [cited by applicant]
US 20240135683A1 · Li · 2024 [cited by examiner]
US 20240171788A1 · Kreis · 2024 [cited by examiner]
US 20240185035A1 · Yu et al. · 2024 [cited by applicant]
US 20240193726A1 · Misra et al. · 2024 [cited by applicant]
US 20240221235A1 · Gafni et al. · 2024 [cited by applicant]
US 20240264718A1 · Benedetto et al. · 2024 [cited by applicant]
US 20240265204A1 · Meeks et al. · 2024 [cited by applicant]
US 20240282016A1 · Liu et al. · 2024 [cited by applicant]
US 20250209712A1 · Svitov · 2025 [cited by examiner]
CN 111062865A · 2020 [cited by applicant]
GB 2621492A1 · 2024 [cited by applicant]
GB 118521474A · 2024 [cited by applicant]
KR 1020210121537A · 2021 [cited by applicant]
WO WO2022156350A1 · 2022 [cited by applicant]
Zbinden, Robin; Implementing and Experimenting with Diffusion Models for Text-to-Image Generation; Sep. 2022; EPFL; p. 21-22; https://arxiv.org/pdf/2209.10948. (Year: 2022). [cited by examiner]
Zhu, Peihao; Improved StyleGAN Embedding: Where are the Good Latents ?; Oct. 2021; p. 1-2; https://arxiv.org/pdf/2012.09036. (Year: 2021). [cited by examiner]
Wang, X.; Deep Multi-Task Learning for Diabetic Retinopathy Grading in Fundus Images; May 2021; Proceedings of the AAAI Conference on Artificial Intelligence; 35(4); p. 2826-2834; https://doi.org/10.1609/aaai.v35i4.1638… [cited by examiner]
Li, Kun; MILI: Multi-person inference from a low-resolution image; Mar. 2023; Fundamental Research; vol. 3, Issue 3; p. 434-441; https://doi.org/10.1016/j.fmre.2023.02.006. (Year: 2023). [cited by examiner]
Office Action dated Jul. 3, 2025 in related U.S. Appl. No. 18/170,963. [cited by applicant]
CN 111062865-A ((Machine Translation on Mar. 6, 2025). [cited by applicant]
Office Action dated Mar. 12, 2025 in related U.S. Appl. No. 18/170,963. [cited by applicant]
Office Action dated May 20, 2025 in related U.S. Appl. No. 18/171,046. [cited by applicant]
Karras, et al., “Analyzing and Improving the Image Quality of StyleGAN”, arXiv preprint arXiv:1912.04958v2 [cs.CV] Mar. 23, 2020, 21 pages. [cited by applicant]
Kim, et al., “The Lipschitz Constant of Self-Attention”, arXiv preprint arXiv:2006.04710v2 [stat.ML] Jun. 9, 2021, 26 pages. [cited by applicant]
Kumari, et al., “Ensembling Off-the-shelf Models for GAN Training”, arXiv preprint arXiv:2112.09130v3 [cs.CV] May 4, 2022, 35 pages. [cited by applicant]
Lee, et al., “ViTGAN: Training GANs with Vision Transformers”, In International Conference on Learning Representations (ICLR), Apr. 2022, 18 pages. [cited by applicant]
Liang, et al., “CPGAN: Content-Parsing Generative Adversarial Networks for Text-to-Image Synthesis”, arXiv preprint arXiv:1912.08562v2 [cs.CV] Jul. 12, 2020, 18 pages. [cited by applicant]
Miyato, et al., “cGANs with Projection Discriminator”, arXiv preprint arXiv:1802.05637v2 [cs.LG] Aug. 15, 2018, 21 pages. [cited by applicant]
Nichol, et al., “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models”, arXiv preprint arXiv:2112.10741v3 [cs.CV] Mar. 8, 2022, 20 pages. [cited by applicant]
Van den Oord, et al., “Representation Learning with Contrastive Predictive Coding”, arXiv preprint arXiv:1807.03748v2 [cs.LG] Jan. 22, 2019, 13 pages. [cited by applicant]
Park, et al., “Contrastive Learning for Unpaired Image-to-Image Translation”, arXiv preprint arXiv:2007.15651v3 [cs.CV] Aug. 20, 2020, 29 pages. [cited by applicant]
Radford, et al., “Learning Transferable Visual Models from Natural Language Supervision”, arXiv preprint arXiv:2103.00020v1 [cs.CV] Feb. 26, 2021, 48 pages. [cited by applicant]
Brock, et al., “Large Scale GAN Training for High Fidelity Natural Image Synthesis”, arXiv preprint arXiv:1809.11096v2 [cs.LG] Feb. 25, 2019, 35 pages. [cited by applicant]
Radford, et al., “Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks”, arXiv preprint arXiv:1511.06434v2 [cs.LG] Jan. 7, 2016, 16 pages. [cited by applicant]
Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv preprint arXiv:2204.06125v1 [cs.CV] Apr. 13, 2022, 27 pages. [cited by applicant]
Reed, et al., “Generative Adversarial Text to Image Synthesis”, arXiv preprint arXiv:1605.05396v2 [cs.NE] Jun. 5, 2016, 10 pages. [cited by applicant]
Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, arXiv preprint arXiv:2112.10752v2 [cs.CV] Apr. 13, 2022, 45 pages. [cited by applicant]
Saharia, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv preprint arXiv:2205.11487v1 [cs.CV] May 23, 2022, 46 pages. [cited by applicant]
Sauer, et al., “StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets”, arXiv preprint arXiv:2202.00273v2 [cs.LG] May 5, 2022, 19 pages. [cited by applicant]
Tanjim, et al., “DynamicRec: A Dynamic Convolutional Network for Next Item Recommendation”, In Proceedings of the 29th ACM International Conference on Information and Knowledge Management (CIKM-2020), pp. 2237-2240, 202… [cited by applicant]
Tao, et al., “Df-Gan: A Simple and Effective Baseline for Text-to-Image Synthesis”, arXiv preprint arXiv:2008.05865v4 [cs.CV] Oct. 15, 2022, 11 pages. [cited by applicant]
Vaswani, et al., “Attention is All You Need”, arXiv preprint arXiv:1706.03762v5 [cs.CL] Dec. 6, 2017, 15 pages. [cited by applicant]
Wang, et al., “Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere”, arXiv preprint arXiv:2005.10242v10 [cs.LG] Aug. 15, 2022, 41 pages. [cited by applicant]
De Brabandere, et al., “Dynamic Filter Networks”, arXiv preprint arXiv:1605.09673v2 [cs.LG] Jun. 6, 2016, 14 pages. [cited by applicant]
Wu, et al., “Pay Less Attention with Lightweight and Dynamic Convolutions”, arXiv preprint arXiv:1901.10430v2 [cs.CL] Feb. 22, 2019, 14 pages. [cited by applicant]
Xu, et al., “AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks”, In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 1316-1324, 9 pages. [cited by applicant]
Yu, et al., “Scaling Autoregressive Models for Content-Rich Text-to-Image Generation”, arXiv preprint arXiv:2206.10789v1 [cs.CV] Jun. 22, 2022, 49 pages. [cited by applicant]
Zhang, et al., “Cross-Modal Contrastive Learning for Text-to-Image Generation”, arXiv preprint arXiv:2101.04702v5 [cs.CV] Apr. 14, 2022, 19 pages. [cited by applicant]
Zhang, et al., “StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks”, arXiv preprint arXiv:1612.03242v2 [cs.CV] Aug. 5, 2017, 14 pages. [cited by applicant]
Zhang, et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”, arXiv preprint arXiv:1801.03924v2 [cs.CV] Apr. 10, 2018, 14 pages. [cited by applicant]
Zhao, et al., “Large Scale Image Completion via Co-Modulated Generative Adversarial Networks”, arXiv preprint arXiv:2103.10428v1 [cs.CV] Mar. 18, 2021, 25 pages. [cited by applicant]
Zhu, et al., “DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-to-Image Synthesis”, arXiv preprint arXiv:1904.01310v1 [cs.CV] Apr. 2, 2019, 9 pages. [cited by applicant]
Deng, et al., “Imagenet: A Large-Scale Hierarchical Image Database”, In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248-255, 2009, 8 pages. [cited by applicant]
Dhariwal, et al., “Diffusion Models beat GANs on Image Synthesis”, arXiv preprint arXiv:2105.05233v4 [cs.LG] Jun. 1, 2021, 44 pages. [cited by applicant]
Goodfellow, et al., “Generative Adversarial Nets”, arXiv preprint arXiv:1406.2661v1 [stat.ML] Jun. 10, 2014, 9 pages. [cited by applicant]
Ha, et al., “HyperNetworks”, arXiv preprint arXiv:1609.09106v4 [cs.LG] Dec. 1, 2016, 29 pages. [cited by applicant]
Ho, et al., “Cascaded Diffusion Models for High Fidelity Image Generation”, in Journal of Machine Learning Research 23, pp. 1-33, Jan. 2022, 33 pages. [cited by applicant]
Ho, et al., “Classifier-Free Diffusion Guidance”, arXiv preprint arXiv:2207.12598v1 [cs.LG] Jul. 26, 2022, 14 pages. [cited by applicant]
Karras, et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, arXiv preprint arXiv:1812.04948v3 [cs.NE] Mar. 29, 2019, 12 pages. [cited by applicant]
Combined Search and Examination Report dated Jun. 18, 2024 in corresponding United Kingdom Patent Application No. 2319189.3 (7 pages). [cited by applicant]
Li, et al., “Text to Realistic Image Generation with Attentional Concatenation Generative Adversarial Networks”, Discrete Dynamics in Nature and Society, vol. 2020, Article ID 6452536, 10 pages. [cited by applicant]
Office Action dated Oct. 1, 2025 in relation U.S. Appl. No. 18/171,046. [cited by applicant]
Lüddecke, et al., “Image segmentation using text and image prompts”, Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 7086-7096, (Year: 2022). [cited by applicant]
Ma, et al., “AI illustrator: Translating raw descriptions into images by prompt-based cross-modal generation”, Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 4282-4290, (Year: 2022). [cited by applicant]
Office Action dated Aug. 13, 2025 in relation U.S. Appl. No. 18/426,763. [cited by applicant]
Office Action dated Sep. 16, 2025 in relation U.S. Appl. No. 18/439,036. [cited by applicant]
Office Action dated Jan. 13, 2026 in relation U.S. Appl. No. 18/439,036. [cited by applicant]
Office Action dated Jan. 26, 2026 in related U.S. Appl. No. 18/171,046. [cited by applicant]
Jain, et al., “Keys to Better Image Inpainting: Structure and Texture Go Hand in Hand”, 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2023, pp. 208-217, doi:10.1109/WACV56… [cited by applicant]
Office Action dated Apr. 2, 2026 in related U.S. Appl. No. 18/487,764. [cited by applicant]