IP Library › Granted Patent US 12,579,606
Granted Patent B2
US 12,579,606 · App. 18/167,324 · Granted Mar 17, 2026

Generative AI inferred prompt outpainting

Inventors: Michael Spencer Cragg (Troy, MI); Edward Christopher Wright (Shorewood, MN); Davis Taylor Brown (Seattle, WA); Dana Michelle Jefferson (Astoria, NY); Andreas Kuefer (Sausalito, CA)
Assignee: ADOBE INC.
G06T3/40G06T3/608G06T7/11G06V10/25G06T2207/20081G06T2207/20132
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,606
App. No.
18/167,324
Granted
Mar 17, 2026
Kind
B2
Abstract

Systems and methods for image processing are provided. Embodiments of the present disclosure obtain an image and a target dimension for expanding the image. The system generates a prompt based on the image using a prompt generation network. A diffusion model generates an expanded image based on the image, the target dimension, and the prompt, where the expanded image includes additional content in an outpainted region that is consistent with content of the image and the prompt.

Claims (60)

1 . A method comprising:

obtaining an image and a target dimension from a user, wherein the image includes metadata indicating a camera setting used for capturing the image, and wherein the target dimension indicates an aspect ratio different from the image;

generating a prompt based on the image using a prompt generation network, wherein the prompt includes a term based on the camera setting; and

generating an expanded image based on the image, the target dimension, and the prompt using a diffusion model, wherein the expanded image includes additional content in an outpainted region that is consistent with the term based on the camera setting, and wherein the expanded image has the aspect ratio indicated by the target dimension.

2 . The method of claim 1 , further comprising:

providing an image cropping interface to the user, wherein the image cropping interface enables the user to increase and decrease a size of the image; and

receiving the target dimension via the image cropping interface.

3 . The method of claim 1 , further comprising:

identifying an expanded region having the target dimension, wherein the expanded region includes the image and the outpainted region.

4 . The method of claim 1 , further comprising:

rotating the image to obtain a rotated image; and

identifying the outpainted region based on the rotated image and the target dimension.

5 . The method of claim 1 , further comprising:

identifying a skew angle corresponding to a perspective of the image;

stretching the image based on the skew angle to obtain a stretched image; and

identifying the outpainted region based on the stretched image and the target dimension.

6 . The method of claim 1 , further comprising:

identifying the metadata of the image, wherein the prompt is generated based on the metadata.

7 . The method of claim 6 , wherein:

the metadata comprises time information, location information, color information, or a combination thereof.

8 . The method of claim 1 , further comprising:

generating an input map for the diffusion model that includes the image in an internal region and noise in the outpainted region, wherein the expanded image is generated based on the input map.

9 . The method of claim 1 , further comprising:

generating a plurality of low-resolution images depicting a plurality of candidate dimensions for the expanded image; and

receiving a user input selecting one of the plurality of candidate dimensions as the target dimension.

10 . The method of claim 1 , further comprising:

identifying a first region including a first portion of the image and a first portion of the outpainted region;

generating a first tile for the expanded image based on the first region using the diffusion model;

identifying a second region including a second portion of the image and a second portion of the outpainted region; and

generating a second tile for the expanded image based on the second region using the diffusion model.

11 . The method of claim 10 , further comprising:

identifying a first prompt based on the first portion of the image, wherein the first tile is generated based on the first prompt; and

identifying a second prompt based on the second portion of the image, wherein the second tile is generated based on the second prompt.

12 . A method comprising:

initializing a diffusion model;

obtaining training data including an input image, a prompt, a target dimension, and a ground-truth expanded image, wherein the input image includes metadata indicating a camera setting used for capturing the input image, and wherein the target dimension indicates an aspect ratio different from the input image, and wherein the prompt includes a term based on the camera setting; and

training the diffusion model to generate an expanded image that includes additional content in an outpainted region that is consistent with the term based on the camera setting and the training data, wherein the expanded image has the aspect ratio indicated by the target dimension.

13 . The method of claim 12 , wherein:

the training data includes the metadata of the input image, and wherein the diffusion model is trained to generate the expanded image based on the metadata.

14 . The method of claim 12 , further comprising:

generating the prompt based on the input image and the metadata.

15 . The method of claim 12 , further comprising:

cropping the ground-truth expanded image to obtain the input image.

16 . The method of claim 12 , further comprising:

initializing a prompt generation network; and

training the prompt generation network to generate the prompt based on the input image.

17 . The method of claim 16 , wherein:

the prompt generation network is trained to generate the prompt based on the metadata of the input image.

18 . The method of claim 12 , further comprising:

performing a forward diffusion process to obtain a plurality of noise maps; and

performing a reverse diffusion process using the diffusion model to obtain a plurality of predicted noise maps, wherein the training is based on the plurality of noise maps and the plurality of predicted noise maps.

19 . An apparatus comprising:

a processor; and

a memory including instructions executable by the processor to:

obtain an image and a target dimension from a user, wherein the image includes metadata indicating a camera setting used for capturing the image, and wherein the target dimension indicates an aspect ratio different from the image;

generate a prompt based on the image using a prompt generation network, wherein the prompt includes a term based on the camera setting; and

generate an expanded image based on the image, the target dimension, and the prompt using a diffusion model, wherein the expanded image includes additional content in an outpainted region that is consistent with the term based on the camera setting, and wherein the expanded image has the aspect ratio indicated by the target dimension.

20 . The apparatus of claim 19 , further comprising instructions executable by the processor to:

train the prompt generation network to generate the prompt; and

train the diffusion model to generate the expanded image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2023
From: CRAGG, MICHAEL SPENCER; WRIGHT, EDWARD CHRISTOPHER; BROWN, DAVIS TAYLOR; JEFFERSON, DANA MICHELLE; KUEFER, ANDREAS
To: ADOBE INC.
Reel/Frame 062654/0814 →
Continuity (1)
Related Publication 20240273670A1 · Aug 15, 2024
References Cited (15)
US 20130107284A1 · Hayashi · 2013 [cited by examiner]
US 20140240357A1 · Hou · 2014 [cited by examiner]
US 20160378788A1 · Panneer · 2016 [cited by examiner]
US 20170115853A1 · Allekotte · 2017 [cited by examiner]
US 20240046422A1 · Song · 2024 [cited by examiner]
Screen captures from YouTube video clip titled “DALL-E 2 Outpainting Feature—Super Simple AI Tutorial,” uploaded on Oct. 19, 2022 by user “Prompt Muse”. retrieved from Internet: <https://www.youtube.com/watch?v=jo--DkGa… [cited by examiner]
Screen captures from YouTube video clip titled “How To Outpaint—Stable Diffusion AI | Draw Outside the Lines,” uploaded on Nov. 7, 2022 by user “Jennifer Doebelin”. retrieved from Internet: <https://www.youtube.com/watc… [cited by examiner]
Sabini, M., & Rusak, G. (2018). Painting outside the box: Image outpainting with gans. arXiv preprint arXiv:1808.08483. (Year: 2018). [cited by examiner]
Weng, L. (2021). What are diffusion models?. lilianweng. github. io, 21. (Year: 2021). [cited by examiner]
Xu, X., Wang, Z., Zhang, G., Wang, K., & Shi, H. (2023). Versatile diffusion: Text, images and variations all in one diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 7754-… [cited by examiner]
Xu, S. (2022). Clip-diffusion-Im: Apply diffusion model on image captioning. arXiv preprint arXiv:2210.04559. (Year: 2022). [cited by examiner]
1Ramesh, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv:2204.06125v1 [cs.CV] Apr. 13, 2022, 27 pages. [cited by applicant]
2Brown, et al. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33, 1877-1901. [cited by applicant]
3Dall⋅E: Creating Images from Text, Jan. 5, 2021, https://openai.com/blog/dall-e/, 14 pages. [cited by applicant]
4Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, In Proceedings of the 38th International Conference on Machine Learning, PMLR 139, pp. 8748-8763, PMLR (Jul. 2021). [cited by applicant]