IP Library Granted Patent US 12,524,839
Granted Patent B2
US 12,524,839 · App. 18/515,378 · Granted Jan 13, 2026

Generative image filling using a reference image

Inventors: Yuqian Zhou (Bellevue, WA); Krishna Kumar Singh (San Jose, CA); Zhe Lin (Clyde Hill, WA); Qing Liu (San Jose, CA); Zhifei Zhang (San Jose, CA); Sohrab Amirghodsi (Seattle, CA); Elya Shechtman (Seattle, WA); Jingwan Lu (Sunnyvale, CA)
Assignee: ADOBE INC.
G06T5/50G06T9/00G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,839
App. No.
18/515,378
Granted
Jan 13, 2026
Kind
B2
Abstract

Embodiments include systems and methods for generative image filling based on text and a reference image. In one aspect, the system obtains an input image, a reference image, and a text prompt. Then, the system encodes the reference image to obtain an image embedding and encodes the text prompt to obtain a text embedding. Subsequently, a composite image is generated based on the input image, the image embedding, and the text embedding.

Claims (51)

1 . A method comprising:

obtaining an input image, a mask, a reference image, and a text prompt, wherein the mask indicates a region of the input image, and the text comprises guidance for editing the region of the input image based on the reference image;

encoding, using an image encoder, the reference image to obtain an image embedding;

encoding, using a text encoder, the text prompt to obtain a text embedding; and

generating, using an image generation model, a composite image based on the input image, the image embedding, and the text embedding, wherein the composite image includes content from the input image outside the region indicated by the mask and content from the reference image within the region indicated by the mask.

2 . The method of claim 1 , wherein generating the composite image comprises:

providing the image embedding and the text embedding as guidance for the image generation model.

3 . The method of claim 1 , wherein:

the composite image includes content corresponding to the reference image in a region corresponding to a mask.

4 . The method of claim 1 , wherein:

the composite image includes content corresponding to the input image in a region outside of a mask.

5 . The method of claim 1 , wherein:

the text prompt describes an object in the reference image.

6 . The method of claim 1 , wherein:

the reference image comprises a portion of the input image.

7 . The method of claim 1 , wherein:

the reference image comprises a style from the input image.

8 . The method of claim 1 , wherein:

the composite image includes an object described by the text prompt with a style from the reference image.

9 . The method of claim 1 , further comprising:

receiving a mask from a user, wherein the mask indicates a region generated by the image generation model based on the reference image.

10 . The method of claim 1 , further comprising:

generating a mask based on the input image, the reference image, or the text prompt, wherein the mask indicates a region generated by the image generation model.

11 . A method comprising:

obtaining an input image, a mask, and a reference image, wherein the mask indicates a region of the input image;

inserting the reference image into the input image based on the mask to obtain a combined input image;

encoding, using an image encoder, the reference image to obtain an image embedding; and

generating, using an image generation model, a composite image based on the combined input image and the image embedding, wherein the composite image includes content from the input image outside the region indicated by the mask and content from the reference image within the region indicated by the mask.

12 . The method of claim 11 , further comprising:

performing, by the image generation model, a self-attention operation on the combined input image, wherein the composite image is generated based on the self-attention operation.

13 . The method of claim 11 , further comprising:

obtaining a text prompt; and

encoding, using a text encoder, the text prompt to obtain a text embedding, wherein the composite image is generated based on the text embedding.

14 . The method of claim 11 , wherein:

the reference image is inserted into the input image in a region outside of a mask.

15 . An apparatus comprising:

at least one memory;

at least one processor coupled to the at least one memory, wherein the processor is configured to execute instructions stored in the at least one memory;

an image encoder configured to encode a reference image to obtain an image embedding;

a text encoder configured to encode a text prompt to obtain a text embedding; and

an image generation model configured to generate a composite image based on an input image, the image embedding, and the text embedding.

16 . The system of aspect 15 , wherein:

the image generation model comprises a diffusion model.

17 . The system of aspect 15 , wherein:

the image generation model comprises a self-attention layer configured to operate on a combination of the input image and the reference image.

18 . The system of aspect 15 , wherein:

the image encoder and the text encoder are components of a multimodal encoder.

19 . The system of aspect 15 , wherein:

the image generation model uses the image embedding and the text embedding for classifier-free guidance.

20 . The system of aspect 15 , further comprising:

a user interface configured to obtain a selection input from a user, wherein a mask is created based on the selection input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2023
From: ZHOU, YUQIAN; SINGH, KRISHNA KUMAR; LIN, ZHE; LIU, QING; ZHANG, ZHIFEI; AMIRGHODSI, SOHRAB; SHECHTMAN, ELYA; LU, JINGWAN
To: ADOBE INC.
Reel/Frame 065630/0204 →
Continuity (2)
Provisional Application 63505902 · Jun 2, 2023
Related Publication 20240404013A1 · Dec 5, 2024
References Cited (9)
US 11983806B1 · Ramesh · 2024 [cited by examiner]
Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, arXiv preprint arXiv:2112.10752v2 [cs.CV] Apr. 13, 2022, 45 pages. [cited by applicant]
Zhang, et al., “Adding Conditional Control to Text-to-Image Diffusion Models”, arXiv preprint arXiv:2302.05543v2 [cs.CV] Sep. 2, 2023, 12 pages. [cited by applicant]
Li, et al., “GLIGEN: Open-Set Grounded Text-to-Image Generation”, arXiv preprint arXiv:2301.07093v2 [cs.CV] Apr. 17, 2023, 21 pages. [cited by applicant]
Mou, et al., “T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models”, arXiv preprint arXiv:2302.08453v2 [cs.CV] Mar. 20, 2023, 10 pages. [cited by applicant]
Chung, et al., “Scaling Instruction-Finetuned Language Models”, arXiv preprint arXiv:2210.11416v5 [cs.LG] Dec. 6, 2022, 54 pages. [cited by applicant]
Song, et al., “ObjectStitch: Object Compositing with Diffusion Model”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18310-18319. [cited by applicant]
Xu, et al., “Prompt-Free Diffusion: Taking “Text” out of Text-to-Image Diffusion Models”, arXiv preprint arXiv:2305.16223v2 [cs.CV] Jun. 1, 2023, 12 pages. [cited by applicant]
Xu, et al., “Reference-based Painterly Inpainting via Diffusion: Crossing the Wild Reference Domain Gap”, arXiv preprint arXiv:2307.10584v1 [cs.CV] Jul. 20, 2023, 15 pages. [cited by applicant]