IP Library Granted Patent US 12,430,727
Granted Patent B2
US 12,430,727 · App. 18/058,027 · Granted Sep 30, 2025

Image and object inpainting with diffusion models

Inventors: Haitian Zheng (Rochester, NY); Zhe Lin (Clyde Hill, WA); Jianming Zhang (Fremont, CA); Connelly Stuart Barnes (Seattle, WA); Elya Shechtman (Seattle, WA); Jingwan Lu (Sunnyvale, CA); Qing Liu (Santa Clara, CA); Sohrab Amirghodsi (Seattle, WA); Yuqian Zhou (Bellevue, WA); Scott Cohen (Cupertino, CA)
Assignee: ADOBE INC.
G06T5/77G06T5/73G06T2207/20081G06T2207/20104
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,727
App. No.
18/058,027
Granted
Sep 30, 2025
Kind
B2
Abstract

Systems and methods for image processing are described. Embodiments of the present disclosure receive an image comprising a first region that includes content and a second region to be inpainted. Noise is then added to the image to obtain a noisy image, and a plurality of intermediate output images are generated based on the noisy image using a diffusion model trained using a perceptual loss. The intermediate output images predict a final output image based on a corresponding intermediate noise level of the diffusion model. The diffusion model then generates the final output image based on the intermediate output image. The final output image includes inpainted content in the second region that is consistent with the content in the first region.

Claims (58)

1. A method comprising:

receiving an image comprising a first region that includes content and a second region to be inpainted;

adding noise to the image to obtain a noisy image;

generating a plurality of intermediate output images based on the noisy image using a diffusion model, wherein the diffusion model is trained using a perceptual loss based on a difference between perceptual features extracted from a predicted image and perceptual features extracted from a clean image, and wherein each of the plurality of intermediate output images comprises an intermediate prediction of a final output image based on a corresponding intermediate noise level of the diffusion model; and

generating the final output image based on the plurality of intermediate output images using the diffusion model, wherein the final output image includes inpainted content in the second region that is consistent with the content in the first region.

2. The method of claim 1 , further comprising:

providing the image as an input to the diffusion model, wherein the intermediate output image is conditioned based on the first region of the image.

3. The method of claim 1 , further comprising:

encoding the noisy image to obtain image features; and

decoding the image features to obtain the intermediate output image.

4. The method of claim 1 , further comprising:

receiving a user input indicating the second region to be inpainted.

5. A method comprising:

initializing a diffusion model;

receiving training data including an image comprising content in a first region and in a second region;

masking the second region to obtain a masked image;

adding noise to the masked image to obtain a noisy image;

identifying a first kernel at a first step;

identifying a second kernel at a second step, wherein a size of the second kernel is different from a size of the first kernel;

computing an adaptively-blurred perceptual loss based on the first kernel and the second kernel; and

training the diffusion model to inpaint the second region based on the noisy image using a perceptual loss, wherein the perceptual loss comprises the adaptively-blurred perceptual loss.

6. The method of claim 5 , further comprising:

computing a predicted noise based on the noisy image; and

computing a reconstruction loss based on the predicted noise and a ground truth noise, wherein the diffusion model is trained based on the reconstruction loss.

7. The method of claim 5 , further comprising:

computing a weighted signal-to-noise-ratio loss, wherein the diffusion model is trained based on the weighted signal-to-noise-ratio loss.

8. The method of claim 5 , further comprising:

generating a predicted output image based on the noisy image using the diffusion model; and

comparing the predicted output image to the image to obtain the perceptual loss.

9. The method of claim 5 , further comprising:

adding intermediate noise to a predicted output image to obtain a noisy output image;

encoding the noisy output image to obtain image features;

identifying an intermediate noisy image between the image and the noisy image;

encoding the intermediate noisy image to obtain intermediate image features; and

comparing the image features and the intermediate image features to obtain a sample-based perceptual loss, wherein the diffusion model is trained based on the sample-based perceptual loss.

10. The method of claim 9 , further comprising:

identifying a plurality of intermediate noisy images including the intermediate noisy image between the image and the noisy image, wherein the sample-based perceptual loss is computed based on the plurality of intermediate noisy images.

11. The method of claim 9 , further comprising:

selecting a plurality of samples of the predicted output image, wherein the sample-based perceptual loss is computed based on the plurality of samples of the predicted output image.

12. The method of claim 5 , further comprising:

identifying a filter of a predetermined kernel size, wherein the adaptively-blurred perceptual loss is computed based on the filter.

13. An apparatus comprising:

a processor; and

a memory including instructions executable by the processor to:

receive an image comprising a first region that includes content and a second region to be inpainted;

add noise to the image to obtain a noisy image;

generate a plurality of intermediate output images based on the noisy image using a diffusion model, wherein the diffusion model is trained using a perceptual loss based on a difference between perceptual features extracted from a predicted image and perceptual features extracted from a clean image, and wherein each of the plurality of intermediate output images comprises an intermediate prediction of a final output image based on a corresponding intermediate noise level of the diffusion model; and

generate the final output image based on the plurality of intermediate output images using the diffusion model, wherein the final output image includes inpainted content in the second region that is consistent with the content in the first region.

14. The apparatus of claim 13 , wherein:

the diffusion model comprises a U-Net architecture.

15. The apparatus of claim 13 , wherein:

the diffusion model comprises a denoising diffusion probabilistic model (DDPM).

16. The apparatus of claim 13 , wherein:

the perceptual loss comprises a sample-based perceptual loss.

17. The apparatus of claim 13 , wherein:

the perceptual loss comprises an adaptively-blurred perceptual loss.

18. The apparatus of claim 13 , further comprising:

a user interface configured to receive a user input indicating the second region to be inpainted.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2022
From: ZHENG, HAITIAN; LIN, ZHE; ZHANG, JIANMING; BARNES, CONNELLY STUART; SHECHTMAN, ELYA; LU, JINGWAN; LIU, QING; AMIRGHODSI, SOHRAB; ZHOU, YUQIAN; COHEN, SCOTT
To: ADOBE INC.
Reel/Frame 061856/0115 →
Continuity (1)
Related Publication 20240169500A1 · May 23, 2024
References Cited (23)
US 11798132B2 · Wang · 2023 [cited by examiner]
US 20220156896A1 · Li · 2022 [cited by examiner]
US 20230067841A1 · Saharia · 2023 [cited by examiner]
US 20230245285A1 · Dinh · 2023 [cited by examiner]
US 20230377214A1 · Kansy · 2023 [cited by examiner]
US 20240135630A1 · Nagano · 2024 [cited by examiner]
WO WO2022108373A1 · 2022 [cited by applicant]
Combined Search and Examination Report dated Mar. 6, 2024 in corresponding United Kingdom Patent Application No. 2314582.4 (5 pages). [cited by applicant]
1Ho, et al., “Denoising Diffusion Probabilistic Models”, Advances in Neural Information Processing Systems, 33, pp. 6840-6851, arXiv preprint arXiv:2006.11239v2 [cs.LG] Dec. 16, 2020. [cited by applicant]
2Nichol, et al., “Improved Denoising Diffusion Probabilistic Models”, In International Conference on Machine Learning (pp. 8162-8171), PMLR, arXiv preprint arXiv:2102.09672v1 [cs.LG] Feb. 18, 2021. [cited by applicant]
3Saharia, et al., “Palette: Image-to-Image Diffusion Models”, In ACM SIGGRAPH 2022 Conference Proceedings (pp. 1-10), arXiv preprint arXiv:2111.05826v2 [cs.CV] May 3, 2022. [cited by applicant]
4Suvorov, et al., “Resolution-robust Large Mask Inpainting with Fourier Convolutions”, In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, (2022), (pp. 2149-2159). [cited by applicant]
5Song, et al., “Denoising Diffusion Implicit Models”, arXiv preprint arXiv:2010.02502v4 [cs.LG] Oct. 5, 2022, 22 pages. [cited by applicant]
6Yu, et al., “Generative Image Inpainting with Contextual Attention”, In Proceedings of the IEEE conference on computer vision and pattern recognition, (2018), (pp. 5505-5514). [cited by applicant]
Yu, et al., “Free-Form Image Inpainting with Gated Convolution”, In Proceedings of the IEEE/CVF international conference on computer vision (pp. 4471-4480); arXiv preprint arXiv:1806.03589v2 [cs.CV] Oct. 22, 2019. [cited by applicant]
8Zhao, et al., “Large Scale Image Completion via Co-Modulated Generative Adversarial Networks”, arXiv preprint arXiv:2103.10428v1 [cs.CV] Mar. 18, 2021, 25 pages. [cited by applicant]
9Zheng, et al., “Image Inpainting with Cascaded Modulation GAN and Object-Aware Training”, arXiv preprint arXiv:2203.11947v3 [cs.CV] Jul. 21, 2022, 32 pages. [cited by applicant]
10Zeng, et al., “High-Resolution Image Inpainting with Iterative Confidence Feedback and Guided Upsampling”, In European conference on computer vision (pp. 1-17), Springer, Cham, arXiv preprint arXiv:2005.11742v2 [cs.CV… [cited by applicant]
11Benny, et al., “Dynamic Dual-Output Diffusion Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 11482-11491); arXiv preprint arXiv:2203.04304v2 [cs.CV] Mar. 15, 2022. [cited by applicant]
12Choi, et al., “Perception Prioritized Training of Diffusion Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2022), (pp. 11472-11481). [cited by applicant]
13Xiao, et al., “Tackling the Generative Learning Trilemma With Denoising Diffusion GANS”, arXiv preprint arXiv:2112.07804v2 [cs.LG] Apr. 4, 2022, 28 pages. [cited by applicant]
14Rombach, et al., “High-Resolution Image Synthesis with Latent Diffusion Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10684-10695), arXiv preprint arXiv:2112.10752v… [cited by applicant]
15Lugmayr, et al., “RePaint: Inpainting using Denoising Diffusion Probabilistic Models”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 11461-11471), arXiv preprint arXiv:2201.… [cited by applicant]