IP Library Granted Patent US 12,499,520
Granted Patent B2
US 12,499,520 · App. 17/815,409 · Granted Dec 16, 2025

Generating neural network based perceptual artifact segmentations in modified portions of a digital image

Inventors: Sohrab Amirghodsi (Seattle, WA); Lingzhi Zhang (Philadelphia, PA); Zhe Lin (Fremont, CA); Elya Shechtman (Seattle, WA); Yuqian Zhou (Urbana, IL); Connelly Barnes (Seattle, WA)
Assignee: Adobe Inc.
G06T5/77G06T7/194G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,520
App. No.
17/815,409
Granted
Dec 16, 2025
Kind
B2
Abstract

Methods, systems, and non-transitory computer readable storage media are disclosed for generating neural network based perceptual artifact segmentations in synthetic digital image content. The disclosed system utilizing neural networks to detect perceptual artifacts in digital images in connection with generating or modifying digital images. The disclosed system determines a digital image including one or more synthetically modified portions. The disclosed system utilizes an artifact segmentation machine-learning model to detect perceptual artifacts in the synthetically modified portion(s). The artifact segmentation machine-learning model is trained to detect perceptual artifacts based on labeled artifact regions of synthetic training digital images. Additionally, the disclosed system utilizes the artifact segmentation machine-learning model in an iterative inpainting process. The disclosed system utilizes one or more digital image inpainting models to inpaint in a digital image. The disclosed system utilizes the artifact segmentation machine-learning model detect perceptual artifacts in the inpainted portions for additional inpainting iterations.

Claims (67)

1 . A computer-implemented method comprising:

determining, by at least one processor, one or more portions of a digital image to fill by generating one or more synthetically modified portions of the digital image utilizing a digital image inpainting model; and

generating one or more artifact segmentations from the digital image by:

determining, utilizing an artifact segmentation machine-learning model, one or more predicted perceptual artifact regions indicating one or more artifacts comprising visible errors in synthetically generated image content, the one or more predicted perceptual artifact regions corresponding to pixels within the one or more synthetically modified portions of the digital image,

wherein the artifact segmentation machine-learning model comprises parameters learned based on labeled artifact regions of synthetic training digital images.

2 . The computer-implemented method of claim 1 , wherein determining the digital image comprises:

generating the one or more synthetically modified portions utilizing an image generation neural network; or

selecting the digital image comprising the one or more synthetically modified portions from a database of digital images.

3 . The computer-implemented method of claim 1 , wherein generating the one or more artifact segmentations comprises determining, utilizing the artifact segmentation machine-learning model, a plurality of predicted perceptual artifact regions corresponding to a plurality of separate artifacts within a synthetically modified portion of the digital image.

4 . The computer-implemented method of claim 1 , further comprising:

determining a combined size of one or more artifacts within a synthetically modified portion of the one or more synthetically modified portions of the digital image;

determining a size of the synthetically modified portion of the digital image; and

generating an artifact ratio metric for the digital image based on the combined size of the one or more artifacts relative to the size of the synthetically modified portion.

5 . The computer-implemented method of claim 4 , further comprising generating, utilizing an image generation neural network, an additional synthetically modified portion replacing an artifact within the synthetically modified portion in response to comparing the artifact ratio metric to a ratio threshold.

6 . The computer-implemented method of claim 4 , further comprising:

determining, based on the artifact ratio metric for the digital image, a first performance of a first image generation neural network utilized to generate the one or more synthetically modified portions of the digital image;

determining, based on an additional artifact ratio metric for an additional version of the digital image, a second performance of a second image generation neural network utilized to generate one or more additional synthetically modified portions of the additional version of the digital image; and

providing, for display within a graphical user interface, a comparison of the first performance of the first image generation neural network and the second performance of the second image generation neural network.

7 . The computer-implemented method of claim 1 , further comprising:

generating a plurality of candidate digital images comprising synthetically modified portions, the plurality of candidate digital images comprising the digital image;

generating a plurality of artifact ratio metrics corresponding to the plurality of candidate digital images based on artifacts relative to sizes of the synthetically modified portions of the plurality of candidate digital images; and

selecting the digital image from the plurality of candidate digital images based on the plurality of artifact ratio metrics.

8 . The computer-implemented method of claim 1 , further comprising:

generating predicted artifact bounding regions for portions of the synthetic training digital images;

generating modified artifact bounding regions by dilating synthetically modified regions corresponding to the predicted artifact bounding regions by a predetermined amount; and

providing, for display at a client device, the synthetic training digital images including the modified artifact bounding regions.

9 . The computer-implemented method of claim 8 , further comprising:

determining a plurality of marked regions of the synthetic training digital images based on user inputs;

determining the labeled artifact regions by intersecting the plurality of marked regions and hole masks corresponding to synthetically modified portions of the synthetic training digital images; and

learning the parameters of the artifact segmentation machine-learning model based on the labeled artifact regions.

10 . A system comprising:

one or more computer memory devices; and

one or more servers configured to cause the system to:

determine one or more portions of a digital image to fill by generating one or more synthetically modified portions of the digital image utilizing a digital image inpainting model;

generate one or more artifact segmentations from the digital image by:

determining, utilizing an artifact segmentation machine-learning model, one or more predicted perceptual artifact regions indicating one or more artifacts comprising visible errors in synthetically generated image content, the one or more predicted perceptual artifact regions corresponding to pixels within the one or more synthetically modified portions of the digital image,

wherein the artifact segmentation machine-learning model comprises parameters learned based on labeled artifact regions of synthetic training digital images; and

generate, for display within a graphical user interface of a client device, one or more indications of the one or more artifact segmentations in the digital image.

11 . The system of claim 10 , wherein the one or more servers are further configured to cause the system to determine the digital image by generating, utilizing the digital image inpainting model, the one or more synthetically modified portions according to a digital image mask associated with a detected object.

12 . The system of claim 10 , wherein the one or more servers are further configured to cause the system to generate the one or more artifact segmentations by:

determining, utilizing the artifact segmentation machine-learning model, a first predicted perceptual artifact region corresponding to a first artifact within a synthetically modified portion of the digital image; and

determining, utilizing the artifact segmentation machine-learning model, a second predicted perceptual artifact region corresponding to a second artifact within the synthetically modified portion of the digital image.

13 . The system of claim 10 , wherein the one or more servers are further configured to cause the system to generate the one or more indications of the one or more predicted perceptual artifact regions by:

generating an artifact ratio metric for the digital image based on a combined size of the one or more predicted perceptual artifact regions relative to a combined size of the one or more synthetically modified portions; and

providing, for display within a graphical user interface of a client device, an indication to further modify the one or more synthetically modified portions of the digital image in response to comparing the artifact ratio metric to a ratio threshold.

14 . The system of claim 10 , wherein the one or more servers are further configured to cause the system to generate the one or more indications of the one or more predicted perceptual artifact regions by:

generating an artifact ratio metric for the digital image based on a combined size of the one or more predicted perceptual artifact regions relative to a combined size of the one or more synthetically modified portions;

determining, based on the artifact ratio metric for the digital image, a performance of an image generation neural network that generated the one or more synthetically modified portions of the digital image; and

providing, for display within a graphical user interface of a client device, an indication of the performance of the image generation neural network.

15 . The system of claim 10 , wherein the one or more servers are further configured to cause the system to determine the labeled artifact regions of the synthetic training digital images by:

generating a predicted artifact bounding region for a portion of a synthetic training digital image of the synthetic training digital image;

generating a modified artifact bounding region by dilating a synthetically modified region corresponding to the predicted artifact bounding region to a rectangle enclosing the synthetically modified region;

providing, for display at a client device, the synthetic training digital image comprising the modified artifact bounding region and a duplicate of the synthetic training digital image; and

determining, based on a user input via the client device, a labeled artifact region indicating an artifact within the modified artifact bounding region.

16 . A non-transitory computer readable medium comprising instructions that, when executed by at least one processor, cause a computing device to:

determining one or more portions of a digital image to fill by generating one or more synthetically modified portions of the digital image utilizing a digital image inpainting model;

generating one or more artifact segmentations from the digital image by:

determining, utilizing an artifact segmentation machine-learning model, one or more predicted perceptual artifact regions indicating one or more artifacts comprising visible errors in synthetically generated image content, the one or more predicted perceptual artifact regions corresponding to pixels within the one or more synthetically modified portions of the digital image,

wherein the artifact segmentation machine-learning model comprises parameters learned based on labeled artifact regions of synthetic training digital images; and

generating, for display within a graphical user interface of a client device, one or more indications of the one or more predicted perceptual artifact regions based on a size of the one or more predicted perceptual artifact regions.

17 . The non-transitory computer readable medium of claim 16 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate an artifact ratio metric for the digital image by determining a ratio of a combined size of the one or more artifacts relative to a combined size of the one or more synthetically modified portions.

18 . The non-transitory computer readable medium of claim 17 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the one or more indications of the one or more predicted perceptual artifact regions by generating, in response to comparing the artifact ratio metric to a ratio threshold, a recommendation to generate an additional synthetic modified portion within a portion of the digital image corresponding to an artifact segmentation of the one or more artifact segmentations.

19 . The non-transitory computer readable medium of claim 17 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the one or more indications of the one or more predicted perceptual artifact regions by generating, based on the artifact ratio metric, a performance comparison of an image generation neural network utilized to generate the digital image relative to an additional image generation neural network.

20 . The non-transitory computer readable medium of claim 17 , further comprising instructions that, when executed by the at least one processor, cause the computing device to:

provide, to a plurality of client devices, the synthetic training digital images with dilated artifact bounding regions corresponding to a plurality of artifacts in the synthetic training digital images;

determine the labeled artifact regions in response to user interactions with the synthetic training digital images; and

learn the parameters of the artifact segmentation machine-learning model based on the labeled artifact regions in the synthetic training digital images and ground-truth training digital images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2022
From: AMIRGHODSI, SOHRAB; ZHANG, LINGZHI; LIN, ZHE; SHECHTMAN, ELYA; ZHOU, YUQIAN; BARNES, CONNELLY
To: ADOBE INC.
Reel/Frame 060644/0240 →
Continuity (1)
Related Publication 20240037717A1 · Feb 1, 2024
References Cited (25)
US 20170193594A1 · Glasgow · 2017 [cited by examiner]
US 20210264591A1 · Park et al. · 2021 [cited by applicant]
US 20220084181A1 · Isken · 2022 [cited by examiner]
US 20230073223A1 · Bergmann · 2023 [cited by examiner]
US 20230169325A1 · Xie et al. · 2023 [cited by applicant]
CN 114387642A · 2022 [cited by examiner]
WO WO2022119870A1 · 2022 [cited by examiner]
Barnes, C., Shechtman, E., Finkelstein, A., Goldman, D.B.: Patchmatch: A ran-452 domized correspondence algorithm for structural image editing. ACM Trans. 453 Graph. 28(3), 24 (2009). [cited by applicant]
Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. International journal of computer vision 88(2), 303-338 (2010). [cited by applicant]
He, K., Gkioxari, G., Dollar, P., Girshick, R.: Mask r-cnn. In: Proceedings of the IEEE international conference on computer vision. pp. 2961-2969 (2017). [cited by applicant]
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770-778 (2016). [cited by applicant]
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S .: Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30 (2017). [cited by applicant]
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll'ar, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European conference on computer vision. pp. 740-755. Springer (2014). [cited by applicant]
Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3431-3-3340 (2015). [cited by applicant]
Nazeri, K., Ng, E., Joseph, T., Qureshi, F.Z., Ebrahimi, M.: Edgeconnect: Generative image inpainting with adversarial edge learning. arXiv preprint arXiv:1901.00212 (2019). [cited by applicant]
Parmar, G., Zhang, R., Zhu, J.Y.: On buggy resizing libraries and surprising subtleties in fid calculation. arXiv preprint arXiv:2104.11222 (2021. [cited by applicant]
Su, S., Yan, Q., Zhu, Y., Zhang, C., Ge, X., Sun, J., Zhang, Y.: Blindly assess image quality in the wild guided by a self-adaptive hyper network. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Patter… [cited by applicant]
Suvorov, R., Logacheva, E., Mashikhin, A., Remizova, A., Ashukha, A., Silvestrov, A., Kong, N., Goka, H., Park, K., Lempitsky, V.: Resolution-robust large mask inpainting with fourier convolutions. arXiv preprint arXiv:… [cited by applicant]
Yu, J., Lin, Z., Yang, J., Shen, X., Lu, X., Huang, T.S.: Free-form image inpainting with gated convolution. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4471-4480 (2019). [cited by applicant]
Zeng, Y., Lin, Z., Yang, J., Zhang, J., Shechtman, E., Lu, H.: High-resolution image inpainting with iterative confidence feedback and guided upsampling. In: European Conference on Computer Vision. pp. 1-17. Springer (2… [cited by applicant]
Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. … [cited by applicant]
Zhao, H., Shi, J., Qi, X., Wang, X., Jia, J.: Pyramid scene parsing network. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 2881-2890 (2017). [cited by applicant]
Zhao, S., Cui, J., Sheng, Y., Dong, Y., Liang, X., Chang, E.I., Xu, Y.: Large scale image completion via co-modulated generative adversarial networks. arXiv preprint arXiv:2103.10428 (2021). [cited by applicant]
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., Torralba, A.: Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence (2017. [cited by applicant]
U.S. Appl. No. 17/815,418, Apr. 28, 2025, Office Action. [cited by applicant]