IP Library Granted Patent US 12,437,375
Granted Patent B2
US 12,437,375 · App. 18/743,497 · Granted Oct 7, 2025

Improving digital image inpainting utilizing plane panoptic segmentation and plane grouping

Inventors: Yuqian Zhou (Urbana, IL); Connelly Barnes (Seattle, WA); Sohrab Amirghodsi (Seattle, WA); Elya Shechtman (Seattle, WA)
Assignee: Adobe Inc.
G06T5/77G06F18/22G06N3/02G06T7/11G06T11/00G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,375
App. No.
18/743,497
Granted
Oct 7, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer readable media for accurately generating inpainted digital images utilizing a guided inpainting model guided by both plane panoptic segmentation and plane grouping. For example, the disclosed systems utilize a guided inpainting model to fill holes of missing pixels of a digital image as informed or guided by an appearance guide and a geometric guide. Specifically, the disclosed systems generate an appearance guide utilizing plane panoptic segmentation and generate a geometric guide by grouping plane panoptic segments. In some embodiments, the disclosed systems generate a modified digital image by implementing an inpainting model guided by both the appearance guide (e.g., a plane panoptic segmentation map) and the geometric guide (e.g., a plane grouping map).

Claims (53)

1. A method comprising:

generating, using an inpainting model, a preliminary inpainted image by inpainting one or more holes of an initial digital image;

generating, from the preliminary inpainted image, an appearance guide from the preliminary inpainted image by:

detecting surface planes within the preliminary inpainted image utilizing a plane detection model; and

detecting panoptic segments within the preliminary inpainted image utilizing a panoptic segmentation model that distinguishes between instances of objects with a shared semantic label; and

generating a final inpainted image by utilizing a guided inpainting model informed by the appearance guide as a guide for inpainting the one or more holes of the initial digital image.

2. The method of claim 1 , wherein generating the appearance guide further comprises:

grouping plane panoptic segments of the appearance guide into a plane grouping map by:

determining normal vectors for plane panoptic segments made up of the panoptic segments and the surface planes; and

grouping the plane panoptic segments into surface plane groups according to the normal vectors.

3. The method of claim 1 , wherein generating the final inpainted image comprises utilizing the guided inpainting model to inpaint the one or more holes of the initial digital image according to an appearance guidance optimization algorithm that constrains sampling based on structure separation indicated by the appearance guide.

4. The method of claim 1 , further comprising generating a geometric guide by grouping the surface planes of the appearance guide based on comparing respective normal vectors of the surface planes;

wherein generating the final inpainted image comprises utilizing the guided inpainting model informed by the geometric guide.

5. The method of claim 4 , wherein generating the final inpainted image comprises utilizing the guided inpainting model informed by the appearance guide and the geometric guide to guide inpainting the one or more holes of the initial digital image.

6. The method of claim 1 , wherein generating the appearance guide comprises generating, for a pixel of the preliminary inpainted image, a triplet label comprising a semantic label, an instance label, and a surface plane identification for content depicted by the pixel.

7. The method of claim 1 , further comprising providing the final inpainted image for display on a client device.

8. A system comprising:

one or more memory components; and

one or more processors coupled to the one or more memory components, wherein the one or more processors are configured to cause the system to perform operations comprising:

generating, using an inpainting model, a preliminary inpainted image by inpainting one or more holes of an initial digital image;

generating, from the preliminary inpainted image, an appearance guide that indicates semantic classes and surface planes from the preliminary inpainted image by:

detecting surface planes within the preliminary inpainted image utilizing a plane detection model; and

detecting panoptic segments within the preliminary inpainted image utilizing a panoptic segmentation model that distinguishes between instances of objects with a shared semantic label; and

generating a final inpainted image by utilizing a guided patch match model informed by the appearance guide for inpainting the one or more holes of the initial digital image.

9. The system of claim 8 , further comprising generating a geometric guide by grouping the surface planes of the appearance guide based on comparing respective normal vectors of the surface planes;

wherein generating the final inpainted image comprises utilizing the guided patch match model informed by the geometric guide.

10. The system of claim 9 , wherein the operations further comprise grouping plane panoptic segments of the appearance guide into a plane grouping map by:

determining normal vectors for plane panoptic segments made up of the panoptic segments and the surface planes; and

grouping the plane panoptic segments into surface plane groups according to the normal vectors.

11. The system of claim 10 , wherein generating the final inpainted image comprises utilizing the guided patch match model to inpaint the one or more holes of the initial digital image as guided by the appearance guide and the plane grouping map.

12. The system of claim 8 , wherein generating the final inpainted image comprises utilizing the guided patch match model to inpaint the one or more holes of the initial digital image according to an appearance guidance optimization algorithm that constrains sampling based on structure separation indicated by the appearance guide.

13. The system of claim 8 , wherein generating the appearance guide comprises generating, for pixels of the preliminary inpainted image, triplet labels that each include a semantic label, an instance label, and a surface plane identification for content depicted by the pixels.

14. The system of claim 8 , wherein:

generating the preliminary inpainted image is at a first resolution; and

generating the final inpainted image is at a second resolution higher than the first resolution.

15. A non-transitory computer readable medium storing instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

generating, using an inpainting model, a preliminary inpainted image by inpainting one or more holes of an initial digital image;

generating, from the preliminary inpainted image, a plane panoptic segmentation map that combines panoptic segments and surface planes from the preliminary inpainted image;

determining a plane grouping map by grouping the surface planes of the plane panoptic segmentation map based on comparing respective normal vectors of the surface planes; and

generating a final inpainted image by utilizing a guided inpainting model informed by the plane panoptic segmentation map and the plane grouping map to guide inpainting the one or more holes of the initial digital image.

16. The non-transitory computer readable medium of claim 15 , wherein generating the plane panoptic segmentation map comprises generating, for a pixel of the preliminary inpainted image, a triplet label comprising a semantic label, an instance label, and a surface plane identification for content depicted by the pixel.

17. The non-transitory computer readable medium of claim 15 , wherein determining the plane grouping map comprises:

performing an early grouping process comprising detecting and clustering lines depicted across all of the preliminary inpainted image according to vanishing points; and

performing a late grouping process comprising detecting lines depicted in individual plane panoptic segments of the plane panoptic segmentation map.

18. The non-transitory computer readable medium of claim 15 , wherein generating the plane panoptic segmentation map comprises:

detecting the surface planes within the preliminary inpainted image utilizing a plane detection model;

detecting the panoptic segments within the preliminary inpainted image utilizing a panoptic segmentation model that distinguishes between instances of objects with a shared semantic label; and

combining the surface planes and the panoptic segments into shared labels.

19. The non-transitory computer readable medium of claim 15 , wherein determining the plane grouping map comprises:

determining normal vectors for the surface planes;

comparing the normal vectors to determine differences among the normal vectors; and

grouping, from comparing the normal vectors, two or more plane panoptic segments into a surface plane group according based on determining that normal vectors of the two or more plane panoptic segments are within a threshold difference.

20. The non-transitory computer readable medium of claim 15 , wherein the operations further comprise providing the final inpainted image for display on a client device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2024
From: ZHOU, YUQIAN; BARNES, CONNELLY; AMIRGHODSI, SOHRAB; SHECHTMAN, ELYA
To: ADOBE INC.
Reel/Frame 067758/0020 →
Continuity (2)
Continuation 17520249 · Nov 5, 2021
Related Publication 20240331114A1 · Oct 3, 2024
References Cited (74)
US 7912255B2 · Rahmes · 2011 [cited by examiner]
US 9406131B2 · Würmlin · 2016 [cited by examiner]
US 10290085B2 · Lin · 2019 [cited by examiner]
US 11023747B2 · Pojman · 2021 [cited by examiner]
US 11210774B2 · Schroers · 2021 [cited by examiner]
US 11282164B2 · Liao · 2022 [cited by examiner]
US 11328392B2 · Bai · 2022 [cited by examiner]
US 11551429B2 · Rong · 2023 [cited by examiner]
US 11580622B2 · Fu · 2023 [cited by examiner]
US 11627318B2 · Danielsson · 2023 [cited by examiner]
US 11887310B2 · Jagadeesh · 2024 [cited by examiner]
US 20180300937A1 · Chien · 2018 [cited by examiner]
US 20210142497A1 · Pugh · 2021 [cited by examiner]
US 20210158043A1 · Hou · 2021 [cited by examiner]
US 20210279866A1 · Svekolkin · 2021 [cited by examiner]
US 20230063150A1 · Tu · 2023 [cited by examiner]
US 20230222671A1 · Kim · 2023 [cited by examiner]
KR 20190101020A · 2019 [cited by examiner]
J.-B. Huang, S. B. Kang, N. Ahuja, and J. Kopf, “Image completion using planar structure guidance,” ACM Transactions on graphics (TOG), vol. 33, No. 4, pp. 1-10, 2014. [cited by applicant]
M. Bertalmio, G. Sapiro, V. Caselles, and C. Ballester, “Image inpainting,” in Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pp. 417-424, 2000. [cited by applicant]
C. Ballester, M. Bertalmio, V. Caselles, G. Sapiro, and J. Verdera, “Filling-in by joint interpolation of vector fields and gray levels,” IEEE transactions on image processing, vol. 10, No. 8, pp. 1200-1211, 2001. [cited by applicant]
Y. Wexler, E. Shechtman, and M. Irani, “Space-time completion of video,” IEEE Transactions on pattern analysis and machine intelligence, vol. 29, No. 3, pp. 463-476, 2007. Part 1. [cited by applicant]
Y. Wexler, E. Shechtman, and M. Irani, “Space-time completion of video,” IEEE Transactions on pattern analysis and machine intelligence, vol. 29, No. 3, pp. 463-476, Part 2. 2007. [cited by applicant]
C. Barnes, E. Shechtman, A. Finkelstein, and D. B. Goldman, “Patchmatch: A randomized correspondence algorithm for structural image editing,” ACM Trans. Graph., vol. 28, No. 3, p. 24, 2009. [cited by applicant]
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2536-2544, 201… [cited by applicant]
Dundar et al“Panoptic-based Image Synthesis”, NVIDIA Corporation, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 8070-8079 (Year: 2020). [cited by applicant]
S. lizuka, E. Simo-Serra, and H. Ishikawa, “Globally and locally consistent image completion,” ACM Transactions on Graphics (ToG), vol. 36, No. 4, pp. 1-14, 2017. [cited by applicant]
Liu et al.“Pan DA: Panoptic Data Augmentation”, California Institute of Technology, Pasadena CA 91125, USA, Apr. 4, 2020 (Year: 2020). [cited by applicant]
Liu et al.“An End-to-End Network for Panoptic Segmentation” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 6172-6181 (Year: 2019). [cited by applicant]
G. Liu, F. A. Reda, K. J. Shih, T.-C. Wang, A. Tao, and B. Catanzaro, “Image inpainting for irregular holes using partial convolutions,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 85-100, 2… [cited by applicant]
J. Yu, Z. Lin, J. Yang, X. Shen, X. Lu, and T. S. Huang, “Free-form image inpainting with gated convolution,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4471-4480, 2019. [cited by applicant]
K. Nazeri, E. Ng, T. Joseph, F. Qureshi, and M. Ebrahimi, “Edgeconnect: Generative image inpainting with adversarial edge learning,” 2019. [cited by applicant]
Y. Song, C. Yang, Y. Shen, P. Wang, Q. Huang, and C.-C. J. Kuo, “Spg-net: Segmentation prediction and guidance network for image inpainting,” arXiv preprint arXiv:1805.03356, 2018. [cited by applicant]
Y. Ren, X. Yu, R. Zhang, T. H. Li, S. Liu, and G. Li, “Structureflow: Image inpainting via structure-aware appearance flow,” in IEEE International Conference on Computer Vision (ICCV), 2019. [cited by applicant]
L. Liao, J. Xiao, Z. Wang, C.-w. Lin, and S. Satoh, “Guidance and evaluation: Semantic-aware image inpainting for mixed scenes,” arXiv preprint arXiv:2003.06877, 2020. [cited by applicant]
C. Zheng, T.-J. Cham, and J. Cai, “Pluralistic image completion,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1438-1447, 2019. [cited by applicant]
C. Yang, X. Lu, Z. Lin, E. Shechtman, O. Wang, and H. Li, “High-resolution image in-painting using multi-scale neural patch synthesis,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p… [cited by applicant]
Y. Zeng, Z. Lin, J. Yang, J. Zhang, E. Shechtman, and H. Lu, “High-resolution image inpainting with iterative confidence feedback and guided upsampling,” in European Conference on Computer Vision, pp. 1-17, Springer, 20… [cited by applicant]
Z. Yi, Q. Tang, S. Azizi, D. Jang, and Z. Xu, “Contextual residual aggregation for ultra high-resolution image inpainting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7508-… [cited by applicant]
Z. Yi, Q. Tang, S. Azizi, D. Jang, and Z. Xu, “Contextual residual aggregation for ultra high-resolution image inpainting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7508-… [cited by applicant]
Z. Yi, Q. Tang, S. Azizi, D. Jang, and Z. Xu, “Contextual residual aggregation for ultra high-resolution image inpainting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7508-… [cited by applicant]
M. Lukac, D. Sykora, K. Sunkavalli, E. Shechtman, O. Jamriska, N. Carr, and T. Pajdla, “Nautilus: Recovering regional symmetry transformations for image editing,” ACM Transactions on Graphics (TOG), vol. 36, No. 4, pp. … [cited by applicant]
M. Lukac, D. Sykora, K. Sunkavalli, E. Shechtman, O. Jamriska, N. Carr, and T. Pajdla, “Nautilus: Recovering regional symmetry transformations for image editing,” ACM Transactions on Graphics (TOG), vol. 36, No. 4, pp. … [cited by applicant]
M. Lukac, D. Sykora, K. Sunkavalli, E. Shechtman, O. Jamriska, N. Carr, and T. Pajdla, “Nautilus: Recovering regional symmetry transformations for image editing,” ACM Transactions on Graphics (TOG), vol. 36, No. 4, pp. … [cited by applicant]
C. Cao and Y. Fu, “Learning a sketch tensor space for image inpainting of man-made scenes,” arXiv preprint arXiv:2103.15087, 2021. [cited by applicant]
Z. Wan, J. Zhang, D. Chen, and J. Liao, “High-fidelity pluralistic image completion with transformers,” arXiv preprint arXiv:2103.14031, 2021. [cited by applicant]
C. Liu, J. Yang, D. Ceylan, E. Yumer, and Y. Furukawa, “Planenet: Piece-wise planar reconstruction from a single rgb image,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2579-258… [cited by applicant]
C. Liu, K. Kim, J. Gu, Y. Furukawa, and J. Kautz, “Planercnn: 3d plane detection and reconstruction from a single image,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4450-44… [cited by applicant]
A. Kirillov, K. He, R. Girshick, C. Rother, and p. Doll'ar, “Panoptic segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9404-9413, 2019. [cited by applicant]
D. Deng, Z. Chen, and B. E. Shi, “Multitask emotion recognition with incomplete labels,” in 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020), pp. 592-599, IEEE, 2020. [cited by applicant]
Y. Zhou, J. Huang, X. Dai, S. Liu, L. Luo, Z. Chen, and Y. Ma, “Holicity: A city-scale data platform for learning holistic 3d structures,” arXiv preprint arXiv:2008.03286, 2020. [cited by applicant]
Y. Zhu, Y. Tian, D. Metaxas, and p. Doll'ar, “Semantic amodal segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1464-1472, 2017. [cited by applicant]
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in Proceedings of the IEEE conference on computer vision and pattern recognition,… [cited by applicant]
Z. Li, T.-W. Yu, S. Sang, S. Wang, M. Song, Y. Liu, Y.-Y. Yeh, R. Zhu, N. Gundavarapu, J. Shi, et al., “Openrooms: An open framework for photorealistic indoor scene datasets,” in Proceedings of the IEEE/CVF Conference o… [cited by applicant]
Z. Li, T.-W. Yu, S. Sang, S. Wang, M. Song, Y. Liu, Y.-Y. Yeh, R. Zhu, N. Gundavarapu, J. Shi, et al., “Openrooms: An open framework for photorealistic indoor scene datasets,” in Proceedings of the IEEE/CVF Conference o… [cited by applicant]
Z. Li, M. Shafiei, R. Ramamoorthi, K. Sunkavalli, and M. Chandraker, “Inverse rendering for complex indoor scenes: Shape, spatially-varying lighting and svbrdf from a single image,” in Proceedings of the IEEE/CVF Confer… [cited by applicant]
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European conference on computer vision, pp. 740-755, Springer, 2014. [cited by applicant]
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 633-641, 2017. [cited by applicant]
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proceedings of the IEEE conference on compute… [cited by applicant]
K. He, G. Gkioxari, P. Dollar, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision, pp. 2961-2969, 2017. [cited by applicant]
Y. Li, H. Zhao, X. Qi, L. Wang, Z. Li, J. Sun, and J. Jia, “Fully convolutional networks for panoptic segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 214-223, 202… [cited by applicant]
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” arXiv preprint arXiv:2105.15203, 2021. [cited by applicant]
B. Cheng, A. G. Schwing, and A. Kirillov, “Per-pixel classification is not all you need for semantic segmentation,” arXiv preprint arXiv:2107.06278, 2021. [cited by applicant]
R. Toldo and A. Fusiello, “Robust multiple structures estimation with j-linkage,” in European conference on computer vision, pp. 537-547, Springer, 2008. [cited by applicant]
Y. Zhou, H. Qi, J. Huang, and Y. Ma, “Neurvps: neural vanishing point scanning via conic convolution,” arXiv preprint arXiv:1910.06316, 2019. [cited by applicant]
J.-B. Huang, A. Singh, and N. Ahuja, “Single image super-resolution from transformed self-exemplars,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5197-5206, 2015. [cited by applicant]
P. Denis, J. H. Elder, and F. J. Estrada, “Efficient edge-based methods for estimating manhattan frames in urban imagery,” in European conference on computer vision, pp. 197-210, Springer, 2008. [cited by applicant]
K. Chaudhury, S. DiVerdi, and S. Ioffe, “Auto-rectification of user photos,” in 2014 IEEE International Conference on Image Processing (ICIP), pp. 3479-3483, IEEE, 2014. [cited by applicant]
S. Zhao, J. Cui, Y. Sheng, Y. Dong, X. Liang, E. I. Chang, and Y. Xu, “Large scale image completion via co-modulated generative adversarial networks,” in International Conference on Learning Representations (ICLR), 2021. [cited by applicant]
C. Zheng, T.-J. Cham, and J. Cai, “Tfill: Image completion via a transformer-based architecture,” arXiv preprint arXiv:2104.00845, 2021. [cited by applicant]
Y. Wu, A. Kirillov, F. Massa, W.-Y. Lo, and R. Girshick, “Detectron2.” https://github.com/facebookresearch/detectron2, 2019. [cited by applicant]
F. Kluger, E. Brachmann, H. Ackermann, C. Rother, M. Y. Yang, and B. Rosenhahn, “Consac: Robust multi-model fitting by conditional sample consensus,” in Proceedings of the IEEE/CVF conference on computer vision and patt… [cited by applicant]
U.S. Appl. No. 17/520,249, Feb. 12, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 17/520,249, May 1, 2024, Notice of Allowance. [cited by applicant]