IP Library › Granted Patent US 12,367,586
Granted Patent B2
US 12,367,586 · App. 17/937,680 · Granted Jul 22, 2025

Learning parameters for neural networks using a semantic discriminator and an object-level discriminator

Inventors: Zhe Lin (Fremont, CA); Haitian Zheng (Rochester, NY); Elya Shechtman (Seattle, WA); Jianming Zhang (Campbell, CA); Jingwan Lu (Santa Clara, CA); Ning Xu (Milpitas, CA); Qing Liu (Santa Clara, CA); Scott Cohen (Sunnyvale, CA); Sohrab Amirghodsi (Seattle, WA)
Assignee: Adobe Inc.
G06T7/11G06T2207/20081G06T2207/20084G06T2207/20132
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,586
App. No.
17/937,680
Granted
Jul 22, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer readable media for panoptically guiding digital image inpainting utilizing a panoptic inpainting neural network. In some embodiments, the disclosed systems utilize a panoptic inpainting neural network to generate an inpainted digital image according to panoptic segmentation map that defines pixel regions corresponding to different panoptic labels. In some cases, the disclosed systems train a neural network utilizing a semantic discriminator that facilitates generation of digital images that are realistic while also conforming to a semantic segmentation. The disclosed systems generate and provide a panoptic inpainting interface to facilitate user interaction for inpainting digital images. In certain embodiments, the disclosed systems iteratively update an inpainted digital image based on changes to a panoptic segmentation map.

Claims (60)

1. A non-transitory computer readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

generating a predicted digital image from a semantic segmentation of a digital image utilizing a neural network;

generating, utilizing a semantic discriminator, a semantic image embedding from the predicted digital image and a panoptic condition combining a binary mask, a normalized semantic embedding, and an edge map;

combining the semantic image embedding with an image embedding of the predicted digital image;

generating a realism prediction, utilizing the semantic discriminator, from the combination of the semantic image embedding and the image embedding; and

modifying parameters of the neural network based on the realism prediction.

2. The non-transitory computer readable medium of claim 1 , wherein generating the realism prediction comprises:

generating the image embedding from the predicted digital image utilizing a first encoder of the semantic discriminator;

generating the semantic image embedding from the predicted digital image and the semantic segmentation utilizing a second encoder of the semantic discriminator; and

determining the realism prediction from a concatenation of the image embedding and the semantic image embedding.

3. The non-transitory computer readable medium of claim 1 , wherein generating the realism prediction comprises utilizing the semantic discriminator to determine realism of the predicted digital image together with conformity of the predicted digital image to the semantic segmentation.

4. The non-transitory computer readable medium of claim 1 , wherein generating the realism prediction comprises utilizing the semantic discriminator as part of an image-level discriminator to determine a realism score for an entirety of the predicted digital image.

5. The non-transitory computer readable medium of claim 1 , wherein generating the realism prediction comprises utilizing the semantic discriminator as part of an object-level discriminator to determine a realism score for a crop of the predicted digital image.

6. The non-transitory computer readable medium of claim 1 , wherein generating the realism prediction comprises utilizing the semantic discriminator to generate a first realism score and utilizing a generative adversarial discriminator to generate a second realism score.

7. The non-transitory computer readable medium of claim 1 , wherein modifying the parameters of the neural network comprises:

determining an adversarial loss utilizing the semantic discriminator;

determining a reconstruction loss by comparing the predicted digital image with the digital image; and

modifying the parameters of the neural network based on the adversarial loss and the reconstruction loss.

8. A system comprising:

one or more memory devices comprising a digital image, a semantic segmentation, a neural network, an image-level semantic discriminator, and an object-level semantic discriminator; and

one or more processors configured to cause the system to learn parameters for the neural network utilizing the image-level semantic discriminator and the object-level semantic discriminator by:

generating a predicted digital image from the digital image and the semantic segmentation utilizing the neural network;

generating, utilizing the object-level semantic discriminator, a semantic image embedding from the predicted digital image and a panoptic condition combining a binary mask, a normalized semantic embedding, and an edge map;

combining the semantic image embedding with an image embedding of the predicted digital image;

generating, utilizing the image-level semantic discriminator and the object-level semantic discriminator, a realism prediction from the combination of the semantic image embedding and the image embedding; and

modifying parameters of the neural network based on the realism prediction.

9. The system of claim 8 , wherein generating the realism prediction comprises:

determining a bounding box for a crop of the predicted digital image; and

utilizing the object-level semantic discriminator to determine a realism score for the crop of the predicted digital image.

10. The system of claim 9 , wherein generating the realism prediction further comprises:

identifying a binary mask indicating background pixels and foreground pixels for the crop of the predicted digital image; and

utilizing the object-level semantic discriminator to determine the realism score for the foreground pixels of the crop of the predicted digital image indicated by the binary mask.

11. The system of claim 8 , wherein combining the semantic image embedding with the image embedding of the predicted digital image comprises:

utilizing an image embedding model to extract the image embedding from the predicted digital image; and

concatenating the semantic image embedding with the image embedding.

12. The system of claim 11 , wherein generating the semantic image embedding comprises determining the panoptic condition by:

identifying the binary mask indicating pixels to replace within a sample digital image;

generating the normalized semantic embedding indicating semantic labels for objects depicted within the sample digital image; and

determining the edge map defining boundaries between the objects depicted within the sample digital image.

13. The system of claim 8 , wherein modifying the parameters of the neural network comprises:

determining an overall adversarial loss by combining a first adversarial loss associated with the image-level semantic discriminator and a second adversarial loss associated with the object-level semantic discriminator; and

modifying the parameters based on the overall adversarial loss.

14. The system of claim 8 , wherein generating the realism prediction comprises:

generating a crop of the predicted digital image;

generating a cropped binary mask, a cropped semantic label map, and a cropped edge map associated with sample digital image data; and

utilizing the object-level semantic discriminator to generate the realism prediction from the crop of the predicted digital image, the cropped binary mask, the cropped semantic label map, and the cropped edge map.

15. A computer-implemented method comprising:

generating a predicted digital image from a semantic segmentation of a digital image utilizing a neural network;

generating, utilizing a first encoder of a semantic discriminator, an image embedding from the predicted digital image;

generating, utilizing a second encoder of the semantic discriminator, a semantic image embedding from the predicted digital image and a panoptic condition combining a binary mask, a normalized semantic embedding, and an edge map;

combining the semantic image embedding with the image embedding;

determining a realism prediction, utilizing the semantic discriminator, from the combination of the image embedding and the semantic image embedding; and

modifying parameters of the neural network based on the realism prediction.

16. The computer-implemented method of claim 15 , wherein determining the realism prediction comprises:

utilizing the semantic discriminator as part of an image-level discriminator to determine a realism score for an entirety of the predicted digital image; and

utilizing an additional semantic discriminator as part of an object-level discriminator to determine a realism score for a crop of the predicted digital image.

17. The computer-implemented method of claim 15 , wherein generating the semantic image embedding from the predicted digital image and the panoptic condition further comprises determining, from the sample digital image data, the binary mask indicating pixels to replace within a sample digital image, the normalized semantic embedding representing semantic labels for objects within the sample digital image, the edge map reflecting boundaries between the objects within the sample digital image.

18. The computer-implemented method of claim 17 , wherein determining the realism prediction comprises utilizing the semantic discriminator to generate a realism score for the predicted digital image based on the panoptic condition.

19. The computer-implemented method of claim 15 , further comprising determining an overall adversarial loss by combining a first adversarial loss associated with an image-level semantic discriminator, a second adversarial loss associated with an object-level semantic discriminator, a third adversarial loss associated with an image-level generative adversarial discriminator, and a fourth adversarial loss associated with an object-level generative adversarial discriminator.

20. The computer-implemented method of claim 19 , wherein modifying the parameters of the neural network comprises modifying the parameters to reduce the overall adversarial loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2022
From: LIN, ZHE; ZHENG, HAITIAN; SHECHTMAN, ELYA; ZHANG, JIANMING; LU, JINGWAN; XU, NING; LIU, QING; COHEN, SCOTT; AMIRGHODSI, SOHRAB
To: ADOBE INC.
Reel/Frame 061293/0001 →
Continuity (1)
Related Publication 20240127452A1 · Apr 18, 2024
References Cited (101)
US 10922793B2 · Baek et al. · 2021 [cited by applicant]
US 12159412B2 · Baruch et al. · 2024 [cited by applicant]
US 20190196698A1 · Cohen et al. · 2019 [cited by applicant]
US 20200242774A1 · Park et al. · 2020 [cited by applicant]
US 20200405242A1 · Kearney et al. · 2020 [cited by applicant]
US 20210097691A1 · Liu · 2021 [cited by applicant]
US 20210118099A1 · Kearney et al. · 2021 [cited by applicant]
US 20210158043A1 · Hou et al. · 2021 [cited by applicant]
US 20210158491A1 · Li et al. · 2021 [cited by applicant]
US 20210342983A1 · Lin et al. · 2021 [cited by applicant]
US 20210357684A1 · Amirghodsi et al. · 2021 [cited by applicant]
US 20220129682A1 · Tang et al. · 2022 [cited by applicant]
US 20220172369A1 · Tang et al. · 2022 [cited by applicant]
US 20220366544A1 · Kudelski et al. · 2022 [cited by applicant]
US 20230368339A1 · Zheng et al. · 2023 [cited by applicant]
US 20240013351A1 · Liu · 2024 [cited by applicant]
CN 111556278A · 2020 [cited by applicant]
CN 115205161A · 2022 [cited by applicant]
GB 2619381A · 2023 [cited by applicant]
Li, Daiqing, et al. “Semantic segmentation with generative models: Semi-supervised learning and strong out-of-domain generalization.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 20… [cited by examiner]
Ang Li, Jianzhong Qi, Rui Zhang, Ramamohanarao Kotagiri, Boosted GAN with Semantically Interpretable Information at least for Image Inpainting, IJCNN, Budapest, Hungary, 2019. [cited by applicant]
Combined Search and Examination Report as received in GB application 2311936.5 dated Feb. 6, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB application 2311866.4 dated Feb. 9, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB application 2311936.5 dated Feb. 9, 2024. [cited by applicant]
Heng Wang, et al., “Region and Object Based Panoptic Image Synthesis Through Conditional GANs”. arXiv preprint arXiv:1912.060840v1, Dec. 14, 2019. [cited by applicant]
Jean-François Aujol, Guy Gilboa, Tony Chan, and Stanley Osher. 2006. Structure-Texture Image Decomposition—Modeling, Algorithms, and Parameter Selection. International journal of computer vision 67, 1 (2006), 111-136. [cited by applicant]
Coloma Ballester, Marcelo Bertalmio, Vicent Caselles, Guillermo Sapiro, and Joan Verdera. 2001. Filling-in by joint interpolation of vector fields and gray levels. IEEE transactions on image processing 10, 8 (2001), 120… [cited by applicant]
Connelly Barnes, Eli Shechtman, Adam Finkelstein, and Dan B Goldman. 2009. PatchMatch: A randomized correspondence algorithm for structural image editing. ACM Trans. Graph. 28, 3 (2009), 24. [cited by applicant]
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba. 2020. Understanding the role of individual units in a deep neural network. Proceedings of the National Academy of Sciences 117… [cited by applicant]
Marcelo Bertalmio, Luminita Vese, Guillermo Sapiro, and Stanley Osher. 2003. Simultaneous structure and texture image inpainting. IEEE transactions on image processing 12, 8 (2003), 882-889. [cited by applicant]
Tony F Chan and Jianhong Shen. 2001. Nontexture inpainting by curvature-driven diffusions. Journal of visual communication and image representation 12, 4 (2001), 436-449. [cited by applicant]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning. PMLR, 1597-1607. [cited by applicant]
Lu Chi, Borui Jiang, and Yadong Mu. 2020. Fast fourier convolution. Advances in Neural Information Processing Systems 33 (2020). [cited by applicant]
Taeg Sang Cho, Moshe Butman, Shai Avidan, and William T Freeman. 2008. The patch transform and its applications to image editing. In 2008 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1-8. [cited by applicant]
Antonio Criminisi, Patrick Perez, and Kentaro Toyama. 2004. Region filling and object removal by exemplar-based image inpainting. IEEE Transactions on image processing 13, 9 (2004), 1200-1212. [cited by applicant]
Soheil Darabi, Eli Shechtman, Connelly Barnes, Dan B Goldman, and Pradeep Sen. 2012. Image melding: Combining inconsistent images using patch-based synthesis. ACM Transactions on graphics (TOG) 31, 4 (2012), 1-10. [cited by applicant]
Alexei A Efros andWilliam T Freeman. 2001. Image quilting for texture synthesis and transfer. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques. ACM, 341-346. [cited by applicant]
Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12873-12883. [cited by applicant]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, DavidWarde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in neural information processing systems. 26… [cited by applicant]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Proce… [cited by applicant]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems 33 (2020), 6840-6851. [cited by applicant]
Seunghoon Hong, Xinchen Yan, Thomas S Huang, and Honglak Lee. 2018. Learning hierarchical semantic image manipulation through structured representations. Advances in Neural Information Processing Systems 31 (2018). [cited by applicant]
Mbing SongWei Huang Hongyu Liu, Bin Jiang and Chao Yang. 2020. Rethinking Image Inpainting via a Mutual Encoder-Decoder with Feature Equalizations. In Proceedings of the European Conference on Computer Vision. [cited by applicant]
Satoshi Iizuka, Edgar Simo-Serra, and Hiroshi Ishikawa. 2017. Globally and locally consistent image completion. ACM Transactions on Graphics (ToG) 36, 4 (2017), 1-14. [cited by applicant]
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition. 112… [cited by applicant]
Jongheon Jeong and Jinwoo Shin. 2021. Training gans with stronger augmentations via contrastive discriminator. arXiv preprint arXiv:2103.09742 (2021). [cited by applicant]
Youngjoo Jo and Jongyoul Park. 2019. Sc-fegan: Face editing generative adversarial network with user's sketch and color. In Proceedings of the IEEE/CVF international conference on computer vision. 1745-1753. [cited by applicant]
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016. Perceptual losses for real-time style transfer and super-resolution. In European conference on computer vision. Springer, 694-711. [cited by applicant]
Minguk Kang and Jaesik Park. 2020. Contragan: Contrastive learning for conditional image generation. Advances in Neural Information Processing Systems 33 (2020), 21357-21369. [cited by applicant]
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2017. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196 (2017). [cited by applicant]
Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4401-4410. [cited by applicant]
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020. Analyzing and Improving the Image Quality of StyleGAN. In Proc. CVPR. [cited by applicant]
Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). [cited by applicant]
Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollár. 2019. Panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9404-9413. [cited by applicant]
Nupur Kumari, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. 2021. Ensembling Off-the-shelf Models for GAN Training. arXiv preprint arXiv:2112.09130 (2021). [cited by applicant]
Vivek Kwatra, Irfan Essa, Aaron Bobick, and Nipun Kwatra. 2005. Texture optimization for example-based synthesis. In ACM SIGGRAPH 2005 Papers. 795-802. [cited by applicant]
Chuan Li and Michael Wand. 2016. Precomputed real-time texture synthesis with markovian generative adversarial networks. In European conference on computer vision. Springer, 702-716. [cited by applicant]
Yanwei Li, Hengshuang Zhao, Xiaojuan Qi, Liwei Wang, Zeming Li, Jian Sun, and Jiaya Jia. 2021. Fully Convolutional Networks for Panoptic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat… [cited by applicant]
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European conference on computer vision. Spr… [cited by applicant]
Guilin Liu, Fitsum A Reda, Kevin J Shih, Ting-Chun Wang, Andrew Tao, and Bryan Catanzaro. 2018. Image inpainting for irregular holes using partial convolutions. In Proceedings of the European Conference on Computer Visi… [cited by applicant]
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022. RePaint: Inpainting using Denoising Diffusion Probabilistic Models. arXiv preprint arXiv:2201.09865 (2022). [cited by applicant]
Liqian Ma, Xu Jia, Qianru Sun, Bernt Schiele, Tinne Tuytelaars, and Luc Van Gool. 2017. Pose guided person image generation. Advances in neural information processing systems 30 (2017). [cited by applicant]
Mehdi Mirza and Simon Osindero. 2014. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784 (2014). [cited by applicant]
Kamyar Nazeri, Eric Ng, Tony Joseph, Faisal Z Qureshi, and Mehran Ebrahimi. 2019. Edgeconnect: Generative image inpainting with adversarial edge learning. arXiv preprint arXiv:1901.00212 (2019). [cited by applicant]
Evangelos Ntavelis, Andres Romero, lason Kastanis, Luc Van Gool, and Radu Timofte. 2020. Sesame: Semantic editing of scenes by adding, manipulating or erasing objects. In European Conference on Computer Vision. Springer… [cited by applicant]
Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. 2019. Semantic Image Synthesis with Spatially-Adaptive Normalization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. [cited by applicant]
Taesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu, Eli Shechtman, Alexei A Efros, and Richard Zhang. 2020. Swapping autoencoder for deep image manipulation. arXiv preprint arXiv:2007.00653 (2020). [cited by applicant]
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros. 2016. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recogniti… [cited by applicant]
Tiziano Portenier, Qiyang Hu, Attila Szabo, Siavash Arjomand Bigdeli, Paolo Favaro, and Matthias Zwicker. 2018. Faceshop: Deep sketch-based face image editing. arXiv preprint arXiv:1804.08972 (2018). [cited by applicant]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language … [cited by applicant]
Chitwan Saharia, William Chan, Huiwen Chang, Chris A Lee, Jonathan Ho, Tim Salimans, David J Fleet, and Mohammad Norouzi. 2021. Palette: Image-to-image diffusion models. arXiv preprint arXiv:2111.05826 (2021). [cited by applicant]
Axel Sauer, Kashyap Chitta, Jens Müller, and Andreas Geiger. 2021. Projected gans converge faster. Advances in Neural Information Processing Systems 34 (2021). [cited by applicant]
Axel Sauer, Katja Schwarz, and Andreas Geiger. 2022. StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets. arXiv preprint arXiv:2202.00273 (2022). [cited by applicant]
Jianhong Shen and Tony F Chan. 2002. Mathematical models for local nontexture inpaintings. SIAM J. Appl. Math. 62, 3 (2002), 1019-1043. [cited by applicant]
Yuhang Song, Chao Yang, Yeji Shen, Peng Wang, Qin Huang, and C-C Jay Kuo. 2018. Spg-net: Segmentation prediction and guidance network for image inpainting. arXiv preprint arXiv:1805.03356 (2018). [cited by applicant]
Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. 2021. Resolution-robust Large Mask Inpainting… [cited by applicant]
Zheng, Haitian, et al. “Image Inpainting with Cascaded Modulation GAN and Object-Aware Training.” Computer Vision—ECCV 2022: 17th European Conference, Tel Aviv, Israel, Oct. 23-27, 2022, Proceedings, Part XVI. Cham: Spr… [cited by applicant]
Ziyu Wan, Jingbo Zhang, Dongdong Chen, and Jing Liao. 2021. High-Fidelity Pluralistic Image Completion with Transformers. arXiv preprint arXiv:2103.14031 (2021). [cited by applicant]
Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros. 2020. Cnn-generated images are surprisingly easy to spot . . . for now. In Proceedings of the IEEE/CVF conference on computer vision and patte… [cited by applicant]
Wei Xiong, Jiahui Yu, Zhe Lin, Jimei Yang, Xin Lu, Connelly Barnes, and JiebovLuo. 2019. Foreground-aware image inpainting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5840-5848. [cited by applicant]
Jie Yang, Zhiquan Qi, and Yong Shi. 2020. Learning to incorporate structure knowledge for image inpainting. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34. 12605-12612. [cited by applicant]
Zili Yi, Qiang Tang, Shekoofeh Azizi, Daesik Jang, and Zhan Xu. 2020. Contextual residual aggregation for ultra high-resolution image inpainting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern … [cited by applicant]
Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. 2018. Generative image inpainting with contextual attention. In Proceedings of the IEEE conference on computer vision and pattern recognition. 55… [cited by applicant]
Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. 2019. Free-form image inpainting with gated convolution. In Proceedings of the IEEE International Conference on Computer Vision. 4471-4480. [cited by applicant]
Ning Yu, Guilin Liu, Aysegul Dundar, Andrew Tao, Bryan Catanzaro, Larry S Davis, and Mario Fritz. 2021. Dual contrastive loss and attention for gans. In Proceedings of the IEEE/CVF International Conference on Computer V… [cited by applicant]
Yu Zeng, Zhe Lin, and Vishal M Patel. 2021. SketchEdit: Mask-Free Local Image Manipulation with Partial Sketches. arXiv preprint arXiv:2111.15078 (2021). [cited by applicant]
Yu Zeng, Zhe Lin, Jimei Yang, Jianming Zhang, Eli Shechtman, and Huchuan Lu. 2020. High-Resolution Image Inpainting with Iterative Confidence Feedback and Guided Upsampling. arXiv preprint arXiv:2005.11742 (2020). [cited by applicant]
Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang. 2021. Cross-modal contrastive learning for text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pa… [cited by applicant]
Shengyu Zhao, Jonathan Cui, Yilun Sheng, Yue Dong, Xiao Liang, Eric I Chang, and Yan Xu. 2021. Large scale image completion via co-modulated generative adversarial networks. arXiv preprint arXiv:2103.10428 (2021). [cited by applicant]
Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai. 2021. Pluralistic Free-From Image Completion. International Journal of Computer Vision (2021), 1-20. [cited by applicant]
Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai, and Dinh Phung. 2021. Bridging Global Context Interactions for High-Fidelity Image Completion. (2021). arXiv:cs.CV/2104.00845. [cited by applicant]
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017. Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence 40, 6 (2017),… [cited by applicant]
Examination Report as received in GB application 2311936.5 dated Oct. 16, 2024. [cited by applicant]
Martin Belan, “Testing The Photoshop 2022 Objection Selection Tool on a Variety of Nature Photographs.” Oct. 30, 2021. Retrieved from the internet: <https://blog.martinbelan.com/2021/10/30/testing-the-photoshop-2023-obj… [cited by applicant]
Berta Bescos, et al. “DynaSLAM: Tracking, Mapping, and Inpainting in Dynamic Scenes”, IEEE Robotics and Automation Letters 3.4 (2018): 4076-4083. (Year: 2018). [cited by applicant]
Hu Zhu et al., “Fusing Panoptic Segmentation and Geometry Information for Robust Visual SLAM in Dynamic Environments”, IEEE 18th International Conference on Automation Science and Engineering (CASE). IEEE, (Year: 2022). [cited by applicant]
Jasper R.R. Uijlings, Mykhaylo Andriluka, Vittorio Ferrari, “Panoptic Image Annotation with a Collaborative Assistant”, Proceedings of the 28th ACM International Conference on Multimedia. (Year: 2020). [cited by applicant]
U.S. Appl. No. 17/937,695, filed Mar. 18, 2025. [cited by applicant]
U.S. Appl. No. 17/937,706, filed Mar. 12, 2025. [cited by applicant]
U.S. Appl. No. 17/937,708, filed Feb. 5, 2025. [cited by applicant]