IP Library › Granted Patent US 12,743,783
Granted Patent B2
US 12,743,783 · App. 18/487,453 · Granted Sep 22, 2026

Generating improved panoptic segmented digital images based on panoptic segmentation neural networks that utilize exemplar unknown object classes

Inventors: Jaedong Hwang (Seoul, KR); Seoung Wug Oh (San Jose, CA); Joon-Young Lee (Miliptas, CA)
Assignee: Adobe Inc.
G06T7/11G06F18/24137G06V10/40G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,783
App. No.
18/487,453
Granted
Sep 22, 2026
Kind
B2
Abstract

This disclosure describes one or more implementations of a panoptic segmentation system that generates panoptic segmented digital images that classify both known and unknown instances of digital images. For example, the panoptic segmentation system builds and utilizes a panoptic segmentation neural network to discover, cluster, and segment new unknown object subclasses for previously unknown object instances. In addition, the panoptic segmentation system can determine additional unknown object instances from additional digital images. Moreover, in some implementations, the panoptic segmentation system utilizes the newly generated unknown object subclasses to refine and tune the panoptic segmentation neural network to improve the detection of unknown object instances in input digital images.

Claims (55)

1 . A computer-implemented method comprising:

receiving, via a client device, a panoptic segmentation request comprising a digital image portraying a plurality of objects;

generating, from the digital image utilizing a panoptic segmentation neural network trained utilizing unknown object instance clusters to classify known object classes and unknown object classes, a panoptic segmented image comprising a semantic map of known labels assigned to a first set of objects and unknown labels assigned to a second set of objects for the plurality of objects; and

providing, for display via the client device, the panoptic segmented image with the known labels and the unknown labels for the plurality of objects.

2 . The computer-implemented method of claim 1 , wherein generating the panoptic segmented image comprises:

generating, utilizing the panoptic segmentation neural network, a first unknown label of a first unknown object class for a first object of the plurality of objects; and

generating, utilizing the panoptic segmentation neural network, a second unknown label of a second unknown object class for a second object of the plurality of objects.

3 . The computer-implemented method of claim 2 , wherein providing the panoptic segmented image for display via the client device further comprises:

providing the first unknown label of the first unknown object class for display with the first object in the panoptic segmented image; and

providing the second unknown label of the second unknown object class for display with the second object in the panoptic segmented image.

4 . The computer-implemented method of claim 2 , wherein the panoptic segmentation neural network is trained to classify the first unknown object class based on a first unknown object class cluster and the second unknown object class based on a second unknown object class cluster.

5 . The computer-implemented method of claim 1 , wherein generating the panoptic segmented image comprises:

generating, utilizing the panoptic segmentation neural network, a first unknown label instance of a first unknown object class for a first object of the plurality of objects; and

generating, utilizing the panoptic segmentation neural network, a second unknown label instance of the first unknown object class for a second object of the plurality of objects.

6 . The computer-implemented method of claim 5 , wherein providing the panoptic segmented image for display via the client device further comprises:

providing the first unknown label instance of the first unknown object class for display with the first object in the panoptic segmented image; and

providing the second unknown label instance of the first unknown object class for display with the second object in the panoptic segmented image.

7 . The computer-implemented method of claim 1 , wherein providing the panoptic segmented image for display comprises providing a semantic map for display with the known labels and the unknown labels, wherein the semantic map comprises pixel classifications for pixels of the digital image.

8 . The computer-implemented method of claim 7 , further comprising editing the digital image by selecting and modifying an object of the plurality of objects having an unknown label of the unknown labels according to the pixel classifications of the semantic map.

9 . A system comprising:

one or more memory devices; and

one or more processors configured to cause the system to:

receive, via a client device, a panoptic segmentation request comprising a digital image portraying a plurality of objects;

generate, from the digital image utilizing a panoptic segmentation neural network trained utilizing unknown object instance clusters to classify known object classes and unknown object classes, a panoptic segmented image comprising a semantic map of known labels assigned to a first set of objects and unknown labels assigned to a second set of objects for the plurality of objects; and

provide, for display via the client device, the panoptic segmented image with the known labels and the unknown labels for the plurality of objects.

10 . The system of claim 9 , wherein the one or more processors are further configured to cause the system to generate the panoptic segmented image by:

generating, utilizing the panoptic segmentation neural network, a first unknown label of a first unknown object class for a first object of the plurality of objects; and

generating, utilizing the panoptic segmentation neural network, a second unknown label of a second unknown object class for a second object of the plurality of objects.

11 . The system of claim 10 , wherein the one or more processors are further configured to cause the system to provide the panoptic segmented image for display via the client device by:

providing the first unknown label of the first unknown object class for display with the first object in the panoptic segmented image; and

providing the second unknown label of the second unknown object class for display with the second object in the panoptic segmented image.

12 . The system of claim 10 , wherein the panoptic segmentation neural network is trained to classify the first unknown object class based on a first unknown object class cluster and the second unknown object class based on a second unknown object class cluster.

13 . The system of claim 9 , wherein the one or more processors are further configured to cause the system to:

generate the panoptic segmented image by generating, utilizing the panoptic segmentation neural network, a first unknown label instance of a first unknown object class for a first object of the plurality of objects and a second unknown label instance of the first unknown object class for a second object of the plurality of objects; and

provide the panoptic segmented image for display by providing, for display, the first unknown label instance of the first unknown object class for display with the first object in the panoptic segmented image and the second unknown label instance of the first unknown object class for display with the second object in the panoptic segmented image.

14 . The system of claim 9 , wherein the one or more processors are further configured to cause the system to:

provide the panoptic segmented image for display by providing a semantic map for display with the known labels and the unknown labels, wherein the semantic map comprises pixel classifications for pixels of the digital image; and

edit the digital image by selecting and modifying an object of the plurality of objects having an unknown label of the unknown labels according to the pixel classifications of the semantic map.

15 . A non-transitory computer readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

receive, via a client device, a panoptic segmentation request comprising a digital image portraying a plurality of objects;

generate, from the digital image utilizing a panoptic segmentation neural network trained utilizing unknown object instance clusters to classify known object classes and unknown object classes, a panoptic segmented image comprising a semantic map of known labels assigned to a first set of objects and unknown labels assigned to a second set of objects for the plurality of objects; and

provide, for display via the client device, the panoptic segmented image with the known labels and the unknown labels for the plurality of objects.

16 . The non-transitory computer readable medium of claim 15 , wherein generating the panoptic segmented image comprises:

generating, utilizing the panoptic segmentation neural network, a first unknown label of a first unknown object class for a first object of the plurality of objects; and

generating, utilizing the panoptic segmentation neural network, a second unknown label of a second unknown object class for a second object of the plurality of objects.

17 . The non-transitory computer readable medium of claim 16 , wherein providing the panoptic segmented image for display via the client device comprises:

providing the first unknown label of the first unknown object class for display with the first object in the panoptic segmented image; and

providing the second unknown label of the second unknown object class for display with the second object in the panoptic segmented image.

18 . The non-transitory computer readable medium of claim 16 , wherein the panoptic segmentation neural network is trained to classify the first unknown object class based on a first unknown object class cluster and the second unknown object class based on a second unknown object class cluster.

19 . The non-transitory computer readable medium of claim 15 , wherein the operations further comprise:

generating the panoptic segmented image by generating, utilizing the panoptic segmentation neural network, a first unknown label instance of a first unknown object class for a first object of the plurality of objects and a second unknown label instance of the first unknown object class for a second object of the plurality of objects; and

providing the panoptic segmented image for display by providing, for display, the first unknown label instance of the first unknown object class for display with the first object in the panoptic segmented image and the second unknown label instance of the first unknown object class for display with the second object in the panoptic segmented image.

20 . The non-transitory computer readable medium of claim 15 , wherein the operations further comprise:

providing the panoptic segmented image for display by providing a semantic map for display with the known labels and the unknown labels, wherein the semantic map comprises pixel classifications for pixels of the digital image; and

editing the digital image by selecting and modifying an object of the plurality of objects having an unknown label of the unknown labels according to the pixel classifications of the semantic map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2023
From: HWANG, JAEDONG; OH, SEOUNG WUG; LEE, JOON-YOUNG
To: ADOBE INC.
Reel/Frame 065233/0989 →
Continuity (2)
Continuation 17319979 · May 13, 2021
Related Publication 20240037750A1 · Feb 1, 2024
References Cited (97)
US 11055566B1 · Pham et al. · 2021 [cited by applicant]
US 11176384B1 · Yang et al. · 2021 [cited by applicant]
US 11417097B2 · Lin et al. · 2022 [cited by applicant]
US 11562490B2 · Homayounfar et al. · 2023 [cited by applicant]
US 11587234B2 · Zhao et al. · 2023 [cited by applicant]
US 20170206431A1 · Sun et al. · 2017 [cited by applicant]
US 20190244060A1 · Dundar · 2019 [cited by examiner]
US 20190347804A1 · Kim · 2019 [cited by examiner]
US 20200082219A1 · Li · 2020 [cited by applicant]
US 20210026355A1 · Chen · 2021 [cited by examiner]
US 20210150230A1 · Smolyanskiy et al. · 2021 [cited by applicant]
US 20210158043A1 · Hou et al. · 2021 [cited by applicant]
US 20210263962A1 · Chang et al. · 2021 [cited by applicant]
US 20210319236A1 · Tang · 2021 [cited by examiner]
US 20210326638A1 · Lee et al. · 2021 [cited by applicant]
US 20210326656A1 · Lee et al. · 2021 [cited by applicant]
US 20210358127A1 · Jagadeesh et al. · 2021 [cited by applicant]
US 20210366128A1 · Kim et al. · 2021 [cited by applicant]
US 20210397876A1 · Hemani et al. · 2021 [cited by applicant]
US 20220230321A1 · Zhao et al. · 2022 [cited by applicant]
US 20220264149A1 · Lee · 2022 [cited by examiner]
US 20220309762A1 · Zhao et al. · 2022 [cited by applicant]
US 20220375090A1 · Hwang · 2022 [cited by examiner]
US 20220404504A1 · Kim · 2022 [cited by examiner]
US 20230063150A1 · Tu · 2023 [cited by examiner]
US 20230072731A1 · Li · 2023 [cited by examiner]
US 20230154005A1 · Borse · 2023 [cited by examiner]
US 20230206525A1 · Harikumar · 2023 [cited by examiner]
US 20230259587A1 · Lin · 2023 [cited by examiner]
US 20230281824A1 · Mei · 2023 [cited by examiner]
US 20230334673A1 · Rajan · 2023 [cited by examiner]
US 20240404064A1 · Nguyen · 2024 [cited by examiner]
US 20250104249A1 · Lee · 2025 [cited by examiner]
EP 3985551A1 · 2022 [cited by examiner]
EP 4138039A2 · 2023 [cited by examiner]
Abhijit Bendale and Terrance E Boult. Towards open set deep networks. In CVPR, 2016. [cited by applicant]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning libra… [cited by applicant]
Akshay Dhamija, Manuel Gunther, Jonathan Ventura, and Terrance Boult. The overlooked elephant of object detection: Open set. In WACV, 2020. [cited by applicant]
Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollar. Panoptic segmentation. In CVPR, 2019. [cited by applicant]
Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollar. Panoptic feature pyramid networks. In CVPR, 2019. [cited by applicant]
Ameya Prabhu, Philip HS Torr, and Puneet K Dokania. Gdumb: A simple approach that questions our progress in continual learning. In ECCV, 2020. [cited by applicant]
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, 2017. [cited by applicant]
Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen. Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In CVPR, 2020. [cited by applicant]
Cheng-Yang Fu, Tamara L Berg, and Alexander C Berg. Imp: Instance mask projection for high accuracy semantic segmentation of things. In ICCV, 2019. [cited by applicant]
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In CVPR, 2015. [cited by applicant]
Daniel Bolya, Chong Zhou, Fanyi Xiao, and Yong Jae Lee. Yolact: Real-time instance segmentation. In CVPR, 2019. [cited by applicant]
Dimity Miller, Lachlan Nicholson, Feras Dayoub, and Niko Sunderhauf. Dropout sampling for robust object detection in open-set conditions. In ICRA, 2018. [cited by applicant]
Dong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh, and Anton van den Hengel. Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly det… [cited by applicant]
Douglas L Medin and Marguerite M Schaffer. Context theory of classification learning. Psychological review, 85(3):207, 1978. [cited by applicant]
Gerhard Neuhold, Tobias Ollmann, Samuel Rota Bulo, and Peter Kontschieder. The mapillary vistas dataset for semantic understanding of street scenes. In ICCV, 2017. [cited by applicant]
Huanyu Liu, Chao Peng, Changqian Yu, Jingbo Wang, Xu Liu, Gang Yu, and Wei Jiang. An end-to-end network for panoptic segmentation. In CVPR, 2019. [cited by applicant]
Huiyu Wang, Yukun Zhu, Bradley Green, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Axial-deeplab: Stand-alone axial-attention for panoptic segmentation. In ECCV, 2020. [cited by applicant]
Hyeonwoo Noh, Seunghoon Hong, and Bohyung Han. Learning deconvolution network for semantic segmentation. In ICCV, 2015. [cited by applicant]
Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. In NeurIPS, 2017. [cited by applicant]
James MacQueen. Some methods for classification and analysis of multivariate observations. In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability, vol. 1, pp. 281-297. Oakland, CA, USA… [cited by applicant]
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. Bilinear attention networks. In NeurIPS, 2018. [cited by applicant]
John Lambert, Zhuang Liu, Ozan Sener, James Hays, and Vladlen Koltun. Mseg: A composite dataset for multi-domain semantic segmentation. In CVPR, 2020. [cited by applicant]
Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. [cited by applicant]
Joseph Redmon and Ali Farhadi. Yolo9000: better, faster, stronger. In CVPR, 2017. [cited by applicant]
Justin Lazarow, Kwonjoon Lee, Kunyu Shi, and Zhuowen Tu. Learning instance occlusion for panoptic segmentation. In CVPR, 2020. [cited by applicant]
Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In ICCV, 2017. [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In CVPR, 2016. [cited by applicant]
Kelsey R Allen, Evan Shelhamer, Hanul Shin, and Joshua B Tenenbaum. Infinite mixture prototypes for few-shot learning. ICML, 2019. [cited by applicant]
Lawrence Neal, Matthew Olson, Xiaoli Fern, Weng-Keen Wong, and Fuxin Li. Open set learning with counterfactual images. In ECCV, 2018. [cited by applicant]
Lei Guo, Gang Xie, Xinying Xu, and Jinchang Ren. Exemplar-supported representation for effective class-incremental learning. IEEE Access, 8:51276-51284, 2020. [cited by applicant]
Li et al. “A Survey on Deep Learning Based Panoptic Segmentation” Digital Signal Processing 120, Oct. 12, 2021, pp. 1-20. [cited by applicant]
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. TPAMI, 40(4):834-8… [cited by applicant]
Lorenzo Porzi, Samuel Rota Bulo, Aleksander Colovic, and Peter Kontschieder. Seamless scene segmentation. In CVPR, 2019. [cited by applicant]
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. [cited by applicant]
Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation, vol. 24, pp. 109-165. Elsevier, 1989. [cited by applicant]
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognitio… [cited by applicant]
Pramuditha Perera, Vlad I Morariu, Rajiv Jain, Varun Manjunatha, Curtis Wigington, Vicente Ordonez, and Vishal M Patel. Generative-discriminative feature representations for open-set recognition. In CVPR, 2020. [cited by applicant]
Priya Goyal, Piotr Dollar, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large mini-batch sgd: Training imagenet in 1 hour. arXiv preprint arXiv… [cited by applicant]
Qizhu Li, Anurag Arnab, and Philip HS Torr. Weakly-and semi-supervised panoptic segmentation. In ECCV, 2018. [cited by applicant]
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced d… [cited by applicant]
Robert M Nosofsky. Attention, similarity, and the identification-categorization relationship. Journal of experimental psychology: General, 115(1):39, 1986. [cited by applicant]
Ross Girshick. Fast r-cnn. In CVPR, 2015. [cited by applicant]
Shaoan Xie, Zibin Zheng, Liang Chen, and Chuan Chen. Learning semantic representations for unsupervised domain adaptation. In ICML, 2018. [cited by applicant]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NeurIPS, 2015. [cited by applicant]
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In CVPR, 2017. [cited by applicant]
Thomas Cover and Peter Hart. Nearest neighbor pattern classification. TIT, 13(1):21-27, 1967. [cited by applicant]
Tien-Ju Yang, Maxwell D Collins, Yukun Zhu, Jyh-Jing Hwang, Ting Liu, Xiao Zhang, Vivienne Sze, George Papandreou, and Liang-Chieh Chen. Deeperlab: Single-shot image parser. arXiv preprint arXiv:1902.05093, 2019. [cited by applicant]
Trung Pham, Vijay BG Kumar, Thanh-Toan Do, Gustavo Carneiro, and Ian Reid. Bayesian semantic instance segmentation in open set world. In ECCV, 2018. [cited by applicant]
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. [cited by applicant]
Tsung-Yi Lin, Piotr Dollar, Ross B Girshick, Kaiming He, Bharath Hariharan, and Serge J Belongie. Feature pyramid networks for object detection. In CVPR, 2017. [cited by applicant]
Walter J Scheirer, Anderson de Rezende Rocha, Archana Sapkota, and Terrance E Boult. Toward open set recognition. TPAMI, 35(7):1757-1772, 2012. [cited by applicant]
Walter J Scheirer, Lalit P Jain, and Terrance E Boult. Probability models for open set recognition. TPAMI, 36(11):2317-2324, 2014. [cited by applicant]
Xu Yang, Kaihua Tang, Hanwang Zhang, and Jianfei Cai. Auto-encoding scene graphs for image captioning. In CVPR, 2019. [cited by applicant]
Yanwei Li, Xinze Chen, Zheng Zhu, Lingxi Xie, Guan Huang, Dalong Du, and Xingang Wang. Attention-guided unified network for panoptic segmentation. In CVPR, 2019. [cited by applicant]
Yuwen Xiong, Renjie Liao, Hengshuang Zhao, Rui Hu, Min Bai, Ersin Yumer, and Raquel Urtasun. Upsnet: A unified panoptic segmentation network. In CVPR, 2019. [cited by applicant]
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019. [cited by applicant]
Zendel, et al. “Unifying Panoptic Segmentation for Autonomous Driving” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 21351-21360. [cited by applicant]
Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In CVPR, 2018. [cited by applicant]
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. Deep modular co-attention networks for visual question answering. In CVPR, 2019. [cited by applicant]
Ziwei Liu, Zhongqi Miao, Xiaohang Zhan, Jiayun Wang, Boqing Gong, and Stella X Yu. Large-scale long-tailed recognition in an open world. In CVPR, 2019. [cited by applicant]
U.S. Appl. No. 17/319,979, filed Mar. 30, 2023, Office Action. [cited by applicant]
U.S. Appl. No. 17/319,979, filed Jun. 14, 2023, Notice of Allowance. [cited by applicant]