IP Library › Granted Patent US 12,725,260
Granted Patent B2
US 12,725,260 · App. 18/118,705 · Granted Sep 1, 2026

Generating panoptic segmentation labels

Inventors: Jieru Mei (Baltimore, MD); Hang Yan (Sunnyvale, CA); Liang-Chieh Chen (Los Angeles, CA); Siyuan Qiao (Santa Clara, CA); Yukun Zhu (Shoreline, WA); Alex Zihao Zhu (Mountain View, CA); Xinchen Yan (San Francisco, CA); Henrik Kretzschmar (Mountain View, CA)
Assignee: Waymo LLC
G06T7/11G01S17/89G06V10/82G06V20/58G06V20/64G06T2207/10016G06T2207/10028G06T2207/20081G06T2207/30252G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,260
App. No.
18/118,705
Granted
Sep 1, 2026
Kind
B2
Abstract

Methods, systems, and apparatus for generating a panoptic segmentation label for a sensor data sample. In one aspect, a system comprises one or more computers configured to obtain a sensor data sample characterizing a scene in an environment. The one or more computers obtain a 3D bounding box annotation at each time point for a point cloud characterizing the scene at the time point. The one or more computers obtain, for each camera image and each time point, annotation data identifying object instances depicted in the camera image, and the one or more computers generate a panoptic segmentation label for the sensor data sample characterizing the scene in the environment.

Claims (76)

1 . A method comprising:

obtaining a sensor data sample characterizing a scene in an environment, the sensor data sample comprising one or more camera images of the scene from each of a plurality of cameras that are each located at a different camera viewpoint within the scene;

obtaining respective three-dimensional (3D) bounding box annotations for one or more point clouds characterizing the scene, wherein each 3D bounding box annotation includes one or more 3D bounding boxes in the one or more point clouds, and wherein each 3D bounding box includes a plurality of LIDAR points;

generating a plurality of projected LIDAR points for each 3D bounding box that are in a camera image of the one or more camera images of the scene;

obtaining, for each camera image, annotation data identifying object instances depicted in the camera image; and

generating a panoptic segmentation label for the sensor data sample characterizing the scene in the environment by modifying the annotation data for one or more of the object instances depicted in the camera image, the generating comprising:

identifying a first object depicted in a first camera image from a first camera at a first camera viewpoint within the scene;

identifying a second object depicted in a second camera image from a second camera at a second camera viewpoint within the scene, wherein the second camera viewpoint is different than the first camera viewpoint;

determining that the identified first object and the identified second object are each depicted in a respective overlapping portion that overlaps between the first camera image and the second camera image, wherein the respective overlapping portions correspond to a region of the scene that is simultaneously captured in both the first camera image and the second camera based on the first camera viewpoint and the second camera viewpoint; and

based on determining that the identified first object and the identified second object correspond to a same object instance, assigning a same object instance identifier to a first set of pixels in the first camera image and to a second set of pixels in the second camera image, wherein the first set of pixels and the second set of pixels are each within the respective overlapping portion.

2 . The method of claim 1 , further comprising:

training a neural network configured to perform panoptic segmentation on input sensor samples on training data that includes a training example that associates the sensor sample with the panoptic segmentation label for the sensor data sample.

3 . The method of claim 1 , wherein modifying the annotation data for one or more of the object instances comprises, for each of the one or more object instances:

identifying, using the 3D bounding box annotations, a corresponding 3D bounding box that corresponds to the same object as the object instance.

4 . The method of claim 3 , wherein identifying, using the 3D bounding box annotations, a corresponding 3D bounding box that corresponds to the same object as the object instance comprises, for each camera image:

for each 3D bounding box in a corresponding point cloud that was generated at a same time point as the camera image, generating a plurality of projected points by projecting the points within the 3D bounding box onto the camera image;

for each object instance detected in the camera image:

computing a respective association score for each 3D bounding box based on a similarity between the object instance and the projected points generated for the 3D bounding box; and

determining whether any of the 3D bounding boxes correspond to the same object as the object instance based on the respective association scores.

5 . The method of claim 4 , wherein the respective association score for each 3D bounding box is based on a similarity between the object instance and a convex hull of the projected points generated for the 3D bounding box.

6 . The method of claim 4 , wherein determining whether any of the 3D bounding boxes correspond to the same object as the object instance based on the respective association scores comprises:

performing a bipartite matching across the object instances and the 3D bounding boxes based on the respective association scores for the object instance-3D bounding box pairs.

7 . The method of claim 1 , further comprising:

obtaining, for each camera image, 2D bounding box annotations for the image, the 2D bounding box annotations comprising (i) one or more 2D bounding boxes in the image and (ii) a respective object instance identifier from a second set of object instance identifiers for each 2D bounding box identifying an object depicted within the 2D bounding box.

8 . The method of claim 7 , wherein generating the panoptic segmentation label for the sensor data sample characterizing the scene in the environment further comprises:

for each of one or more of the object instances identified in the annotation data for the camera images:

identifying, using the 3D bounding box annotations and the 2D bounding box annotations, a corresponding 3D bounding box that corresponds to the same object as the object instance; and

associating, with the object instance, the respective object instance identifier for the corresponding 3D bounding box.

9 . The method of claim 7 , wherein generating the panoptic segmentation label for the sensor data sample characterizing the scene in the environment further comprises:

for each of one or more of the object instances identified in the annotation data for the camera images:

identifying, using the 2D bounding box annotations, a corresponding 2D bounding box that corresponds to the same object as the object instance; and

associating, with the object instance, the respective object instance identifier for the corresponding 2D bounding box.

10 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by one or more computers cause the one more computers to perform operations comprising:

obtaining a sensor data sample characterizing a scene in an environment, the sensor data sample comprising one or more camera images of the scene from each of a plurality of cameras that are each located at a different camera viewpoint within the scene;

obtaining respective three-dimensional (3D) bounding box annotations for one or more point clouds characterizing the scene, wherein each 3D bounding box annotation includes one or more 3D bounding boxes in the one or more point clouds, and wherein each 3D bounding box includes a plurality of LIDAR points;

generating a plurality of projected LIDAR points for each 3D bounding box that are in a camera image of the one or more camera images of the scene;

obtaining, for each camera image, annotation data identifying object instances depicted in the camera image; and

generating a panoptic segmentation label for the sensor data sample characterizing the scene in the environment by modifying the annotation data for one or more of the object instances depicted in the camera image, the generating comprising:

identifying a first object depicted in a first camera image from a first camera at a first camera viewpoint within the scene;

identifying a second object depicted in a second camera image from a second camera at a second camera viewpoint within the scene, wherein the second camera viewpoint is different than the first camera viewpoint;

determining that the identified first object and the identified second object are each depicted in a respective overlapping portion that overlaps between the first camera image and the second camera image, wherein the overlapping portions correspond to a region of the scene that is simultaneously captured in both the first camera image and the second camera based on the first camera viewpoint and the second camera viewpoint; and

based on determining that the identified first object and the identified second object correspond to a same object instance, assigning a same object instance identifier to a first set of pixels in the first camera image and to a second set of pixels in the second camera image, wherein the first set of pixels and the second set of pixels are each within the respective overlapping portion.

11 . The system of claim 10 , the operations further comprising:

training a neural network configured to perform panoptic segmentation on input sensor samples on training data that includes a training example that associates the sensor sample with the panoptic segmentation label for the sensor data sample.

12 . The system of claim 10 , wherein modifying the annotation data for one or more of the object instances comprises, for each of the one or more object instances:

identifying, using the 3D bounding box annotations, a corresponding 3D bounding box that corresponds to the same object as the object instance.

13 . The system of claim 12 , wherein identifying, using the 3D bounding box annotations, a corresponding 3D bounding box that corresponds to the same object as the object instance comprises, for each camera image:

for each 3D bounding box in a corresponding point cloud that was generated at a same time point as the camera image, generating a plurality of projected points by projecting the points within the 3D bounding box onto the camera image;

for each object instance detected in the camera image:

computing a respective association score for each 3D bounding box based on a similarity between the object instance and the projected points generated for the 3D bounding box; and

determining whether any of the 3D bounding boxes correspond to the same object as the object instance based on the respective association scores.

14 . The system of claim 13 , wherein the respective association score for each 3D bounding box is based on a similarity between the object instance and a convex hull of the projected points generated for the 3D bounding box.

15 . The system of claim 13 , wherein determining whether any of the 3D bounding boxes correspond to the same object as the object instance based on the respective association scores comprises:

performing a bipartite matching across the object instances and the 3D bounding boxes based on the respective association scores for the object instance-3D bounding box pairs.

16 . The system of claim 10 , the operations further comprising:

obtaining, for each camera image, 2D bounding box annotations for the image, the 2D bounding box annotations comprising (i) one or more 2D bounding boxes in the image and (ii) a respective object instance identifier from a second set of object instance identifiers for each 2D bounding box identifying an object depicted within the 2D bounding box.

17 . The system of claim 16 , wherein generating the panoptic segmentation label for the sensor data sample characterizing the scene in the environment further comprises:

for each of one or more of the object instances identified in the annotation data for the camera images:

identifying, using the 3D bounding box annotations and the 2D bounding annotations, a corresponding 3D bounding box that corresponds to the same object as the object instance; and

associating, with the object instance, the respective object instance identifier for the corresponding 3D bounding box.

18 . The system of claim 16 , wherein generating the panoptic segmentation label for the sensor data sample characterizing the scene in the environment further comprises:

for each of one or more of the object instances identified in the annotation data for the camera images:

identifying, using the 2D bounding box annotations, a corresponding 2D bounding box that corresponds to the same object as the object instance; and

associating, with the object instance, the respective object instance identifier for the corresponding 2D bounding box.

19 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one more computers to perform operations comprising:

obtaining a sensor data sample characterizing a scene in an environment, the sensor data sample comprising one or more camera images of the scene from each of a plurality of cameras that are each located at a different camera viewpoint within the scene;

obtaining respective three-dimensional (3D) bounding box annotations for one or more point clouds characterizing the scene, wherein each 3D bounding box annotation includes one or more 3D bounding boxes in the one or more point clouds, and wherein each 3D bounding box includes a plurality of LIDAR points;

generating a plurality of projected LIDAR points for each 3D bounding box that are in a camera image of the one or more camera images of the scene;

obtaining, for each camera image, annotation data identifying object instances depicted in the camera image; and

generating a panoptic segmentation label for the sensor data sample characterizing the scene in the environment by modifying the annotation data for one or more of the object instances depicted in the camera image, the generating comprising:

identifying a first object depicted in a first camera image from a first camera at a first camera viewpoint within the scene;

identifying a second object depicted in a second camera image from a second camera at a second camera viewpoint within the scene, wherein the second camera viewpoint is different than the first camera viewpoint;

determining that the identified first object and the identified second object are each depicted in a respective overlapping portion that overlap between the first camera image and the second camera image, wherein the respective overlapping portions correspond to a region of the scene that is simultaneously captured in both the first camera image and the second camera based on the first camera viewpoint and the second camera viewpoint; and

based on determining that the identified first object and the identified second object correspond to a same object instance, assigning a same object instance identifier to a first set of pixels in the first camera image and to a second set of pixels in the second camera image, wherein the first set of pixels and the second set of pixels are each within the respective overlapping portion.

20 . The non-transitory computer-readable storage media of claim 19 , the operations further comprising:

training a neural network configured to perform panoptic segmentation on input sensor samples on training data that includes a training example that associates the sensor sample with the panoptic segmentation label for the sensor data sample.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2024
From: MEI, JIERU; YAN, HANG; CHEN, LIANG-CHIEH; QIAO, SIYUAN; ZHU, YUKUN; ZHU, ALEX ZIHAO; YAN, XINCHEN; KRETZSCHMAR, HENRIK
To: WAYMO LLC
Reel/Frame 067347/0552 →
Continuity (2)
Provisional Application 63317534 · Mar 7, 2022
Related Publication 20230281824A1 · Sep 7, 2023
References Cited (95)
US 20180314921A1 · Mercep · 2018 [cited by examiner]
US 20190197778A1 · Sachdeva · 2019 [cited by examiner]
US 20210150230A1 · Smolyanskiy · 2021 [cited by examiner]
US 20210181758A1 · Das · 2021 [cited by examiner]
US 20220237799A1 · Price · 2022 [cited by examiner]
US 20220284666A1 · Chandler · 2022 [cited by examiner]
US 20230252638A1 · Hotson · 2023 [cited by examiner]
US 20230267615A1 · Agia · 2023 [cited by examiner]
Li, Peiliang, and Tong Qin. “Stereo vision-based semantic 3d object and ego-motion tracking for autonomous driving.” Proceedings of the European Conference on Computer Vision (ECCV). 2018. (Year: 2018). [cited by examiner]
Khan, Sohaib, and Mubarak Shah. “Consistent labeling of tracked objects in multiple cameras with overlapping fields of view.” IEEE transactions on pattern analysis and machine intelligence 25.10 (2003): 1355-1360. (Year… [cited by examiner]
Simon, Martin, et al. “Complexer-yolo: Real-time 3d object detection and tracking on semantic point clouds.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops. 2019. (Year: 2019… [cited by examiner]
Baque et al., “Deep occlusion reasoning for multi-camera multitarget detection,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 271-279. [cited by applicant]
Behley et al., “SemanticKITTI: A dataset for semantic scene understanding of lidar sequences,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 9297-9307. [cited by applicant]
Berclaz et al., “Multiple object tracking using k-shortest paths optimization,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Sep. 2011, 33(9):1806-1819. [cited by applicant]
Brostow et al., “Semantic object classes in video: A high definition ground truth database,” Pattern Recognition Letters, Jan. 15, 2009, 30(2):88-97. [cited by applicant]
Caesar et al., “nuScenes: A multimodal dataset for autonomous driving,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 11621-11631. [cited by applicant]
Chang et al., “Argoverse: 3D tracking and forecasting with rich maps,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 8748-8757. [cited by applicant]
Chavdarova et al., “Wildtrack: A multi-camera HD dataset for dense unscripted pedestrian detection,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 5030-5039. [cited by applicant]
Chen et al., “DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Apr. 1, 2018, 40(4):834-848. [cited by applicant]
Chen et al., “Encoder-decoder with atrous separable convolution for semantic image segmentation,” Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 801-818. [cited by applicant]
Chen et al., “Geosim: Realistic video simulation via geometry-aware composition for self-driving,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 7230-7240. [cited by applicant]
Chen et al., “Semantic image segmentation with deep convolutional nets and fully connected CRFs,” CoRR, Dec. 2014, arxiv.org/abs/1412.7062, 14 pages. [cited by applicant]
Cheng et al., “Panoptic-Deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 12475-12485. [cited by applicant]
Cheng et al., “Segow: Joint learning for video object segmentation and optical Flow,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 686-695. [cited by applicant]
Cordts et al., “The Cityscapes Dataset for Semantic Urban Scene Understanding,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 3213-3223. [cited by applicant]
Dehghan et al., “GMMCP Tracker: Globally optimal generalized maximum multi clique problem for multiple object tracking,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 40… [cited by applicant]
Dendorfer et al., “MOTChallenge: A Benchmark for Single-camera Multiple Target Tracking,” International Journal of Computer Vision, Dec. 23, 2020, 129:845-881. [cited by applicant]
Eigen et al., “Depth map prediction from a single image using a multi-scale deep network,” Advances in Neural Information Processing Systems 27, 2014, 9 pages. [cited by applicant]
Eshel et al., “Homography based multiple camera detection and tracking of people in a dense crowd,” 2008 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 23-28, 2008, 8 pages. [cited by applicant]
Ettinger et al., “Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 9710-9719. [cited by applicant]
Everingham et al., “The Pascal Visual Object Classes (VOC) Challenge,” IJCV, Sep. 9, 2009, 88:303-338. [cited by applicant]
Felzenszwalb et al., “Efficient graph-based image segmentation,” IJCV, Sep. 2004, 59:167-181. [cited by applicant]
Ferryman et al., “Pets2009: Dataset and Challenge,” 2009 Twelfth IEEE international workshop on performance evaluation of tracking and surveillance, Dec. 7-9, 2009, 6 pages. [cited by applicant]
Fleuret et al., “Multicamera people tracking with a probabilistic occupancy map,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Feb. 2008, 30(2):267-282. [cited by applicant]
Gao et al., “SSAP: Single-Shot Instance Segmentation With A nity Pyramid,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 642-651. [cited by applicant]
Geiger et al., “Are we ready for autonomous driving? the kitti vision benchmark suite,” 2012 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 16-21, 2012, pp. 3354-3361. [cited by applicant]
Han et al., “MMPTRACK: Large-scale densely annotated multi-camera multiple people tracking benchmark,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023, pp. 4860-4869. [cited by applicant]
Hariharan et al., “Simultaneous detection and segmentation,” Computer Vision—ECCV 2014, 2014, pp. 297-312. [cited by applicant]
He et al., “Deep residual learning for image recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778. [cited by applicant]
He et al., “Multiscale conditional random elds for image labeling,” Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Jun. 27, 2004, 8 pages. [cited by applicant]
Hofmann et al., “Hypergraphs for joint multi-view reconstruction and multi-object tracking,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2013, pp. 3650-3657. [cited by applicant]
Huang et al., “The ApolloScape open dataset for autonomous driving and its application,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Oct. 1, 2020, 42(10)2702-2719. [cited by applicant]
Kendall et al., “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7482-7491. [cited by applicant]
Kim et al., “Video Panoptic Segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 9859-9868. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization,” CoRR, Dec. 22, 2014, arxiv.org/abs/1412.6980, 15 pages. [cited by applicant]
Kirillov et al., “Panoptic Feature Pyramid Networks,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 6399-6408. [cited by applicant]
Kirillov et al., “Panoptic Segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 9404-9413. [cited by applicant]
Kretzschmar et al., “Scalability in Perception for autonomous driving: Waymo Open Dataset,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2446-2454. [cited by applicant]
Kuo et al., “Inter-camera association of multi-target tracks by on-line learned appearance a nity models,” Computer Vision—ECCV 2010, Sep. 2010, 6311:383-396. [cited by applicant]
Ladicky et al., “What, where and how many? combining object detectors and CRFs,” Computer Vision—ECCV 2010, Sep. 2010, 6314:424-437. [cited by applicant]
Li et al., “Attention-guided unifed network for panoptic segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 7026-7035. [cited by applicant]
Liang et al., “Polytransform: Deep polygon transformer for instance segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 9131-9140. [cited by applicant]
Liao et al., “Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2D and 3D,” CoRR, Sep. 28, 2021, arxiv.org/abs/2109.13410, 32 pages. [cited by applicant]
Lin et al., “Microsoft COCO: common objects in context,” Computer Vision—ECCV 2010, Sep. 2014, 8693:740-755. [cited by applicant]
Ling et al., “Variational amodal object completion,” Advances in Neural Information Processing Systems, 2020, 33:16246-57. [cited by applicant]
Liu et al., “An End-to-End Network for Panoptic Segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 6172-6181. [cited by applicant]
Long et al., “Fully convolutional networks for semantic segmentation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 3431-3440. [cited by applicant]
Luiten et al., “HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking,” IJCV, Oct. 8, 2020, 129:548-578. [cited by applicant]
Mallya et al., “World-consistent video-to-video synthesis,” Computer Vision—ECCV 2020, Nov. 2020, 12353:359-378. [cited by applicant]
Miao et al., “VSPW: A large-scale dataset for video scene parsing in the wild,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 4133-4143. [cited by applicant]
Neuhold et al., “The mapillary vistas dataset for semantic understanding of street scenes,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 4990-4999. [cited by applicant]
Papandreou et al., “PersonLab: Person pose estimation and instance segmentation with a bottom-up, part-based, geometric embedding model,” Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 269-2… [cited by applicant]
Philion et al., “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3D,” Computer Vision—ECCV 2020, Nov. 13, 2020, pp. 194-210. [cited by applicant]
Porzi et al., “Seamless Scene Segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 8277-8286. [cited by applicant]
Qi et al., “Offboard 3D object detection from point cloud sequences,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 6134-6144. [cited by applicant]
Qiao et al., “VIP-Deeplab: Learning visual perception with depth-aware video panoptic segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 3997-4008. [cited by applicant]
Ristani et al., “Features for multi-target multi-camera tracking and re-identification,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 6036-6046. [cited by applicant]
Ristani et al., “Performance measures and a data set for multi-target, multi-camera tracking,” Computer Vision—ECCV 2016 Workshops, Nov. 3, 2016, pp. 17-35. [cited by applicant]
Roddick et al., “Predicting semantic map representations from images using pyramid occupancy networks,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 11138-11147. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision, Apr. 11, 2015, 115:211-252. [cited by applicant]
Shi et al., “Normalized cuts and image segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Aug. 2000, 22(8):888-905. [cited by applicant]
Song et al., “Im2pano3d: Extrapolating 360 structure and semantics beyond the eld of view,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 3847-3856. [cited by applicant]
Tang et al., “CityFlow: A city-scale benchmark for multi-target multi-camera vehicle tracking and re-identification,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 8… [cited by applicant]
Tateno et al., “Distortion-aware convolutional filters for dense prediction in panoramic images,” Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 707-722. [cited by applicant]
Tu et al., “Image parsing: Unifying segmentation, detection, and recognition,” International Journal of Computer Vision, Feb. 1, 2005, 63:113-140. [cited by applicant]
Voigtlaender et al., “MOTS: Multi-object tracking and segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 7942-7951. [cited by applicant]
Wang et al., “Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation,” Computer Vision—ECCV 2020, Oct. 29, 2020, pp. 108-126. [cited by applicant]
Wang et al., “Pixel Consensus Voting for Panoptic Segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 9464-9473. [cited by applicant]
Weber et al., “DeepLab2: A TensorFlow Library for Deep Labeling,” CoRR, Jun. 17, 2021, arXiv: 2106.09748, 7 pages. [cited by applicant]
Weber et al., “Single-shot Panoptic Segmentation,” 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 24, 2020, pp. 8476-8483. [cited by applicant]
Weber et al., “STEP: Segmenting and tracking every pixel,” 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmark, 2021, 13 pages. [cited by applicant]
Wu et al., “Online object tracking: A benchmark,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2013, pp. 2411-2418. [cited by applicant]
Xiong et al., “UPSNet: A Unified Panoptic Segmentation Network,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 8818-8826. [cited by applicant]
Xu et al., “Multi-view people tracking via hierarchical trajectory composition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4256-4265. [cited by applicant]
Xu et al., “Streaming hierarchical video segmentation,” Computer Vision—ECCV 2012, Oct. 7-13, 2012, pp. 626-639. [cited by applicant]
Yang et al., “Auto4D: Learning to label 4D objects from sequential point clouds,” CORR, Jan. 17, 2021, arXiv:2101.06586, 8 pages. [cited by applicant]
Yang et al., “Capturing omni-range context for omnidirectional segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 1376-1386. [cited by applicant]
Yang et al., “DeeperLab: Single-Shot Image Parser,” CoRR, Feb. 13, 2019, arXiv:1902.05093, 20 pages. [cited by applicant]
Yang et al., “Video Instance Segmentation,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 5188-5197. [cited by applicant]
Yao et al., “Describing the scene as a whole: Joint object detection, scene classification and semantic segmentation,” 2012 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 16-21, 2012, pp. 702-709. [cited by applicant]
Yogamani et al., “Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving,” IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 9308-9318. [cited by applicant]
Yu et al., “BDD100k: A diverse driving dataset for heterogeneous multitask learning,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2636-2645. [cited by applicant]
Zakharov et al., “Autolabeling 3D objects with differentiable rendering of SDF shape priors,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 12224-12233. [cited by applicant]
Zamir et al., “GMCP-tracker: Global multi-object tracking using generalized minimum clique graphs,” Computer Vision—ECCV 2012, Oct. 7-13, 2012, pp. 343-356. [cited by applicant]
Zhang et al., “Orientation-aware semantic segmentation on icosahedron spheres,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 3533-3541. [cited by applicant]