IP Library Granted Patent US 12,437,412
Granted Patent B2
US 12,437,412 · App. 18/397,921 · Granted Oct 7, 2025

Deep neural network for segmentation of road scenes and animate object instances for autonomous driving applications

Inventors: Ke Chen (Sunnvale, CA); Nikolai Smolyanskiy (Seattle, WA); Alexey Kamenev (Bellevue, WA); Ryan Oldja (Redmond, WA); Tilman Wekel (Sunnyvale, CA); David Nister (Bellevue, WA); Joachim Pehserl (Lynnwood, WA); Ibrahim Eden (Redmond, WA); Sangmin Oh (San Jose, CA); Ruchi Bhargava (Redmond, WA)
Assignee: NVIDIA Corporation
G06T7/11G05D1/0088G05D1/81G06F18/22G06F18/23G06T5/50G06T7/10G06V10/82G06V20/56G06V20/58G06T2207/10028G06T2207/20084G06T2207/30252G06V10/454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,412
App. No.
18/397,921
Filed
Dec 27, 2023
Granted
Oct 7, 2025
Kind
B2
Art Unit
3662
USPC
701/28
Abstract

A deep neural network(s) (DNN) may be used to perform panoptic segmentation by performing pixel-level class and instance segmentation of a scene using a single pass of the DNN. Generally, one or more images and/or other sensor data may be stitched together, stacked, and/or combined, and fed into a DNN that includes a common trunk and several heads that predict different outputs. The DNN may include a class confidence head that predicts a confidence map representing pixels that belong to particular classes, an instance regression head that predicts object instance data for detected objects, an instance clustering head that predicts a confidence map of pixels that belong to particular instances, and/or a depth head that predicts range values. These outputs may be decoded to identify bounding shapes, class labels, instance labels, and/or range values for detected objects, and used to enable safe path planning and control of an autonomous vehicle.

Claims (60)

1. One or more processors comprising processing circuitry to:

generate, using a neural network and based at least on a representation of image data corresponding to an environment of an ego-object, one or more classifications of one or more pixels, the one or more classifications indicating associations of the one or more pixels with one or more unique instances corresponding to one or more respective channels of the neural network;

generate, based at least on the one or more classifications, one or more bounding shapes of the one or more unique instances of one or more detected objects in the environment; and

execute one or more operations of the ego-object based at least on the one or more bounding shapes.

2. The one or more processors of claim 1 , wherein the neural network comprises an instance clustering head comprising a respective classification channel for each of a plurality of detectable unique instances.

3. The one or more processors of claim 1 , wherein the one or more classifications comprise, for each channel of at least one of the one or more respective channels, a respective confidence map representing pixels that belong to a respective instance of the one or more unique instances.

4. The one or more processors of claim 1 , wherein the processing circuitry is further to generate the one or more bounding shapes using a connected components analysis to detect one or more boundaries of one or more clusters of connected or occluded unique instances represented by the one or more classifications.

5. The one or more processors of claim 1 , wherein the processing circuitry is further to identify a plurality of globally unique instances from each cluster of one or more clusters of connected or occluded unique instances detected from each classification channel of at least one of the one or more respective channels.

6. The one or more processors of claim 1 , wherein the processing circuitry is further to identify a plurality of globally unique instances based at least on: using a connected components analysis to assign a plurality of disconnected regions to a first instance, and determining that the plurality of disconnected regions correspond to the plurality of globally unique instances based at least on a minimum separation between the disconnected regions.

7. The one or more processors of claim 1 , wherein the one or more classifications comprise a depth-wise probability distribution per pixel representing a predicted likelihood, for each channel of a plurality of channels of the neural network, that each pixel of at least one of the one or more pixels belongs to a respective unique instance corresponding to the channel.

8. The one or more processors of claim 1 , wherein the processing circuitry is further to identify at least one unique instance of the one or more unique instances based at least on joining distinct connected regions of the one or more classifications, that are separated by less than a threshold gap, into a composite region representing the at least one unique instance.

9. The one or more processors of claim 1 , wherein using the neural network performs panoptic segmentation comprising class segmentation and instance regression in a single pass of the neural network.

10. The one or more processors of claim 1 , wherein the one or more processors are comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system for performing remote operations;

a system for performing real-time streaming;

a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content;

a system implemented using a robot;

a system for generating synthetic data;

a system for generating synthetic data using AI; or

a system implemented at least partially using cloud computing resources.

11. A system comprising one or more processors to:

generate, using a neural network and based at least on a representation of sensor data corresponding to an environment of an ego-object, one or more classifications associating one or more pixels with one or more unique instances corresponding to one or more respective channels of the neural network; and

execute one or more operations of the ego-object based at least on the one or more classifications.

12. The system of claim 11 , wherein the neural network comprises an instance clustering head comprising a respective classification channel for each of a plurality of detectable unique instances.

13. The system of claim 11 , wherein the one or more classifications comprise, for each channel of at least one of the one or more respective channels, a respective confidence map representing pixels that belong to a respective instance of the one or more unique instances.

14. The system of claim 11 , wherein the one or more processors are further to generate one or more bounding shapes of the one or more unique instances using a connected components analysis to detect one or more boundaries of one or more clusters of connected or occluded unique instances represented by the one or more classifications.

15. The system of claim 11 , wherein the one or more processors are further to identify a plurality of globally unique instances from each cluster of one or more clusters of connected or occluded unique instances detected from each classification channel of at least one of the one or more respective channels.

16. The system of claim 11 , wherein the one or more processors are further to identify a plurality of globally unique instances based at least on: using a connected components analysis to assign a plurality of disconnected regions to a first instance, and determining that the plurality of disconnected regions correspond to the plurality of globally unique instances based at least on a minimum separation between the disconnected regions.

17. The system of claim 11 , wherein the one or more classifications comprise a depth-wise probability distribution per pixel representing a predicted likelihood, for each channel of a plurality of channels of the neural network, that each pixel of at least one of the one or more pixels belongs to a respective unique instance corresponding to the channel.

18. The system of claim 11 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system for performing remote operations;

a system for performing real-time streaming;

a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content;

a system implemented using a robot;

a system for generating synthetic data;

a system for generating synthetic data using AI; or

a system implemented at least partially using cloud computing resources.

19. A method comprising:

generate, based at least on using a neural network to process a representation of sensor data corresponding to an environment of an ego-object, one or more classifications of one or more pixels into one or more unique instances corresponding to one or more respective channels of the neural network; and

execute one or more operations of the ego-object based at least on the one or more classifications.

20. The method of claim 19 , wherein the method is performed by at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system for performing remote operations;

a system for performing real-time streaming;

a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content;

a system implemented using a robot;

a system for generating synthetic data;

a system for generating synthetic data using AI; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2024
From: CHEN, KE; SMOLYANSKIY, NIKOLAI; KAMENEV, ALEXEY; OLDJA, RYAN; WEKEL, TILMAN; NISTER, DAVID; PEHSERL, JOACHIM; EDEN, IBRAHIM; OH, SANGMIN; BHARGAVA, RUCHI
To: NVIDIA CORPORATION
Reel/Frame 067019/0907 →
Continuity (3)
Continuation 16938706 · Jul 24, 2020
Provisional Application 62878659 · Jul 25, 2019
Related Publication 20250014186A1 · Jan 9, 2025
References Cited (122)
US 7035463B1 · Monobe · 2006 [cited by examiner]
US 9098754B1 · Stout et al. · 2015 [cited by applicant]
US 9286538B1 · Chen et al. · 2016 [cited by applicant]
US 10593042B1 · Douillard et al. · 2020 [cited by applicant]
US 10809361B2 · Vallespi-Gonzalez et al. · 2020 [cited by applicant]
US 10824862B2 · Qi et al. · 2020 [cited by applicant]
US 10825188B1 · Tan et al. · 2020 [cited by applicant]
US 10860034B1 · Ziyaee et al. · 2020 [cited by applicant]
US 10885698B2 · Muthler et al. · 2021 [cited by applicant]
US 10915793B2 · Corral-Soto et al. · 2021 [cited by applicant]
US 10970871B2 · Nezhadarya et al. · 2021 [cited by applicant]
US 11062454B1 · Cohen et al. · 2021 [cited by applicant]
US 11380108B1 · Cai et al. · 2022 [cited by applicant]
US 11532168B2 · Smolyanskiy et al. · 2022 [cited by applicant]
US 11704572B1 · Pronovost et al. · 2023 [cited by applicant]
US 11727601B2 · Marschner et al. · 2023 [cited by applicant]
US 11762094B2 · Laddha et al. · 2023 [cited by applicant]
US 11768292B2 · Liang et al. · 2023 [cited by applicant]
US 11885907B2 · Popov et al. · 2024 [cited by applicant]
US 20020147694A1 · Dempsey et al. · 2002 [cited by applicant]
US 20150036870A1 · Mundhenk et al. · 2015 [cited by applicant]
US 20160073080A1 · Wagner et al. · 2016 [cited by applicant]
US 20170023473A1 · Wegner et al. · 2017 [cited by applicant]
US 20170293837A1 · Cosatto et al. · 2017 [cited by applicant]
US 20170307735A1 · Rohani et al. · 2017 [cited by applicant]
US 20180101720A1 · Liu · 2018 [cited by applicant]
US 20180108134A1 · Venable et al. · 2018 [cited by applicant]
US 20180173971A1 · Jia et al. · 2018 [cited by applicant]
US 20180211403A1 · Hotson et al. · 2018 [cited by applicant]
US 20180247160A1 · Rohani et al. · 2018 [cited by applicant]
US 20180276845A1 · Bjorgvinsdottir et al. · 2018 [cited by applicant]
US 20180314253A1 · Mercep et al. · 2018 [cited by applicant]
US 20180349746A1 · Vallespi-Gonzalez · 2018 [cited by applicant]
US 20190026571A1 · Ryan · 2019 [cited by applicant]
US 20190026588A1 · Ryan · 2019 [cited by examiner]
US 20190026597A1 · Zeng et al. · 2019 [cited by applicant]
US 20190137287A1 · Pazhayampallil et al. · 2019 [cited by applicant]
US 20190145765A1 · Luo et al. · 2019 [cited by applicant]
US 20190147260A1 · May · 2019 [cited by applicant]
US 20190147331A1 · Arditi · 2019 [cited by applicant]
US 20190147610A1 · Frossard et al. · 2019 [cited by applicant]
US 20190258878A1 · Koivisto et al. · 2019 [cited by applicant]
US 20190279366A1 · Sick et al. · 2019 [cited by applicant]
US 20190286153A1 · Rankawat et al. · 2019 [cited by applicant]
US 20190324148A1 · Kim et al. · 2019 [cited by applicant]
US 20190361454A1 · Zeng et al. · 2019 [cited by applicant]
US 20200013219A1 · Dhua · 2020 [cited by examiner]
US 20200104584A1 · Zheng et al. · 2020 [cited by applicant]
US 20200174132A1 · Nezhadarya et al. · 2020 [cited by applicant]
US 20200175326A1 · Shen · 2020 [cited by examiner]
US 20200193606A1 · Douillard et al. · 2020 [cited by applicant]
US 20200210721A1 · Goel et al. · 2020 [cited by applicant]
US 20200301013A1 · Banerjee · 2020 [cited by examiner]
US 20210026355A1 · Chen et al. · 2021 [cited by applicant]
US 20210082181A1 · Shi et al. · 2021 [cited by applicant]
US 20210096241A1 · Bongio et al. · 2021 [cited by applicant]
US 20210109523A1 · Zou et al. · 2021 [cited by applicant]
US 20210146952A1 · Vora et al. · 2021 [cited by applicant]
US 20210149051A1 · Ding et al. · 2021 [cited by applicant]
US 20210166426A1 · Mccormac · 2021 [cited by examiner]
US 20210181758A1 · Das et al. · 2021 [cited by applicant]
US 20220327743A1 · Oh et al. · 2022 [cited by applicant]
CA 2934636A1 · 2017 [cited by applicant]
CN 106796718A · 2017 [cited by applicant]
CN 108171217A · 2018 [cited by applicant]
CN 108334081A · 2018 [cited by applicant]
CN 108596058A · 2018 [cited by applicant]
CN 109284764A · 2019 [cited by applicant]
CN 109291929A · 2019 [cited by applicant]
CN 109814130A · 2019 [cited by applicant]
CN 110032949A · 2019 [cited by applicant]
CN 110366710A · 2019 [cited by applicant]
EP 3525000A1 · 2019 [cited by applicant]
WO 2019178548A1 · 2019 [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/895,940, mailed on Mar. 20, 2024, 9 pages. [cited by applicant]
Object Detection and Classification by Decision-Level Fusion for Intelligent Vehicle Systems. Oh et al. (Year: 2016). [cited by applicant]
Office Action received for Chinese Patent Application No. 202011294650.3, mailed on Mar. 1, 2024, 24 pages (14 pages of Original OA and 10 pages of English Translation). [cited by applicant]
Office action received for Chinese Patent Application No. 202011297922.5, mailed on Mar. 16, 2024, 7 pages (2 pages English Translation and 5 pages of Original Copy). [cited by applicant]
Office Action received for European Application No. 20204403.8, mailed on Nov. 22, 2023, 8 pages. [cited by applicant]
Office Action received for European Application No. 20205868.1, mailed on Nov. 23, 2023, 9 pages. [cited by applicant]
Non-Final office action received for U.S. Appl. No. 17/377,064, mailed on Jan. 31, 2024, 9 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/377,053, mailed on Feb. 14, 2024, 11 pages. [cited by applicant]
Office Action Appendix received for U.S. Appl. No. 17/377,053, mailed on Jan. 8, 2024, 2 pages. [cited by applicant]
European Office Action dated Nov. 23, 2023 in Application No. 20205868.1, 9 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/493,452, Notification Date: Dec. 10, 2024, 51 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 17/377,064, Notification Date: Jun. 5, 2024, 10 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/976,581, Notification Date: May 31, 2024, 7 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/938,706, Notification Date: Jun. 12, 2024, 8 pages. [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Society of Automotive Engineers (SAE), Standard No. J3016-201609, pp. 1-30 (Sep. 30, 2016). [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Society of Automotive Engineers (SAE), Standard No. J3016-201806, pp. 1-35 (Jun. 15, 2018). [cited by applicant]
Chen, X., et al., “Multi-View 3D Object Detection Network for Autonomous Driving”, Cornell University Library, pp. 1-9 {Nov. 23, 2016). [cited by applicant]
Daan, D. G., et al., “Single Network Panoptic Segmentation for Street Scene Understanding”, 2019 IEEE Intelligent Vehicles Symposium (IV), Jun. 9, 2019, pp. 709-715. [cited by applicant]
Erhan, D., et al., “Scalable Object Detection using Deep Neural Networks”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, pp. 8 (2014). [cited by applicant]
Furukawa, H., “Deep learning for end-lo-end automatic target recognition from synthetic aperture radar imagery”, IEICE, pp. 35-40 (2018). [cited by applicant]
He, K., et al., “Deep residual learning for image recognition”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778 (2016). [cited by applicant]
IEC 61508, “Functional Safety of Electrical/Electronic/Programmable Electronic Safety-related Systems,” Retrieved from Internet URL: hllps://en.wikipedia.org/wiki/IEC_61508, accessed on Apr. 1, 2022, 7 pages. [cited by applicant]
ISO 26262, “Road vehicle—Functional safety,” International Standard for Functional Safety of Electronic System, Retrieved from Internet URL: hllps://en.wikipedia.org/wiki/ISO_26262, accessed on Sep. 13, 2021, 8 pages. [cited by applicant]
Jayakrishnan Unnikrishan, et al. “Resolving Elevation Ambiguity in 1-D Radar Array Measurements Using Deep Learning”, International Conference on Intelligent Robots and Systems, Macau, China, Nov. 4-8, 2019, 6 pages. [cited by applicant]
Kendall, A, et al., “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7482-7491 (2018). [cited by applicant]
Kendall, A., & Gal, Y. (2017). What uncertainties do we need in bayesian deep learning for computer vision?. In Advances in neural information processing systems (pp. 5574-5584). [cited by applicant]
Kirillov, A, et al., “Panoptic Segmentation”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1-10 (2019). [cited by applicant]
Krizhevsky, A, et al., “Imagenet classification with deep convolutional neural networks”, In Advances in Neural Information Processing Systems, pp. 1-9 (2012). [cited by applicant]
Ku, J., et al., “Joint 3D Proposal Generation and Object Detection from View Aggregation”, IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 1-8 {Oct. 2018). [cited by applicant]
Liu, H., et al., “An End-To-End Network for Panoptic Segmentation”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6165-6174 (2019). [cited by applicant]
Luo, W., et al., “Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net”, In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, … [cited by applicant]
Nezhadarya Ehsan et al: “BoxNet: A Deep Learning Method for 2D Bounding Box Estimation from Bird's-Eye View Point Cloud”, 2019 IEEE Intelligent Vehicles Symposium (IV), IEEE, Jun. 9, 2019, pp. 1557-1564, XP033606092, DO… [cited by applicant]
Office Action received for Chinese Patent Application No. 202011272919.8, mailed on Dec. 29, 2023, 24 pages (12 pages of English Translation and 12 pages of Office Action). [cited by applicant]
Qi, C.R., et al., “PointNet: Deep Learning on Point Sets for 30 Classification and Segmentation”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 652-660 (2017). [cited by applicant]
Ronneberger, 0., et al., “U-net: Convolutional networks for biomedical image segmentation”, In International Conference on Medical image computing and computer-assisted intervention, pp. 1-8 (2015). [cited by applicant]
Szegedy, C., et al., “Going Deeper with Convolutions”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1-9 (2015). [cited by applicant]
Wenquan, Z., et al., “LiSeg: Lightweight Road-object Semantic Segmentation in 3D LiDAR Scans For Autonomous Driving”, 2018 IEEE Intelligent Vehicles Symposium (IV), IEEE, Jun. 26, 2018, pp. 1021-1026. [cited by applicant]
Xiong, Y., et al., “UPSNet: A Unified Panoptic Segmentation Network”, IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8810-8818 (2019). [cited by applicant]
Zhou, Y., et al. “End-lo-End Multi-View Fusion for 3D Object Detection in LiDAR Point Clouds” Cornell University Library, pp. 1-10 {Oct. 2019). [cited by applicant]
Office Action received for Chinese Patent Application No. 202011272919.8, mailed on Jul. 16, 2024, 2024, 7 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/377,064, Notification Date: Aug. 7, 2024, 8 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 18/493,452, Notification Date: Feb. 25, 2025, 24 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/482,183, Notification Date: Mar. 13, 2025, 8 pages. [cited by applicant]
Castorena, Juan, and Siddharth Agarwal. “Ground-edge-based LIDAR localization without a reflectivity calibration for autonomous driving.” IEEE Robotics and Automation Letters 3.1 (2017): 344-351, 14 pages. [cited by applicant]
Lu, Weixin, et al. “L3-net: Towards learning based lidar localization for autonomous driving.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019, 10 pages. [cited by applicant]
Non-Final Office Action, European Application No. 20 206 733.6-1207, Notification Date: Feb. 14, 2025, 7pages. [cited by applicant]
Shen, Xiaotong, Seong-Woo Kim, and Marcelo H. Ang. “Spatio-temporal motion features for laser-based moving objects detection and tracking.” 2014 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE,… [cited by applicant]
Notice of Allowance, U.S. Appl. No. 18/493,452, Notification Date: Jun. 13, 2025, 13 pages. [cited by applicant]