IP Library Granted Patent US 12,651,465
Granted Patent B2
US 12,651,465 · App. 18/647,415 · Granted Jun 9, 2026

Multi-view deep neural network for LiDAR perception

Inventors: Nikolai Smolyanskiy (Seattle, WA); Ryan Oldja (Redmond, WA); Ke Chen (Sunnyvale, CA); Alexander Popov (Kirkland, WA); Joachim Pehserl (Lynnwood, WA); Ibrahim Eden (Redmond, WA); Tilman Wekel (Sunnyvale, CA); David Wehr (Redmond, WA); Ruchi Bhargava (Redmond, WA); David Nister (Bellevue, WA)
Assignee: NVIDIA Corporation
G06V20/584B60W60/0011B60W60/0016B60W60/0027G01S7/4802G01S17/89G01S17/931G05D1/0088G05D1/81G06N3/045G06T19/006G06V10/25G06V10/26G06V10/454G06V10/764G06V10/774G06V10/803G06V10/82G06V20/56G06V20/58B60W2420/403G06T2207/10028G06T2207/20081G06T2207/20084G06T2207/30261G06V10/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,465
App. No.
18/647,415
Filed
Apr 26, 2024
Granted
Jun 9, 2026
Kind
B2
Art Unit
2661
USPC
382/104
Abstract

A deep neural network(s) (DNN) may be used to detect objects from sensor data of a three dimensional (3D) environment. For example, a multi-view perception DNN may include multiple constituent DNNs or stages chained together that sequentially process different views of the 3D environment. An example DNN may include a first stage that performs class segmentation in a first view (e.g., perspective view) and a second stage that performs class segmentation and/or regresses instance geometry in a second view (e.g., top-down). The DNN outputs may be processed to generate 2D and/or 3D bounding boxes and class labels for detected objects in the 3D environment. As such, the techniques described herein may be used to detect and classify animate objects and/or parts of an environment, and these detections and classifications may be provided to an autonomous vehicle drive stack to enable safe planning and control of the autonomous vehicle.

Claims (53)

1 . One or more processors comprising one or more circuits to:

receive initial classification data generated using a first representation of sensor data and representing one or more classifications;

generate refined classification data based at least on a neural network processing a projected representation of the initial classification data and a second representation of the sensor data different from the first representation of the sensor data; and

cause performance of one or more operations corresponding to a machine based at least on the refined classification data.

2 . The one or more processors of claim 1 , wherein the one or more circuits are further to generate the projected representation of the initial classification data based at least on projecting one or more detected three-dimensional locations labeled with one or more classifications extracted based at least on a range image corresponding to the first representation of the sensor data.

3 . The one or more processors of claim 1 , wherein the one or more circuits are further to generate the projected representation of the initial classification data based at least on projecting a representation of labeled object geometry generated based at least on associating the first representation of the sensor data with the initial classification data.

4 . The one or more processors of claim 1 , wherein the one or more circuits are further to generate the projected representation of the initial classification data based at least on transforming a semantically labeled range image associating the first representation of the sensor data with the initial classification data.

5 . The one or more processors of claim 1 , wherein the one or more circuits are further to generate the refined classification data based at least on the neural network processing one or more first channels comprising one or more height values represented by the second representation of the sensor data and one or more second channels comprising one or more view transformed classifications represented by the projected representation of the initial classification data.

6 . The one or more processors of claim 1 , wherein the one or more circuits are further to generate the refined classification data based at least on the neural network processing at least one of one or more view transformed confidence maps or one or more view transformed segmentation masks corresponding to the projected representation of the initial classification data.

7 . The one or more processors of claim 1 , wherein the first representation of the sensor data comprises one or more range image images corresponding to a point cloud, and the second representation of the sensor data comprises one or more height maps corresponding to the point cloud.

8 . The one or more processors of claim 1 , wherein the one or more processors are comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

9 . A method comprising:

receiving initial classification data generated using sensor data and representing one or more classifications;

generating refined classification data based at least on a neural network processing a projected representation of one or more labeled points that associate the sensor data with the initial classification data; and

causing performance of one or more operations corresponding to a machine based at least on the refined classification data.

10 . The method of claim 9 , further comprising generating the one or more labeled points based at least on labeling one or more detected three-dimensional locations with one or more classifications extracted based at least on a range image corresponding to the sensor data.

11 . The method of claim 9 , further comprising generating the one or more labeled points based at least on associating one or more per-pixel classifications corresponding to the initial classification data with one or more corresponding pixels of a range image corresponding to the sensor data.

12 . The method of claim 9 , further comprising generating the projected representation of the one or more labeled points based at least on transforming a semantically labeled range image associating the sensor data with the initial classification data.

13 . The method of claim 9 , further comprising generating the refined classification data based at least on the neural network processing one or more first channels comprising one or more height values represented by the sensor data and one or more second channels comprising one or more view transformed classifications represented by the projected representation of the one or more labeled points.

14 . The method of claim 9 , further comprising generating the refined classification data based at least on the neural network processing at least one of one or more view transformed confidence maps or one or more view transformed segmentation masks corresponding to the initial classification data.

15 . The method of claim 9 , the initial classification data generated using a first representation of the sensor data comprising one or more range images corresponding to a point cloud, further comprising generating the refined classification data based at least on the neural network further processing a second representation of the sensor data comprising one or more height maps corresponding to the point cloud.

16 . The method of claim 9 , wherein the method is performed by at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

17 . A system comprising one or more processing units to cause performance of one or more operations corresponding to a machine based at least on refined classification data, the refined classification data generated based at least on a neural network processing initial classification data and a first representation of sensor data, the initial classification data generated using a second representation of the sensor data different from the first representation of the sensor data.

18 . The system of claim 17 , wherein the one or more processing units are further to generate the refined classification data based at least on the neural network processing a projected representation of the initial classification data.

19 . The system of claim 17 , wherein the one or more processing units are further to generate the refined classification data based at least on the neural network processing a projected representation of the initial classification data generated based at least on projecting one or more detected three-dimensional locations labeled with the initial classification data.

20 . The system of claim 17 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2024
From: OLDJA, RYAN; CHEN, KE; POPOV, ALEXANDER; PEHSERL, JOACHIM; EDEN, IBRAHIM; WEKEL, TILMAN; WEHR, DAVID; BHARGAVA, RUCHI; NISTER, DAVID; SMOLYANSKIY, NIKOLAI
To: NVIDIA CORPORATION
Reel/Frame 068166/0832 →
Continuity (7)
Continuation 17895940 · Aug 25, 2022
Continuation 16915346 · Jun 29, 2020
Continuation 16836583 · Mar 31, 2020
Continuation 16836618 · Mar 31, 2020
Provisional Application 62938852 · Nov 21, 2019
Provisional Application 62936080 · Nov 15, 2019
Related Publication 20240273919A1 · Aug 15, 2024
References Cited (128)
US 7035463B1 · Monobe et al. · 2006 [cited by applicant]
US 9098754B1 · Stout et al. · 2015 [cited by applicant]
US 9286538B1 · Chen et al. · 2016 [cited by applicant]
US 10593042B1 · Douillard et al. · 2020 [cited by applicant]
US 10809361B2 · Vallespi-Gonzalez et al. · 2020 [cited by applicant]
US 10824862B2 · Qi et al. · 2020 [cited by applicant]
US 10825188B1 · Tan et al. · 2020 [cited by applicant]
US 10860034B1 · Ziyaee et al. · 2020 [cited by applicant]
US 10884409B2 · Mercep · 2021 [cited by examiner]
US 10885698B2 · Muthler et al. · 2021 [cited by applicant]
US 10915793B2 · Corral-Soto et al. · 2021 [cited by applicant]
US 10921817B1 · Kangaspunta · 2021 [cited by examiner]
US 10970871B2 · Nezhadarya et al. · 2021 [cited by applicant]
US 11062454B1 · Cohen et al. · 2021 [cited by applicant]
US 11380108B1 · Cai et al. · 2022 [cited by applicant]
US 11532168B2 · Smolyanskiy et al. · 2022 [cited by applicant]
US 11704572B1 · Pronovost et al. · 2023 [cited by applicant]
US 11727601B2 · Marschner et al. · 2023 [cited by applicant]
US 11762094B2 · Laddha et al. · 2023 [cited by applicant]
US 11768292B2 · Liang et al. · 2023 [cited by applicant]
US 11885907B2 · Popov et al. · 2024 [cited by applicant]
US 20020147694A1 · Dempsey et al. · 2002 [cited by applicant]
US 20150036870A1 · Mundhenk et al. · 2015 [cited by applicant]
US 20160073080A1 · Wagner et al. · 2016 [cited by applicant]
US 20170023473A1 · Wegner et al. · 2017 [cited by applicant]
US 20170293837A1 · Cosatto et al. · 2017 [cited by applicant]
US 20170307735A1 · Rohani et al. · 2017 [cited by applicant]
US 20180074506A1 · Branson · 2018 [cited by examiner]
US 20180101720A1 · Liu · 2018 [cited by applicant]
US 20180108134A1 · Venable et al. · 2018 [cited by applicant]
US 20180173971A1 · Jia et al. · 2018 [cited by applicant]
US 20180211403A1 · Hotson et al. · 2018 [cited by applicant]
US 20180247160A1 · Rohani et al. · 2018 [cited by applicant]
US 20180276845A1 · Bjorgvinsdottir et al. · 2018 [cited by applicant]
US 20180314253A1 · Mercep et al. · 2018 [cited by applicant]
US 20180314921A1 · Mercep · 2018 [cited by examiner]
US 20180349746A1 · Vallespi-Gonzalez · 2018 [cited by applicant]
US 20190026571A1 · Ryan · 2019 [cited by applicant]
US 20190026588A1 · Ryan · 2019 [cited by examiner]
US 20190026597A1 · Zeng et al. · 2019 [cited by applicant]
US 20190137287A1 · Pazhayampallil et al. · 2019 [cited by applicant]
US 20190145765A1 · Luo et al. · 2019 [cited by applicant]
US 20190147253A1 · Bai · 2019 [cited by examiner]
US 20190147254A1 · Bai · 2019 [cited by examiner]
US 20190147255A1 · Homayounfar · 2019 [cited by examiner]
US 20190147260A1 · May · 2019 [cited by applicant]
US 20190147331A1 · Arditi · 2019 [cited by applicant]
US 20190147610A1 · Frossard et al. · 2019 [cited by applicant]
US 20190220013A1 · Bradley · 2019 [cited by examiner]
US 20190258878A1 · Koivisto et al. · 2019 [cited by applicant]
US 20190279366A1 · Sick et al. · 2019 [cited by applicant]
US 20190286153A1 · Rankawat et al. · 2019 [cited by applicant]
US 20190324148A1 · Kim et al. · 2019 [cited by applicant]
US 20190361454A1 · Zeng et al. · 2019 [cited by applicant]
US 20200013219A1 · Dhua et al. · 2020 [cited by applicant]
US 20200104584A1 · Zheng et al. · 2020 [cited by applicant]
US 20200174132A1 · Nezhadarya et al. · 2020 [cited by applicant]
US 20200175326A1 · Shen et al. · 2020 [cited by applicant]
US 20200193606A1 · Douillard et al. · 2020 [cited by applicant]
US 20200210721A1 · Goel et al. · 2020 [cited by applicant]
US 20200272148A1 · Karasev · 2020 [cited by examiner]
US 20200301013A1 · Banerjee · 2020 [cited by examiner]
US 20210026355A1 · Chen et al. · 2021 [cited by applicant]
US 20210082181A1 · Shi et al. · 2021 [cited by applicant]
US 20210096241A1 · Bongio Karrman et al. · 2021 [cited by applicant]
US 20210109523A1 · Zou · 2021 [cited by examiner]
US 20210146952A1 · Vora et al. · 2021 [cited by applicant]
US 20210149051A1 · Ding et al. · 2021 [cited by applicant]
US 20210166426A1 · Mccormac et al. · 2021 [cited by applicant]
US 20210181758A1 · Das et al. · 2021 [cited by applicant]
US 20220327743A1 · Oh et al. · 2022 [cited by applicant]
CA 2934636A1 · 2017 [cited by applicant]
CN 106796718A · 2017 [cited by applicant]
CN 108171217A · 2018 [cited by applicant]
CN 108334081A · 2018 [cited by applicant]
CN 108596058A · 2018 [cited by applicant]
CN 109284764A · 2019 [cited by applicant]
CN 109291929A · 2019 [cited by applicant]
CN 109814130A · 2019 [cited by applicant]
CN 110032949A · 2019 [cited by applicant]
CN 110366710A · 2019 [cited by applicant]
WO 2019178548A1 · 2019 [cited by applicant]
Notice of Allowance, U.S. Appl. No. 18/482,183, Notification Date: Jul. 25, 2025, 8 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 18/493,452, Notification Date: Feb. 25, 2025, 24 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/397,921, Notification Date: Mar. 13, 2025, 17 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/482,183, Notification Date: Mar. 13, 2025, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 18/397,921, Notification Date: Aug. 12, 2025, 9 pages. [cited by applicant]
Office Action received for Chinese Patent Application No. 202011272919.8, mailed on Jul. 16, 2024, 2024, 7 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/377,064, Notification Date: Aug. 7, 2024, 8 pages. [cited by applicant]
Castorena, Juan, and Siddharth Agarwal. “Ground-edge-based LIDAR localization without a reflectivity calibration for autonomous driving.” IEEE Robotics and Automation Letters 3.1 (2017): 344-351, 14 pages. [cited by applicant]
Lu, Weixin, et al. “L3-net: Towards learning based lidar localization for autonomous driving.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019, 10 pages. [cited by applicant]
Non-Final Office Action, European Application No. 20 206 733.6-1207, Notification Date: Feb. 14, 2025, 7pages. [cited by applicant]
Shen, Xiaotong, Seong-Woo Kim, and Marcelo H. Ang. “Spatio-temporal motion features for laser-based moving objects detection and tracking.” 2014 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE,… [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Society of Automotive Engineers (SAE), Standard No. J3016-201609, pp. 1-30 (Sep. 30, 2016). [cited by applicant]
Chen, X., et al., “Multi-View 3D Object Detection Network for Autonomous Driving”, Cornell University Library, pp. 1-9 {Nov. 23, 2016). [cited by applicant]
Erhan, D., et al., “Scalable Object Detection using Deep Neural Networks”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, pp. 8 (2014). [cited by applicant]
European Office Action dated Nov. 23, 2023 in Application No. 20205868.1, 9 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 17/377,064, Notification Date: Jun. 5, 2024, 10 pages. [cited by applicant]
Furukawa, H., “Deep learning for end-lo-end automatic target recognition from synthetic aperture radar imagery”, IEICE, pp. 35-40 (2018). [cited by applicant]
Geus. D. D., et al., “Single Network Panoptic Segmentation for Street Scene Understanding”, 2019 IEEE Intelligent Vehicles Symposium (IV), Jun. 9, 2019, pp. 709-715. [cited by applicant]
He, K., et al., “Deep residual learning for image recognition”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778 (2016). [cited by applicant]
IEC 61508, “Functional Safety of Electrical/Electronic/Programmable Electronic Safety-related Systems,” Retrieved from Internet URL: hllps://en.wikipedia.org/wiki/IEC_61508, accessed on Apr. 1, 2022, 7 pages. [cited by applicant]
ISO 26262, “Road vehicle—Functional safety,” International Standard for Functional Safety of Electronic System, Retrieved from Internet URL: hllps://en.wikipedia.org/wiki/ISO_26262, accessed on Sep. 13, 2021, 8 pages. [cited by applicant]
Jayakrishnan Unnikrishan, et al. “Resolving Elevation Ambiguity in 1-D Radar Array Measurements Using Deep Learning”, International Conference on Intelligent Robots and Systems, Macau, China, Nov. 4-8, 2019, 6 pages. [cited by applicant]
Kendall, A, et al., “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7482-7491 (2018). [cited by applicant]
Kendall, A., et al. “What uncertainties do we need in bayesian deep learning for computer vision?”, In Advances in neural information processing systems pp. 1-11 (2017). [cited by applicant]
Kirillov, A, et al., “Panoptic Segmentation”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1-10 (2019). [cited by applicant]
Krizhevsky, A, et al., “Imagenet classification with deep convolutional neural networks”, In Advances in Neural Information Processing Systems, pp. 1-9 (2012). [cited by applicant]
Ku, J., et al., “Joint 3D Proposal Generation and Object Detection from View Aggregation”, IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 1-8 {Oct. 2018). [cited by applicant]
Liu, H., et al., “An End-To-End Network for Panoptic Segmentation”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6165-6174 (2019). [cited by applicant]
Luo, W., et al., “Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net”, In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, … [cited by applicant]
Nezhadarya Ehsan et al: “BoxNet: A Deep Learning Method for 2D Bounding Box Estimation from Bird's-Eye View Point Cloud”, 2019 IEEE Intelligent Vehiclessymposium (IV), IEEE, Jun. 9, 2019, pp. 1557-1564, XP033606092, DOI… [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/976,581, Notification Date: May 31, 2024, 7 pages. [cited by applicant]
Object Detection and Classification by Decision-Level Fusion for Intelligent Vehicle Systems. Oh et al. (Year: 2016). [cited by applicant]
Office Action received for Chinese Patent Application No. 202011272919.8, mailed on Dec. 29, 2023, 24 pages (12 pages of English Translation and 12 pages of Office Action). [cited by applicant]
Office Action received for Chinese Patent Application No. 202011294650.3, mailed on Mar. 1, 2024, 24 pages (14 pages of Original OA and 10 pages of English Translation). [cited by applicant]
Office action received for Chinese Patent Application No. 202011297922.5, mailed on Mar. 16, 2024, 7 pages (2 pages English Translation and 5 pages of Original Copy). [cited by applicant]
Office Action received for European Application No. 20204403.8, mailed on Nov. 22, 2023, 8 pages. [cited by applicant]
Qi, C.R., et al., “PointNet: Deep Learning on Point Sets for 30 Classification and Segmentation”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 652-660 (2017). [cited by applicant]
Ronneberger, 0., et al., “U-net: Convolutional networks for biomedical image segmentation”, In International Conference on Medical image computing and computer-assisted intervention, pp. 1-8 (2015). [cited by applicant]
Szegedy, C., et al., “Going Deeper with Convolutions”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1-9 (2015). [cited by applicant]
Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles, Society of Automotive Engineers (SAE), Standard No. J3016-201806, pp. 1-35 (Jun. 15, 2018). [cited by applicant]
Xiong, Y., et al., “UPSNet: A Unified Panoptic Segmentation Network”, IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8810-8818(2019). [cited by applicant]
Wenquan, Z., et al., ““LiSeg: LightweightRoad-object Semantic Segmentation In 3DLiDAR Scans for Autonomous Driving””,2018 IEEE Intelligent Vehicles Symposium(IV), IEEE, Jun. 26, 2018, pp. 1021-1026. [cited by applicant]
Zhou, Y., et al. “End-lo-End Multi-View Fusion for 3D Object Detection in LiDAR Point Clouds” Cornell University Library, pp. 1-10 {Oct. 2019). [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/938,706, Notification Date: Jun. 12, 2024, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 18/493,452, Notification Date: Jun. 13, 2025, 13 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/493,452, Notification Date: Dec. 10, 2024, 51 pages. [cited by applicant]