IP Library › Granted Patent US 12,730,190
Granted Patent B2
US 12,730,190 · App. 18/531,103 · Granted Sep 8, 2026

Object detection and classification

Inventors: Tilman Wekel (Sunnyvale, CA); Sangmin Oh (San Jose, CA); David Nister (Bellevue, WA); Joachim Pehserl (Lynnwood, WA); Neda Cvijetic (East Palo Alto, CA); Ibrahim Eden (Redmond, WA)
Assignee: NVIDIA Corporation
G01S7/4802G01S7/481G01S17/894G01S17/931G06V10/764G06V10/80G06V10/82G06V20/58G01S7/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,190
App. No.
18/531,103
Granted
Sep 8, 2026
Kind
B2
Abstract

In various examples, a deep neural network (DNN) may be used to detect and classify animate objects and/or parts of an environment. The DNN may be trained using camera-to-LiDAR cross injection to generate reliable ground truth data for LiDAR range images. For example, annotations generated in the image domain may be propagated to the LiDAR domain to increase the accuracy of the ground truth data in the LiDAR domain—e.g., without requiring manual annotation in the LiDAR domain. Once trained, the DNN may output instance segmentation masks, class segmentation masks, and/or bounding shape proposals corresponding to two-dimensional (2D) LiDAR range images, and the outputs may be fused together to project the outputs into three-dimensional (3D) LiDAR point clouds. This 2D and/or 3D information output by the DNN may be provided to an autonomous vehicle drive stack to enable safe planning and control of the autonomous vehicle.

Claims (88)

1 . A method comprising:

processing, using a neural network, a converted representation of sensor data generated from an initial representation of the sensor data to determine:

one or more segmentation masks representing one or more portions of the converted representation of the sensor data that correspond to one or more object classes; and

one or more first locations within the converted representation of the sensor data of one or more first bounding shapes corresponding to one or more objects;

correlating the one or more first bounding shapes with the one or more segmentation masks;

generating, based at least on the correlating and using the one or more first bounding shapes, one or more second locations within the initial representation of the sensor data of one or more second bounding shapes corresponding to the one or more objects; and

performing one or more operations using a machine based at least on the one or more second locations of the one or more second bounding shapes.

2 . The method of claim 1 , further comprising:

obtaining the initial representation of the sensor data obtained using one or more sensors of the machine; and

generating, based at least on the sensor data, the converted representation of the sensor data.

3 . The method of claim 1 , wherein the generating the one or more second locations of the one or more second bounding shapes associated with the initial representation comprises:

projecting, based at least on the correlating, the one or more first locations of the one or more first bounding shapes from the converted representation to the initial representation; and

generating, based at least on the projecting, the one or more second locations of the one or more second bounding shapes.

4 . The method of claim 1 , further comprising:

processing, using the neural network, the converted representation of the sensor data to determine one or more second segmentation masks corresponding to one or more instances of the one or more objects in the converted representation,

wherein the generating the one or more second locations of the one or more second bounding shapes is further based at least on the one or more second segmentation masks.

5 . The method of claim 1 , wherein the converted representation of the sensor data represents information associated with one or more points, the information including at least one of:

one or more intensity values associated with the one or more points;

one or more elevation values associated with the one or more points; or

one or more distance values associated with the one or more points.

6 . The method of claim 1 , wherein:

the converted representation of the sensor data comprises a range image; and

the initial representation of the sensor data comprises a point cloud.

7 . The method of claim 1 , wherein:

the one or more first locations of the one or more first bounding shapes indicate one or more points of the converted representation that correspond to the one or more objects in the converted representation; and

the one or more segmentation masks indicate the one or more object classes associated with at least the one or more points.

8 . The method of claim 1 , further comprising:

generating, based at least on the one or more segmentation masks corresponding to the one or more object classes, one or more class labels associated with the one or more second bounding shapes,

wherein the performing the one or more operations is further based at least on the one or more class labels.

9 . The method of claim 1 , wherein the correlating at least a first bounding shape of the one or more first bounding shapes with a segmentation mask of the one or more segmentation masks comprises:

determining that one or more first points associated with the first bounding shape correspond to one or more second points associated with the segmentation mask; and

correlating the first bounding shape with the segmentation mask based at least on the one or more first points corresponding to the one or more second points.

10 . A system comprising:

one or more processors to:

process, using a neural network, a converted representation of sensor data generated from an initial representation of the sensor data to determine:

one or more segmentation masks corresponding to one or more objects represented by the converted representation of the sensor data, and

one or more first locations within the converted representation of the sensor data of one or more first bounding shapes corresponding to one or more objects;

correlate the one or more first bounding shapes with the one or more segmentation masks;

generate, based at least on the correlation and using the one or more first bounding shapes, one or more second locations within the initial representation of the sensor data of one or more second bounding shapes corresponding to the one or more objects; and

perform one or more operations using a machine based at least on the one or more second locations of the one or more second bounding shapes.

11 . The system of claim 10 , wherein the one or more processors are further to:

obtain the sensor data obtained using one or more sensors of the machine; and

generate, based at least on the sensor data, the converted representation of the sensor data.

12 . The system of claim 10 , wherein the generation of the one or more second locations of the one or more second bounding shapes associated with the initial representation comprises:

projecting, based at least on the correlation, the one or more first locations of the one or more first bounding shapes from the converted representation to the initial representation; and

generating, based at least on the projecting, the one or more second locations of the one or more second bounding shapes.

13 . The system of claim 10 , wherein the one or more segmentation masks comprise at least one of:

one or more instance segmentation masks indicating whether one or more points correspond to the one or more objects; or

one or more semantic segmentation masks indicating one or more object classes associated with the one or more points.

14 . The system of claim 10 , wherein the converted representation of the sensor data represents information associated with one or more points, the information including at least one of:

one or more intensity values associated with the one or more points;

one or more elevation values associated with the one or more points; or

one or more distance values associated with the one or more points.

15 . The system of claim 10 , wherein:

the converted representation of the sensor data comprises a range image; and

the initial representation of the sensor data comprises a point cloud.

16 . The system of claim 10 , wherein the one or more processors are further to:

generate, based at least on the correlation, one or more class labels associated with the one or more second bounding shapes,

wherein the one or more operations are further to be performed based at least on the one or more class labels.

17 . The system of claim 10 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

18 . One or more processors comprising processing circuitry to:

correlate one or more first locations within a projection image of one or more first bounding shapes associated with one or more objects with one or more segmentation masks associated with the one or more objects as represented in the projection image;

project, based at least on the correlation, the one or more first locations of the one or more first bounding shapes from the projection image to one or more second locations of one or more second bounding shapes within a point cloud; and

cause a machine to perform one or more operations based at least on the one or more second locations of the one or more second bounding shapes.

19 . The one or more processors of claim 18 , wherein:

the one or more first locations of the one or more first bounding shapes indicate at least one or more first points of the projection image that correspond to the one or more objects; and

the one or more segmentation masks indicate at least one or more second points of the projection image that correspond to the one or more objects.

20 . The one or more processors of claim 18 , wherein the one or more processors is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2024
From: WEKEL, TILMAN; OH, SANGMIN; NISTER, DAVID; PEHSERL, JOACHIM; CVIJETIC, NEDA; EDEN, IBRAHIM
To: NVIDIA CORPORATION
Reel/Frame 066008/0979 →
Continuity (3)
Continuation 17005788 · Aug 28, 2020
Provisional Application 62893814 · Aug 30, 2019
Related Publication 20240111025A1 · Apr 4, 2024
References Cited (30)
US 9098754B1 · Stout · 2015 [cited by examiner]
US 10885698B2 · Stout et al. · 2021 [cited by applicant]
US 11906660B2 · Wekel et al. · 2024 [cited by applicant]
US 20180161986A1 · Kee et al. · 2018 [cited by applicant]
US 20180348346A1 · Vallespi-Gonzalez · 2018 [cited by examiner]
US 20190057507A1 · El-Khamy · 2019 [cited by examiner]
He et al, Mask R-CNN, ICCV (Year: 2017). [cited by examiner]
Vaquero et al, Deconvolutional Networks for Point-Cloud Vehicle Detection and Tracking in Driving Scenarios, European Conference on Mobile Robots (ECMR) (Year: 2017). [cited by examiner]
IEC 61508, “Functional Safety of Electrical/Electronic/Programmable Electronic Safety-related Systems,” https://en.wikipedia.org/wiki/IEC_61508, accessed on Apr. 1, 2022, 7 pgs. [cited by applicant]
ISO 26262, “Road vehicle—Functional safety,” International standard for functional safety of electronic system, https://en.wikipedia.org/wiki/ISO_26262, accessed on Sep. 13, 2021, 8 pgs. [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Society of Automotive Engineers (SAE), Standard No. J3016-201609, pp. 30 (Sep. 30, 2016). [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Society of Automotive Engineers (SAE), Standard No. J3016-201806, pp. 35 (Jun. 15, 2018). [cited by applicant]
Wekel, Tilman; Requirement for Restriction/Election for U.S. Appl. No. 17/005,788, filed Aug. 28, 2020, mailed Mar. 21, 2023, 6 pgs. [cited by applicant]
Wekel, et al.; Non-Final Office Action for U.S. Appl. No. 17/005,788, filed Aug. 28, 2020, mailed Jun. 29, 2023, 28 pgs. [cited by applicant]
Wikipedia, Image Segmentation, 2023; 22 pgs. [cited by applicant]
Wekel, Tilman; Final Office Action for U.S. Appl. No. 17/005,788, filed Aug. 28, 2020, mailed Sep. 20, 2023, 16 pgs. [cited by applicant]
Wekel, Tilman; Notice of Allowance for U.S. Appl. No. 17/005,788, filed Aug. 28, 2020; mailed Nov. 7, 2023, 13 pgs. [cited by applicant]
Wekel, Tilman; Invitation to pay additional fees for PCT Application No. PCT/US2020/048466, filed Aug. 28, 2020, mailed Nov. 27, 2020, 11 pgs. [cited by applicant]
Wekel, Tilman; Preliminary Report on Patentability for PCT Application No. PCT/US2020/048466, filed Aug. 28, 2020, mailed Mar. 10, 2022, 12 pgs. [cited by applicant]
Radi, et al.; “VolMap: A Real-time Model for Semantic Segmentation of a LiDAR 360 surrounding view”, Proceedings of the 36th International Conference on Machine Learning, Long Beach CA, PMLR 97, arXiv: 1906.11873v1 [cs.… [cited by applicant]
Wekel, Tilman; International Search Report and Written Opinion for PCT Application No. PCT/US2020/049466, Filed Aug. 28, 2020, mailed Feb. 1, 2021, 16 pgs. [cited by applicant]
Luo, et al.; (2018). Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition… [cited by applicant]
Qi, et al.; (2017). Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 652-660). [cited by applicant]
Kendall, et al. (2017). What uncertainties do we need in bayesian deep learning for computer vision?. In Advances in neural information processing systems (pp. 5574-5584). [cited by applicant]
Furukawa, H. (2018). Deep learning for end-to-end automatic target recognition from synthetic aperture radar imagery. arXiv preprint arXiv:1801.08558. [cited by applicant]
Kendall, et al. (2018). Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7482-7491). [cited by applicant]
Ronneberger, et al. (Oct. 2015). U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention (pp. 234-241). Springer, Cham. [cited by applicant]
Szegedy, et al. (2015). Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1-9). [cited by applicant]
He, et al.; (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778). [cited by applicant]
Krizhevsky, et al.; (2012). Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems (pp. 1097-1105). [cited by applicant]