Object detection and classification
In various examples, a deep neural network (DNN) may be used to detect and classify animate objects and/or parts of an environment. The DNN may be trained using camera-to-LiDAR cross injection to generate reliable ground truth data for LiDAR range images. For example, annotations generated in the image domain may be propagated to the LiDAR domain to increase the accuracy of the ground truth data in the LiDAR domain—e.g., without requiring manual annotation in the LiDAR domain. Once trained, the DNN may output instance segmentation masks, class segmentation masks, and/or bounding shape proposals corresponding to two-dimensional (2D) LiDAR range images, and the outputs may be fused together to project the outputs into three-dimensional (3D) LiDAR point clouds. This 2D and/or 3D information output by the DNN may be provided to an autonomous vehicle drive stack to enable safe planning and control of the autonomous vehicle.
1 . A method comprising:
processing, using a neural network, a converted representation of sensor data generated from an initial representation of the sensor data to determine:
one or more segmentation masks representing one or more portions of the converted representation of the sensor data that correspond to one or more object classes; and
one or more first locations within the converted representation of the sensor data of one or more first bounding shapes corresponding to one or more objects;
correlating the one or more first bounding shapes with the one or more segmentation masks;
generating, based at least on the correlating and using the one or more first bounding shapes, one or more second locations within the initial representation of the sensor data of one or more second bounding shapes corresponding to the one or more objects; and
performing one or more operations using a machine based at least on the one or more second locations of the one or more second bounding shapes.
2 . The method of claim 1 , further comprising:
obtaining the initial representation of the sensor data obtained using one or more sensors of the machine; and
generating, based at least on the sensor data, the converted representation of the sensor data.
3 . The method of claim 1 , wherein the generating the one or more second locations of the one or more second bounding shapes associated with the initial representation comprises:
projecting, based at least on the correlating, the one or more first locations of the one or more first bounding shapes from the converted representation to the initial representation; and
generating, based at least on the projecting, the one or more second locations of the one or more second bounding shapes.
4 . The method of claim 1 , further comprising:
processing, using the neural network, the converted representation of the sensor data to determine one or more second segmentation masks corresponding to one or more instances of the one or more objects in the converted representation,
wherein the generating the one or more second locations of the one or more second bounding shapes is further based at least on the one or more second segmentation masks.
5 . The method of claim 1 , wherein the converted representation of the sensor data represents information associated with one or more points, the information including at least one of:
one or more intensity values associated with the one or more points;
one or more elevation values associated with the one or more points; or
one or more distance values associated with the one or more points.
6 . The method of claim 1 , wherein:
the converted representation of the sensor data comprises a range image; and
the initial representation of the sensor data comprises a point cloud.
7 . The method of claim 1 , wherein:
the one or more first locations of the one or more first bounding shapes indicate one or more points of the converted representation that correspond to the one or more objects in the converted representation; and
the one or more segmentation masks indicate the one or more object classes associated with at least the one or more points.
8 . The method of claim 1 , further comprising:
generating, based at least on the one or more segmentation masks corresponding to the one or more object classes, one or more class labels associated with the one or more second bounding shapes,
wherein the performing the one or more operations is further based at least on the one or more class labels.
9 . The method of claim 1 , wherein the correlating at least a first bounding shape of the one or more first bounding shapes with a segmentation mask of the one or more segmentation masks comprises:
determining that one or more first points associated with the first bounding shape correspond to one or more second points associated with the segmentation mask; and
correlating the first bounding shape with the segmentation mask based at least on the one or more first points corresponding to the one or more second points.
10 . A system comprising:
one or more processors to:
process, using a neural network, a converted representation of sensor data generated from an initial representation of the sensor data to determine:
one or more segmentation masks corresponding to one or more objects represented by the converted representation of the sensor data, and
one or more first locations within the converted representation of the sensor data of one or more first bounding shapes corresponding to one or more objects;
correlate the one or more first bounding shapes with the one or more segmentation masks;
generate, based at least on the correlation and using the one or more first bounding shapes, one or more second locations within the initial representation of the sensor data of one or more second bounding shapes corresponding to the one or more objects; and
perform one or more operations using a machine based at least on the one or more second locations of the one or more second bounding shapes.
11 . The system of claim 10 , wherein the one or more processors are further to:
obtain the sensor data obtained using one or more sensors of the machine; and
generate, based at least on the sensor data, the converted representation of the sensor data.
12 . The system of claim 10 , wherein the generation of the one or more second locations of the one or more second bounding shapes associated with the initial representation comprises:
projecting, based at least on the correlation, the one or more first locations of the one or more first bounding shapes from the converted representation to the initial representation; and
generating, based at least on the projecting, the one or more second locations of the one or more second bounding shapes.
13 . The system of claim 10 , wherein the one or more segmentation masks comprise at least one of:
one or more instance segmentation masks indicating whether one or more points correspond to the one or more objects; or
one or more semantic segmentation masks indicating one or more object classes associated with the one or more points.
14 . The system of claim 10 , wherein the converted representation of the sensor data represents information associated with one or more points, the information including at least one of:
one or more intensity values associated with the one or more points;
one or more elevation values associated with the one or more points; or
one or more distance values associated with the one or more points.
15 . The system of claim 10 , wherein:
the converted representation of the sensor data comprises a range image; and
the initial representation of the sensor data comprises a point cloud.
16 . The system of claim 10 , wherein the one or more processors are further to:
generate, based at least on the correlation, one or more class labels associated with the one or more second bounding shapes,
wherein the one or more operations are further to be performed based at least on the one or more class labels.
17 . The system of claim 10 , wherein the system is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for an autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for performing deep learning operations;
a system implemented using an edge device;
a system implemented using a robot;
a system for generating synthetic data;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources.
18 . One or more processors comprising processing circuitry to:
correlate one or more first locations within a projection image of one or more first bounding shapes associated with one or more objects with one or more segmentation masks associated with the one or more objects as represented in the projection image;
project, based at least on the correlation, the one or more first locations of the one or more first bounding shapes from the projection image to one or more second locations of one or more second bounding shapes within a point cloud; and
cause a machine to perform one or more operations based at least on the one or more second locations of the one or more second bounding shapes.
19 . The one or more processors of claim 18 , wherein:
the one or more first locations of the one or more first bounding shapes indicate at least one or more first points of the projection image that correspond to the one or more objects; and
the one or more segmentation masks indicate at least one or more second points of the projection image that correspond to the one or more objects.
20 . The one or more processors of claim 18 , wherein the one or more processors is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for an autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for performing deep learning operations;
a system implemented using an edge device;
a system implemented using a robot;
a system for generating synthetic data;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources.