Machine learning-based point cloud alignment classification
Provided are methods, systems, and computer program products for machine-learning based point cloud alignment classification. An example method may include: obtaining at least two light detection and ranging (LiDAR) point clouds; processing the at least two LiDAR point clouds using at least one classifier network; obtaining at least one output dataset from the at least one classifier network; determining that the at least two LiDAR point clouds are misaligned based on the at least one output dataset; and performing a first action based on the determining that the at least two LiDAR point clouds are misaligned.
1 . A method, comprising:
generating a plurality of merged point cloud images based on a plurality of sets of LiDAR point clouds, wherein generating a first merged point cloud image of the plurality of merged point cloud images comprises merging at least two LiDAR point clouds of a first set of LiDAR point clouds of the plurality of sets of LiDAR point clouds, wherein the first merged point cloud image includes respective LiDAR points of the at least two LiDAR point clouds of the first set of LiDAR point clouds;
indicating that the at least two LiDAR point clouds of the first set of LiDAR point clouds are aligned based on the generating the first merged point cloud image;
indicating that at least two LiDAR point clouds of a second set of LiDAR point clouds of the plurality of sets of LiDAR point clouds are aligned based on generating a second merged point cloud image;
processing the plurality of merged point cloud images using at least one neural network, wherein the at least one neural network was trained by:
obtaining training data comprising at least one first set of merged point cloud images comprising misaligned pairs of LiDAR point clouds and at least one second set of merged point cloud images comprising aligned pairs of LiDAR point clouds, and
training, using the training data, the at least one neural network to classify merged point cloud images as having aligned or misaligned LiDAR point clouds;
receiving, from the at least one neural network, a first probability corresponding to a likelihood that the at least two LiDAR point clouds of the first set of LiDAR point clouds included in the first merged point cloud image are misaligned;
receiving, from the at least one neural network, a second probability corresponding to a likelihood that the at least two LiDAR point clouds of the second set of LiDAR point clouds included in the second merged point cloud image are aligned;
determining that the at least two LiDAR point clouds of the first set of LiDAR point clouds are misaligned based on the first probability;
determining that the at least two LiDAR point clouds of the second set of LiDAR point clouds are aligned based on the second probability; and
generating a map using the second merged point cloud image based on determining that the at least two LiDAR point clouds of the second set of LiDAR point clouds are aligned.
2 . The method of claim 1 , wherein obtaining at least one set of LiDAR point clouds of the plurality of sets of LiDAR point clouds comprises:
obtaining the at least one set of LiDAR point clouds from a first plurality of LiDAR point clouds for a point cloud registration process to map a locality of a map;
obtaining a first LiDAR point cloud of the at least one set of LiDAR point clouds from a LiDAR system onboard a vehicle and a second LiDAR point cloud of the at least one set of LiDAR point clouds from a second plurality of LiDAR point clouds for a map correction process;
obtaining a first LiDAR point cloud of the at least one set of LiDAR point clouds from a LiDAR system onboard a vehicle and a second LiDAR point cloud of the at least one set of LiDAR point clouds from a third plurality of LiDAR point clouds for a localization process; or obtaining a first LiDAR point cloud of the at least one set of LiDAR point clouds from a LiDAR system onboard the vehicle and a second LiDAR point cloud of the at least one set of LiDAR point clouds from a fourth plurality of LiDAR point clouds for a calibration process.
3 . The method of claim 1 , wherein the at least one neural network comprises at least one of: a pillar-based network or a kernel point convolution-based network.
4 . The method of claim 1 , wherein the at least one neural network comprises a pillar-based network and a kernel point convolution-based network, wherein receiving the first probability from the at least one neural network comprises receiving a third probability from the pillar-based network and a fourth probability from the kernel point convolution-based network, and wherein determining that the at least two LiDAR point clouds of the first set of LiDAR point clouds are misaligned comprises:
weighting the third probability from the pillar-based network and the fourth probability from the kernel point convolution-based network based on a confidence value associated with the third probability and a confidence value associated with the fourth probability,
fusing the third probability from the pillar-based network and the fourth probability from the kernel point convolution-based network based on the weighting the third probability and the fourth probability, and
determining that the at least two LiDAR point clouds of the first set of LiDAR point clouds are misaligned based on the fusing the third probability and the fourth probability.
5 . The method of claim 1 , wherein a first neural network of the at least one neural network is a pillar-based network, wherein the pillar-based network comprises:
a feature network that receives at least one LiDAR point cloud and outputs at least one feature map,
at least one functional network that receives the at least one feature map and outputs a feature vector, and
a fully connected layer that receives the feature vector and outputs the first probability or the second probability.
6 . The method of claim 5 , wherein the feature network includes:
a pillar encoder that receives the at least one LiDAR point cloud and outputs at least one pseudo-image, and
a feature backbone that receives the at least one pseudo-image and outputs the at least one feature map.
7 . The method of claim 5 , wherein the at least one functional network comprises at least one of:
a concatenation network, at least one convolutional network, or a flatten network.
8 . The method of claim 1 , wherein a first neural network of the at least one neural network is a pillar-based network, wherein the pillar-based network comprises:
a first feature network that receives a first LiDAR point cloud and outputs a first feature map,
a second feature network that receives a second LiDAR point cloud and outputs a second feature map,
at least one functional network that receives the first feature map and the second feature map, and outputs a feature vector, and
a fully connected layer that receives the feature vector and outputs the first probability or the second probability.
9 . The method of claim 1 , wherein a first neural network of the at least one neural network is a kernel point convolution-based network, wherein the kernel point convolution-based network comprises:
a kernel point convolution-based encoder that receives at least one merged point cloud image and outputs a plurality of feature vectors,
an aggregation function that receives the plurality of feature vectors and aggregates the plurality of feature vectors into a single feature vector, and
a fully connected layer that receives the single feature vector and outputs the first probability or the second probability.
10 . The method of claim 1 , wherein at least one merged point cloud image of the plurality of merged point cloud images is source-labeled to each of at least two LiDAR point clouds.
11 . The method of claim 9 , wherein the aggregation function comprises a max pooling function a random choice function, a global average function, a mean value function, or a non-parametric aggregation function.
12 . The method of claim 1 , further comprising:
performing a first action based on the determining that the at least two LiDAR point clouds of the second set of LiDAR point clouds are aligned, wherein performing the first action comprises:
labeling the at least two LiDAR point clouds of the second set of LiDAR point clouds as aligned, and/or
updating a locality of a map based on a previous labeling of the at least two LiDAR point clouds of the second set of LiDAR point clouds as aligned.
13 . The method of claim 1 , further comprising:
performing a first action based on the determining that the at least two LiDAR point clouds of the first set of LiDAR point clouds are misaligned, wherein the first action comprises:
labeling the at least two LiDAR point clouds of the first set of LiDAR point clouds as misaligned, and/or
updating a locality of a map based on a previous labeling of the at least two LiDAR point clouds of the first set of LiDAR point clouds as misaligned.
14 . A system, comprising:
at least one processor, and
at least one non-transitory storage media storing instructions that, when executed by the at least one processor, cause the at least one processor to:
generate a plurality of merged point cloud images based on a plurality of sets of LiDAR point clouds, wherein to generate a first merged point cloud image of the plurality of merged point cloud images, the instructions, when executed by the at least one processor, cause the at least one processor to merge at least two LiDAR point clouds of a first set of LiDAR point clouds of the plurality of sets of LiDAR point clouds, wherein the first merged point cloud image includes the respective plurality of LiDAR points of the at least two LiDAR point clouds of the first set of LiDAR point clouds;
indicate that the at least two LiDAR point clouds of the first set of LiDAR point clouds are aligned based on the generation of the first merged point cloud image;
indicate that at least two LiDAR point clouds of a second set of LiDAR point clouds of the plurality of sets of LiDAR point clouds are aligned based on generation of a second merged point cloud image;
process the plurality of merged point cloud images using at least one neural network, wherein the at least one neural network was trained by:
obtaining training data comprising at least one first set of merged point cloud images comprising misaligned pairs of LiDAR point clouds and at least one second set of merged point cloud images comprising aligned pairs of LiDAR point clouds, and
training, using the training data, the at least one neural network to classify merged point cloud images as having aligned or misaligned LiDAR point clouds;
receive, from the at least one neural network, a first probability corresponding to a likelihood that the at least two LiDAR point clouds of the first set of LiDAR point clouds included in the first merged point cloud image are misaligned;
receive, from the at least one neural network, a second probability corresponding to a likelihood that the at least two LiDAR point clouds of the second set of LiDAR point clouds included in the second merged point cloud image are aligned;
determine that the at least two LiDAR point clouds of the first set of LiDAR point clouds are misaligned based on the first probability;
determine that the at least two LiDAR point clouds of the second set of LiDAR point clouds that are aligned based on the second probability of the plurality of probabilities, the second probability corresponding to the second merged point cloud image; and
generate a map using the second merged point cloud image based on the determination that the at least two LiDAR point clouds of the second set of LiDAR point clouds are aligned.
15 . The system of claim 14 , wherein the at least one neural network comprises at least one of: a pillar-based network or a kernel point convolution-based network.
16 . The system of claim 14 , wherein the at least one neural network comprises a pillar-based network and a kernel point convolution-based network, wherein to receive the first probability from the at least one neural network, the instructions, when executed by the at least one processor, further cause the at least one processor to: receive a third probability from the pillar-based network and a fourth probability from the kernel point convolution-based network, and wherein to determine that the at least two LiDAR point clouds of the first set of LiDAR point clouds are misaligned the instructions, when executed by the at least one processor, cause the at least one processor to:
weight the third probability from the pillar-based network and the fourth probability from the kernel point convolution-based network based on a confidence value associated with the third probability and a confidence value associated with the fourth probability,
fuse the third probability from the pillar-based network and the fourth probability from the kernel point convolution-based network based on the weighting the third probability and the fourth probability, and
determining that the at least two LiDAR point clouds of the first set of LiDAR point clouds are misaligned based on fusing the third probability and the fourth probability.
17 . The system of claim 14 , wherein a first neural network of the at least one neural network is a pillar-based network, wherein the pillar-based network comprises:
a feature network that receives at least one LiDAR point cloud and outputs at least one feature map,
at least one functional network that receives the at least one feature map and outputs a feature vector, and
a fully connected layer that receives the feature vector and outputs the first probability or the second probability.
18 . The system of claim 14 , wherein a first neural network of the at least one neural network is a kernel point convolution-based network, wherein the kernel point convolution-based network comprises:
a kernel point convolution-based encoder that receives at least one merged point cloud image and outputs a plurality of feature vectors,
an aggregation function that receives the plurality of feature vectors and aggregates the plurality of feature vectors into a single feature vector, and
a fully connected layer that receives the single feature vector and outputs the first probability or the second probability.
19 . At least one non-transitory storage media storing instructions that, when executed by at least one processor, cause the at least one processor to:
generate a plurality of merged point cloud images based on a plurality of sets of LiDAR point clouds, wherein to generate a first merged point cloud image of the plurality of merged point cloud images, the instructions, when executed by the at least one processor, cause the at least one processor to merge at least two LiDAR point clouds of a first set of LiDAR point clouds of the plurality of sets of LiDAR point clouds, wherein the first merged point cloud image includes respective LiDAR points of the at least two LiDAR point clouds of the first set of LiDAR point clouds;
indicate that the at least two LiDAR point clouds of the first set of LiDAR point clouds are aligned based on the generation of the first merged point cloud image;
indicate that at least two LiDAR point clouds of a second set of LiDAR point clouds of the plurality of sets of LiDAR point clouds are aligned based on generation of a second merged point cloud image;
process the plurality of merged point cloud images using at least one neural network, wherein the at least one neural network was trained by:
obtaining training data comprising at least one first set of merged point cloud images comprising misaligned pairs of LiDAR point clouds and at least one second set of merged point cloud images comprising aligned pairs of LiDAR point clouds, and
training, using the training data, the at least one neural network to classify merged point cloud images as having aligned or misaligned LiDAR point clouds;
receive, from the at least one neural network, a first probability corresponding to a likelihood that the at least two LiDAR point clouds of the first set of LiDAR point clouds included in the first merged point cloud image are misaligned;
receive, from the at least one neural network, a second probability corresponding to a likelihood that the at least two LiDAR point clouds of the second set of LiDAR point clouds included in the second merged point cloud image are aligned;
determine that the at least two LiDAR point clouds of the first set of LiDAR point clouds are misaligned based on the first probability;
determine that the at least two LiDAR point clouds of the second set of LiDAR point clouds are aligned based on the second probability; and
generate a map using the second merged point cloud image based on the determination that the at least two LiDAR point clouds of the second set of LiDAR point clouds are aligned.