Matching between 2D and 3D for direct localization
Determining a location of an entity comprises: receiving a query comprising a 2D image depicting an environment of the entity; searching for a match between the query and a 3D map of the environment. The 3D map comprising a 3D point cloud, the match indicating the location of the entity in the environment. Searching for the match comprises: extracting descriptors from the 2D image referred to as image descriptors; extracting descriptors from the 3D point cloud referred to as point cloud descriptors; correlating the image descriptors with the point cloud descriptors to produce correspondences, wherein a correspondence is an image descriptor corresponding to a point cloud descriptor; estimating, using the correspondences, the location of the entity.
1 . A method of determining a pose of an entity, the pose comprising a 3D position and orientation of the entity, the method comprising:
receiving a query comprising a single 2D image depicting an environment of the entity; and
searching for a match between the query and a 3D map of the environment, the 3D map comprising a 3D point cloud omitting visual imagery of the environment, searching for the match further comprising:
predicting a coarse image feature map comprising a feature vector at a flattened spatial location;
down-sampling the 3D point cloud to a coarsest resolution to reduce one or more of: a pointwise correspondence density, a correspondence clustering, and redundant correspondence of the 3D point cloud;
extracting an image descriptor from the single 2D image, the image descriptor based on the predicted coarse image feature map;
extracting a point cloud descriptor from the down-sampled 3D point cloud;
using the single 2D image, directly localizing the entity with respect to the 3D map, directly localizing comprising:
correlating the image descriptor with the point cloud descriptor to produce a correspondence,
the correspondence being a pair comprising the image descriptor and the point cloud descriptor, the pair having a numerical similarity value above a threshold representing a likelihood that the pair is depicting a same element of the environment; and
estimating, using the correspondence, the pose of the entity with respect to the 3D map.
2 . The method of claim 1 wherein the estimated pose of the entity is a relative pose between the entity and the 3D map of the environment.
3 . The method of claim 1 comprising, prior to the correlating, refining the image descriptor using the point cloud descriptor and refining the point cloud descriptor using the image descriptor, refining the image descriptor and refining the point cloud descriptor facilitating the correlating by making the image descriptor similar to the point cloud descriptor.
4 . The method of claim 3 wherein the refining comprises using a trained machine learning model having a cross-attention layer.
5 . The method of claim 4 wherein the refining comprises using a trained machine learning model having a first self-attention layer for the image descriptor and a second self-attention layer for the point cloud descriptor.
6 . The method of claim 1 wherein the correlating comprises computing similarity between point cloud descriptor and image descriptor.
7 . The method of claim 1 wherein:
extracting the image descriptor is done using a machine learning model, and extracting the point cloud descriptor is done using a machine learning model;
omitting visual imagery of the environment further comprises increasing a security of the single 2D image and securing the direct localization of the entity; and
correlating the image descriptor with the point cloud descriptor to produce a correspondence further comprises:
computing an output cost volume as dense scalar products, the output cost volume encoding a deep feature similarity between a coarse point-cloud location and a coarse image feature map location;
converting the output cost volume into a soft assignment matrix by applying a softmax operator over a flattened image dimension,
a row of the soft assignment matrix being a predicted probability distribution of where a point on the 3D point cloud projects in the single 2D image, and
an entry of the soft assignment matrix encoding a predicted confidence of a candidate match;
extracting a predicted correspondence using mutual top-β selection; and
selecting the candidate match as a match when:
the candidate match is among a β largest entries of the row and a column of the soft assignment matrix where the candidate match is located, and
a corresponding confidence of the candidate match is above the threshold.
8 . An apparatus comprising:
a processor;
a memory storing instructions that, when executed by the processor, perform a method of determining a pose of an entity, the pose comprising a 3D position and orientation of the entity, the method comprising:
receiving a query comprising a single 2D image depicting an environment of the entity; and
searching for a match between the query and a 3D map of the environment, the 3D map comprising a 3D point cloud omitting visual imagery of the environment, searching for the match further comprising:
predicting a coarse image feature map comprising a feature vector at a flattened spatial location;
down-sampling the 3D point cloud to a coarsest resolution to reduce one or more of: a pointwise correspondence density, reduce a correspondence clustering, and reduce a redundant correspondence;
extracting an image descriptor from the single 2D image;
extracting a point cloud descriptor from the 3D point cloud;
using the single 2D image, directly localizing the entity with respect to the 3D map, directly localizing comprising:
correlating the image descriptor with the point cloud descriptor to produce a correspondence,
the correspondence being a pair comprising the image descriptor and the point cloud descriptor, the pair having a numerical similarity value above a threshold representing a likelihood that the pair is depicting a same element of the environment; and
estimating, using the correspondence, the pose of the entity with respect to the 3D map.
9 . The apparatus of claim 8 comprising a machine learning model configured to, prior to the correlating, refine the image descriptor using the point cloud descriptor and, refine the point cloud descriptor using the image descriptor, refining the image descriptor and refining the point cloud descriptor facilitating the correlating by making the image descriptor similar to the point cloud descriptor.
10 . The apparatus of claim 9 wherein the machine learning model has been trained using training data comprising a plurality of pairs, a pair of the plurality of pairs comprising a point cloud and a corresponding image.
11 . The apparatus of claim 10 wherein a ground truth pose of the pair of the plurality of pairs is known.
12 . The apparatus of claim 10 wherein the training data is computed from an RGB-D image of a plurality of RGB-D images by:
selecting the RGB-D image as a query image,
finding a plurality of reference images distinct from the query image in the plurality of RGB-D images which are covisible with the query image, and
projecting a plurality of pixels of the plurality of reference images to 3D to form the point cloud, and
storing the query image and the point cloud as a training data item.
13 . The apparatus of claim 12 wherein the training data is computed by augmenting the point cloud with at least one of: rotation, scaling, or noise.
14 . The apparatus of claim 12 wherein computing the training data further comprises applying a rotation, the rotation accounting for a gravity direction.
15 . The apparatus of claim 14 wherein a ground truth pose relating the point cloud to the image is modified according to the rotation.
16 . The apparatus of claim 10 wherein the processor executes further instructions stored in the memory, further comprising: computing the training data from a plurality of RGB-D images by:
selecting a first RGB-D image and a second RGB-D image of the plurality of RGB-D images as a first query image and a second query image,
finding a plurality of reference images distinct from the first query image and the second query image in the plurality of RGB-D images which are covisible with first query image or the second query image,
projecting a plurality of pixels of the plurality of reference images to 3D to form a single point cloud, and
storing the first query image and the second query image and the single point cloud as a training data item.
17 . The apparatus of claim 16 further comprising a 2D-2D matching network which shares a component for extracting the image descriptor from the single 2D image.
18 . The apparatus of claim 17 wherein training the machine learning model further comprises training the 2D-2D matching network.
19 . The apparatus of claim 9 wherein the machine learning model is trained using a training loss being a focal loss.
20 . A computer program embodied on a non-transitory computer-readable storage and configured to, when executed on a processor, perform a method to determine a pose of an entity, the pose comprising a 3D position and orientation of the entity, the method comprising:
receiving a query comprising a single 2D image depicting an environment of the entity; and
searching for a match between the query and a 3D map of the environment, the 3D map comprising a 3D point cloud omitting visual imagery of the environment, searching for the match further comprising:
predicting a coarse image feature map comprising a feature vector at a flattened spatial location;
down-sampling the 3D point cloud to a coarsest resolution to reduce one or more of: a pointwise correspondence density, a correspondence clustering, and redundant correspondence of the 3D point cloud;
extracting an image descriptor from the single 2D image, the image descriptor based on the predicted coarse image feature map;
extracting a point cloud descriptor from the down-sampled 3D point cloud;
using the single 2D image, directly localizing the entity with respect to the 3D map, directly localizing comprising:
refining the image descriptor or the point cloud descriptor using a trained machine learning model having a cross-attention layer,
correlating the image descriptor with the point cloud descriptor to produce a correspondence,
the correspondence being a pair comprising the image descriptor and the point cloud descriptor, the pair having a numerical similarity value above a threshold representing a likelihood that the pair is depicting a same element of the environment; and
estimating, using the correspondence, the pose of the entity with respect to the 3D map.