IP Library › Granted Patent US 12,657,764
Granted Patent B2
US 12,657,764 · App. 18/055,722 · Granted Jun 16, 2026

Matching between 2D and 3D for direct localization

Inventors: Johannes Lutz Schönberger (Zurich, CH); Rui Wang (Zurich, CH); Prune Solange Garance Truong (Zurich, CH); Marc André Léon Pollefeys (Zurich, CH)
Assignee: Microsoft Technology Licensing, LLC.
G06T7/74G06F16/535G06V20/647G06T2207/10024G06T2207/10028G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,764
App. No.
18/055,722
Granted
Jun 16, 2026
Kind
B2
Abstract

Determining a location of an entity comprises: receiving a query comprising a 2D image depicting an environment of the entity; searching for a match between the query and a 3D map of the environment. The 3D map comprising a 3D point cloud, the match indicating the location of the entity in the environment. Searching for the match comprises: extracting descriptors from the 2D image referred to as image descriptors; extracting descriptors from the 3D point cloud referred to as point cloud descriptors; correlating the image descriptors with the point cloud descriptors to produce correspondences, wherein a correspondence is an image descriptor corresponding to a point cloud descriptor; estimating, using the correspondences, the location of the entity.

Claims (72)

1 . A method of determining a pose of an entity, the pose comprising a 3D position and orientation of the entity, the method comprising:

receiving a query comprising a single 2D image depicting an environment of the entity; and

searching for a match between the query and a 3D map of the environment, the 3D map comprising a 3D point cloud omitting visual imagery of the environment, searching for the match further comprising:

predicting a coarse image feature map comprising a feature vector at a flattened spatial location;

down-sampling the 3D point cloud to a coarsest resolution to reduce one or more of: a pointwise correspondence density, a correspondence clustering, and redundant correspondence of the 3D point cloud;

extracting an image descriptor from the single 2D image, the image descriptor based on the predicted coarse image feature map;

extracting a point cloud descriptor from the down-sampled 3D point cloud;

using the single 2D image, directly localizing the entity with respect to the 3D map, directly localizing comprising:

correlating the image descriptor with the point cloud descriptor to produce a correspondence,

the correspondence being a pair comprising the image descriptor and the point cloud descriptor, the pair having a numerical similarity value above a threshold representing a likelihood that the pair is depicting a same element of the environment; and

estimating, using the correspondence, the pose of the entity with respect to the 3D map.

2 . The method of claim 1 wherein the estimated pose of the entity is a relative pose between the entity and the 3D map of the environment.

3 . The method of claim 1 comprising, prior to the correlating, refining the image descriptor using the point cloud descriptor and refining the point cloud descriptor using the image descriptor, refining the image descriptor and refining the point cloud descriptor facilitating the correlating by making the image descriptor similar to the point cloud descriptor.

4 . The method of claim 3 wherein the refining comprises using a trained machine learning model having a cross-attention layer.

5 . The method of claim 4 wherein the refining comprises using a trained machine learning model having a first self-attention layer for the image descriptor and a second self-attention layer for the point cloud descriptor.

6 . The method of claim 1 wherein the correlating comprises computing similarity between point cloud descriptor and image descriptor.

7 . The method of claim 1 wherein:

extracting the image descriptor is done using a machine learning model, and extracting the point cloud descriptor is done using a machine learning model;

omitting visual imagery of the environment further comprises increasing a security of the single 2D image and securing the direct localization of the entity; and

correlating the image descriptor with the point cloud descriptor to produce a correspondence further comprises:

computing an output cost volume as dense scalar products, the output cost volume encoding a deep feature similarity between a coarse point-cloud location and a coarse image feature map location;

converting the output cost volume into a soft assignment matrix by applying a softmax operator over a flattened image dimension,

a row of the soft assignment matrix being a predicted probability distribution of where a point on the 3D point cloud projects in the single 2D image, and

an entry of the soft assignment matrix encoding a predicted confidence of a candidate match;

extracting a predicted correspondence using mutual top-β selection; and

selecting the candidate match as a match when:

the candidate match is among a β largest entries of the row and a column of the soft assignment matrix where the candidate match is located, and

a corresponding confidence of the candidate match is above the threshold.

8 . An apparatus comprising:

a processor;

a memory storing instructions that, when executed by the processor, perform a method of determining a pose of an entity, the pose comprising a 3D position and orientation of the entity, the method comprising:

receiving a query comprising a single 2D image depicting an environment of the entity; and

searching for a match between the query and a 3D map of the environment, the 3D map comprising a 3D point cloud omitting visual imagery of the environment, searching for the match further comprising:

predicting a coarse image feature map comprising a feature vector at a flattened spatial location;

down-sampling the 3D point cloud to a coarsest resolution to reduce one or more of: a pointwise correspondence density, reduce a correspondence clustering, and reduce a redundant correspondence;

extracting an image descriptor from the single 2D image;

extracting a point cloud descriptor from the 3D point cloud;

using the single 2D image, directly localizing the entity with respect to the 3D map, directly localizing comprising:

correlating the image descriptor with the point cloud descriptor to produce a correspondence,

the correspondence being a pair comprising the image descriptor and the point cloud descriptor, the pair having a numerical similarity value above a threshold representing a likelihood that the pair is depicting a same element of the environment; and

estimating, using the correspondence, the pose of the entity with respect to the 3D map.

9 . The apparatus of claim 8 comprising a machine learning model configured to, prior to the correlating, refine the image descriptor using the point cloud descriptor and, refine the point cloud descriptor using the image descriptor, refining the image descriptor and refining the point cloud descriptor facilitating the correlating by making the image descriptor similar to the point cloud descriptor.

10 . The apparatus of claim 9 wherein the machine learning model has been trained using training data comprising a plurality of pairs, a pair of the plurality of pairs comprising a point cloud and a corresponding image.

11 . The apparatus of claim 10 wherein a ground truth pose of the pair of the plurality of pairs is known.

12 . The apparatus of claim 10 wherein the training data is computed from an RGB-D image of a plurality of RGB-D images by:

selecting the RGB-D image as a query image,

finding a plurality of reference images distinct from the query image in the plurality of RGB-D images which are covisible with the query image, and

projecting a plurality of pixels of the plurality of reference images to 3D to form the point cloud, and

storing the query image and the point cloud as a training data item.

13 . The apparatus of claim 12 wherein the training data is computed by augmenting the point cloud with at least one of: rotation, scaling, or noise.

14 . The apparatus of claim 12 wherein computing the training data further comprises applying a rotation, the rotation accounting for a gravity direction.

15 . The apparatus of claim 14 wherein a ground truth pose relating the point cloud to the image is modified according to the rotation.

16 . The apparatus of claim 10 wherein the processor executes further instructions stored in the memory, further comprising: computing the training data from a plurality of RGB-D images by:

selecting a first RGB-D image and a second RGB-D image of the plurality of RGB-D images as a first query image and a second query image,

finding a plurality of reference images distinct from the first query image and the second query image in the plurality of RGB-D images which are covisible with first query image or the second query image,

projecting a plurality of pixels of the plurality of reference images to 3D to form a single point cloud, and

storing the first query image and the second query image and the single point cloud as a training data item.

17 . The apparatus of claim 16 further comprising a 2D-2D matching network which shares a component for extracting the image descriptor from the single 2D image.

18 . The apparatus of claim 17 wherein training the machine learning model further comprises training the 2D-2D matching network.

19 . The apparatus of claim 9 wherein the machine learning model is trained using a training loss being a focal loss.

20 . A computer program embodied on a non-transitory computer-readable storage and configured to, when executed on a processor, perform a method to determine a pose of an entity, the pose comprising a 3D position and orientation of the entity, the method comprising:

receiving a query comprising a single 2D image depicting an environment of the entity; and

searching for a match between the query and a 3D map of the environment, the 3D map comprising a 3D point cloud omitting visual imagery of the environment, searching for the match further comprising:

predicting a coarse image feature map comprising a feature vector at a flattened spatial location;

down-sampling the 3D point cloud to a coarsest resolution to reduce one or more of: a pointwise correspondence density, a correspondence clustering, and redundant correspondence of the 3D point cloud;

extracting an image descriptor from the single 2D image, the image descriptor based on the predicted coarse image feature map;

extracting a point cloud descriptor from the down-sampled 3D point cloud;

using the single 2D image, directly localizing the entity with respect to the 3D map, directly localizing comprising:

refining the image descriptor or the point cloud descriptor using a trained machine learning model having a cross-attention layer,

correlating the image descriptor with the point cloud descriptor to produce a correspondence,

the correspondence being a pair comprising the image descriptor and the point cloud descriptor, the pair having a numerical similarity value above a threshold representing a likelihood that the pair is depicting a same element of the environment; and

estimating, using the correspondence, the pose of the entity with respect to the 3D map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2023
From: SCHÖNBERGER, JOHANNES LUTZ; WANG, RUI; TRUONG, PRUNE SOLANGE GARANCE; POLLEFEYS, MARC ANDRÉ LÉON
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062886/0928 →
Continuity (1)
Related Publication 20240161337A1 · May 16, 2024
References Cited (40)
US 9251417B1 · Xu · 2016 [cited by examiner]
US 9633483B1 · Xu · 2017 [cited by examiner]
US 11189049B1 · Chakravarty · 2021 [cited by examiner]
US 11704841B2 · Mashita · 2023 [cited by examiner]
US 12377549B2 · Moreno Noguer · 2025 [cited by examiner]
US 20180158235A1 · Wu · 2018 [cited by examiner]
US 20190206116A1 · Xu · 2019 [cited by examiner]
US 20210335033A1 · Meng · 2021 [cited by examiner]
US 20220351465A1 · Pantpratinidhi · 2022 [cited by examiner]
US 20230126333A1 · Ali · 2023 [cited by examiner]
US 20230298307A1 · Mao · 2023 [cited by examiner]
US 20240029295A1 · You · 2024 [cited by examiner]
US 20240265657A1 · Pei · 2024 [cited by examiner]
US 20250052590A1 · Yoshida · 2025 [cited by examiner]
US 20250316036A1 · Evangelidis · 2025 [cited by examiner]
CN 114882106A · 2022 [cited by applicant]
Aiello E, Valsesia D, Magli E. Cross-modal Learning for Image-Guided Point Cloud Shape Completion. InAdvances in Neural Information Processing Systems Oct. 31, 2022. (Year: 2022). [cited by examiner]
Yun P, Tai L, Wang Y, Liu C, Liu M. Focal loss in 3d object detection. IEEE Robotics and Automation Letters. Jan. 23, 2019;4(2): 1263-70. (Year: 2019). [cited by examiner]
Zhao C, Yang J, Xiong X, Zhu A, Cao Z, Li X. Rotation invariant point cloud analysis: Where local geometry meets global topology. Pattern Recognition. Jul. 1, 2022;127:108626. (Year: 2022). [cited by examiner]
Shi C, Wang C, Liu X, Sun S, Xiao B, Li X, Li G. Three-dimensional point cloud denoising via a gravitational feature function. Applied Optics. Feb. 10, 2022;61(6):1331-43. (Year: 2022). [cited by examiner]
Kim J, Choi C, Jang H, Kim YM. Piccolo: Point cloud-centric omnidirectional localization. InProceedings of the IEEE/CVF International Conference on Computer Vision 2021 (pp. 3313-3323). (Year: 2021). [cited by examiner]
Li, Minhao, et al. “2d3d-matr: 2d-3d matching transformer for detection-free registration between images and point clouds.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023. (Year: 2023). [cited by examiner]
Dong, Chuanxiang, et al. “A Novel Point Cloud Coarse Registration Method Combining 2D image information.” 2023 International Conference on Image Processing, Computer Vision and Machine Learning (ICICML). IEEE, 2023. (Ye… [cited by examiner]
Feng, et al., “2D3D-Matchnet: Learning to Match Keypoints Across 2D Image and 3D Point Cloud” 2019 International Conference on Robotics and Automation, IEEE, May 20, 2019, pp. 4790-4796. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US23/033800, Feb. 7, 2024, 18 pages. [cited by applicant]
Pham, et al., “LCD: Learned Cross-Domain Descriptors for 2D-3D Matching” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, Issue 07, Nov. 21, 2019, pp. 11856-11864. [cited by applicant]
Wang, et al., “P2-Net: Joint Description and Detection of Local Features for Pixel and Point Matching”, IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 10, 2021, pp. 16004-16013. [cited by applicant]
Yu, et al., “CoFiNet: Reliable Coarse-to-fine Correspondences for Robust Point Cloud Registration”, Oct. 26, 2021, pp. 1-13. [cited by applicant]
Zeng, et al., “3DMatch: Learning Local Geometric Descriptors From RGB-D Reconstructions”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21, 2017, pp. 1802-1811. [cited by applicant]
Campbell, et al., “Solving the Blind Perspective-n-Point Problem End-to-End with Robust Differentiable Geometric Optimization”, In Proceedings of 16th European Conference on Computer Vision, Aug. 23, 2020, pp. 244-261. [cited by applicant]
Dai, et al., “ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes”, In Proceedings of Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp. 2432-2443. [cited by applicant]
Feng, et al., “2D3D-Matchnet: Learning to Match Keypoints Across 2D Image and 3D Point Cloud”, In Proceedings of International Conference on Robotics and Automation, May 20, 2019, pp. 4790-4796. [cited by applicant]
Li, et al., “DeepI2P: Image-to-Point Cloud Registration via Deep Classification”, In Proceedings of Conference on Computer Vision and Pattern Recognition, Jun. 19, 2021, pp. 15960-15969. [cited by applicant]
Li, et al., “MegaDepth: Learning Single-View Depth Prediction from Internet Photos”, In Proceedings of Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, pp. 2041-2050. [cited by applicant]
Liu, et al., “Learning 2D-3D Correspondences to Solve The Blind Perspective-n-Point Problem”, In Repository of arXiv:2003.06752v1, Mar. 15, 2020, 23 Pages. [cited by applicant]
Pham, et al., “LCD: Learned Cross-Domain Descriptors for 2D-3D Matching”, In Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence, vol. 34, Issue 7, Apr. 3, 2020, pp. 11856-11864. [cited by applicant]
Schonberger, et al., “Structure-from-Motion Revisited”, In Proceedings of Conference on Computer Vision and Pattern Recognition, Jun. 27, 2016, pp. 4104-4113. [cited by applicant]
Wang, et al., “P2-Net: Joint Description and Detection of Local Features for Pixel and Point Matching”, In Proceedings of International Conference on Computer Vision, Oct. 10, 2021, pp. 15984-15993. [cited by applicant]
Zhou, et al., “Is Geometry Enough for Matching in Visual Localization?”, In Repository of arXiv:2203.12979v1, Mar. 24, 2022, 16 Pages. [cited by applicant]
International preliminary report on patentability Received in European Patent Application No. PCT/US23/033800, mailed on May 22, 2025, 13 pages. [cited by applicant]