IP Library › Granted Patent US 12,322,129
Granted Patent B2
US 12,322,129 · App. 17/592,500 · Granted Jun 3, 2025

Camera localization

Inventors: Sudipta Narayan Sinha (Kirkland, WA); Ondrej Miksik (Zurich, CH); Joseph Michael Degol (Seattle, WA); Tien Do (Minneapolis, MN)
Assignee: Microsoft Technology Licensing, LLC
G06T7/70G06V10/7747G06V10/7784G06V20/64G06T2207/10016G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,129
App. No.
17/592,500
Granted
Jun 3, 2025
Kind
B2
Abstract

In various embodiments there is a method for camera localization within a scene. An image of a scene captured by the camera is input to a machine learning model, which has been trained for the particular scene to detect a plurality of 3D scene landmarks. The 3D scene landmarks are pre-specified in a pre-built map of the scene. The machine learning model outputs a plurality of predictions, each prediction comprising: either a 2D location in the image which is predicted to depict one of the 3D scene landmarks, or a 3D bearing vector, being a vector originating at the camera and pointing towards a predicted 3D location of one of the 3D scene landmarks. Using the predictions, an estimate of a position and orientation of the camera in the pre-built map of the scene is computed.

Claims (29)

1. A method for camera localization within a scene, the method comprising:

receiving, at a processor, an image of a scene, the image captured by the camera;

inputting the image to a machine learning model which has been trained for the scene to detect a plurality of 3D scene landmarks, wherein the 3D scene landmarks are specified in a map of the scene, and receiving as output from the machine learning model a plurality of predictions, each prediction comprising a predicted 3D bearing vector originating at the camera and pointing towards a predicted 3D location of one of the 3D scene landmarks; and

using the predictions, computing an estimate of a position and orientation of the camera in the map of the scene.

2. The method of claim 1 wherein the camera is a mobile camera moving in the scene.

3. The method of claim 1 wherein the machine learning model has been trained by denoting for a 2D image point corresponding to a landmark, an associated unitized bearing vector in the camera coordinate as a point on a unit sphere.

4. The method of claim 1 wherein the method further comprises inputting the image to a second machine learning model, the second machine learning model configured to output predicted 2D locations in the image.

5. The method of claim 4 comprising determining pairs of the predicted 2D locations and the predicted 3D bearing vectors which relate to a same one of the 3D scene landmarks, and discarding the predicted 3D bearing vectors from the determined pairs prior to using the predictions to compute the estimate of the position and orientation of the camera.

6. The method of claim 1 wherein a number of the specified 3D landmarks is less than 250.

7. The method of claim 1 comprising automatically selecting the specified 3D landmarks in advance by selecting a specified number of 3D points in the scene which are the most robust, repeatedly detectable, unique in appearance, and generalizable.

8. The method of claim 7 comprising computing a saliency score to determine which are the most robust, repeatedly detectable, unique in appearance and generalizable 3D points in the scene.

9. The method of claim 1 wherein the machine learning model comprises a convolutional neural network which has been trained using supervised learning to predict the predicted 3D bearing vectors.

10. The method of claim 4 wherein the second machine learning model comprises a convolutional neural network and the input image is input into the convolutional neural network which computes feature maps, and wherein the feature maps are used to predict one heatmap for each 3D landmark, and a transposed convolution layer is used for performing up-sampling and generating a second set of heatmaps for each 3D landmark.

11. The method of claim 1 wherein the machine learning model computes the 3D bearing vectors using a plurality of multi-layered perceptrons, one multi-layered perceptron per 3D landmark.

12. The method of claim 11 wherein the machine learning model has been trained using supervised learning.

13. The method of claim 1 wherein the machine learning model has been trained using training data comprising observed videos of the scene and exhibiting appearance, illumination and geometric changes.

14. The method of claim 13 wherein the observed videos are captured over a time frame during which one or more objects in the scene move or during which ambient light in the scene changes.

15. The method of claim 13 wherein the training data is augmented with synthetic videos computed by applying homography and/or intensity changes to the observed videos of the scene.

16. The method of claim 1 comprising automatically selecting the specified 3D scene landmarks in advance by carrying out object recognition on images or videos depicting the scene, and using rules to place the 3D scene landmarks on recognized objects which are likely to be static.

17. The method of claim 1 further comprising:

removing one of the specified 3D scene landmarks in response to disappearance of a corresponding pattern in the scene or adding a new specified 3D scene landmark in response to a new salient object appearing in the scene; and

retraining the machine learning model.

18. The method of claim 1 wherein the machine learning model comprises a first neural network trained using data from multiple scenes and which feeds into additional neural network layers which are scene-specific.

19. The method of claim 1 wherein the scene is a large scene divided into a plurality of overlapping subregions, and where the method comprises using a different plurality of 3D scene landmarks for each subregion and using a different machine learning model for each subregion, using a classifier to classify the image into one of the subregions and inputting the image to the machine learning model associated with the subregion.

20. An apparatus for camera localization within a scene, the apparatus comprising a processor executing instructions from a non-transitory machine readable storage medium, the instructions being for:

receiving an image of a scene, the image captured by the camera;

executing first and second machine learning models, each of the machine learning models receiving the image as input and each of the machine learning models having been trained for the scene to detect a plurality of 3D scene landmarks, wherein the 3D scene landmarks are specified in a map of the scene;

receiving as output from the first machine learning model a first plurality of predictions and from the second machine learning model a second plurality of predictions, wherein each of the predictions of the first plurality of predictions and the second plurality of predictions indicates a location of one of the 3D scene landmarks; and

using the first plurality of predictions and the second plurality of predictions, computing an estimate of a position and orientation of the camera in the map of the scene.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2022
From: SINHA, SUDIPTA NARAYAN; MIKSIK, ONDREJ; DEGOL, JOSEPH MICHAEL; DO, TIEN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 058884/0069 →
Continuity (2)
Provisional Application 63279614 · Nov 15, 2021
Related Publication 20230154032A1 · May 18, 2023
References Cited (108)
US 20180189974A1 · Clark · 2018 [cited by examiner]
US 20190108651A1 · Gu · 2019 [cited by examiner]
US 20200301015A1 · Siddiqui · 2020 [cited by examiner]
US 20220414932A1 · Chen · 2022 [cited by examiner]
US 20230032888A1 · Li · 2023 [cited by examiner]
US 20240104773A1 · Altillawi · 2024 [cited by examiner]
Moolan-Feroze, Oliver, and Andrew Calway. “Predicting out-of-view feature points for model-based camera pose estimation.” 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2018. (Yea… [cited by examiner]
Lee, Timothy E., et al. “Camera-to-robot pose estimation from a single image.” 2020 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2020. (Year: 2020). [cited by examiner]
Tremblay, Jonathan, et al. “Indirect object-to-robot pose estimation from an external monocular rgb camera. In 2020 IEEE.” RSJ International Conference on Intelligent Robots and Systems (IROS). 2020. (Year: 2020). [cited by examiner]
Arandjelovic, et al., “NetVLAD: CNN Architecture for Weakly Supervised Place Recognition”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, pp. 5297-5307. [cited by applicant]
Balntas, et al., “RelocNet: Continuous Metric Learning Relocalisation using Neural Nets”, In Proceedings of the European Conference on Computer Vision, Sep. 2018, pp. 1-17. [cited by applicant]
Bay, et al., “Speeded-Up Robust Features (SURF)”, In Journal of Computer Vision and Image Understanding, vol. 110, Issue 3, Jun. 2008, pp. 346-359. [cited by applicant]
Bergamo, et al., “Leveraging Structure from Motion to Learn Discriminative Codebooks for Scalable Landmark Classification”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2013, pp… [cited by applicant]
Brachmann, et al., “DSAC—Differentiable RANSAC for Camera Localization”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, pp. 6684-6692. [cited by applicant]
Brachmann, et al., “Expert Sample Consensus Applied to Camera Re-Localization”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 7525-7534. [cited by applicant]
Brachmann, et al., “Learning Less Is More—6D Camera Localization via 3D Surface Regression”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 4654-4662. [cited by applicant]
Brachmann, et al., “On the Limits of Pseudo Ground Truth in Visual Camera Re-Localisation”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2021, pp. 6218-6228. [cited by applicant]
Brachmann, et al., “Visual Camera Re-Localization from RGB and RGB-D Images Using DSAC”, In Journal of IEEE Transactions on Pattern Analysis and Machine Intelligence, Apr. 2, 2021, pp. 1-18. [cited by applicant]
Calonder, et al., “Fast Keypoint Recognition Using Random Ferns”, In Journal of IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 32, Issue 3, Mar. 2010, pp. 448-461. [cited by applicant]
Cao, et al., “Minimal Scene Descriptions from Structure from Motion Models”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2014, 8 Pages. [cited by applicant]
Cao, et al., “Unifying Deep Local and Global Features for Image Search”, In Proceedings of the European Conference on Computer Vision, Nov. 12, 2020, pp. 726-743. [cited by applicant]
Cheng, et al., “HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020, pp. 5386-5395. [cited by applicant]
Chum, et al., “Matching with PROSAC—Progressive Sample Consensus”, In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Jun. 20, 2005, 7 Pages. [cited by applicant]
Dai, et al., “ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, pp. 5828-5839. [cited by applicant]
Detone, et al., “SuperPoint: Self-Supervised Interest Point Detection and Description”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Jun. 2018, pp. 337-349. [cited by applicant]
Do, et al., “Surface Normal Estimation of Tilted Images via Spatial Rectifier”, In Proceedings of the European Conference on Computer Vision, Oct. 29, 2020, pp. 265-280. [cited by applicant]
Dong, et al., “Supervision by Registration and Triangulation for Landmark Detection”, In Journal of IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, Issue 10, Oct. 1, 2021, pp. 3681-3694. [cited by applicant]
Douze, et al., “Aggregating Local Descriptors into a Compact Image Representation”, In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Jun. 13, 2010, pp. 3304-3311. [cited by applicant]
Dusmanu, et al., “D2-Net: A Trainable CNN for Joint Description and Detection of Local Features”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 8084-8093. [cited by applicant]
Fischler, et al., “Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography”, In Journal of Communications of the ACM, vol. 24, Issue 6, Jun. 1, 1981, pp. 381-3… [cited by applicant]
Hays, et al., “IM2GPS: Estimating Geographic Information from a Single Image”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 23, 2008, 8 Pages. [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, pp. 770-778. [cited by applicant]
Hyeon, et al., “Pose Correction for Highly Accurate Visual Localization in Large-Scale Indoor Spaces”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2021, pp. 15974-15983. [cited by applicant]
Jin, et al., “Image Matching across Wide Baselines: From Paper to Practice”, In International Journal of Computer Vision, vol. 129, Issue 2, Oct. 7, 2020, pp. 517-547. [cited by applicant]
Johnson, et al., “Billion-Scale Similarity Search with GPUs”, In Journal of IEEE Transactions on Big Data, vol. 7, Issue 3, Jun. 7, 2019, pp. 535-547. [cited by applicant]
Kendall, et al., “Geometric Loss Functions for Camera Pose Regression With Deep Learning”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, pp. 5974-5983. [cited by applicant]
Kendall, et al., “PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization”, In Proceedings of the IEEE International Conference on Computer Vision, Dec. 2015, pp. 2938-2946. [cited by applicant]
Kingma, et al., “Adam: A Method for Stochastic Optimization”, In Repository of arXiv:1412.6980v1, Dec. 22, 2014, pp. 1-9. [cited by applicant]
Laskar, et al., “Camera Relocalization by Computing Pairwise Relative Poses Using Convolutional Neural Network”, In Proceedings of the IEEE International Conference on Computer Vision, Oct. 2017, pp. 929-938. [cited by applicant]
Lee, et al., “Large-Scale Localization Datasets in Crowded Indoor Spaces”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 3227-3236. [cited by applicant]
Lepetit, et al., “Keypoint Recognition using Randomized Trees”, In Journal of IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 28, Issue 9, Sep. 2006, pp. 1465-1479. [cited by applicant]
Li, et al., “Full-Frame Scene Coordinate Regression for Image-Based Localization”, In Proceedings of the Robotics: Science and Systems XIV, Jun. 26, 2018, 9 Pages. [cited by applicant]
Li, et al., “Scene Coordinate Regression with Angle-Based Reprojection Loss for Camera Relocalization”, In Proceedings of the European Conference on Computer Vision, Sep. 2018, pp. 1-17. [cited by applicant]
Li, et al., “Worldwide Pose Estimation Using 3D Point Clouds”, In Proceedings of the European Conference on Computer Vision, Oct. 7, 2012, pp. 15-29. [cited by applicant]
Lowe, David G. , “Distinctive Image Features from Scale-Invariant Keypoints”, In International Journal of Computer Vision, vol. 60, Issue 2, Nov. 2004, pp. 91-110. [cited by applicant]
Lynen, et al., “Get Out of My Lab: Large-Scale, Real-Time VVisual-Inertial Localization”, In Journal of Robotics: Science and Systems, vol. 1, Jul. 13, 2015, 10 Pages. [cited by applicant]
Maddern, et al., “1 Year, 1000km: The Oxford RobotCar Dataset”, In the International Journal of Robotics Research, vol. 36, Issue 1, Nov. 29, 2016, 8 Pages. [cited by applicant]
Massiceti, et al., “Random forests versus Neural Networks—What's Best for Camera Localization?”, In Proceedings of the IEEE International Conference on Robotics and Automation, May 29, 2017, pp. 5118-5125. [cited by applicant]
Meng, et al., “Backtracking Regression Forests for Accurate Camera Relocalization”, In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, Sep. 24, 2017, pp. 6886-6893. [cited by applicant]
Mikolajczyk, et al., “Scale and Affine Invariant Interest Point Detectors”, In International Journal of Computer Vision, vol. 60, Oct. 2004, pp. 63-86. [cited by applicant]
Mishchuk, et al., “Working Hard to Know your Neighbor's Margins: Local Descriptor Learning Loss”, In Proceedings of the 31st Neural Information Processing Systems, 2017, 12 Pages. [cited by applicant]
Muja, et al., “Fast Approximate Nearest Neighbors with Automatic Algorithm Configuration”, In Journal of VISAPP, vol. 2, 2009, 10 Pages. [cited by applicant]
Newcombe, et al., “KinectFusion: Real-Time Dense Surface Mapping and Tracking”, In Proceedings of the 10th IEEE International Symposium on Mixed and Augmented Reality, Oct. 26, 2011, pp. 127-136. [cited by applicant]
Newell, et al., “Stacked Hourglass Networks for Human Pose Estimation”, In Proceedings of the European Conference on Computer Vision, Sep. 17, 2016, pp. 483-499. [cited by applicant]
Ng, et al., “Reassessing the Limitations of CNN Methods for Camera Pose Regression”, In the Repository of arXiv:2108.07260v1, Aug. 16, 2021, pp. 1-15. [cited by applicant]
Park, et al., “3D Point Cloud Reduction Using Mixed-Integer Quadratic Programming”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Jun. 2013, pp. 229-236. [cited by applicant]
Paszke, et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library”, In Proceedings of the 33rd Neural Information Processing Systems, 2019, 12 Pages. [cited by applicant]
Pavlakos, et al., “6-DoF Object Pose from Semantic Keypoints”, In Proceedings of the IEEE International Conference on Robotics and Automation, May 29, 2017, pp. 2011-2018. [cited by applicant]
Peng, et al., “PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 4561-4570. [cited by applicant]
Do, et al., “Learning to Detect Scene Landmarks for Camera Localization”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 18, 2022, pp. 11122-11132. [cited by applicant]
Lee, et al., “Camera-to-Robot Pose Estimation from a Single Image”, In Proceedings of IEEE International Conference on Robotics and Automation (ICRA), May 31, 2020, pp. 9426-9432. [cited by applicant]
Moolan-Feroze, et al., “Predicting Out-of-View Feature Points for Model-Based Camera Pose Estimation”, In Proceedings of International Conference on Intelligent Robots and Systems (IROS), Oct. 1, 2018, pp. 82-88. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/043562”, Mailed Date: Jan. 5, 2023, 13 Pages. [cited by applicant]
Tremblay, et al., “Indirect Object-to-Robot Pose Estimation from an External Monocular RGB Camera”, In Repository of arXiv:2008.11822, Aug. 26, 2020, 8 Pages. [cited by applicant]
Zadeh, et al., “Convolutional Experts Constrained Local Model for 3D Facial Landmark Detection”, In Proceedings of the IEEE International Conference on Computer Vision, Oct. 2017, pp. 2519-2528. [cited by applicant]
Persson, et al., “Lambda Twist: An Accurate Fast Robust Perspective Three Point (P3P) Solver”, In Proceedings of the European Conference on Computer Vision, Sep. 2018, 15 Pages. [cited by applicant]
Pion, et al., “Benchmarking Image Retrieval for Visual Localization”, In Proceedings of International Conference on 3D Vision, Nov. 25, 2020, pp. 483-494. [cited by applicant]
Pittaluga, et al., “Revealing Scenes by Inverting Structure From Motion Reconstructions”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 145-154. [cited by applicant]
Raguram, et al., “USAC: A Universal Framework for Random Sample Consensus”, In Journal of IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, Issue 8, Aug. 2013, pp. 2022-2038. [cited by applicant]
Rublee, et al., “ORB: An Efficient Alternative to SIFT or SURF”, In Proceedings of the International Conference on Computer Vision, Nov. 6, 2011, pp. 2564-2571. [cited by applicant]
Saha, et al., “Improved Visual Relocalization by Discovering Anchor Points”, In Proceedings of the British Machine Vision Conference, Sep. 3, 2018, 11 Pages. [cited by applicant]
Sarlin, et al., “Back to the Feature: Learning Robust Camera Localization From Pixels to Pose”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 3247-3257. [cited by applicant]
Sarlin, et al., “From Coarse to Fine: Robust Hierarchical Localization at Large Scale”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 12716-12725. [cited by applicant]
Sarlin, et al., “SuperGlue: Learning Feature Matching With Graph Neural Networks”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020, pp. 4938-4947. [cited by applicant]
Sattler, et al., “Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 8601-8610. [cited by applicant]
Sattler, et al., “Image Retrieval for Image-Based Localization Revisited”, In Proceedings of the British Machine Vision Conference, Sep. 3, 2012, pp. 1-12. [cited by applicant]
Sattler, et al., “Improving Image-Based Localization by Active Correspondence Search”, In Proceedings of the European Conference on Computer Vision, Oct. 7, 2012, pp. 752-765. [cited by applicant]
Sattler, et al., “Understanding the Limitations of CNN-Based Absolute Camera Pose Regression”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 3302-3312. [cited by applicant]
Schonberger, et al., “Comparative Evaluation of Hand-Crafted and Learned Local Features”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, pp. 1482-1491. [cited by applicant]
Schonberger, et al., “Structure-From-Motion Revisited”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, pp. 4104-4113. [cited by applicant]
Shavit, et al., “Learning Multi-Scene Absolute Pose Regression With Transformer”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2021, pp. 2733-2742. [cited by applicant]
Shotton, et al., “Scene Coordinate Regression Forests for Camera Relocalization in RGB-D Images”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2013, pp. 2930-2937. [cited by applicant]
Simon, et al., “Hand Keypoint Detection in Single Images Using Multiview Bootstrapping”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, pp. 1145-1153. [cited by applicant]
Skydes, “ETH-MS Localization Dataset”, Retrieved from: https://github.com/cvg/visloc-iccv2021, Aug. 31, 2021, 6 Pages. [cited by applicant]
Speciale, et al., “Privacy Preserving Image Queries for Camera Localization”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 1486-1496. [cited by applicant]
Speciale, et al., “Privacy Preserving Image-Based Localization”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 5493-5503. [cited by applicant]
Sun, et al., “Deep High-Resolution Representation Learning for Human Pose Estimation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 5693-5703. [cited by applicant]
Taira, et al., “InLoc: Indoor Visual Localization With Dense Matching and View Synthesis”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 7199-7209. [cited by applicant]
Tekin, et al., “Real-Time Seamless Single Shot 6D Object Pose Prediction”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 292-301. [cited by applicant]
Tian, et al., “SOSNet: Second Order Similarity Regularization for Local Descriptor Learning”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 11016-11025. [cited by applicant]
Tolias, et al., “To Aggregate or Not to Aggregate: Selective Match Kernels for Image Search”, In Proceedings of the IEEE International Conference on Computer Vision, Dec. 2013, pp. 1401-1408. [cited by applicant]
Torii, et al., “24/7 Place Recognition by View Synthesis”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2015, pp. 1808-1817. [cited by applicant]
Toshev, et al., “DeepPose: Human Pose Estimation via Deep Neural Networks”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2014, 8 Pages. [cited by applicant]
Valentin, et al., “Exploiting Uncertainty in Regression Forests for Accurate Camera Relocalization”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2015, pp. 4400-4408. [cited by applicant]
Wald, et al., “Beyond Controlled Environments: 3D Camera Re-Localization in Changing Indoor Scenes”, In Proceedings of the European Conference on Computer Vision, Nov. 9, 2020, pp. 467-487. [cited by applicant]
Wang, et al., “AtLoc: Attention Guided Camera Localization”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, Issue 06, Apr. 3, 2020, pp. 10393-10401. [cited by applicant]
Wang, et al., “Continual Learning for Image-Based Camera Localization”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2021, pp. 3252-3262. [cited by applicant]
Wei, et al., “Convolutional Pose Machines”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, pp. 4724-4732. [cited by applicant]
Weinzaepfel, et al., “Visual Localization by Learning Objects-Of-Interest Dense Match Regression”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 5634-5643. [cited by applicant]
Wenzel, et al., “4Seasons: A Cross-Season Dataset for Multi-Weather SLAM in Autonomous Driving”, In Proceedings of the DAGM German Conference on Pattern Recognition, Sep. 28, 2020, pp. 404-417. [cited by applicant]
Weyand, et al., “PlaNet—Photo Geolocation with Convolutional Neural Networks”, In Proceedings of the European Conference on Computer Vision, Sep. 17, 2016, pp. 37-55. [cited by applicant]
Xiao, et al., “Simple Baselines for Human Pose Estimation and Tracking”, In Proceedings of the European Conference on Computer Vision, Sep. 2018, 16 Pages. [cited by applicant]
Xu, et al., “AnchorFace: An Anchor-based Facial Landmark Detector Across Large Poses”, In Proceedings of the Association for the Advancement of Artificial Intelligence, Jan. 13, 2021, pp. 3092-3100. [cited by applicant]
Xue, et al., “Learning Multi-View Camera Relocalization With Graph Neural Networks”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 11372-11381. [cited by applicant]
Yang, et al., “SANet: Scene Agnostic Network for Camera Localization”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 42-51. [cited by applicant]
Yi, et al., “LIFT: Learned Invariant Feature Transform”, In Proceedings of European Conference on Computer Vision, Sep. 17, 2016, pp. 467-483. [cited by applicant]
Yu, et al., “Multi-Scale Context Aggregation by Dilated Convolutions”, In Repository of arXiv:1511.07122v1, Nov. 23, 2015, pp. 1-9. [cited by applicant]
Badino, et al., “Visual Localization Data Set”, Retrieved from: https://web.archive.org/web/20150219124350/http://3dvis.ri.cmu.edu:80/data-sets/localization/, 2011, 8 Pages. [cited by applicant]