IP Library Granted Patent US 12,688,607
Granted Patent B2
US 12,688,607 · App. 18/645,542 · Granted Jul 21, 2026

System and method for model-free, one-shot object pose estimation via coordinate regression

Inventors: Jérome Revaud (Meylan, FR); Romain Brégier (Meylan, FR); Yohann Cabon (Meylan, FR); Philippe Weinzaepfel (Meylan, FR); JongMin Lee (Meylan, FR)
Assignee: Naver Corporation
G06T7/74G06T7/50G06V10/774G06V10/806G06V10/82G06V20/70G06T2207/20081G06T2207/20084G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,607
App. No.
18/645,542
Granted
Jul 21, 2026
Kind
B2
Abstract

A computer implemented method and system using an object-agnostic model for predicting a pose of an object in an image receives a query image having a target object therein; receives a set of reference images of the target object from different viewpoints; encodes, using a vision transformer, the received query image and the received set of reference images to generate a set of token features for the received query image and a set of token features for the received set of reference images; extracts, using a transformer decoder, information from the set of token features for the encoded reference images with respect to a set of token features for the received query image; processes, using a prediction head, the combined set of token features to generate a 2D-3D mapping and a confidence map of the query image; and processes the 2D-3D mapping and confidence map to determine the pose of the target object in the query image.

Claims (197)

1 . A computer-implemented method of training a machine learning model for regression on pixel-level annotations in images, the machine learning model comprising an image encoder, a feature mixer, and a decoder, the method comprising:

pre-training the image encoder and the decoder for cross-view completion on images;

constructing training tuples, each training tuple comprising a first image and one or more second images, wherein the first image is associated with dense pixel-level annotations, and wherein each of the one or more second images is associated with sparse pixel-level annotations;

generating, by the image encoder, a set of first image tokens from the first image;

generating, by the image encoder, one or more sets of second image tokens from the one or more second images;

generating, by the feature mixer, one or more sets of augmented second image tokens by augmenting each set of second image tokens with encodings of the respective sparse pixel-level annotations associated with the respective set of second image tokens, wherein the augmenting includes mixing the respective set of second image tokens and the encodings of the respective sparse pixel-level annotations;

processing, by the decoder, the set of first image tokens and the one or more sets of augmented second image tokens to generate prediction data of the machine learning model for the first image, wherein the processing comprises receiving a set of augmented second image tokens of the one or more sets of augmented second image tokens, wherein the prediction data comprises predictions for each image pixel of the first image and confidences for the predictions; and

fine-tuning the machine learning model, wherein the fine-tuning comprises adjusting parameters of the image encoder, the feature mixer, and the decoder to minimize a loss function, wherein the loss function is based on the prediction data and the dense pixel-level annotations.

2 . The method as claimed in claim 1 , wherein the feature mixer comprises a first pipeline of image-level decoders for processing the respective second set of tokens, and a second pipeline of point-level decoders for processing the respective encodings, wherein each image-level decoder comprises a first cross-attention layer, wherein each point-level decoder comprises a second cross-attention layer, wherein the mixing the respective set of second image tokens and the respective encodings comprises each first cross-attention layer receiving information from the second pipeline and each second cross-attention block receiving information from the first pipeline.

3 . The method as claimed in claim 2 , wherein the feature mixer further comprises first linear projection modules configured for processing the information from the second pipeline, and second linear projection modules configured for processing the information from the first pipeline, to account for different set sizes of the encodings and the set of second image tokens.

4 . The method as claimed in claim 3 , wherein the decoder comprises a plurality of cross-attention layers, wherein the number of cross-attention layers matches the number of second images, wherein the processing, by the decoder, the set of first image tokens and the one or more sets of augmented second image tokens comprises processing the set of first image tokens as input of the decoder and providing each set of augmented second image tokens to a respective cross-attention layer of the plurality of cross-attention layers, or

wherein the decoder comprises a single cross-attention layer, wherein the processing, by the decoder, the set of first image tokens and the one or more sets of augmented second image tokens comprises providing one set of augmented second image tokens to the single cross-attention layer at a time, wherein for each set of augmented second image tokens, intermediate prediction data are generated and the prediction data of the machine learning model is based on selecting, as the prediction data for an image pixel, an intermediate prediction for the image pixel having a highest confidence value of the confidences.

5 . The method as claimed in claim 4 , wherein the sparse pixel-level annotations and the dense pixel-level annotations each comprise annotations of image pixels with 3D coordinates of real-world object features corresponding to the image pixels, wherein the sparse pixel-level annotations are scattered over the respective image when compared with the dense pixel-level annotations, and wherein the prediction data for each image pixel of the first image comprises a 3D coordinate for the image pixel.

6 . The method as claimed in claim 5 , further comprising generating the encodings of the pixel-level annotations by a trainable pixel annotation encoder, wherein the pixel annotation encoder is fed with representations of the pixel-level annotations based on an embedding of 3D coordinates in the hypercube [−1,1] d .

7 . The method as claimed in claim 6 , wherein the embedding is an injective projection and has an inverse, wherein the inverse of the embedding is well-defined over the hypercube [−1,1] d .

8 . The method as claimed in claim 7 , wherein the embedding is defined by φ(x))=(ψ(x), (ψ(y), ψ(z)) with

ψ

(

x

)

=

[

cos

(

f

1

x

)

,

sin

(

f

1

x

)

,

cos

(

f

2

x

)

,

sin

(

f

2

x

)

,

]

,

ψ

(

y

)

=

[

cos

(

f

1

y

)

,

sin

(

f

1

y

)

,

cos

(

f

2

y

)

,

sin

(

f

2

y

)

,

]

,

ψ

(

z

)

=

[

cos

(

f

1

z

)

,

sin

(

f

1

z

)

,

cos

(

f

2

z

)

,

sin

(

f

2

z

)

,

]

,

wherein

f

i

=

f

0

γ

i

-

1

,

i

{

1

,

,

d

6

}

,

f

0

>

0

,

and

γ

>

0

.

9 . The method as claimed in claim 4 , wherein the pixel-level annotation relates to optical flow, relates to information for identifying instances of objects, or relates to information for segmentation of views.

10 . The method as claimed in claim 9 , wherein the constructing training tuples comprises selecting the one or more second images from an image database based on selecting easy inliers, hard inliers, and hard outliers, wherein the easy and hard inliers are determined based on a viewpoint angle between camera poses employed to capture the respective images, and the hard outliers are selected as being images most similar to the first image.

11 . The method as claimed in claim 9 , wherein the pre-training of the image encoder along with the decoder for cross-view completion comprises pre-training a pipeline comprising an encoder and a decoder on a pair of pre-training images comprising a first pre-training image and a second pre-training image, wherein the pipeline is trained to reconstruct masked portions of the first pre-training image based on the second pre-training image.

12 . The method as claimed in claim 1 , further comprising using the trained machine learning model for performing one or more of:

inferring a camera pose of a camera used for capturing an unannotated query image;

generating a 3D-reconstruction of a scene depicted in the unannotated query image;

performing 3D completion of a sparsely annotated query image, wherein the sparsely annotated query image is employed as the second image; and

performing dense depth prediction for an unannotated query image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2025
From: REVAUD, JEROME; CABON, YOHANN; BREGIER, ROMAIN; WEINZAEPFEL, PHILIPPE; LEE, JONGMIN
To: NAVER CORPORATION
Reel/Frame 070978/0413 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2024
From: NAVERS LAB CORPORATION
To: NAVER CORPORATION
Reel/Frame 068789/0417 →
Priority Claims (1)
EP 23305867 · Jun 1, 2023 · regional
Continuity (2)
Provisional Application 63542114 · Oct 3, 2023
Related Publication 20240404104A1 · Dec 5, 2024
References Cited (162)
US 12154301B2 · Portman · 2024 [cited by examiner]
US 20230196955A1 · Jang · 2023 [cited by examiner]
US 20240096072A1 · He · 2024 [cited by examiner]
US 20250112088A1 · Lee · 2025 [cited by examiner]
US 20250307552A1 · Ebrahimi · 2025 [cited by examiner]
US 20250342604A1 · Chen · 2025 [cited by examiner]
US 20250378591A1 · Yu · 2025 [cited by examiner]
L. Yang, Z. Bai, C. Tang, H. Li, Y. Furukawa and P. Tan, “SANet: Scene Agnostic Network for Camera Localization,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South), 2019, pp. 42-51, … [cited by examiner]
Arnold, Eduardo, Jamie Wynn, Sara Vicente, Guillermo Garcia-Hernando, Áron Monszpart, Victor Prisacariu, Daniyar Turmukhambetov, and Eric Brachmann. ‘Map-Free Visual Relocalization: Metric Pose Relative to a Single Imag… [cited by applicant]
Bachmann, Roman, David Mizrahi, Andrei Atanov, and Amir Zamir. ‘MultiMAE: Multi-Modal Multi-Task Masked Autoencoders’. In Computer Vision—ECCV 2022, edited by Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Mari… [cited by applicant]
Balntas, Vassileios, Shuda Li, and Victor Prisacariu. ‘RelocNet: Continuous Metric Learning Relocalisation Using Neural Nets’. In Computer Vision—ECCV 2018, edited by Vittorio Ferrari, Martial Hebert, Cristian Sminchise… [cited by applicant]
Baruch, Gilad, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, et al. ARKitScenes: A Diverse Real-World Dataset for 3D Indoor Scene Understanding Using Mobile RGB-D Data. [cited by applicant]
Belghit, Hayet, Abdelkader Bellarbi, Nadia Zenati, and Samir Otmane. ‘Vision-Based Pose Estimation for Augmented Reality□: A Comparison Study’, 35th Conference on Neural Information Processing Systems (NeurIPS 2021). [cited by applicant]
Blanton, Hunter, Connor Greenwell, Scott Workman, and Nathan Jacobs. ‘Extending Absolute Pose Regression to Multiple Scenes’. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 170… [cited by applicant]
Blender Online Community. Blender—a 3D modelling and rendering package. Blender Foundation, 2018. https://www.blender.org/download/. [cited by applicant]
Brachmann, Eric, Alexander Krull, Frank Michel, Stefan Gumhold, Jamie Shotton, and Carsten Rother. ‘Learning 6D Object Pose Estimation Using 3D Object Coordinates’. In Computer Vision—ECCV 2014, edited by David Fleet, T… [cited by applicant]
Brachmann, Eric, Alexander Krull, Sebastian Nowozin, Jamie Shotton, Frank Michel, Stefan Gumhold, and Carsten Rother. ‘DSAC—Differentiable RANSAC for Camera Localization’. In 2017 IEEE Conference on Computer Vision and … [cited by applicant]
Brachmann, Eric, Frank Michel, Alexander Krull, Michael Ying Yang, Stefan Gumhold, and Carsten Rother. ‘Uncertainty-Driven 6D Pose Estimation of Objects and Scenes from a Single RGB Image’. In 2016 IEEE Conference on Co… [cited by applicant]
Brachmann, Eric, and Carsten Rother. ‘Expert Sample Consensus Applied to Camera Re-Localization’. arXiv, Aug. 7, 2019. http://arxiv.org/abs/1908.02484. [cited by applicant]
Brachmann, Eric et al. ‘Learning Less Is More—6D Camera Localization via 3D Surface Regression’. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4654-62. Salt Lake City, UT: IEEE, 2018. https://d… [cited by applicant]
Brachmann, Eric et al. ‘Visual Camera Re-Localization from RGB and RGB-D Images Using DSAC’. arXiv, Oct. 9, 2020. http://arxiv.org/abs/2002.12324. [cited by applicant]
Brahmbhatt, Samarth, Jinwei Gu, Kihwan Kim, James Hays, and Jan Kautz. ‘Geometry-Aware Learning of Maps for Camera Localization’. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2616-25. Salt Lak… [cited by applicant]
Brégier, Romain. ‘Deep Regression on Manifolds: A 3D Rotation Case Study’. arXiv, Oct. 12, 2021. https://doi.org/10.48550/arXiv.2103.16317. [cited by applicant]
Cai, Dingding, Janne Heikkilä, and Esa Rahtu. ‘OVE6D: Object Viewpoint Encoding for Depth-Based 6D Object Pose Estimation’. arXiv, Apr. 7, 2022. https://doi.org/10.48550/arXiv.2203.01072. [cited by applicant]
Camposeco, Federico, Andrea Cohen, Marc Pollefeys, and Torsten Sattler. ‘Hybrid Scene Compression for Visual Localization’. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7645-54. Long Be… [cited by applicant]
Cao, Song, and Noah Snavely. ‘Minimal Scene Descriptions from Structure from Motion Models’. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, 461-68. Columbus, OH, USA: IEEE, 2014. https://doi.org/10.… [cited by applicant]
Chen, Dengsheng, Jun Li, Zheng Wang, and Kai Xu. ‘Learning Canonical Shape Space for Category-Level 6D Object Pose and Size Estimation’. arXiv, Nov. 21, 2021. https://doi.org/10.48550/arXiv.2001.09322. [cited by applicant]
Chen, Hansheng, Pichao Wang, Fan Wang, Wei Tian, Lu Xiong, and Hao Li. ‘EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose Estimation’. arXiv, Aug. 11, 2022. https://doi.org/10… [cited by applicant]
Chen, Kai, and Qi Dou. ‘SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose Estimation’. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2753-62. Montreal, QC, Canada: IEEE, 20… [cited by applicant]
Chen, Wei, Xi Jia, Hyung Jin Chang, Jinming Duan, Linlin Shen, and Ales Leonardis. ‘FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation with Decoupled Rotation Mechanism’. arXiv, Jun. 6, 2021. … [cited by applicant]
Chen, Xu, Zijian Dong, Jie Song, Andreas Geiger, and Otmar Hilliges. ‘Category Level Object Pose Estimation via Neural Analysis-by-Synthesis’. arXiv, Aug. 18, 2020. https://doi.org/10.48550/arXiv.2008.08145. [cited by applicant]
Cheng, Wentao, Weisi Lin, Kan Chen, and Xinfeng Zhang. ‘Cascaded Parallel Filtering for Memory-Efficient Image-Based Localization’. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 1032-41. Seoul, Ko… [cited by applicant]
Cheng, Wentao, Weisi Lin, Xinfeng Zhang, Michael Goesele, and Ming-Ting Sun. ‘A Data-Driven Point Cloud Simplification Framework for City-Scale Image-Based Localization’. IEEE Transactions on Image Processing 26, No. 1 … [cited by applicant]
Cheng, Xinjing, Peng Wang, Chenye Guan, and Ruigang Yang. ‘CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth Completion’. Proceedings of the AAAI Conference on Artificial I… [cited by applicant]
Cipolla, Roberto, Yarin Gal, and Alex Kendall. ‘Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics’. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7482-91. S… [cited by applicant]
Collins, Jasmine, Shubham Goel, Kenan Deng, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Xi Zhang, et al. ‘ABO: Dataset and Benchmarks for Real-World 3D Object Understanding’. In 2022 IEEE/CVF Conference on Computer Visi… [cited by applicant]
Dai, Angela, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Niessner. ‘ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes’. In 2017 IEEE Conference on Computer Vision and Patter… [cited by applicant]
Deng, Xinke, Yu Xiang, Arsalan Mousavian, Clemens Eppner, Timothy Bretl, and Dieter Fox. ‘Self-Supervised 6D Object Pose Estimation for Robot Manipulation’. arXiv, Mar. 7, 2020. https://doi.org/10.48550/arXiv.1909.10159. [cited by applicant]
DeTone, Daniel, Tomasz Malisiewicz, and Andrew Rabinovich. ‘SuperPoint: Self-Supervised Interest Point Detection and Description’. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)… [cited by applicant]
Ding, Mingyu, Zhe Wang, Jiankai Sun, Jianping Shi, and Ping Luo. ‘CamNet: Coarse-to-Fine Retrieval for Camera Re-Localization’. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2871-80. Seoul, Korea … [cited by applicant]
Do, Thanh-Toan. ‘Real-Time Monocular Object Instance 6D Pose Estimation’, 29th British Machine Vision Conference, BMVC 2018. British Machine Vision Association and Society for Pattern Recognition. British MachineVision … [cited by applicant]
Dong, Siyan, Shuzhe Wang, Yixin Zhuang, Juho Kannala, Marc Pollefeys, and Baoquan Chen. ‘Visual Localization via Few-Shot Scene Region Classification’. arXiv, Aug. 14, 2022. http://arxiv.org/abs/2208.06933. [cited by applicant]
Dosovitskiy, Alexey, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, et al. ‘An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale’, 2021. ht… [cited by applicant]
Dusmanu, Mihai, Ignacio Rocco, Tomas Pajdla, Marc Pollefeys, Josef Sivic, Akihiko Torii, and Torsten Sattler. ‘D2-Net: A Trainable CNN for Joint Detection and Description of Local Features’. arXiv, May 9, 2019. http://a… [cited by applicant]
Marcin Dymczyk, Simon Lynen, Michael Bosse, and Roland Siegwart. Keep it brief: Scalable creation of compressed localization maps. In IROS, 2015. [cited by applicant]
Eldar, Yuval, Michael Lindenbaum, and Moshe Porat. ‘The Farthest Point Strategy for Progressive Image Sampling’. IEEE Transactions on Image Processing 6, No. 9 (1997). [cited by applicant]
Fischler, Martin A, and Robert C Bolles. ‘Random Sample Consensus’ 24, No. 6 (1981). Communications of the ACM, 1981 https://dl.acm.org/doi/pdf/10.1145/358669.358692. [cited by applicant]
Gou, Minghao, Haolin Pan, Hao-Shu Fang, Ziyuan Liu, Cewu Lu, and Ping Tan. ‘Unseen Object 6D Pose Estimation: A Benchmark and Baselines’. arXiv, Jun. 23, 2022. https://doi.org/10.48550/arXiv.2206.11808. [cited by applicant]
Robert M. Gray and David L. Neuhoff. Quantization. IEEE Trans. Inf. Theory, 1998. [cited by applicant]
Guzman-Rivera, Abner, Pushmeet Kohli, Ben Glocker, Jamie Shotton, Toby Sharp, Andrew Fitzgibbon, and Shahram Izadi. ‘Multi-Output Learning for Camera Relocalization’. In 2014 IEEE Conference on Computer Vision and Patte… [cited by applicant]
He, Kaiming, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. ‘Masked Autoencoders Are Scalable Vision Learners’. arXiv, Dec. 19, 2021. https://doi.org/10.48550/arXiv.2111.06377. [cited by applicant]
He, Xingyi, Jiaming Sun, Yuang Wang, Di Huang, Hujun Bao, and Xiaowei Zhou. ‘OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD Models’. arXiv, Jan. 18, 2023. https://doi.org/10.48550/arXiv.2301.07673. [cited by applicant]
He, Yisheng, Yao Wang, Haoqiang Fan, Jian Sun, and Qifeng Chen. ‘FS6D: Few-Shot 6D Pose Estimation of Novel Objects’. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6804-14. New Orleans, … [cited by applicant]
Heinly, Jared, Johannes L. Schonberger, Enrique Dunn, and Jan-Michael Frahm. ‘Reconstructing the World in Six Days’. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3287-95. Boston, MA, USA: I… [cited by applicant]
Henriques, João F, Pedro Martins, Rui F Caseiro, and Jorge Batista. ‘Fast Training of Pose Detectors in the Fourier Domain’. Advances in Neural Information Processing Systems, vol. 27.https://openaccess.thecvf.com/conte… [cited by applicant]
Hietanen, Antti, Jyrki Latokartano, Alessandro Foi, Roel Pieters, Ville Kyrki, Minna Lanz, and Joni-Kristian Kämäräinen. ‘Benchmarking Pose Estimation for Robot Manipulation’. Robotics and Autonomous Systems 143 (Sep. 2… [cited by applicant]
Hietanen, Antti et al. ‘Object Pose Estimation in Robotics Revisited’. arXiv, May 21, 2020. https://doi.org/10.48550/arXiv.1906.02783. [cited by applicant]
Hinterstoisser, S., C. Cagniart, S. Ilic, P. Sturm, N. Navab, P. Fua, and V. Lepetit. ‘Gradient Response Maps for Real-Time Detection of Textureless Objects’. IEEE Transactions on Pattern Analysis and Machine Intelligen… [cited by applicant]
Zhang, Zichao, Torsten Sattler, and Davide Scaramuzza. ‘Reference Pose Generation for Long-Term Visual Localization via Learned Features and View Synthesis’. International Journal of Computer Vision 129, No. 4 (Apr. 202… [cited by applicant]
Zhou, Lei, Zixin Luo, Tianwei Shen, Jiahui Zhang, Mingmin Zhen, Yao Yao, Tian Fang, and Long Quan. ‘KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering’. In 2020 IEEE/CVF Conference on Computer Vision … [cited by applicant]
Zhou, Qunjie, Torsten Sattler, Marc Pollefeys, and Laura Leal-Taixe. ‘To Learn or Not to Learn: Visual Localization from Essential Matrices’. arXiv, Mar. 9, 2020. https://doi.org/10.48550/arXiv.1908.01293. [cited by applicant]
Ultralytics/yolov5: v3.1—Bug Fixes and Performance Improvements https://zenodo.org/records/4154370. [cited by applicant]
Revaud, Jerome, Philippe Weinzaepfel, César De Souza, Noe Pion, Gabriela Csurka, Yohann Cabon, and Martin Humenberger. ‘R2D2: Repeatable and Reliable Detector and Descriptor’. ArXiv:1906.06195 [Cs], Jun. 17, 2019. http:… [cited by applicant]
Rios-Cabrera, Reyes, and Tinne Tuytelaars. ‘Discriminatively Trained Templates for 3D Object Detection: A Real Time Scalable Approach’. In 2013 IEEE International Conference on Computer Vision, 2048-55. Sydney, Australi… [cited by applicant]
Rublee, Ethan, Vincent Rabaud, Kurt Konolige, and Gary Bradski. ‘ORB: An Efficient Alternative to SIFT or SURF’. In 2011 International Conference on Computer Vision, 2564-71. Barcelona, Spain: IEEE, 2011. https://doi.or… [cited by applicant]
Sarlin, Paul-Edouard, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. ‘From Coarse to Fine: Robust Hierarchical Localization at Large Scale’. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV… [cited by applicant]
Sarlin, Paul-Edouard, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. ‘SuperGlue: Learning Feature Matching with Graph Neural Networks’. arXiv, Mar. 28, 2020. https://doi.org/10.48550/arXiv.1911.11763. [cited by applicant]
Sattler, Torsten, Michal Havlena, Filip Radenovic, Konrad Schindler, and Marc Pollefeys. ‘Hyperpoints and Fine Vocabularies for Large-Scale Location Recognition’. In 2015 IEEE International Conference on Computer Vision… [cited by applicant]
Sattler, Torsten, Bastian Leibe, and Leif Kobbelt. ‘Efficient & Effective Prioritized Matching for Large-Scale Image-Based Localization’. IEEE Transactions on Pattern Analysis and Machine Intelligence 39, No. 9 (Sep. 1,… [cited by applicant]
Sattler, Torsten, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, et al. ‘Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions’. In 2018 IEEE/CVF Conference on Co… [cited by applicant]
Savva, Manolis, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, et al. ‘Habitat: A Platform for Embodied AI Research’. In 2019 IEEE/CVF International Conference on Computer Vi… [cited by applicant]
Schonberger, Johannes L., and Jan-Michael Frahm. ‘Structure-from-Motion Revisited’. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4104-13. Las Vegas, NV, USA: IEEE, 2016. https://doi.org/10.… [cited by applicant]
Shi, Wenzhe, Jose Caballero, Ferenc Huszár, Johannes Totz, Andrew P. Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. ‘Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neu… [cited by applicant]
Shotton, Jamie, Ben Glocker, Christopher Zach, Shahram Izadi, Antonio Criminisi, and Andrew Fitzgibbon. ‘Scene Coordinate Regression Forests for Camera Relocalization in RGB-D Images’. In 2013 IEEE Conference on Compute… [cited by applicant]
Shugurov, Ivan, Fu Li, Benjamin Busam, and Slobodan Ilic. ‘OSOP: A Multi-Stage One Shot Object Pose Estimation Framework’. arXiv, Mar. 30, 2022. https://doi.org/10.48550/arXiv.2203.15533. [cited by applicant]
Snavely, Noah, Steven M. Seitz, and Richard Szeliski. ‘Modeling the World from Internet Photo Collections’. International Journal of Computer Vision 80, No. 2 (Nov. 2008): 189-210. https://doi.org/10.1007/s11263-007-010… [cited by applicant]
Song, Jiaru. ‘Sliding Window Filter Based Unknown Object Pose Estimation’. In 2017 IEEE International Conference on Image Processing (ICIP), 2642-46. Beijing: IEEE, 2017. https://doi.org/10.1109/ICIP.2017.8296761. [cited by applicant]
Straub, Julian, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, et al. ‘The Replica Dataset: A Digital Replica of Indoor Spaces’. arXiv, Jun. 13, 2019. http://arxiv.org/abs/1906.05797. [cited by applicant]
Su, Jianlin, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. ‘RoFormer: Enhanced Transformer with Rotary Position Embedding’. arXiv, Nov. 8, 2023. http://arxiv.org/abs/2104.09864. [cited by applicant]
Sun, Jiaming, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. ‘LoFTR: Detector-Free Local Feature Matching with Transformers’. arXiv, Apr. 1, 2021. https://doi.org/10.48550/arXiv.2104.00680. [cited by applicant]
Sun, Jiaming, Zihao Wang, Siyu Zhang, Xingyi He, Hongcheng Zhao, Guofeng Zhang, and Xiaowei Zhou. ‘OnePose: One-Shot Object Pose Estimation without CAD Models’, https://arxiv.org/pdf/2205.12257. [cited by applicant]
Szot, Andrew, Alex Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, et al. ‘Habitat 2.0: Training Home Assistants to Rearrange Their Habitat’, In NeurIPS, 2021 . https://proceedings.neurips.c… [cited by applicant]
Tang, Shitao, Chengzhou Tang, Rui Huang, Siyu Zhu, and Ping Tan. ‘Learning Camera Localization via Dense Scene Matching’. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1831-41. Nashville… [cited by applicant]
Tang et al. ‘NeuMap: Neural Coordinate Mapping by Auto-Transdecoder for Camera Localization’. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 929-39. Vancouver, BC, Canada: IEEE, 2023. htt… [cited by applicant]
Tejani, Alykhan, Danhang Tang, Rigas Kouskouridas, and Tae-Kyun Kim. ‘Latent-Class Hough Forests for 3D Object Detection and Pose Estimation’, Computer Vision—ECCV 2014, 462-477.Cham: Springer International Publishing h… [cited by applicant]
Terzakis, George, and Manolis Lourakis. ‘A Consistently Fast and Globally Optimal Solution to the Perspective-n-Point Problem’. In Computer Vision—ECCV 2020, edited by Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan… [cited by applicant]
Tian, Meng, Marcelo H. Ang Jr, and Gim Hee Lee. ‘Shape Prior Deformation for Categorical 6D Object Pose and Size Estimation’. arXiv, Jul. 16, 2020. https://doi.org/10.48550/arXiv.2007.08454. [cited by applicant]
Tian, Yurun, Xin Yu, Bin Fan, Fuchao Wu, Huub Heijnen, and Vassileios Balntas. ‘SOSNet: Second Order Similarity Regularization for Local Descriptor Learning’. In 2019 IEEE/CVF Conference on Computer Vision and Pattern R… [cited by applicant]
Tolias, Giorgos, Tomas Jenicek, and Ondřej Chum. ‘Learning and Aggregating Deep Local Descriptors for Instance-Level Recognition’. In Computer Vision—ECCV 2020, edited by Andrea Vedaldi, Horst Bischof, Thomas Brox, and … [cited by applicant]
Tyszkiewicz, Michał J, Pascal Fua, and Eduard Trulls. DISK: Learning Local Features with Policy Gradient https://proceedings.neurips.cc/paper_files/paper/2020/file/a42a596fc71e17828440030074d15e74-Paper.pdf. [cited by applicant]
Valentin, Julien, Matthias Niebner, Jamie Shotton, Andrew Fitzgibbon, Shahram Izadi, and Philip Torr. ‘Exploiting Uncertainty in Regression Forests for Accurate Camera Relocalization’. In 2015 IEEE Conference on Compute… [cited by applicant]
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. ‘Attention Is All You Need’. In Advances in Neural Information Processing Systems 30, edited … [cited by applicant]
Walch, F., C. Hazirbas, L. Leal-Taixe, T. Sattler, S. Hilsenbeck, and D. Cremers. ‘Image-Based Localization Using LSTMs for Structured Feature Correlation’. In 2017 IEEE International Conference on Computer Vision (ICCV… [cited by applicant]
Wang, Angtian, Shenxiao Mei, Alan Yuille, and Adam Kortylewski. ‘Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D Pose’. arXiv, Oct. 27, 2021. https://doi.org/10.48550/arXiv.2110.14213. [cited by applicant]
Wang, Bing, Changhao Chen, Chris Xiaoxuan Lu, Peijun Zhao, Niki Trigoni, and Andrew Markham. ‘AtLoc: Attention Guided Camera Localization’. Proceedings of the AAAI Conference on Artificial Intelligence 34, No. 06 (Apr. … [cited by applicant]
Wang, Gu, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. ‘GDR-Net: Geometry-Guided Direct Regression Network for Monocular 6D Object Pose Estimation’. arXiv, Mar. 9, 2021. https://doi.org/10.48550/arXiv.2102.12145. [cited by applicant]
Wang, He, Srinath Sridhar, Jingwei Huang, Julien Valentin, Shuran Song, and Leonidas J. Guibas. ‘Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation’. arXiv, Jun. 23, 2019. https://d… [cited by applicant]
Wang, Jiaze, Kai Chen, and Qi Dou. ‘Category-Level 6D Object Pose Estimation via Cascaded Relation and Recurrent Reconstruction Networks’. arXiv, Aug. 19, 2021. https://doi.org/10.48550/arXiv.2108.08755. [cited by applicant]
Wang, Qianqian, Xiaowei Zhou, Bharath Hariharan, and Noah Snavely. ‘Learning Feature Descriptors Using Camera Pose Supervision’. In Computer Vision—ECCV 2020, edited by Andrea Vedaldi, Horst Bischof, Thomas Brox, and Ja… [cited by applicant]
Weinzaepfel, Philippe, Vincent Leroy, Thomas Lucas, Romain Brégier, Yohann Cabon, Vaibhav Arora, Leonid Antsfeld, Boris Chidlovskii, Gabriela Csurka, and Jérôme Revaud. ‘CroCo: Self-Supervised Pre-Training for 3D Vision… [cited by applicant]
Weinzaepfel, Philippe, Thomas Lucas, Diane Larlus, and Yannis Kalantidis. ‘Learning Super-Features for Image Retrieval’, ICLR 2022. https://openreview.net/pdf?id=wogsFPHwftY. [cited by applicant]
Weinzaepfel, Philippe, Thomas Lucas, Vincent Leroy, Yohann Cabon, Vaibhav Arora, Romain Brégier, Gabriela Csurka, Leonid Antsfeld, Boris Chidlovskii, and Jerome Revaud. ‘CroCo v2: Improved Cross-View Completion Pre-Trai… [cited by applicant]
Weinzaepfel, Philippe et al. ‘CroCo v2: Improved Cross-View Completion Pre-Training for Stereo Matching and Optical Flow’. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 17923-34. Paris, France: IE… [cited by applicant]
Wen, Bowen, and Kostas Bekris. ‘BundleTrack: 6D Pose Tracking for Novel Objects without Instance or Category-Level 3D Models’. arXiv, Aug. 1, 2021. https://doi.org/10.48550/arXiv.2108.00516. [cited by applicant]
Wu, Xin, Hao Zhao, Shunkai Li, Yingdian Cao, and Hongbin Zha. ‘SC-WLS: Towards Interpretable Feed-Forward Camera Re-Localization’. In Computer Vision—ECCV 2022, edited by Shai Avidan, Gabriel Brostow, Moustapha Cissé, G… [cited by applicant]
Xiang, Yu, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. ‘PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes’. arXiv, May 26, 2018. https://doi.org/10.48550/arXiv.1711.001… [cited by applicant]
Xie, Tao, Kun Dai, Ke Wang, Ruifeng Li, and Lijun Zhao. ‘DeepMatcher: A Deep Transformer-Based Network for Robust and Accurate Local Feature Matching’. arXiv, Jan. 8, 2023. https://doi.org/10.48550/arXiv.2301.02993. [cited by applicant]
Yang, Luwei, Ziqian Bai, Chengzhou Tang, Honghua Li, Yasutaka Furukawa, and Ping Tan. ‘SANet: Scene Agnostic Network for Camera Localization’. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 42-51. … [cited by applicant]
Yang, Luwei, Rakesh Shrestha, Wenbo Li, Shuaicheng Liu, Guofeng Zhang, Zhaopeng Cui, and Ping Tan. ‘SceneSqueezer: Learning to Compress Scene for Camera Relocalization’. In 2022 IEEE/CVF Conference on Computer Vision an… [cited by applicant]
Yen-Chen, Lin, Pete Florence, Jonathan T. Barron, Alberto Rodriguez, Phillip Isola, and Tsung-Yi Lin. ‘INeRF: Inverting Neural Radiance Fields for Pose Estimation’. arXiv, Aug. 10, 2021. https://doi.org/10.48550/arXiv.2… [cited by applicant]
Yi, Kwang Moo, Eduard Trulls, Vincent Lepetit, and Pascal Fua. ‘LIFT: Learned Invariant Feature Transform’. arXiv, Jul. 29, 2016. https://doi.org/10.48550/arXiv.1603.09114. [cited by applicant]
Zakharov, Sergey, Ivan Shugurov, and Slobodan Ilic. ‘DPOD: 6D Pose Object Detector and Refiner’. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 1941-50. Seoul, Korea (South): IEEE, 2019. https://do… [cited by applicant]
Hinterstoisser, Stefan, Stefan Holzer, Cedric Cagniart, Slobodan Ilic, Kurt Konolige, Nassir Navab, and Vincent Lepetit. ‘Multimodal Templates for Real-Time Detection of Texture-Less Objects in Heavily Cluttered Scenes’… [cited by applicant]
Hinterstoisser, Stefan, Vincent Lepetit, Slobodan Ilic, Stefan Holzer, Gary Bradski, Kurt Konolige, and Nassir Navab. ‘Model Based Training, Detection and Pose Estimation of Texture-Less 3D Objects in Heavily Cluttered … [cited by applicant]
Hodan, Tomas, Daniel Barath, and Jiri Matas. ‘EPOS: Estimating 6D Pose of Objects with Symmetries’. arXiv, Apr. 1, 2020. https://doi.org/10.48550/arXiv.2004.00605. [cited by applicant]
Hodan, Tomas, Frank Michel, Eric Brachmann, Wadim Kehl, Anders Glent Buch, Dirk Kraft, Bertram Drost, et al. ‘BOP: Benchmark for 6D Object Pose Estimation’. arXiv, Aug. 24, 2018. https://doi.org/10.48550/arXiv.1808.0831… [cited by applicant]
Hodan et al. EPOS: Estimating 6D Pose of Objects with Symmetries https://arxiv.org/pdf/2004.00605. [cited by applicant]
Hu, Mu, Shuling Wang, Bin Li, Shiyu Ning, Li Fan, and Xiaojin Gong. ‘PENet: Towards Precise and Efficient Image Guided Depth Completion’. arXiv, Mar. 18, 2021. https://doi.org/10.48550/arXiv.2103.00783. [cited by applicant]
Huang, Zhaoyang, Han Zhou, Yijin Li, Bangbang Yang, Yan Xu, Xiaowei Zhou, Hujun Bao, Guofeng Zhang, and Hongsheng Li. ‘VS-Net: Voting with Segmentation for Visual Localization’. In 2021 IEEE/CVF Conference on Computer V… [cited by applicant]
Humenberger, Martin, Yohann Cabon, Nicolas Guerin, Julien Morat, Vincent Leroy, Jerome Revaud, Philippe Rerole, Noé Pion, Cesar de Souza, and Gabriela Csurka. ‘Robust Image Retrieval-Based Visual Localization Using Kapt… [cited by applicant]
Hutchison, David, Takeo Kanade, Josef Kittler, Jon M. Kleinberg, Friedemann Mattern, John C. Mitchell, Moni Naor, et al. ‘Location Recognition Using Prioritized Feature Matching’. In Computer Vision—ECCV 2010, edited by… [cited by applicant]
Iwase, Shun, Xingyu Liu, Rawal Khirodkar, Rio Yokota, and Kris M. Kitani. ‘RePOSE: Fast 6D Object Pose Refinement via Deep Texture Rendering’. arXiv, Aug. 19, 2021. https://doi.org/10.48550/arXiv.2104.00633. [cited by applicant]
Jégou, H, M Douze, and C Schmid. ‘Product Quantization for Nearest Neighbor Search’. IEEE Transactions on Pattern Analysis and Machine Intelligence 33, No. 1 (Jan. 2011): 117-28. https://doi.org/10.1109/TPAMI.2010.57. [cited by applicant]
Kehl, Wadim, Fabian Manhardt, Federico Tombari, Slobodan Ilic, and Nassir Navab. ‘SSD-6D: Making RGB-Based 3D Detection and 6D Pose Estimation Great Again’. arXiv, Nov. 27, 2017. https://doi.org/10.48550/arXiv.1711.1000… [cited by applicant]
Kehl, Wadim, Federico Tombari, Nassir Navab, Slobodan Ilic, and Vincent Lepetit. ‘Hashmod: A Hashing Method for Scalable 3D Object Detection’. arXiv, Jul. 20, 2016. https://doi.org/10.48550/arXiv.1607.06062. [cited by applicant]
Kendall, Alex, and Roberto Cipolla. ‘Geometric Loss Functions for Camera Pose Regression with Deep Learning’. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6555-64. Honolulu, HI: IEEE, 2017.… [cited by applicant]
Kendall, Alex, Matthew Grimes, and Roberto Cipolla. ‘PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization’. In 2015 IEEE International Conference on Computer Vision (ICCV), 2938-46. Santiago, Chile… [cited by applicant]
Lee, JongMin, Yohann Cabon, Romain Brégier, Sungjoo Yoo, and Jerome Revaud. ‘MFOS: Model-Free & One-Shot Object Pose Estimation’. arXiv, Oct. 3, 2023. https://doi.org/10.48550/arXiv.2310.01897. [cited by applicant]
Lee et al. MFOS: Model-Free & One-Shot Object Pose Estimation. Proceedings of the AAAI Conference on Artificial Intelligence 38, No. 4 (Mar. 24, 2024): 2911-19. https://doi.org/10.1609/aaai.v38i4.28072. [cited by applicant]
Lee, Taeyeop, Byeong-Uk Lee, Myungchul Kim, and In So Kweon. ‘Category-Level Metric Scale Object Shape and Pose Estimation’. IEEE Robotics and Automation Letters 6, No. 4 (Oct. 2021): 8575-82. https://doi.org/10.1109/LR… [cited by applicant]
Lepetit, Vincent, Francesc Moreno-Noguer, and Pascal Fua. ‘EPnP: An Accurate O(n) Solution to the PnP Problem’. International Journal of Computer Vision 81, No. 2 (Feb. 2009): 155-66. https://doi.org/10.1007/s11263-008-… [cited by applicant]
Li, Kunhong, Longguang Wang, Li Liu, Qing Ran, Kai Xu, and Yulan Guo. ‘Decoupling Makes Weakly Supervised Local Feature Better’. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15817-27. N… [cited by applicant]
Li, Xiaotian, Shuzhe Wang, Yi Zhao, Jakob Verbeek, and Juho Kannala. ‘Hierarchical Scene Coordinate Classification and Regression for Visual Localization’. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Reco… [cited by applicant]
Li, Zhengqi, and Noah Snavely. ‘MegaDepth: Learning Single-View Depth Prediction from Internet Photos’. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2041-50. Salt Lake City, UT, USA: IEEE, 201… [cited by applicant]
Li, Zhigang, and Xiangyang Ji. ‘Pose-Guided Auto-Encoder and Feature-Based Refinement for 6-DoF Object Pose Regression’. In 2020 IEEE International Conference on Robotics and Automation (ICRA), 8397-8403. Paris, France:… [cited by applicant]
Li, Zhigang, Gu Wang, and Xiangyang Ji. ‘CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose Estimation’. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 7677… [cited by applicant]
Liu, Yuan, Yilin Wen, Sida Peng, Cheng Lin, Xiaoxiao Long, Taku Komura, and Wenping Wang. ‘Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images’. arXiv, Jan. 27, 2023. https://doi.org/10.48550/ar… [cited by applicant]
Liu, Zhuang, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. ‘A ConvNet for the 2020s’, CVPR, 2022 https://openaccess.thecvf.com/content/CVPR2022/papers/Liu_A_ConvNet_for_the_2020s_CVP… [cited by applicant]
Loshchilov, Ilya, and Frank Hutter. ‘Decoupled Weight Decay Regularization’, . ICLR 2019. [cited by applicant]
Lowe, David G. ‘Distinctive Image Features from Scale-Invariant Keypoints’. International Journal of Computer Vision 60, No. 2 (Nov. 2004): 91-110. https://doi.org/10.1023/B:VISI.0000029664.99615.94. [cited by applicant]
Luo, Zixin, Tianwei Shen, Lei Zhou, Siyu Zhu, Runze Zhang, Yao Yao, Tian Fang, and Long Quan. ‘GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints’. In Computer Vision—ECCV 2018, edited by Vittorio F… [cited by applicant]
Luo, Zixin, Lei Zhou, Xuyang Bai, Hongkai Chen, Jiahui Zhang, Yao Yao, Shiwei Li, Tian Fang, and Long Quan. ‘ASLFeat: Learning Local Features of Accurate Shape and Localization’. In 2020 IEEE/CVF Conference on Computer … [cited by applicant]
Macqueen, J. ‘Some Methods for Classification and Analysis of Multivariate Observations’. Multivariate Observations In Proc. of the fifth Berkeley Symposium on Mathematical Statistics and Probability, 1967. Replaced by … [cited by applicant]
Marchand, Eric, Hideaki Uchiyama, and Fabien Spindler. ‘Pose Estimation for Augmented Reality: A Hands-On Survey’. IEEE Transactions on Visualization and Computer Graphics 22, No. 12 (Dec. 1, 2016): 2633-51. https://doi… [cited by applicant]
Mera-Trujillo, Marcela, Benjamin Smith, and Victor Fragoso. ‘Efficient Scene Compression for Visual-Based Localization’. In 2020 International Conference on 3D Vision (3DV), 1-10. Fukuoka, Japan: IEEE, 2020. https://doi… [cited by applicant]
Mildenhall, Ben, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. ‘NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis’. arXiv, Aug. 3, 2020. https://doi.org/10.… [cited by applicant]
Nazir, Danish, Alain Pagani, Marcus Liwicki, Didier Stricker, and Muhammad Zeshan Afzal. ‘SemAttNet: Toward Attention-Based Semantic Aware Guided Depth Completion’. IEEE Access 10 (2022): 120781-91. https://doi.org/10.1… [cited by applicant]
Olson, Edwin. ‘AprilTag: A Robust and Flexible Visual Fiducial System’. In 2011 IEEE International Conference on Robotics and Automation, 3400-3407. Shanghai, China: IEEE, 2011. https://doi.org/10.1109/ICRA.2011.5979561. [cited by applicant]
Ono, Yuki, Eduard Trulls, Pascal Fua, and Kwang Moo Yi. ‘LF-Net: Learning Local Features from Images’. In NeurIPS, 2018 https://papers.nips.cc/paper_files/paper/2018/file/f5496252609c43eb8a3d147ab9b9c006-Paper.pdf. [cited by applicant]
Park, Hyun Soo, Yu Wang, Eriko Nurvitadhi, James C. Hoe, Yaser Sheikh, and Mei Chen. ‘3D Point Cloud Reduction Using Mixed-Integer Quadratic Programming’. In 2013 IEEE Conference on Computer Vision and Pattern Recogniti… [cited by applicant]
Park, Jinsun, Kyungdon Joo, Zhe Hu, Chi-Kuei Liu, and In So Kweon. ‘Non-Local Spatial Propagation Network for Depth Completion’. In Computer Vision—ECCV 2020, edited by Andrea Vedaldi, Horst Bischof, Thomas Brox, and Ja… [cited by applicant]
Park, Keunhong, Arsalan Mousavian, Yu Xiang, and Dieter Fox. ‘LatentFusion: End-to-End Differentiable Reconstruction and Rendering for Unseen Object Pose Estimation’. arXiv, Jun. 12, 2020. https://doi.org/10.48550/arXiv… [cited by applicant]
Park, Kiru, Timothy Patten, and Markus Vincze. ‘Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose Estimation’. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 7667-76. Seoul, Korea (… [cited by applicant]
Pavllo, Dario, David Joseph Tan, Marie-Julie Rakotosaona, and Federico Tombari. ‘Shape, Pose, and Appearance from a Single Image via Bootstrapped Radiance Field Inversion’. arXiv, Mar. 20, 2023. https://doi.org/10.48550… [cited by applicant]
Peng, Sida, Yuan Liu, Qixing Huang, Hujun Bao, and Xiaowei Zhou. ‘PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation’. arXiv, Dec. 31, 2018. https://doi.org/10.48550/arXiv.1812.11788. [cited by applicant]
Rad, Mahdi, and Vincent Lepetit. ‘BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth’. arXiv, Mar. 26, 2018. https://doi.org/10.48550/arX… [cited by applicant]
Ramakrishnan, Santhosh K, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alex Clegg, John Turner, Eric Undersander, et al. ‘Habitat-Matterport 3D Dataset (HM3D): 1000 Large-Scale 3D Environments for Embodied AI’, In… [cited by applicant]
Ranftl, Rene, Alexey Bochkovskiy, and Vladlen Koltun. ‘Vision Transformers for Dense Prediction’. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 12159-68. Montreal, QC, Canada: IEEE, 2021. https://… [cited by applicant]
Revaud, J., G. Lavoue, and A. Baskurt. ‘Improving Zernike Moments Comparison for Optimal Similarity and Rotation Angle Retrieval’. IEEE Transactions on Pattern Analysis and Machine Intelligence 31, No. 4 (Apr. 2009): 62… [cited by applicant]
Revaud, Jerome, Jon Almazan, Rafael Rezende, and Cesar De Souza. ‘Learning With Average Precision: Training Image Retrieval With a Listwise Loss’. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 510… [cited by applicant]
Revaud, Jerome, Yohann Cabon, Romain Brégier, JongMin Lee, and Philippe Weinzaepfel. ‘SACReg: Scene-Agnostic Coordinate Regression for Visual Localization’. arXiv, Nov. 30, 2023. http://arxiv.org/abs/2307.11702. [cited by applicant]
Revaud and al. ‘SACReg: Scene-Agnostic Coordinate Regression for Visual Localization’. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 688-98. Seattle, WA, USA: IEEE, 2024. http… [cited by applicant]