IP Library Granted Patent US 12,400,346
Granted Patent B2
US 12,400,346 · App. 17/855,330 · Granted Aug 26, 2025

Methods and apparatus for metric depth estimation using a monocular visual-inertial system

Inventors: Diana Wofk (Santa Clara, CA); Rene Ranftl (Munich, DE); Matthias Mueller (Munich, DE); Vladlen Koltun (Santa Clara, CA)
Assignee: Intel Corporation
G06T7/50G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,346
App. No.
17/855,330
Granted
Aug 26, 2025
Kind
B2
Abstract

Methods, apparatus, systems, and articles of manufacture are disclosed for metric depth estimation using a monocular visual-inertial system. An example apparatus for metric depth estimation includes at least one memory, instructions in the apparatus, and processor circuitry to execute the instructions to access a globally-aligned depth prediction, the globally-aligned depth prediction generated based on a monocular depth estimator, access a dense scale map scaffolding, the dense scale map scaffolding generated based on visual-inertial odometry, regress a dense scale residual map determined using the globally-aligned depth prediction and the dense scale map scaffolding, and apply the dense scale residual map to the globally-aligned depth prediction.

Claims (44)

1. An apparatus for metric depth estimation, the apparatus comprising:

at least one memory;

instructions; and

processor circuitry to execute the instructions to:

access a globally-aligned depth prediction, the globally-aligned depth prediction generated based on a monocular depth estimator;

access a dense scale map scaffolding, the dense scale map scaffolding generated based on visual-inertial odometry;

regress a dense scale residual map determined using the globally-aligned depth prediction and the dense scale map scaffolding; and

apply the dense scale residual map to the globally-aligned depth prediction.

2. The apparatus of claim 1 , wherein the visual-inertial odometry determines the dense scale map scaffolding based on inertial measurement unit (IMU) data and visual data.

3. The apparatus of claim 2 , wherein the visual-inertial odometry generates a sequence of sparse maps based on the IMU data and the visual data, the sequence of sparse maps including metric depth values, the dense scale map scaffolding based on the metric depth values.

4. The apparatus of claim 3 , wherein the globally-aligned depth prediction is based on an alignment of monocular depth estimates to the metric depth values.

5. The apparatus of claim 1 , wherein the globally-aligned depth prediction is based on a least-squares estimation for global scale and global shift.

6. The apparatus of claim 1 , wherein the processor circuitry is to train a scale map learner (SML) neural network to resolve the scale ambiguity in monocular depth estimates.

7. The apparatus of claim 6 , wherein the SML neural network is to fill a region within a convex hull via linear interpolation of anchor values and fill a region outside the convex hull with an identify scale value of one.

8. A method for metric depth estimation, the method comprising:

accessing a globally-aligned depth prediction, the globally-aligned depth prediction generated based on a monocular depth estimator;

accessing a dense scale map scaffolding, the dense scale map scaffolding generated based on visual-inertial odometry;

regressing, by executing an instruction with at least one processor circuit, a dense scale residual map determined using the globally-aligned depth prediction and the dense scale map scaffolding; and

applying, by executing an instruction with one or more of the at least one processor circuit, the dense scale residual map to the globally-aligned depth prediction.

9. The method of claim 8 , wherein the visual-inertial odometry determines the dense scale map scaffolding based on inertial measurement unit (IMU) data and visual data.

10. The method of claim 9 , wherein the visual-inertial odometry generates a sequence of sparse maps based on the IMU data and the visual data, the sequence of sparse maps including metric depth values, the dense scale map scaffolding based on the metric depth values.

11. The method of claim 10 , wherein the globally-aligned depth prediction is based on an alignment of monocular depth estimates to the metric depth values.

12. The method of claim 8 , wherein the globally-aligned depth prediction is determined based on a least-squares estimation for global scale and global shift.

13. The method of claim 8 , further including training a scale map learner (SML) neural network to resolve the scale ambiguity in monocular depth estimates.

14. The method of claim 13 , wherein the SML neural network is to fill a region within a convex hull via linear interpolation of anchor values and fill a region outside the convex hull with an identify scale value of one.

15. A non-transitory computer readable storage medium comprising instructions that, when executed, cause at least one processor circuit to at least:

access a globally-aligned depth prediction, the globally-aligned depth prediction generated based on a monocular depth estimator;

access a dense scale map scaffolding, the dense scale map scaffolding generated based on visual-inertial odometry;

regresses a dense scale residual map determined using the globally-aligned depth prediction and the dense scale map scaffolding; and

apply the dense scale residual map to the globally-aligned depth prediction.

16. The non-transitory computer readable storage medium of claim 15 , wherein the visual-inertial odometry determines the dense scale map scaffolding based on inertial measurement unit (IMU) data and visual data.

17. The non-transitory computer readable storage medium of claim 16 , wherein the visual-inertial odometry generates a sequence of sparse maps based on the IMU data and the visual data, the sequence of sparse maps including metric depth values, the dense scale map scaffolding based on the metric depth values.

18. The non-transitory computer readable storage medium of claim 17 , wherein globally-aligned depth prediction is based on an alignment of monocular depth estimates to the metric depth values.

19. The non-transitory computer readable storage medium of claim 15 , wherein the instructions, when executed, cause one or more of the at least one processor circuit to train a scale map learner (SML) neural network to resolve the scale ambiguity in monocular depth estimates.

20. The non-transitory computer readable storage medium of claim 19 , wherein the SML neural network is to fill a region within a convex hull via linear interpolation of anchor values and fill a region outside the convex hull with an identify scale value of one.

21. An apparatus for metric depth estimation, the apparatus comprising:

means for accessing a globally-aligned depth prediction, the globally-aligned depth prediction generated based on means for estimating using monocular depth estimation;

means for accessing a dense scale map scaffolding, the dense scale map scaffolding generated based on means for estimating using visual-inertial odometry;

means for regressing a dense scale residual map determined using the globally-aligned depth prediction and the dense scale map scaffolding; and

means for applying the dense scale residual map to the globally-aligned depth prediction.

22. The apparatus of claim 21 , wherein the means for estimating using visual-inertial odometry is to determine the dense scale map scaffolding based on inertial measurement unit (IMU) data and visual data.

23. The apparatus of claim 22 , wherein the means for estimating using visual-inertial odometry is to generate a sequence of sparse maps based on the IMU data and the visual data, the sequence of sparse maps including metric depth values, the dense scale map scaffolding based on the metric depth values.

24. The apparatus of claim 23 , wherein the globally-aligned depth prediction is based on an alignment of monocular depth estimates to the metric depth values.

25. The apparatus of claim 21 , wherein the globally-aligned depth prediction is based on a least-squares estimation for global scale and global shift.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2023
From: RANFTL, RENE
To: INTEL CORPORATION
Reel/Frame 063143/0911 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 18, 2023
From: WOFK, DIANA; MUELLER, MATTHIAS; KOLTUN, VLADLEN
To: INTEL CORPORATION
Reel/Frame 063117/0332 →
Continuity (2)
Provisional Application 63314121 · Feb 25, 2022
Related Publication 20220343521A1 · Oct 27, 2022
References Cited (28)
US 11232583B2 · Narasimha · 2022 [cited by examiner]
US 20210142497A1 · Pugh · 2021 [cited by examiner]
Engel, J. et al., “Lsd-slam: Large-scale direct monocular slam”, In ECCV, 2014, 16 pages. Retrieved from <https://jakobengel.github.io/pdf/engel14eccv.pdf> on Mar. 27, 2023. [cited by applicant]
Mur-Artal, R. et al., “Orb-slam: A versatile and accurate monocular slam system”, IEEE Transactions on Robotics, 2015, 18 pages, Retrieved from <https://arxiv.org/abs/1502.00956> on Mar. 27, 2023. [cited by applicant]
Mur-Artal, R. et al., “Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras”, IEEE Transactions on Robotics, 2017, 9 pages. Retrieved from <https://arxiv.org/pdf/1610.06475.pdf > on Mar. 27, 20… [cited by applicant]
Ranftl, R. et al., “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer”, IEEE TPAMI, 2020, 14 pages. Retrieved from <https://arxiv.org/pdf/1907.01341.pdf > on Mar. 27, 2023. [cited by applicant]
Ranftl, R. et al., “Vision transformers for dense prediction”, ArXiv, 2021, 15 pages. Retrieved from <https://arxiv.org/pdf/2103.13413.pdf > on Mar. 27, 2023. [cited by applicant]
Wong, A. et al., “Unsupervised depth completion from visual inertial odometry”, IEEE Robotics and Automation Letters, 2020, 16 pages. Retrieved from <https://arxiv.org/pdf/1905.08616.pdf> on Mar. 27, 2023. [cited by applicant]
Wong, A. et al., “Unsupervised depth completion with calibrated back projection layers”, In ICCV, 2021, 19 pages. Retrieved from <https://arxiv.org/pdf/2108.10531.pdf> on Mar. 27, 2023. [cited by applicant]
Fei, X. et al., “Geo-supervised visual depth prediction”, IEEE Robotics and Automation Letters, 2019, 8 pages. Retrieved from <https://arxiv.org/pdf/1807.11130.pdf> on Mar. 27, 2023. [cited by applicant]
Almalioglu, Y. et al., “SelfVIO: Self-Supervised Deep Monocular Visual-Inertial Odometry and Depth estimation”, ArXiv, 2019, 18 pages. Retrieved from <https://arxiv.org/pdf/1911.09968.pdf> on Mar. 27, 2023. [cited by applicant]
Patil, V. et al., “Don't Forget The Past: Recurrent Depth Estimation from Monocular Video”, IEEE Robotics and Automation Letters, 2020, 8 pages. Retrieved from <https://arxiv.org/pdf/2001.02613.pdf> on Mar. 27, 2023. [cited by applicant]
Teed, Z. et al., “Deepv2d: Video to depth with differentiable structure from motion”, In ICLR, 2020, 20 pages. Retrieved from <https://arxiv.org/pdf/1812.04605.pdf> on Mar. 27, 2023. [cited by applicant]
Liu, C. et al., “Neural RGB-D Sensing: Depth and Uncertainty from a Video Camera”, CVPR, 2019, 13 pages. Retrieved from <https://arxiv.org/pdf/1901.02571.pdf> on Mar. 27, 2023. [cited by applicant]
Xie, J. et al., “Video Depth Estimation by Fusing Flow-to-Depth Proposals”, IROS, 2020, 8 pages, Retrieved from <https://arxiv.org/pdf/1912.12874.pdf> on Mar. 27, 2023. [cited by applicant]
Sartipi, K. et al., “Deep Depth Estimation from Visual-Inertial SLAM”, IROS, 2020, 9 pages. Retrieved from <https://arxiv.org/pdf/2008.00092.pdf> on Mar. 27, 2023. [cited by applicant]
Merrill, N. et al., “Robust Monocular Visual-Inertial Depth Completion for Embedded Systems”, ICRA, 2021, 7 pages. Retrieved from <https://pgeneva.com/downloads/papers/Merrill2021ICRA.pdf> on Mar. 27, 2023. [cited by applicant]
Luo, X. et al., “Consistent Video Depth Estimation”, ACM Transactions on Graphics, 2020, 13 pages. Retrieved from <https://arxiv.org/pdf/2004.15021.pdf> on Mar. 27, 2023. [cited by applicant]
Schönberger, J. et al., “Structure-from-Motion Revisited”, In CVPR, 2016, 10 pages. Retrieved from <https://openaccess.thecvf.com/content_cvpr_2016/papers/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.pdf>… [cited by applicant]
Kopf, J. et al. “Robust Consistent Video Depth Estimation”, In CVPR, 2021, 11 pages. Retrieved from <https://arxiv.org/pdf/2012.05901.pdf> on Mar. 27, 2023. [cited by applicant]
Qin, T. et al., “VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State Estimator”, IEEE Transactions on Robotics, 2018, 17 pages. Retrieved from <https://arxiv.org/pdf/1708.03852.pdf> on Mar. 27, 2023. [cited by applicant]
Li, Z. et al., “Megadepth: Learning Single-View Depth Prediction from Internet Photos”, In CVPR, 2018, 10 pages. Retrieved from <https://arxiv.org/pdf/1804.00607.pdf> on Mar. 27, 2023. [cited by applicant]
Shah, S. et al., “AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles”, In Field and Service Robotics, 2017, 14 pages. Retrieved from <https://arxiv.org/pdf/1705.05065.pdf> on Mar. 27, 2023. [cited by applicant]
Wang, W. et al., “TartanAir: A Dataset to Push the Limits of Visual SLAM”, IROS, 2020, 8 pages. Retrieved from <https://arxiv.org/pdf/2003.14338.pdf> on Mar. 27, 2023. [cited by applicant]
Lusk, P. & Sudhakar, S., “Anticipated VINS-Mono”, 2018, 4 pages. Retrieved from <https://github.com/plusk01/Anticipated-VINS-Mono> on Mar. 27, 2023. [cited by applicant]
Fei, X. & Soatto, S., “XIVO: X Inertial-aided Visual Odometry and Sparse Mapping”, 2019, 4 pages. Retrieved from <https://github.com/ucla-vision/xivo> on Mar. 27, 2023. [cited by applicant]
Deng, J. et al., “Imagenet: A Large-Scale Hierarchical Image Database”, In CVPR, 2009, 8 pages. Retrieved from <https://www.image-net.org/static_files/papers/imagenet_cvpr09.pdf> on Mar. 27, 2023. [cited by applicant]
Loshchilov, I. et al., “Decoupled Weight Decay Regularization”, In ICLR, 2019, 19 pages. Retrieved from <https://arxiv.org/pdf/1711.05101.pdf> on Mar. 27, 2023. [cited by applicant]