IP Library › Granted Patent US 12,333,744
Granted Patent B2
US 12,333,744 · App. 17/578,830 · Granted Jun 17, 2025

Scale-aware self-supervised monocular depth with sparse radar supervision

Inventors: Vitor Guizilini (Santa Clara, CA); Charles Christopher Ochoa (San Francisco, CA)
Assignee: TOYOTA RESEARCH INSTITUTE, INC.
G06T7/50G01B15/00G01S13/89G06T3/40G06T2207/10028G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,744
App. No.
17/578,830
Granted
Jun 17, 2025
Kind
B2
Abstract

Systems and methods are provided for training a depth model to recover scale factor for self-supervised depth estimation in monocular images. Examples include deriving a depth map for an image based on a depth model. The depth map comprises depth values for pixels of the image. A first scale for the image can be estimated based on the depth values, and depth data captured by a range sensor can be received. The depth data comprises a point cloud comprising depth measures. A second scale for the point cloud can be determined based on the depth measures and a scale factor can be determined based the second scale and the first scale. The depth model can be updated based on the scale factor, wherein the depth model generates metrically accurate depth estimates based on the scale factor.

Claims (55)

1. A method for depth estimation from monocular images, comprising:

receiving an image captured by an image sensor, the image comprising pixels representing a scene of an environment;

deriving a depth map for the image based on a depth model, the depth map comprising a plurality of predicted depth values for a plurality of the pixels of the image;

estimating a first scale comprising a pixel distance between depth values for the image based the plurality of predicted depth values;

receiving depth data captured by a range sensor, the depth data comprising a point cloud representing the scene of the environment, the point cloud comprising depth measures for a plurality of points of the point cloud;

determining a second scale comprising a measure of central tendency for the first scale based on the depth measures, wherein the measure of central tendency, for the first scale, comprises a measure of a central value for probability distribution of the first scale based on the depth measure;

determining a scale factor based on a ratio between the second scale and the first scale; and

updating the depth model based on the scale factor, wherein the depth model generates metrically accurate depth estimates based on the scale factor.

2. The method of claim 1 , wherein the image is a monocular image.

3. The method of claim 1 wherein the range sensor produces a sparse point cloud, wherein the depth measure for each point comprises an error.

4. The method of claim 3 , wherein the range sensor is a radar sensor.

5. The method of claim 1 , further comprising:

generating a point cloud from the derived depth map, the point cloud comprises points based on the depth values for the plurality of pixels,

wherein the estimated first scale is based on the point cloud of the derived depth map.

6. The method of claim 1 , wherein the first scale is a single first scale for the image.

7. The method of claim 6 , further comprising:

determining a pixel-wise scale for each of the plurality of the pixels of the image based on a depth value of each respective pixel; and

determining a measure of central tendency of the pixel-wise scales, wherein the first scale is estimated based on the determined measure of central tendency, wherein the central tendency indicates an arithmetic average for probability distribution.

8. The method of claim 1 , wherein the second scale is a single second scale for the depth data.

9. The method of claim 8 , further comprising:

determining a point-wise scale for each of the plurality of the points of the point cloud based on a depth measure of each respective point; and

determining a measure of central tendency of the point-wise scales, wherein the second scale is estimated based on the determined measure of central tendency, wherein the central tendency indicates a weighted average for probability distribution.

10. The method of claim 1 , wherein the depth measure for each point comprises an error, the method further comprising:

determining that the error is greater than a threshold; and

in response to the determination that the error is greater than the threshold,

determining a plurality of scale factors for a plurality of images captured by the image sensor and a plurality of depth data captured by a radar sensor; and

determining an aggregate scale factor by aggregating the plurality of scale factors, wherein updating the depth model is based on the aggregate scale factor.

11. A system with a method, comprising:

a memory; and one or more processors that are configured to execute machine readable instructions stored in the memory for performing a method comprising:

training a depth model at a first stage according to self-supervised photometric losses generated from at least a first monocular image;

determining a single scale factor comprising a ratio between a pixel distance between depth values from a depth map of the first monocular image and a measure of central tendency based on the pixel distance from a sparse point cloud generated by a range sensor, wherein the measure of central tendency comprises of a measure of a central value for probability distribution across pixel-wise scales; and

training the depth model at a second stage according to a supervised loss based on the single scale factor,

wherein the depth model trained according to the second stage generates metrically accurate depth estimates of monocular images based on the single scale factor.

12. The system of claim 11 , wherein the range sensor produces a sparse point cloud, wherein the measure of central tendency for each point comprises an error.

13. The system of claim 12 , wherein the range sensor is a radar sensor.

14. The system of claim 11 , wherein the method further comprises:

generating a point cloud from the first monocular image, the point cloud comprises points based on the depth values for a plurality of pixels derived from the depth model,

wherein the depth map is based on the point cloud of the first monocular image.

15. The system of claim 11 , wherein determining the single scale factor from the depth map of the first monocular image and the sparse point cloud generated by a range sensor comprises:

estimating a first scale for the first monocular image based on predicted depth values from the depth map, the first scale being a single first scale for the image.

16. The system of claim 15 , wherein the method further comprises:

determining a pixel-wise scale for each of the plurality of the pixels of the first monocular image based on a depth value of each respective pixel; and

determining a measure of central tendency of the pixel-wise scales, wherein the first scale is estimated based on the determined measure of central tendency.

17. The system of claim 15 , wherein determining the single scale factor from the depth map of the first monocular image and the sparse point cloud generated by a range sensor comprises:

receiving depth data captured by a range sensor, the depth data comprising the sparse point cloud representing a scene of an environment of the first monocular image, the point cloud comprising depth measures for a plurality of points of the point cloud; and

determining a second scale for the point cloud based on the depth measures, the second scale is a single second scale for the depth data.

18. The system of claim 17 , wherein the method further comprises:

determining a point-wise scale for each of the plurality of the points of the point cloud based on a depth measure of each respective point; and

determining a measure of central tendency of the point-wise scales, wherein the second scale is estimated based on the determined measure of central tendency.

19. The system of claim 17 , wherein the depth measure for each point comprises an error, the method further comprising:

determining that the error is greater than a threshold; and

in response to the determination that the error is greater than the threshold,

determining a plurality of scale factors for a plurality of images captured by an image sensor and a plurality of depth data captured by a radar sensor; and

determining an aggregate scale factor by aggregating the plurality of scale factors, wherein updating the depth model is based on the aggregate scale factor.

20. The system of claim 17 , wherein the single scale factor is based on a comparison of the second scale with the first scale.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2025
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 071701/0133 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2022
From: GUIZILINI, VITOR; OCHOA, CHARLES CHRISTOPHER
To: TOYOTA RESEARCH INSTITUTE, INC.
Reel/Frame 058694/0353 →
Continuity (1)
Related Publication 20230230264A1 · Jul 20, 2023
References Cited (13)
US 20200167941A1 · Tong · 2020 [cited by applicant]
US 20210004660A1 · Ambrus · 2021 [cited by applicant]
US 20210004974A1 · Guizilini · 2021 [cited by applicant]
US 20210183083A1 · Yan · 2021 [cited by applicant]
US 20210237764A1 · Tang · 2021 [cited by applicant]
CN 111753961A · 2020 [cited by applicant]
CN 113205549A · 2021 [cited by applicant]
“Yevhen Kuznietsov et. al., Semi-Supervised Deep Learning for Monocular Depth Map Prediction, 2017, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition CVPR, pp. 6647-6655” (Year: 2017). [cited by examiner]
“Sumyeong Lee et. al., Robust 3-Dimensional Point Cloud Mapping in Dynamic Environment Using Point-wise static Probability-Based NDT Scan-Matching, Sep. 2020, IEEE Access, vol. 8” (Year: 2020). [cited by examiner]
“Lingfei Ma et. al., Multi-Scale Point-Wise Convolutional Neural Networks for 3D Object Segmentation From LiDAR Point Clouds in Large-Scale Environments, Dec. 2019, IEEE Transactions on Intelligent Transportation System… [cited by examiner]
“Qian Shi et. al., A Deeply Supervised Attention Metric-Based Network and an Open Aerial Image Dataset for Remote Sensing Change Detection, Jun. 2021, IEEE Transactions on Geoscience and Remote Sensing, vol. 60” (Year: … [cited by examiner]
Patil et al., “Don't Forget the Past: Recurrent Depth Estimation from Monocular Video,” 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 2020, 8 pages (http://ras.papercept.net/image… [cited by applicant]
Mitra, “Monocular Depth Estimation using Adversarial Training,” Master's thesis, University of Minnesota, Jun. 2020, 72 pages (https://irvlab.cs.umn.edu/sites/irvlab.dl.umn.edu/files/mitra_msc_thesis_monogan.pdf). [cited by applicant]
Cited By (1)
US 12,641,211