IP Library › Granted Patent US 12,651,360
Granted Patent B2
US 12,651,360 · App. 18/142,445 · Granted Jun 9, 2026

Monocular depth estimation system

Inventors: Wenwen Tu (San Jose, CA); Zuoguan Wang (Los Gatos, CA); Qun Gu (San Jose, CA)
Assignee: Black Sesame Technologies Inc.
G06T7/50G06T2207/10028G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,360
App. No.
18/142,445
Granted
Jun 9, 2026
Kind
B2
Abstract

Systems and methods are directed to monocular depth estimation. A method includes capturing an image; encoding the image to form an encoded image; masking one or more dynamic objects in the encoded image; predicting a final prediction performance based on one or more unmasked static objects from the encoded image; decoding and projecting the final prediction performance; and producing an information based on Global Positioning System (GPS), Inertial Measurement Unit (IMU), and wheel encoder fusion.

Claims (34)

1 . A method for estimating monocular depth, comprising:

capturing an image;

encoding the image to form an encoded image;

comparing a projected mask in one or more dynamic objects with a pre-defined mask in the encoded image to predict an overlap ratio;

updating a pre-defined one or more dynamic objects as a static object or a dynamic object based on the overlap ratio;

masking one or more dynamic objects in the encoded image;

predicting a final prediction performance based on one or more unmasked static objects from the encoded image;

decoding and projecting the final prediction performance;

producing an information based on Global Positioning System (GPS), Inertial Measurement Unit (IMU), and wheel encoder fusion; and

estimating the monocular depth based on the final prediction performance and the information.

2 . A method for estimating monocular depth, comprising:

capturing an image;

encoding the image to form an encoded image;

comparing a projected mask in one or more dynamic objects with a pre-defined mask in the encoded image to predict an overlap ratio;

updating a pre-defined one or more dynamic objects to either a static object or a dynamic object based on the overlap ratio;

predicting a final prediction performance based on one or more unmasked static objects from the encoded image; and

decoding and projecting the final prediction performance;

producing an information based on Global Positioning System (GPS), Inertial Measurement Unit (IMU), and wheel encoder fusion; and

estimating the monocular depth based on the final prediction performance and the information.

3 . The method of claim 1 , wherein masking one or more dynamic objects in the encoded image comprises masking the one or more dynamic objects with a self-detected mask on calculating loss in training.

4 . The method of claim 3 , wherein the self-detected mask is formed by comparing an estimated depth and a projected depth map.

5 . The method of claim 1 , wherein the one or more dynamic objects is masked from a final loss function.

6 . The method of claim 5 , wherein the final loss function is calculated to improve the final prediction performance.

7 . The method of claim 6 , wherein the final loss function comprises a re-projection loss, a smoothness loss, and a geometry consistency loss.

8 . The method of claim 7 , wherein the re-projection loss is summation of photometric loss and structural similarity (SSIM) difference.

9 . The method of claim 7 , wherein the geometry consistency loss comprises one or more re-projection weights.

10 . A system for estimating monocular depth, comprising:

a memory and a processor, wherein the memory is connected to the processor;

the memory stores a computer program; and

the processor implements the method of claim 1 when executing the computer program.

11 . A system for estimating monocular depth, comprising:

a processor, wherein the memory is connected to the processor;

the memory stores a computer program; and

the processor implements the method of claim 2 when executing the computer program.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE INVENTORS INCLUDE WENWNE TU, ZUOGUAN WANG, AND QUN GU. PREVIOUSLY RECORDED AT REEL: 063512 FRAME: 0287. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 12, 2023
From: TU, WENWEN; WANG, ZUOGUAN; GU, QUN
To: BLACK SESAME TECHNOLOGIES INC.
Reel/Frame 063981/0350 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2023
From: TU, WENWEN; WANG, ZUOGUAN
To: BLACK SESAME TECHNOLOGIES INC.
Reel/Frame 063512/0287 →
Continuity (1)
Related Publication 20240371014A1 · Nov 7, 2024
References Cited (35)
US 11144818B2 · Ambrus · 2021 [cited by examiner]
US 20190356905A1 · Godard · 2019 [cited by examiner]
US 20200084427A1 · Sun · 2020 [cited by examiner]
US 20200193623A1 · Liu · 2020 [cited by examiner]
US 20210019892A1 · Zhou · 2021 [cited by examiner]
US 20210150227A1 · Hu · 2021 [cited by examiner]
US 20210181757A1 · Goel · 2021 [cited by examiner]
US 20220156525A1 · Guizilini · 2022 [cited by examiner]
US 20220215568A1 · Dekel · 2022 [cited by examiner]
US 20220245843A1 · Guizilini · 2022 [cited by examiner]
US 20230394843A1 · Lee · 2023 [cited by examiner]
CN 114170286B · 2023 [cited by examiner]
JP 7110502B2 · 2022 [cited by examiner]
WO WO2023165093A1 · 2023 [cited by examiner]
Lei et al. “Attention based multilayer feature fusion convolutional neural network for unsupervised monocular depth estimation” (Year: 2020). [cited by examiner]
Li et al. “Self-supervised coarse-to-fine monocular depth estimation using a lightweight attention module” (Year: 2022). [cited by examiner]
Wang et al. “Self-supervised monocular depth estimation with direct methods” (Year: 2020). [cited by examiner]
“Unsupervised depth estimation from monocular videos with hybrid geometric-refined loss and contextual attention”. [cited by examiner]
Dai et al. “Unsupervised learning of depth estimation based on attention model and global pose optimization” (Year: 2019). [cited by examiner]
Godard, C., Mac Aodha, O., Firman, M., & Brostow, G. J. “Digging into self-supervised monocular depth estimation”. In Proceedings of the IEEE/CVF International Conference on Computer Vision: pp. 3828-3838, 2019. [cited by applicant]
Zhou, T., Brown, M., Snavely, N., & Lowe, D. G. “Unsupervised learning of depth and ego-motion from video”. In Proceedings of the IEEE conference on computer vision and pattern recognition: pp. 1851-1858, 2017. [cited by applicant]
Guizilini, V., Hou, R., Li, J., Ambrus, R., & Gaidon, A. “Semantically-guided representation learning for self-supervised monocular depth”. arXiv preprint arXiv:2002.12319. 2020. [cited by applicant]
Klingner, M., Termohlen, J. A., Mikolajczyk, J., & Fingscheidt, T. “Self-supervised monocular depth estimation: Solving the dynamic object problem by semantic guidance”. In European Conference on Computer Vision: pp. 58… [cited by applicant]
Kuznietsov, Y., Stuckler, J., & Leibe, B. “Semi-supervised deep learning for monocular depth map prediction”. In Proceedings of the IEEE conference on computer vision and pattern recognition: pp. 6647-6655. 2017. [cited by applicant]
Amiri, A. J., Loo, S. Y., & Zhang, H. “Semi-supervised monocular depth estimation with left-right consistency using deep neural network”. In 2019 IEEE International Conference on Robotics and Biomimetics (ROBIO): pp. 60… [cited by applicant]
Guizilini, V., Li, J., Ambrus, R., Pillai, S., & Gaidon, A. “Robust semi-supervised monocular depth estimation with reprojected distances”. In Conference on robot learning: pp. 503-512. 2020. [cited by applicant]
Johnston, A., & Carneiro, G. “Self-supervised monocular trained depth estimation using self-attention and discrete disparity volume”. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition:… [cited by applicant]
Chen, Y., Zhao, H., Hu, Z., & Peng, J. Attention-based context aggregation network for monocular depth estimation. International Journal of Machine Learning and Cybernetics: vol. 12, No. 6, pp. 1583-1596. 2021. [cited by applicant]
Lowe, D. G. “Distinctive image features from scale-invariant keypoints”. International journal of computer vision: vol. 60, No. 2, pp. 91-110. 2004. [cited by applicant]
Casser, V., Pirk, S., Mahjourian, R., & Angelova, A. “Unsupervised learning of depth and ego-motion: A structured approach”. In Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19): vol. 2, p. 7. 2019. [cited by applicant]
Garg, R., Bg, V. K., Carneiro, G., & Reid, I. “Unsupervised cnn for single view depth estimation: Geometry to the rescue”. In European conference on computer vision: pp. 740-756. 2016. [cited by applicant]
Godard, C., Mac Aodha, O., & Brostow, G. J. “Unsupervised monocular depth estimation with left-right consistency”. In Proceedings of the IEEE conference on computer vision and pattern recognition: pp. 270-279. 2017. [cited by applicant]
Brunet, D., Vrscay, E. R., & Wang, Z. “On the mathematical properties of the structural similarity index”. IEEE Transactions on Image Processing: vol. 21, No. 4, pp. 1488-1499. 2011. [cited by applicant]
Wagstaff, B., & Kelly, J. “Self-supervised scale recovery for monocular depth and egomotion estimation”. arXiv preprint arXiv:2009.03787, 2020. [cited by applicant]
Bian, J., Li, Z., Wang, N., Zhan, H., Shen, C., Cheng, M. M., & Reid, I. “Unsupervised scale-consistent depth and ego-motion learning from monocular video”. Advances in neural information processing systems: pp. 35-45. … [cited by applicant]