IP Library Granted Patent US 12,700,113
Granted Patent B2
US 12,700,113 · App. 18/415,047 · Granted Aug 4, 2026

Language-based learning for monocular depth estimation

Inventors: Adrien David Gaidon (San Jose, CA); Vitor Campagnolo Guizilini (Santa Clara, CA)
Assignees: Toyota Research Institute, Inc.; Toyota Jidosha Kabushiki Kaisha
G06T7/50G06F40/30G06V20/56G06T2207/10028G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,700,113
App. No.
18/415,047
Filed
Jan 17, 2024
Granted
Aug 4, 2026
Kind
B2
Art Unit
2669
USPC
382/154
Abstract

Systems, methods, and other embodiments described herein relate to using a language model to facilitate training a depth model for monocular depth estimation. In one embodiment, a method includes acquiring an image depicting surrounding objects present in an environment. The method includes generating a depth map from the image using a depth model that performs monocular depth estimation. The method includes analyzing the depth map to derive a semantic loss according to a language model. The method includes training the depth model according to at least the semantic loss.

Claims (37)

1 . A depth system, comprising:

one or more processors;

a memory communicably coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the one or more processors to:

acquire an image depicting surrounding objects present in an environment;

generate a depth map from the image using a depth model that performs monocular depth estimation;

analyze the depth map to derive a semantic loss according to a language model, including querying the language model with a query including the depth map and a text string to identify anomalies in the depth map, wherein the text string defines at least one characteristic of the query; and

train the depth model according to at least the semantic loss.

2 . The depth system of claim 1 , wherein the depth model performs monocular depth estimation and is trained according to the semantic loss and a depth loss from self-supervised structure-from-motion (SfM) training.

3 . The depth system of claim 1 , wherein the language model is one of a large language model (LLM) and a visual language model (VLM), and

wherein the instructions to train the depth model include instructions to combine the semantic loss with a depth loss that is a self-supervised loss.

4 . The depth system of claim 1 , wherein the instructions to analyze the depth map include instructions to generate the semantic loss according to an output of the language model identifying a presence of one or more anomalies in the depth map.

5 . The depth system of claim 1 , wherein the instructions to analyze the depth map include instructions to generate a mask to segment anomalies within the depth map, and

wherein the instructions to train the depth model include instructions to filter the image from a training data set if the image causes anomalies in the depth map.

6 . The depth system of claim 1 , wherein the instructions include instructions to:

provide the depth model, including integrating the depth model in a perception pipeline of an autonomous vehicle to facilitate control of the autonomous vehicle.

7 . The depth system of claim 1 , wherein the depth system is embedded within a vehicle to perceive depth in the environment.

8 . A non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to:

acquire an image depicting surrounding objects present in an environment;

generate a depth map from the image using a depth model that performs monocular depth estimation;

analyze the depth map to derive a semantic loss according to a language model, including querying the language model with a query including the depth map and a text string to identify anomalies in the depth map, wherein the text string defines at least one characteristic of the query; and

train the depth model according to at least the semantic loss.

9 . The non-transitory computer-readable medium of claim 8 , wherein the depth model performs monocular depth estimation and is trained according to the semantic loss and a depth loss from self-supervised structure-from-motion (SfM) training.

10 . The non-transitory computer-readable medium of claim 8 , wherein the language model is one of a large language model (LLM) and a visual language model (VLM), and

wherein the instructions to train the depth model include instructions to combine the semantic loss with a depth loss that is a self-supervised loss.

11 . The non-transitory computer-readable medium of claim 8 , wherein the instructions to analyze the depth map include instructions to generate the semantic loss according to an output of the language model identifying a presence of one or more anomalies in the depth map.

12 . A method, comprising:

acquiring an image depicting surrounding objects present in an environment;

generating a depth map from the image using a depth model that performs monocular depth estimation;

analyzing the depth map to derive a semantic loss according to a language model, including querying the language model with a query including the depth map and a text string to identify anomalies in the depth map, wherein the text string defines at least one characteristic of the query; and

training the depth model according to at least the semantic loss.

13 . The method of claim 12 , wherein the depth model performs monocular depth estimation and is trained according to the semantic loss and a depth loss from self-supervised structure-from-motion (SfM) training.

14 . The method of claim 12 , wherein the language model is one of a large language model (LLM) and a visual language model (VLM), and

wherein training the depth model includes combining the semantic loss with a depth loss that is a self-supervised loss.

15 . The method of claim 12 , wherein analyzing the depth map includes generating the semantic loss according to an output of the language model identifying a presence of one or more anomalies in the depth map.

16 . The method of claim 12 , wherein analyzing the depth map includes generating a mask to segment anomalies within the depth map, and wherein training the depth model includes filtering the image from a training data set if the image causes anomalies in the depth map.

17 . The method of claim 12 , further comprising:

providing the depth model, including integrating the depth model in a perception pipeline of an autonomous vehicle to facilitate control of the autonomous vehicle.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2024
From: GAIDON, ADRIEN DAVID; CAMPAGNOLO GUIZILINI, VITOR
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 066426/0920 →
Continuity (1)
Related Publication 20250232461A1 · Jul 17, 2025
References Cited (14)
US 10410351B2 · Lin et al. · 2019 [cited by applicant]
US 11263753B2 · Larlus-Larrondo et al. · 2022 [cited by applicant]
US 20050038650A1 · Bellegarda et al. · 2005 [cited by applicant]
US 20170193009A1 · Rapantzikos et al. · 2017 [cited by applicant]
US 20210090277A1 · Guizilini et al. · 2021 [cited by applicant]
US 20220172390A1 · Redford et al. · 2022 [cited by applicant]
US 20230023126A1 · Ansari · 2023 [cited by examiner]
US 20230274086A1 · Tunstall-Pedoe et al. · 2023 [cited by applicant]
US 20250200773A1 · Slobodyanyuk · 2025 [cited by examiner]
CN 116523987 · 2023 [cited by examiner]
Look Deeper into Depth: Monocular Depth Estimation with Semantic Booster and Attention-Driven Loss (Year: 2018). [cited by examiner]
Prospective Role of Foundation Models in Advancing Autonomous Vehicles (Year: 2023). [cited by examiner]
Talker et al., “Mind The Edge: Refining Depth Edges in Sparsely-Supervised Monocular Depth Estimation”, 2023, 18 pages, retrieved Jan. 17, 2024 from the arXiv database at: https://doi.org/10.48550/arXiv.2212.05315. [cited by applicant]
Hornauer et al., “Out-of-Distribution Detection for Monocular Depth Estimation”, Oct. 2023, 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 11 pages, retrieved Jan. 17, 2024 from the arXiv database at:… [cited by applicant]