IP Library › Granted Patent US 12,430,840
Granted Patent B2
US 12,430,840 · App. 18/156,958 · Granted Sep 30, 2025

Systems and methods for depth synthesis with transformer architectures

Inventors: Vitor Guizilini (Santa Clara, CA); Igor Vasiljevic (Pacifica, CA); Adrien D. Gaidon (San Jose, CA); Greg Shakhnarovich (Chicago, IL); Matthew Walter (Chicago, IL); Jiading Fang (Chicago, IL); Rares A. Ambrus (San Francisco, CA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA; TOYOTA TECHNOLOGICAL INSTITUTE AT CHICAGO
G06T15/10B60W60/001G06T7/50B60W2420/403B60W2556/40G06T2207/10024G06T2207/10028G06T2207/20212G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,840
App. No.
18/156,958
Granted
Sep 30, 2025
Kind
B2
Abstract

Systems and methods for enhanced computer vision capabilities, particularly including depth synthesis, which may be applicable to autonomous vehicle operation are described. A vehicle may be equipped with a geometric scene representation (GSR) architecture for synthesizing depth views at arbitrary viewpoints. The GSR architecture synthesizes depth views enable advanced functions, including depth interpolation and depth extrapolation. The GSR architecture implements functions (i.e., depth interpolation, depth extrapolation) that are useful for various computer vision applications for autonomous vehicles, such as predicting depth maps from unseen locations. For example, a vehicle includes a processor device synthesizing depth views at multiple viewpoints, where the multiple viewpoints are from image data of a surrounding environment for the vehicle. Further, the vehicle can have a controller device that receives depth views from the processor device and performs autonomous operations in response to analysis of the depth views.

Claims (19)

1. A vehicle, comprising:

one or more processors; and

memory storing machine-executable instructions in non-transitory memory, which when executed by the one or more processors, cause the vehicle to:

synthesize depth views at multiple viewpoints, wherein synthesizing the depth views at the multiple viewpoints comprises:

encoding images of an environment surrounding the vehicle into image embeddings, wherein the images are obtained from the multiple viewpoints,

encoding intrinsic parameters and relatives poses of one or more cameras into camera embeddings, wherein the one or more cameras captured the images,

projecting the image embeddings and the camera embeddings onto a latent representation for the environment using cross-attention layers of a neural network,

conditioning the latent representation using self-attention layers of the neural network,

generating second camera encodings for arbitrary cameras at arbitrary relative poses, and

querying the conditioned latent representation using the second camera embeddings to synthesize the depth views; and

perform one or more autonomous operations in response to analysis of the synthesized depth views.

2. The vehicle of claim 1 , wherein the one or more processors comprise a geometric scene representation (GSR) component.

3. The vehicle of claim 2 , wherein synthesizing depth views comprises depth estimations, depth interpolations, and depth extrapolations.

4. The vehicle of claim 3 , wherein the depth extrapolations comprise completed unseen portions of a scene including the surrounding environment for the vehicle.

5. The vehicle of claim 4 , wherein the depth extrapolations comprise generated dense depth maps from the multiple viewpoints.

6. The vehicle of claim 5 , wherein performing the one or more autonomous operations is in response to the dense depth maps.

7. The vehicle of claim 1 , wherein the one or more processors comprise a computer vision component performing one or more computer visual capabilities for the one or more autonomous operations.

8. The vehicle of claim 7 , wherein the one or more computer visual capabilities comprise object detection.

9. The vehicle of claim 1 , wherein the vehicle comprises an autonomous vehicle.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2025
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 073194/0269 →
CORRECTIVE ASSIGNMENT TO CORRECT THE 2ND ASSIGNOR'S NAME PREVIOUSLY RECORDED AT REEL: 62428 FRAME: 443. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 9, 2025
From: VASILJEVIC, IGOR; SHAKHNAROVICH, GREGORY; WALTER, MATTHEW; FANG, JIADING
To: TOYOTA TECHNOLOGICAL INSTITUTE AT CHICAGO
Reel/Frame 073063/0636 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2023
From: GUIZILINI, VITOR; GAIDON, ADRIEN D.; AMBRUS, RARES A.
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 062428/0348 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2023
From: VASILJEVIC, IGOR; SHAKHNAROVICH, GREG; WALTER, MATTHEW; FANG, JIADING
To: TOYOTA TECHNOLOGICAL INSTITUTE AT CHICAGO
Reel/Frame 062428/0443 →
Continuity (1)
Related Publication 20240249465A1 · Jul 25, 2024
References Cited (17)
US 10984543B1 · Srinivasan · 2021 [cited by examiner]
US 20220012848A1 · Ranftl · 2022 [cited by applicant]
CN 112907641A · 2021 [cited by applicant]
KR 102110690B1 · 2020 [cited by applicant]
Mahmud et al. (“ViewSynth: Learning Local Features from Depth Using View Synthesis”) Computer Vision and Pattern Recognition [Submitted on Nov. 22, 2019 (v1), last revised Sep. 1, 2020 (this version, v4)] (Year: 2020). [cited by examiner]
Jaegle et al., Perceiver IO: A General Architecture for Structured Inputs & Outputs, 10th International Conference on Learning Representations (ICLR 2022), Apr. 25, 2022, 29 pages (https://doi.org/10.48550/arXiv.2107.14… [cited by applicant]
Liu et al., “Deep Convolutional Neural Fields for Depth Estimation from a Single Image,” 2015 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 10, 2015, pp. 5162-5170 (https://openaccess.thecv… [cited by applicant]
Wang et al., “MVDepthNet: Real-time Multiview Depth Estimation Neural Network,” 2018 International Conference on 3D Vision (3DV), Sep. 5, 2018, 18 pages (https://doi.org/10.1109/3DV.2018.00037). [cited by applicant]
Varma et al., “Transformers in Self-Supervised Monocular Depth Estimation with Unknown Camera Intrinsics,” 17th International Conference on Computer Vision Theory and Applications (VISAPP 2022), Feb. 7, 2022, 12 pages (… [cited by applicant]
Rombach et al., “Geometry-Free View Synthesis: Transformers and no 3D Priors,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 13, 2021, pp. 14356-14366 (https://openaccess.thecvf.com/content/ICCV… [cited by applicant]
Yifan et al., “Input-Level Inductive Biases for 3D Reconstruction,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Jun. 22, 2022, 13 pages (https://openaccess.thecvf.com/content/CVPR2022/pape… [cited by applicant]
Mildenhall et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” 16th European Conference on Computer Vision (ECCV), Aug. 24, 2020, 17 pages (https://www.ecva.net/papers/eccv_2020/papers_ECCV… [cited by applicant]
Ranftl et al., “Vision Transformers for Dense Prediction,” 2021 IEEE/CVF International Conference on Computer Vision, Oct. 13, 2021, pp. 12179-12188 (https://openaccess.thecvf.com/content/ICCV2021/papers/Ranftl_Vision_T… [cited by applicant]
Long et al., “Multi-View Depth Estimation using Epipolar Spatio-Temporal Networks,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8258-8267 (https://openaccess.thecvf.com/content/CVPR20… [cited by applicant]
Saxena et al., “Learning Depth from Single Monocular Images,” 19th Annual Conference on Neural Information Processing Systems (NIPS 2005), Dec. 2005, 8 pages (https://papers.nips.cc/paper/2005/file/17d8da815fa21c57af982… [cited by applicant]
Godard et al., “Digging into Self-Supervised Monocular Depth Estimation,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 30, 2019, pp. 3828-3838 (https://openaccess.thecvf.com/content_ICCV_2019/p… [cited by applicant]
Guizilini et al., “3D Packing for Self-Supervised Monocular Depth Estimation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 16, 2020. pp. 2485-2494 (https://openaccess.t… [cited by applicant]