IP Library › Granted Patent US 12,488,483
Granted Patent B2
US 12,488,483 · App. 18/110,421 · Granted Dec 2, 2025

Geometric 3D augmentations for transformer architectures

Inventors: Vitor Guizilini (Santa Clara, CA); Igor Vasiljevic (Pacifica, CA); Adrien D. Gaidon (San Jose, CA); Jiading Fang (Chicago, IL); Gregory Shakhnarovich (Chicago, IL); Matthew R. Walter (Chicago, IL); Rares A. Ambrus (San Francisco, CA)
Assignees: Toyota Research Institute, Inc.; Toyota Jidosha Kabushiki Kaisha; Toyota Technological Institute at Chicago
G06T7/593G06T7/85G06T2207/10024G06T2207/10028G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,483
App. No.
18/110,421
Granted
Dec 2, 2025
Kind
B2
Abstract

A method of generating additional supervision data to improve learning of a geometrically-consistent latent scene representation with a geometric scene representation architecture is provided. The method includes receiving, with a computing device, a latent scene representation encoding a pointcloud from images of a scene captured by a plurality of cameras each with known intrinsics and poses, generating a virtual camera having a viewpoint different from viewpoints of the plurality of cameras, projecting information from the pointcloud onto the viewpoint of the virtual camera, and decoding the latent scene representation based on the virtual camera thereby generating an RGB image and depth map corresponding to the viewpoint of the virtual camera for implementation as additional supervision data.

Claims (63)

1 . A method of generating additional supervision data to improve learning of a geometrically-consistent latent scene representation with a geometric scene representation architecture, the method comprising:

receiving, with a computing device, a latent scene representation encoding a pointcloud from images of a scene captured by a plurality of cameras each with known intrinsics and poses;

selecting one of the plurality of cameras;

translating a pose of the selected camera;

adjusting a viewing angle of the translated selected camera toward a center of the pointcloud;

generating a virtual camera having a viewpoint different from viewpoints of the plurality of cameras;

projecting information from the pointcloud onto the viewpoint of the virtual camera; and

decoding the latent scene representation based on the virtual camera thereby generating an RGB image and depth map corresponding to the viewpoint of the virtual camera for implementation as additional supervision data.

2 . The method of claim 1 , wherein translating the pose of the selected camera comprises adding translation noise to the pose of the selected camera.

3 . The method of claim 1 , wherein the virtual camera is generated by:

selecting one of the plurality of cameras as a canonical camera,

applying a rotation matrix to the canonical camera, and

propagating a rotation and a translation offset of the canonical camera resulting from the rotation matrix to other ones of the plurality of cameras.

4 . The method of claim 1 , further comprising:

implementing the geometric scene representation architecture with the computing device;

inputting the images of the scene captured by the plurality of cameras into the geometric scene representation architecture, wherein each camera of the plurality of cameras includes known embeddings; and

encoding the images of the scene captured by the plurality of cameras, with the geometric scene representation architecture, into the latent scene representation.

5 . The method of claim 1 , further comprising training the geometric scene representation architecture by inputting the RGB image and the depth map corresponding to the viewpoint of the virtual camera.

6 . The method of claim 1 , further comprising:

querying the latent scene representation with a camera embedding; and

decoding the latent scene representation based on the camera embedding, thereby generating an estimated depth map and an estimated RGB image based on the camera embedding.

7 . A system for generating additional supervision data to improve learning of a geometrically-consistent latent scene representation with a geometric scene representation architecture, the system comprising:

one or more processors; and

a non-transitory, computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to:

receive a latent scene representation encoding a pointcloud from images of a scene captured by a plurality of cameras each with known intrinsics and poses;

select one of the plurality of cameras;

translate a pose of the selected camera;

adjust a viewing angle of the translated selected camera toward a center of the pointcloud;

generate a virtual camera having a viewpoint different from viewpoints of the plurality of cameras;

project information from the pointcloud onto the viewpoint of the virtual camera; and

decode the latent scene representation based on the virtual camera thereby generating an RGB image and depth map corresponding to the viewpoint of the virtual camera for implementation as additional supervision data.

8 . The system of claim 7 , wherein translating the pose of the selected camera comprises adding translation noise to the pose of the selected camera.

9 . The system of claim 7 , wherein the virtual camera is generated by:

selecting one of the plurality of cameras as a canonical camera,

applying a rotation matrix to the canonical camera, and

propagating a rotation and a translation offset of the canonical camera resulting from the rotation matrix to other ones of the plurality of cameras.

10 . The system of claim 7 , wherein the instructions further cause the one or more processors to:

implement the geometric scene representation architecture with the system;

input the images of the scene captured by the plurality of cameras into the geometric scene representation architecture, wherein each camera of the plurality of cameras includes known embeddings; and

encode the images of the scene captured by the plurality of cameras, with the geometric scene representation architecture, into the latent scene representation.

11 . The system of claim 7 , wherein the instructions further cause the one or more processors to train the geometric scene representation architecture by inputting the RGB image and the depth map corresponding to the viewpoint of the virtual camera.

12 . The system of claim 7 , wherein the instructions further cause the one or more processors to:

query the latent scene representation with a camera embedding; and

decode the latent scene representation based on the camera embedding, thereby generating an estimated depth map and an estimated RGB image based on the camera embedding.

13 . A computing program product for generating additional supervision data to improve learning of a geometrically-consistent latent scene representation with a geometric scene representation architecture, the computing program product comprising machine-readable instructions stored on a non-transitory computer readable memory, which when executed by a computing device, causes the computing device to carry out steps comprising:

receiving, with the computing device, a latent scene representation encoding a pointcloud from images of a scene captured by a plurality of cameras each with known intrinsics and poses;

selecting one of the plurality of cameras;

translating a pose of the selected camera;

adjusting a viewing angle of the translated selected camera toward a center of the pointcloud;

generating a virtual camera having a viewpoint different from viewpoints of the plurality of cameras;

projecting information from the pointcloud onto the viewpoint of the virtual camera; and

decoding the latent scene representation based on the virtual camera thereby generating an RGB image and depth map corresponding to the viewpoint of the virtual camera for implementation as additional supervision data.

14 . The computing program product of claim 13 , wherein translating the pose of the selected camera comprises adding translation noise to the pose of the selected camera.

15 . The computing program product of claim 13 , wherein the virtual camera is generated by:

selecting one of the plurality of cameras as a canonical camera,

applying a rotation matrix to the canonical camera, and

propagating a rotation and a translation offset of the canonical camera resulting from the rotation matrix to other ones of the plurality of cameras.

16 . The computing program product of claim 13 , the steps caused to be carried out by

the computing device further comprising:

implementing the geometric scene representation architecture with the computing device;

inputting the images of the scene captured by the plurality of cameras into the geometric scene representation architecture, wherein each camera of the plurality of cameras includes known embeddings; and

encoding the images of the scene captured by the plurality of cameras, with the geometric scene representation architecture, into the latent scene representation.

17 . The computing program product of claim 13 , the steps caused to be carried out by the computing device further comprising training the geometric scene representation architecture by inputting the RGB image and the depth map corresponding to the viewpoint of the virtual camera.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2026
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 074439/0545 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2023
From: GUIZILINI, VITOR; GAIDON, ADRIEN D.; AMBRUS, RARES A.
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 062717/0429 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2023
From: VASILJEVIC, IGOR; FANG, JIADING; SHAKHNAROVICH, GREGORY; WALTER, MATTHEW
To: TOYOTA TECHNOLOGICAL INSTITUTE AT CHICAGO
Reel/Frame 062717/0967 →
Continuity (2)
Provisional Application 63392114 · Jul 25, 2022
Related Publication 20240029286A1 · Jan 25, 2024
References Cited (30)
US 8619082B1 · Ciurea et al. · 2013 [cited by applicant]
US 10390005B2 · Nisenzon et al. · 2019 [cited by applicant]
US 10452927B2 · Stojanović · 2019 [cited by examiner]
US 10826786B2 · Eckart · 2020 [cited by examiner]
US 10915793B2 · Corral-Soto · 2021 [cited by examiner]
US 11107228B1 · Shrivastava · 2021 [cited by examiner]
US 11238650B2 · Li · 2022 [cited by examiner]
US 11250616B2 · Kaplan · 2022 [cited by examiner]
US 11288857B2 · Meshry et al. · 2022 [cited by applicant]
US 11605151B2 · Holzer · 2023 [cited by examiner]
US 11625864B2 · Engelland-Gay · 2023 [cited by examiner]
US 11663778B2 · Hosfield · 2023 [cited by examiner]
US 11670088B2 · Voodarla · 2023 [cited by examiner]
US 11721044B2 · Fleureau · 2023 [cited by examiner]
US 11975738B2 · Singh · 2024 [cited by examiner]
US 11995749B2 · Borer · 2024 [cited by examiner]
US 12014446B2 · Koh · 2024 [cited by examiner]
US 12052408B2 · Tauber · 2024 [cited by examiner]
US 12100230B2 · Wang · 2024 [cited by examiner]
US 12243273B2 · Lv · 2025 [cited by examiner]
US 20200041276A1 · Chakravarty · 2020 [cited by examiner]
US 20200294194A1 · Sun · 2020 [cited by examiner]
US 20200380762A1 · Karafin et al. · 2020 [cited by applicant]
US 20210133990A1 · Eckart · 2021 [cited by examiner]
US 20210358193A1 · Oz et al. · 2021 [cited by applicant]
US 20220079510A1 · Robillard et al. · 2022 [cited by applicant]
US 20230334764A1 · Lee · 2023 [cited by examiner]
US 20230342507A1 · Beltrand · 2023 [cited by examiner]
EP 3428875A1 · 2019 [cited by examiner]
WO WO2022221267A2 · 2022 [cited by examiner]