IP Library Granted Patent US 12,530,835
Granted Patent B2
US 12,530,835 · App. 18/364,853 · Granted Jan 20, 2026

Self-supervised depth for volumetric rendering regularization

Inventors: Vitor Guizilini (Santa Clara, CA); Rares A. Ambrus (San Francisco, CA); Jiading Fang (Chicago, IL); Sergey Zakharov (San Francisco, CA); Vincent Sitzmann (Cambridge, MA); Igor Vasiljevic (Pacifica, CA); Adrien Gaidon (San Jose, CA)
Assignees: Toyota Research Institute, Inc.; Toyota Jidosha Kabushiki Kaisha; Massachusetts Institute of Technology
G06T15/08G06T3/18G06T15/20G06V10/25G06V10/7747G06V20/41G06V20/56G06V20/64
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,835
App. No.
18/364,853
Granted
Jan 20, 2026
Kind
B2
Abstract

An example method includes generating embeddings of image data that includes multiple images, where each image has a different viewpoints of a scene, generating a latent space and a decoder, wherein the decoder receives embeddings as input to generate an output viewpoint, for each viewpoint in the image data, determining a volumetric rendering view synthesis loss and a multi-view photometric loss, and applying an optimization algorithm to the latent space and the decoder over a number of epochs until the volumetric rendering view synthesis loss is within a volumetric threshold and the multi-view photometric loss is within a multi-view threshold.

Claims (40)

1 . A method of training a latent space trainer for volumetric rendering, the method comprising:

generating embeddings of image data that includes multiple images by sampling values along a viewing ray to generate 3D points and Fourier encoding the sampled points, where each image has a different viewpoint of a scene;

generating a latent space and a decoder, wherein the decoder receives embeddings as input to generate an output viewpoint;

for each viewpoint in the image data, determining a volumetric rendering view synthesis loss and a multi-view photometric loss; and

applying an optimization algorithm to the latent space and the decoder over a number of epochs until the volumetric rendering view synthesis loss is within a volumetric threshold and the multi-view photometric loss is within a multi-view threshold.

2 . The method of claim 1 , wherein the optimization algorithm uses a Mean Square Error objective for the volumetric rendering view synthesis loss.

3 . The method of claim 1 , wherein the optimization algorithm is a gradient descent algorithm.

4 . The method of claim 1 , wherein the multi-view photometric loss is determined using a photometric objective.

5 . The method of claim 4 , wherein the photometric objective is determined by:

for each pixel of a target image of the image data, with a predicted depth ({circumflex over (d)}), generating, by a warping operation, projected coordinates with a predicted depth ({circumflex over (d)}′) in a context image;

generating a synthesized target image from the context image; and

determining a difference between the target image and the synthesized target image.

6 . The method of claim 5 , wherein the context image is generated by a transformation matrix.

7 . The method of claim 5 , wherein the difference between the target image and the synthesized target image is determined by a weighted structural similarity index.

8 . The method of claim 5 , wherein the synthesized target image is generated by applying grid sampling with bilinear interpolation to place information from the context image onto each target pixel of the synthesized target image based on the projected coordinates.

9 . The method of claim 5 , wherein pixels of the target image for determining the photometric objective are determined using strided ray sampling.

10 . A system for training a latent space trainer for volumetric rendering, the system comprising:

one or more processors;

a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to:

generate embeddings of image data that includes multiple images by sampling values along a viewing ray to generate 3D points and Fourier encoding the sampled points, where each image has a different viewpoint of a scene;

generate a latent space and a decoder, wherein the decoder receives embeddings as input to generate an output viewpoint;

for each viewpoint in the image data, determine a volumetric rendering view synthesis loss and a multi-view photometric loss; and

apply an optimization algorithm to the latent space and the decoder over a number of epochs until the volumetric rendering view synthesis loss is within a volumetric threshold and the multi-view photometric loss is within a multi-view threshold.

11 . The system of claim 10 , wherein the optimization algorithm uses a Mean Square Error objective for the volumetric rendering view synthesis loss.

12 . The system of claim 10 , wherein the optimization algorithm is a gradient descent algorithm.

13 . The system of claim 10 , wherein the multi-view photometric loss is determined using a photometric objective.

14 . The system of claim 13 , wherein the photometric objective is determined by:

for each pixel of a target image of the image data, with a predicted depth ({circumflex over (d)}), generating, by a warping operation, projected coordinates with a predicted depth ({circumflex over (d)}′) in a context image;

generating a synthesized target image from the context image; and

determining a difference between the target image and the synthesized target image.

15 . The system of claim 14 , wherein the context image is generated by a transformation matrix.

16 . The system of claim 14 , wherein the difference between the target image and the synthesized target image is determined by a weighted structural similarity index.

17 . The system of claim 14 , wherein the synthesized target image is generated by applying grid sampling with bilinear interpolation to place information from the context image onto each target pixel of the synthesized target image based on the projected coordinates.

18 . The system of claim 14 , wherein pixels of the target image for determining the photometric objective are determined using strided ray sampling.

19 . A tangible computer-readable medium comprising instructions that, when executed, cause a system to:

generate embeddings of image data that includes multiple images by sampling values along a viewing ray to generate 3D points and Fourier encoding the sampled points, where each image has a different viewpoint of a scene;

generate a latent space and a decoder, wherein the decoder receives embeddings as input to generate an output viewpoint;

for each viewpoint in the image data, determine a volumetric rendering view synthesis loss and a multi-view photometric loss; and

apply an optimization algorithm to the latent space and the decoder over a number of epochs until the volumetric rendering view synthesis loss is within a volumetric threshold and the multi-view photometric loss is within a multi-view threshold.

20 . The system of claim 19 , wherein the multi-view photometric loss is determined using a photometric objective.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2026
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 074440/0174 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2023
From: GUIZILINI, VITOR; AMBRUS, RARES A.; FANG, JIADING; ZAKHAROV, SERGEY; VASILJEVIC, IGOR; GAIDON, ADRIEN
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 064485/0484 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2023
From: SITZMANN, VINCENT
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 064485/0568 →
Continuity (2)
Provisional Application 63382776 · Nov 8, 2022
Related Publication 20240153197A1 · May 9, 2024
References Cited (15)
US 11403807B2 · Kim et al. · 2022 [cited by applicant]
US 11967015B2 · Shan · 2024 [cited by examiner]
US 20190244107A1 · Murez · 2019 [cited by examiner]
US 20200349772A1 · Tkach · 2020 [cited by examiner]
US 20210158561A1 · Park · 2021 [cited by examiner]
US 20210248811A1 · Shan · 2021 [cited by examiner]
US 20210279952A1 · Chen et al. · 2021 [cited by applicant]
US 20220014723A1 · Pandey · 2022 [cited by examiner]
US 20220292781A1 · Bautista et al. · 2022 [cited by applicant]
US 20230281913A1 · Rematas · 2023 [cited by examiner]
US 20240087214A1 · Martin Brualla · 2024 [cited by examiner]
WO 2019073267A1 · 2019 [cited by applicant]
WO 2022167602A2 · 2022 [cited by applicant]
Ke Qiu, Yawen Lai, Shiyi Liu, and Ronggang Wang; Self-supervised Multi-view Stereo via Inter and Intra Network Pseudo Depth; Proceedings of the 30th ACM International Conference on Multimedia, 2022; pp. 2305-2313; Oct. … [cited by examiner]
Depth field networks for generalizable multi-view scene representation (https://arxiv.org/pdf/2207.14287v1.pdf), dated Jul. 28, 2022. [cited by applicant]