IP Library Granted Patent US 11,652,972
Granted Patent B2
US 11,652,972 · App. 16/899,906 · Granted May 16, 2023

Systems and methods for self-supervised depth estimation according to an arbitrary camera

Inventors: Vitor Guizilini (Santa Clara, CA); Igor Vasiljevic (Chicago, IL); Rares A. Ambrus (San Francisco, CA); Sudeep Pillai (Santa Clara, CA); Adrien David Gaidon (Mountain View, CA)
Assignee: Toyota Research Institute, Inc.
H04N13/128G06N3/08G06T9/002G06T15/06H04N13/261H04N13/271H04N2013/0081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,652,972
App. No.
16/899,906
Granted
May 16, 2023
Kind
B2
Abstract

System, methods, and other embodiments described herein relate to improving depth estimates for monocular images using a neural camera model that is independent of a camera type. In one embodiment, a method includes receiving a monocular image from a pair of training images derived from a monocular video. The method includes generating, using a ray surface network, a ray surface that approximates an image character of the monocular image as produced by a camera having the camera type. The method includes creating a synthesized image according to at least the ray surface and a depth map associated with the monocular image.

Claims (46)

1. A depth system for improving depth estimates for monocular images using a neural camera model that is independent of a camera type, comprising:

one or more processors;

a memory communicably coupled to the one or more processors and storing:

a ray module including instructions that, when executed by the one or more processors, cause the one or more processors to:

receive a monocular image from a pair of training images derived from a monocular video, and

generate, using a ray surface network, a ray surface that approximates an image character of the monocular image as produced by a camera with a defined type; and

a training module including instructions that, when executed by the one or more processors, cause the one or more processors to train a depth model by creating a synthesized image according to at least the ray surface and a depth map associated with the monocular image and using the synthesized image to derive a loss according to a loss function that updates at least the depth model, and the ray surface network, wherein the training module includes the instructions to train according to a self-supervised structure from motion (SfM) process.

2. The depth system of claim 1 , wherein the training module includes instructions to create the synthesized image including instructions to apply the neural camera model by:

lifting pixels to produce three-dimensional points using the ray surface, the depth map, and a camera offset, and

projecting the three-dimensional points onto a context image to create the synthesized image.

3. The depth system of claim 2 , wherein the training module includes instructions to lift the pixels including instructions to scale predicted ray vectors from the ray surface using the depth map and adjust the predicted ray vectors according to the camera offset that is an origin of a reference coordinate system.

4. The depth system of claim 2 , wherein the training module includes instructions to project including instructions to apply a softmax approximation to derive each pixel in the synthesized image by identifying a predicted ray vector from the ray surface that corresponds with a direction associated with each of the three-dimensional points as defined relative to the camera offset, and

wherein the ray module includes instructions to generate the ray surface including instructions to learn the camera type to provide the ray surface as part of the neural camera model that approximates the camera type for a set of training data including the pair of training images.

5. The depth system of claim 2 , wherein the training module includes instructions to project includes determining a patch-based data association for searching each pixel in the synthesized image by defining search grids for target pixels according to coordinates of respective ones of the target pixels and a defined grid size, and

wherein the training module includes instructions to project the three-dimensional points into the synthesized image includes applying a softmax approximation with an annealing temperature to search over the search grids.

6. The depth system of claim 1 , wherein the ray surface is comprised of a residual component and a fixed component, wherein the residual component is learned by the ray surface network and the fixed component is a geometric prior for the camera type,

wherein the image character associated with the camera type includes at least a format of the monocular image and lens distortion, and

wherein the ray surface associates pixels within the monocular image with directions in an environment from which light that generates the pixels originate.

7. The depth system of claim 1 , wherein the training module further includes instructions to provide the synthesized image as part of training a depth model for generating the depth estimates.

8. A non-transitory computer-readable medium for improving depth estimates for monocular images using a neural camera model that is independent of a camera type and including instructions that when executed by one or more processors cause the one or more processors to:

receive a monocular image from a pair of training images derived from a monocular video;

generate, using a ray surface network, a ray surface that approximates an image character of the monocular image as produced by a camera with a defined type; and

train a depth model by creating a synthesized image according to at least the ray surface and a depth map associated with the monocular image and using the synthesized image to derive a loss according to a loss function that updates at least the depth model, and the ray surface network, wherein the training occurs according to a self-supervised structure from motion (SfM) process.

9. The non-transitory computer-readable medium of claim 8 , wherein the instructions to create the synthesized image include instructions to apply the neural camera model by:

lifting pixels to produce three-dimensional points using the ray surface, the depth map, and a camera offset, and

projecting the three-dimensional points onto a context image to create the synthesized image.

10. The non-transitory computer-readable medium of claim 9 , wherein the instructions to lift the pixels include instructions to scale predicted ray vectors from the ray surface using the depth map and adjust the predicted ray vectors according to the camera offset that is an origin of a reference coordinate system.

11. The non-transitory computer-readable medium of claim 9 , wherein the instructions to project include instructions to apply a softmax approximation to derive each pixel in the synthesized image by identifying a predicted ray vector from the ray surface that corresponds with a direction associated with each of the three-dimensional points as defined relative to the camera offset, and

wherein the instructions to generate the ray surface include instructions to learn the camera type to provide the ray surface as part of the neural camera model that approximates the camera type for a set of training data including the pair of training images.

12. The non-transitory computer-readable medium of claim 9 , wherein the instructions to project include determining a patch-based data association for searching each pixel in the synthesized image by defining search grids for target pixels according to coordinates of respective ones of the target pixels and a defined grid size, and

wherein the instructions to project the three-dimensional points into the synthesized image includes applying a softmax approximation with an annealing temperature to search over the search grids.

13. A method of improving depth estimates for monocular images using a neural camera model that is independent of a camera type, comprising:

receiving a monocular image from a pair of training images derived from a monocular video;

generating, using a ray surface network, a ray surface that approximates an image character of the monocular image as produced by a camera having the camera type; and

training a depth model by creating a synthesized image according to at least the ray surface and a depth map associated with the monocular image and using the synthesized image to derive a loss according to a loss function that updates at least the depth model, and the ray surface network, wherein the training occurs according to a self-supervised structure from motion (SfM) process.

14. The method of claim 13 , wherein creating the synthesized image includes applying the neural camera model by:

lifting pixels to produce three-dimensional points using the ray surface, the depth map, and a camera offset, and

projecting the three-dimensional points onto a context image to create the synthesized image.

15. The method of claim 14 , wherein lifting the pixels includes scaling predicted ray vectors from the ray surface using the depth map and adjusting the predicted ray vectors according to the camera offset that is an origin of a reference coordinate system.

16. The method of claim 14 , wherein projecting includes applying a softmax approximation to derive each pixel in the synthesized image by identifying a predicted ray vector from the ray surface that corresponds with a direction associated with each of the three-dimensional points as defined relative to the camera offset, and

wherein generating the ray surface includes learning the camera type to provide the ray surface as part of the neural camera model that approximates the camera type for a set of training data including the pair of training images.

17. The method of claim 14 , wherein projecting includes determining a patch-based data association for searching each pixel in the synthesized image by defining search grids for target pixels according to coordinates of respective ones of the target pixels and a defined grid size, and

wherein projecting the three-dimensional points into the synthesized image includes applying a softmax approximation with an annealing temperature to search over the search grids.

18. The method of claim 13 , wherein the ray surface is comprised of a residual component and a fixed component, wherein the residual component is learned by the ray surface network and the fixed component is a geometric prior for the camera type,

wherein the image character associated with the camera type includes at least a format of the monocular image and lens distortion, and

wherein the ray surface associates pixels within the monocular image with directions in an environment from which light that generates the pixels originate.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2023
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 064150/0962 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2020
From: GUIZILINI, VITOR; VASILJEVIC, IGOR; AMBRUS, RARES A.; PILLAI, SUDEEP; GAIDON, ADRIEN DAVID
To: TOYOTA RESEARCH INSTITUTE, INC.
Reel/Frame 052958/0821 →
Continuity (2)
Provisional Application 62984903 · Mar 4, 2020
Related Publication 20210281814A1 · Sep 9, 2021