IP Library › Granted Patent US 11,315,266
Granted Patent B2
US 11,315,266 · App. 16/716,077 · Granted Apr 26, 2022

Self-supervised depth estimation method and system

Inventors: Zhixin Yan (Sunnyvale, CA); Liang Mi (Emeryville, CA); Liu Ren (Cupertino, CA)
Assignee: Robert Bosch GmbH
G06T7/50G06N3/0454G06T5/002G06T7/11G06T2207/10028G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,315,266
App. No.
16/716,077
Granted
Apr 26, 2022
Kind
B2
Abstract

Depth perception has become of increased interest in the image community due to the increasing usage of deep neural networks for the generation of dense depth maps. The applications of depth perception estimation, however, may still be limited due to the needs of a large amount of dense ground-truth depth for training. It is contemplated that a self-supervised control strategy may be developed for estimating depth maps using color images and data provided by a sensor system (e.g., sparse LiDAR data). Such a self-supervised control strategy may leverage superpixels (i.e., group of pixels that share common characteristics, for instance, pixel intensity) as local planar regions to regularize surface normal derivatives from estimated depth together with the photometric loss. The control strategy may be operable to produce a dense depth map that does not require a dense ground-truth supervision.

Claims (40)

1. A method for self-supervised depth estimation, comprising:

receiving a digital image of an environment;

extracting one or more deep superpixel segmentations from the digital image, wherein the one or more deep superpixel segmentations are partitioned to represent a homogenous area of the digital image, and wherein the one or more deep superpixel segmentations are operable as local planar regions that constrain a local normal direction and a secondary derivative of depth within the one or more deep superpixel segmentations;

generating a dense depth map using the one or more deep superpixel segmentations and

smoothing and suppressing inconsistencies within the dense depth map by minimizing a depth secondary derivative within the one or more deep superpixel segmentations.

2. The method of claim 1 , further comprising:

receiving a sparse depth map sample; and

deriving a surface normal map using a depth regression neural network that regresses a full resolution depth map from the digital image and the sparse depth map sample.

3. The method of claim 2 , wherein the depth regression neural network is designed using an encoder-decoder structure having an encoding layer, a decoding layer, and a plurality of skip connections.

4. The method of claim 3 , wherein the encoding layer includes one or more convolutional layers, one or more rectified linear unit (ReLU) layers, one or more residual neural networks (ResNet), and one or more pooling layers.

5. The method of claim 3 , wherein the decoding layer includes one or more deconvolutional layers, one or more unpooling layers, one or more residual neural networks (ResNet) layers, and one or more rectified linear unit (ReLU) layers.

6. The method of claim 5 , wherein a final convolution layer operates to produce a non-negative gray-scale depth image that is used to derive the surface normal map.

7. The method of claim 2 , further comprising:

computing a gradient of the sparse depth map sample in four directions;

converting the sparse depth map sample into one or more 3-dimensional vectors; and

averaging one or more normalized cross products of the one or more 3-dimensional vectors to determine a vertex normal.

8. The method of claim 1 , further comprising: determining a relative transformation between the digital image and a related image using a simultaneous localization and mapping system.

9. The method of claim 8 , further comprising: determining a photometric loss using the relative transformation, the digital image, and the related image.

10. The method of claim 1 , further comprising negating a boundary and an edge within the one or more deep superpixel segmentations.

11. The method of claim 1 , wherein the local normal direction is derived using an estimated depth.

12. The method of claim 1 , further comprising applying a consistency of normal direction within each of the one or more deep superpixel segmentations.

13. A method for self-supervised depth estimation, comprising:

receiving a digital image of an environment;

extracting one or more deep superpixel segmentations from the digital image, wherein the one or more deep superpixel segmentations are partitioned to represent a homogenous area of the digital image;

generating a dense depth map using the one or more deep superpixel segmentations;

receiving a sparse depth map sample; and

deriving a surface normal map using a depth regression neural network that regresses a full resolution depth map from the digital image and the sparse depth map sample, wherein a final convolution layer operates to produce a non-negative gray-scale depth image that is used to derive the surface normal map.

14. A system for self-supervised depth estimation, comprising:

a sensor operable to receive a digital image of an environment;

a controller operable to:

extract one or more deep superpixel segmentations from the digital image, wherein the one or more deep superpixel segmentations are partitioned to represent a homogenous area of the digital image, and wherein the one or more deep superpixel segmentations are operable as local planar regions that constrain a local normal direction and secondary derivative of depth within the one or more deep superpixel segmentations;

generate a dense depth map using the one or more deep superpixel segmentations; and

smooth and suppress inconsistencies within the dense depth map by minimizing a depth secondary derivative within the one or more deep superpixel segmentations.

15. The system of claim 14 , further comprising:

a depth sensor operable to receive a sparse depth map sample; and

the controller further being operable to:

derive a surface normal map using a depth regression neural network that regresses a full resolution depth map from the digital image and the sparse depth map sample.

16. The system of claim 15 , wherein the sensor is a digital camera and the depth sensor is a LiDAR sensor.

17. The system of claim 15 , the controller further being operable to: negate a boundary and an edge within the one or more deep superpixel segmentations.

18. The system of claim 15 , the controller further being operable to: determine a relative transformation between the digital image and a related image using a simultaneous localization and mapping system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2019
From: YAN, ZHIXIN; MI, LIANG; REN, LIU
To: ROBERT BOSCH GMBH
Reel/Frame 051297/0704 →
Continuity (1)
Related Publication 20210183083A1 · Jun 17, 2021