IP Library › Granted Patent US 10,671,082
Granted Patent B2
US 10,671,082 · App. 15/641,113 · Granted Jun 2, 2020

High resolution 3D point clouds generation based on CNN and CRF models

Inventors: Yu Huang (Sunnyvale, CA); Hsien-Ting Cheng (Sunnyvale, CA); Jun Zhu (Sunnyvale, CA); Weide Zhang (Sunnyvale, CA)
Assignee: BAIDU USA LLC
G05D1/0251G01S7/4808G01S17/86G01S17/89G01S17/931G05D1/00G05D1/0088G06N3/04G06N3/0454G06N7/005G06T3/40G06T3/4046G06T5/003G06T5/50G06T7/521G06T7/55G06T11/60H04N5/2258H04N5/23238H04N7/181H04N13/00H04N13/128H04N13/254H04N13/271G06N3/0445G06N3/0481G06N3/082G06N3/088G06T2207/10024G06T2207/10028G06T2207/20081G06T2207/20084G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,671,082
App. No.
15/641,113
Filed
Jul 3, 2017
Granted
Jun 2, 2020
Kind
B2
Art Unit
3665
USPC
701/28
Abstract

In one embodiment, a method or system generates a high resolution 3-D point cloud to operate an autonomous driving vehicle (ADV) from a low resolution 3-D point cloud and camera-captured image(s). The system receives a first image captured by a camera for a driving environment. The system receives a second image representing a first depth map of a first point cloud corresponding to the driving environment. The system determines a second depth map by applying a convolutional neural network model to the first image. The system generates a third depth map by applying a conditional random fields model to the first image, the second image and the second depth map, the third depth map having a higher resolution than the first depth map such that the third depth map represents a second point cloud perceiving the driving environment surrounding the ADV.

Claims (45)

1. A computer-implemented method for operating an autonomous driving vehicle (ADV), the method comprising:

receiving a first image captured by a first camera, the first image capturing a portion of a driving environment of the ADV;

receiving a second image representing a first depth map of a first point cloud corresponding to the portion of the driving environment produced by a light detection and ranging (LiDAR) device;

determining a second depth map by applying a convolutional neural network (CNN) model to the first image;

generating a third depth map by applying a conditional random fields (CRF) model to the first image, the second image, and the second depth map, including optimizing, using the CRF model, an overall cost in generating the third depth map, wherein the overall cost comprises at least an image matching cost for estimating image pixel smoothness, and a smoothness pairwise term for estimating image pixel discontinuities, wherein the third depth map has a higher resolution than the first depth map; and

operating the ADV using the third depth map, which represents a second point cloud utilized to perceive the driving environment surrounding the ADV.

2. The method of claim 1 , further comprising:

receiving a third image captured by a second camera; and

determining the second depth map by applying the CNN model to the first and the third images.

3. The method of claim 1 , wherein the first image comprises a cylindrical panorama image or a spherical panorama image.

4. The method of claim 3 , wherein the cylindrical panorama image or the spherical panorama image is generated based on a plurality of images captured by a plurality of camera devices.

5. The method of claim 3 , further comprising reconstructing the second point cloud by projecting the second depth map into a 3-D space based on the cylindrical panorama image or the spherical panorama image.

6. The method of claim 1 , further comprising mapping the third image onto an image plane of the first image.

7. The method of claim 6 , wherein the third depth map is generated by blending one or more generated depth maps, wherein the third depth map is a panorama map.

8. The method of claim 1 , wherein the CNN model comprises:

a plurality of contractive layers, wherein each contractive layer includes an encoder to downsample a respective input; and

a plurality of expansive layers coupled to the plurality of contractive layers, wherein each expansive layer includes a decoder to upsample a respective input.

9. The method of claim 8 , wherein information of the plurality of contractive layers are fed forward to the plurality of expansive layers.

10. The method of claim 8 , wherein each of the plurality of expansive layers includes a prediction layer to predict a depth map for a subsequent layer.

11. A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations, the operations comprising:

receiving a first image captured by a first camera, the first image capturing a portion of a driving environment of an autonomous driving vehicle (ADV);

receiving a second image representing a first depth map of a first point cloud corresponding to the portion of the driving environment produced by a light detection and ranging (LiDAR) device;

determining a second depth map by applying a convolutional neural network (CNN) model to the first image; and

generating a third depth map by applying a conditional random fields (CRF) model to the first image, the second image, and the second depth map, including optimizing, using the CRF model, an overall cost in generating the third depth map, wherein the overall cost comprises at least an image matching cost for estimating image pixel smoothness, and a smoothness pairwise term for estimating image pixel discontinuities, wherein the third depth map has a higher resolution than the first depth map; and

operating the ADV using the third depth map, which represents a second point cloud utilized to perceive the driving environment surrounding the ADV.

12. The non-transitory machine-readable medium of claim 11 , further comprising:

receiving a third image captured by a second camera; and

determining the second depth map by applying the CNN model to the first and the third images.

13. The non-transitory machine-readable medium of claim 11 , wherein the first image comprises a cylindrical panorama image or a spherical panorama image.

14. The non-transitory machine-readable medium of claim 13 , wherein the cylindrical panorama image or the spherical panorama image is generated based on a plurality of images captured by a plurality of camera devices.

15. The non-transitory machine-readable medium of claim 13 , further comprising reconstructing the second point cloud by projecting the second depth map into a 3-D space based on the cylindrical panorama image or the spherical panorama image.

16. A data processing system, comprising:

a processor; and

a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations, the operations including

receiving a first image captured by a first camera, the first image capturing a portion of a driving environment of an autonomous driving vehicle (ADV);

receiving a second image representing a first depth map of a first point cloud corresponding to the portion of the driving environment produced by a light detection and ranging (LiDAR) device;

determining a second depth map by applying a convolutional neural network (CNN) model to the first image; and

generating a third depth map by applying a conditional random fields (CRF) model to the first image, the second image, and the second depth map, including optimizing, using the CRF model, an overall cost in generating the third depth map, wherein the overall cost comprises at least an image matching cost for estimating image pixel smoothness, and a smoothness pairwise term for estimating image pixel discontinuities, wherein the third depth map has a higher resolution than the first depth map; and

operating the ADV using the third depth map, which represents a second point cloud utilized to perceive the driving environment surrounding the ADV.

17. The system of claim 16 , further comprising:

receiving a third image captured by a second camera; and

determining the second depth map by applying the CNN model to the first and the third images.

18. The system of claim 16 , wherein the first image comprises a cylindrical panorama image or a spherical panorama image.

19. The system of claim 18 , wherein the cylindrical panorama image or the spherical panorama image is generated based on a plurality of images captured by a plurality of camera devices.

20. The system of claim 18 , further comprising reconstructing the second point cloud by projecting the second depth map into a 3-D space based on the cylindrical panorama image or the spherical panorama image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 3, 2017
From: HUANG, YU; CHENG, HSIEN-TING; ZHU, JUN; ZHANG, WEIDE
To: BAIDU USA LLC
Reel/Frame 042884/0485 →
Continuity (1)
Related Publication 20190004535A1 · Jan 3, 2019
Cited By (21)
US 12,198,396 US 12,216,610 US 12,223,428 US 12,236,689 US 12,243,240 US 12,266,122 US 12,307,350 US 12,346,816 US 12,367,405 US 12,384,410 US 12,387,354 US 12,455,739 US 12,462,575 US 12,522,243 US 12,536,131 US 12,554,467 US 12,591,240 US 12,618,976 US 12,623,691 US 12,651,168 US 12,709,294