IP Library Granted Patent US 10,824,862
Granted Patent B2
US 10,824,862 · App. 16/119,939 · Granted Nov 3, 2020

Three-dimensional object detection for autonomous robotic systems using image proposals

Inventors: Ruizhongtai Qi (Stanford, CA); Wei Liu (Mountain View, CA); Chenxia Wu (Menlo Park, CA)
Assignee: Nuro, Inc.
G06K9/00664G06K9/00208G06K9/00791G06K9/3233G06K9/6268G06K9/6273G06K9/66
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,824,862
App. No.
16/119,939
Granted
Nov 3, 2020
Kind
B2
Abstract

Provided herein are methods and systems for implementing three-dimensional perception in an autonomous robotic system comprising an end-to-end neural network architecture that directly consumes large-scale raw sparse point cloud data and performs such tasks as object localization, boundary estimation, object classification, and segmentation of individual shapes or fused complete point cloud shapes.

Claims (42)

1. A computer-implemented method of implementing three-dimensional perception in an autonomous robotic system, the method comprising:

a) receiving, at a processor, two-dimensional image data from an optical camera;

b) generating, by the processor, an attention region in the two-dimensional image data, the attention region marking an object of interest;

c) receiving, at the processor, three-dimensional depth data from a depth sensor, the depth data corresponding to the image data;

d) extracting, by the processor, a three-dimensional frustum from the depth data corresponding to the attention region; and

e) applying, by the processor, a deep learning model to the frustum to:

i) generate and regress an oriented three-dimensional boundary for the object of interest; and

ii) classify the object of interest based on a combination of features from the attention region of the two-dimensional image data and the three-dimensional depth data within and around the regressed boundary.

2. The method of claim 1 , wherein the classification the object of interest is further based on at least one of the two-dimensional image data and the three-dimensional depth data.

3. The method of claim 1 , wherein the autonomous robotic system is an autonomous vehicle.

4. The method of claim 1 , wherein the two-dimensional image data is RGB image data.

5. The method of claim 1 , wherein the two-dimensional image data is IR image data.

6. The method of claim 1 , wherein the depth sensor comprises a LiDAR.

7. The method of claim 6 , wherein the three-dimensional depth data comprises a sparse point cloud.

8. The method of claim 1 , wherein the depth sensor comprises a stereo camera or a time-of-flight sensor.

9. The method of claim 8 , wherein the three-dimensional depth data comprises a dense depth map.

10. The method of claim 1 , wherein the deep learning model comprises a PointNet.

11. The method of claim 1 , wherein the deep learning model comprises a three-dimensional convolutional neural network on voxelized volumetric grids of the point cloud in frustum.

12. The method of claim 1 , wherein the deep learning model comprises a two-dimensional convolutional neural network on bird's eye view projection of the point cloud in frustum.

13. The method of claim 1 , wherein the deep learning model comprises a recurrent neural network on the sequence of the three-dimensional points from close to distant.

14. The method of claim 1 , wherein the classifying comprises semantic classification to apply a category label to the object of interest.

15. An autonomous robotic system comprising: an optical camera, a depth sensor, a memory, and at least one processor configured to:

a) receive two-dimensional image data from the optical camera;

b) generate an attention region in the two-dimensional image data, the attention region marking an object of interest;

c) receive three-dimensional depth data from the depth sensor, the depth data corresponding to the image data;

d) extract a three-dimensional frustum from the depth data corresponding to the attention region; and

e) apply a deep learning model to the frustum to:

i) generate and regress an oriented three-dimensional boundary for the object of interest; and

ii) classify the object of interest based on a combination of features from the attention region of the two-dimensional image data and the three-dimensional depth data within and around the regressed boundary.

16. The system of claim 15 , wherein the classification the object of interest is further based on at least one of the two-dimensional image data and the three-dimensional depth data.

17. The system of claim 15 , wherein the autonomous robotic system is an autonomous vehicle.

18. The system of claim 15 , wherein the two-dimensional image data is RGB image data.

19. The system of claim 15 , wherein the two-dimensional image data is IR image data.

20. The system of claim 15 , wherein the depth sensor comprises a LiDAR.

21. The system of claim 20 , wherein the three-dimensional depth data comprises a sparse point cloud.

22. The system of claim 15 , wherein the depth sensor comprises a stereo camera or a time-of-flight sensor.

23. The system of claim 22 , wherein the three-dimensional depth data comprises a dense depth map.

24. The system of claim 15 , wherein the deep learning model comprises a PointNet.

25. The system of claim 15 , wherein the deep learning model comprises a three-dimensional convolutional neural network on voxelized volumetric grids of the point cloud in frustum.

26. The system of claim 15 , wherein the deep learning model comprises a two-dimensional convolutional neural network on bird's eye view projection of the point cloud in frustum.

27. The system of claim 15 , wherein the deep learning model comprises a recurrent neural network on the sequence of the three-dimensional points from close to distant.

28. The system of claim 15 , wherein the classifying comprises semantic classification to apply a category label to the object of interest.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2019
From: QI, RUIZHONGTAI; LIU, WEI; WU, CHENXIA
To: NURO, INC.
Reel/Frame 048285/0573 →
Continuity (3)
Provisional Application 62588194 · Nov 17, 2017
Provisional Application 62586009 · Nov 14, 2017
Related Publication 20190147245A1 · May 16, 2019
Cited By (26)
US 12,198,396 US 12,216,610 US 12,223,428 US 12,236,689 US 12,283,120 US 12,307,350 US 12,346,816 US 12,367,405 US 12,399,253 US 12,437,412 US 12,455,739 US 12,462,575 US 12,494,034 US 12,522,243 US 12,525,031 US 12,536,131 US 12,554,467 US 12,567,171 US 12,591,240 US 12,618,976 US 12,623,691 US 12,632,978 US 12,651,465 US 12,700,129 US 12,709,294 US 12,722,296