IP Library › Granted Patent US 12,223,705
Granted Patent B2
US 12,223,705 · App. 18/299,970 · Granted Feb 11, 2025

Semantic segmentation of three-dimensional data

Inventors: Chris Jia-Han Zhang (Mississauga, CA); Wenjie Luo (Toronto, CA); Raquel Urtasun (Toronto, CA)
Assignee: AURORA OPERATIONS, INC.
G06V10/82G01S7/4808G06T3/06G06V10/764G06V10/776G06V20/10G06V20/41G06V20/56G06V20/64G01S17/89
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,705
App. No.
18/299,970
Granted
Feb 11, 2025
Kind
B2
Abstract

Systems and methods for performing semantic segmentation of three-dimensional data are provided. In one example embodiment, a computing system can be configured to obtain sensor data including three-dimensional data associated with an environment. The three-dimensional data can include a plurality of points and can be associated with one or more times. The computing system can be configured to determine data indicative of a two-dimensional voxel representation associated with the environment based at least in part on the three-dimensional data. The computing system can be configured to determine a classification for each point of the plurality of points within the three-dimensional data based at least in part on the two-dimensional voxel representation associated with the environment and a machine-learned semantic segmentation model. The computing system can be configured to initiate one or more actions based at least in part on the per-point classifications.

Claims (50)

1. A computer-implemented method of semantic segmentation, the method comprising:

obtaining sensor data comprising three-dimensional data associated with an environment;

determining data indicative of a two-dimensional voxel representation associated with the environment based at least in part on the three-dimensional data and a voxel grid, wherein the two-dimensional voxel representation is associated with one or more voxels;

determining a voxel classification for respective voxels of the one or more voxels associated with the two-dimensional voxel representation, based at least in part on the two-dimensional voxel representation; and

determining a classification for respective points of a plurality of points within the three-dimensional data based at least in part on the voxel classification for the respective voxels.

2. The computer-implemented method of claim 1 , wherein determining a voxel classification for a respective voxel comprises:

determining that a voxel of the one or more voxels is associated with an object in the environment; and

determining a classification of the object based on the voxel classification.

3. The computer-implemented method of claim 2 , wherein the classification is indicative of (i) a static object or (ii) a dynamic object.

4. The computer-implemented method of claim 1 , further comprising:

determining that the sensor data is associated with a plurality of times; and

determining the classification for the respective points across the plurality of times.

5. The computer-implemented method of claim 4 , wherein the classification for the respective points across the plurality of times is associated with motion of one or more objects in the environment.

6. The computer-implemented method of claim 1 , further comprising determining a classification of one or more portions of the environment based on the voxel classification for the respective voxels.

7. The computer-implemented method of claim 1 , further comprising:

determining a probability distribution across a plurality of classes for the respective voxels; and

predicting the voxel classification for the respective voxels based on the probability distribution.

8. The computer-implemented method of claim 1 , wherein the voxel grid comprises a plurality of voxel cells.

9. A vehicle computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions executable by the one or more processors to perform operations, the operations comprising:

obtaining sensor data comprising three-dimensional data associated with an environment;

determining data indicative of a two-dimensional voxel representation associated with the environment based at least in part on the three-dimensional data and a voxel grid, wherein the two-dimensional voxel representation is associated with one or more voxels;

determining a voxel classification for respective voxels of the one or more voxels associated with the two-dimensional voxel representation, based at least in part on the two-dimensional voxel representation; and

determining a classification for respective points of a plurality of points within the three-dimensional data based at least in part on the voxel classification for the respective voxels.

10. The vehicle computing system of claim 9 , wherein determining a voxel classification for a respective voxel comprises:

determining that a voxel of the one or more voxels is associated with an object in the environment; and

determining a classification of the object based on the voxel classification.

11. The vehicle computing system of claim 10 , wherein the classification is indicative of (i) a static object or (ii) a dynamic object.

12. The vehicle computing system of claim 9 , further comprising:

determining that the sensor data is associated with a plurality of times; and

determining the classification for the respective points across the plurality of times.

13. The vehicle computing system of claim 12 , wherein the classification for the respective points across the plurality of times is associated with motion of one or more objects in the environment.

14. The vehicle computing system of claim 9 , further comprising determining a classification of one or more portions of the environment based on the voxel classification for the respective voxels.

15. The vehicle computing system of claim 9 , further comprising:

determining a probability distribution across a plurality of classes for the respective voxels; and

predicting the voxel classification for the respective voxels based on the probability distribution.

16. The vehicle computing system of claim 9 , wherein the voxel grid comprises a plurality of voxel cells.

17. One or more non-transitory computer-readable media storing instructions executable by one or more processors to cause the one or more processors to perform operations, the operations comprising:

obtaining sensor data comprising three-dimensional data associated with an environment;

determining data indicative of a two-dimensional voxel representation associated with the environment based at least in part on the three-dimensional data and a voxel grid, wherein the two-dimensional voxel representation is associated with one or more voxels;

determining a voxel classification for respective voxels of the one or more voxels associated with the two-dimensional voxel representation, based at least in part on the two-dimensional voxel representation; and

determining a classification for respective points of a plurality of points within the three-dimensional data based at least in part on the voxel classification for the respective voxels.

18. The one or more non-transitory computer-readable media of claim 17 , wherein determining a voxel classification for a respective voxel comprises:

determining that a voxel of the one or more voxels is associated with an object in the environment; and

determining a classification of the object based on the voxel classification.

19. The one or more non-transitory computer-readable media of claim 18 , wherein the classification is indicative of (i) a static object or (ii) a dynamic object.

20. The one or more non-transitory computer-readable media of claim 17 , further comprising:

determining that the sensor data is associated with a plurality of times; and

determining the classification for the respective points across the plurality of times.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 067733/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2023
From: UBER TECHNOLOGIES, INC.
To: UATC, LLC
Reel/Frame 065449/0620 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2023
From: ZHANG, CHRIS JIA-HAN; LUO, WENJIE; URTASUN, RAQUEL
To: UBER TECHNOLOGIES, INC.
Reel/Frame 065424/0470 →
Continuity (4)
Continuation 17208509 · Mar 22, 2021
Continuation 16123233 · Sep 6, 2018
Provisional Application 62586777 · Nov 15, 2017
Related Publication 20230252777A1 · Aug 10, 2023
References Cited (52)
US 6400831B2 · Lee · 2002 [cited by examiner]
US 9734455B2 · Levinson · 2017 [cited by examiner]
US 9811756B2 · Liu · 2017 [cited by examiner]
US 10203210B1 · Tagawa · 2019 [cited by examiner]
US 10303956B2 · Huang · 2019 [cited by examiner]
US 10445928B2 · Nehmadi · 2019 [cited by examiner]
US 10471955B2 · Kouri · 2019 [cited by examiner]
US 10546387B2 · Hirzer · 2020 [cited by examiner]
US 10678244B2 · Iandola · 2020 [cited by examiner]
US 20120148162A1 · Zhang · 2012 [cited by examiner]
US 20170124476A1 · Levinson · 2017 [cited by examiner]
US 20180275658A1 · Iandola · 2018 [cited by examiner]
US 20190051056A1 · Chiu · 2019 [cited by examiner]
US 20190080467A1 · Hirzer · 2019 [cited by examiner]
US 20190147220A1 · McCormac · 2019 [cited by examiner]
US 20190188477A1 · Mair · 2019 [cited by examiner]
I. Armeni, et al., “3D Semantic Parsing Of Largescale Indoor Spaces”, In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition, 2016. [cited by applicant]
V. Badrinarayanan, et al., “Segnet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation”, arXiv preprint arXiv:1511.00561, 2015. [cited by applicant]
L. C. Chen, et al., “Deeplab: Semantic Image Segmentation With Deep Convolutional Nets, Atrous Convolution, And Fully Connected CRFS”, arXiv preprint arXiv:1606.00915, 2016. [cited by applicant]
X. Chen, et al., “Multi-view 3D Object Detection Network for Autonomous Driving”, In IEEE CVPR, 2017. [cited by applicant]
O. Cicek, et al., “3D u-net: Learning Dense Volumetric Segmentation from Sparse Annotation”, In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 424-432. Springer, 2016. [cited by applicant]
A. Dai, et al., “Scannet: Richly-Annotated 3D Reconstructions of Indoor Scenes”, In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2017. [cited by applicant]
J. Dai, et al., “R-fcn: Object Detection Via Region-Based Fully Convolutional Networks”, In Advances in neural information processing systems, pp. 379-387, 2016. [cited by applicant]
D. Eigen, et al., “Predicting Depth, Surface Normal and Semantic Labels with a Common Multi-Scale Convolutional Architecture”, In Proceedings of the IEEE International Conference on Computer Vision, pp. 2650-2658, 2015. [cited by applicant]
A. Geiger, et al., “Are We Ready for Autonomous Driving? The KITTI Vision Benchmark Suite.” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012. [cited by applicant]
R. Girshick. Fast r-cnn. In Proceedings of the IEEE international conference on computer vision, pp. 1440-1448, 2015. [cited by applicant]
T. Hackel, et al., “Fast Semantic Segmentation of 3D Point Clouds with Strongly Varying Density”, ISPRS Annals of Photogrammetry, Remote Sensing & Spatial Information Sciences, 3(3), 2016. [cited by applicant]
K. He, et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016. [cited by applicant]
J. Huang, et al., “Point Cloud Labeling Using 3D Convolutional Neural Network”, In 2016 23rd International Conference on Pattern Recognition (ICPR), pp. 2670-2675, Dec. 2016. [cited by applicant]
S. Ioffe, et al. “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, In International Conference on Machine Learning, pp. 448-456, 2015. [cited by applicant]
D. Kingma, et al. “Adam: A method for Stochastic Optimization”, arXiv preprint arXiv:1412.6980, 2014. [cited by applicant]
A. Krizhevsky, et al., “Imagenet Classification With Deep Convolutional Neural Networks”, In Advances in neural information processing systems, pp. 1097-1105, 2012. [cited by applicant]
J. Long, et al., “Fully Convolutional Networks For Semantic Segmentation”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3431-3440, 2015. [cited by applicant]
D. Maturana, et al., “VoxNet: A 3D Convolutional Neural Network for Real-Time Object Recognition”, In IROS, 2015. [cited by applicant]
B. A. Olshausen, et al., Sparse Coding With an Over-Complete Basis Set: A Strategy Employed by vl? Vision research, 37(23):3311-3325, 1997. [cited by applicant]
C. R. Qi, et al., “Pointnet: Deep Learning on Point Sets for 3D Classification and Segmentation”, In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017. [cited by applicant]
C. R. Qi, et al., “Volumetric and Multi-View cnns for Object Classification On 3D Data”, In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2016. [cited by applicant]
C. R. Qi, et al., “Pointnet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space”, 2017. [cited by applicant]
X. Qi, et al., “3D Graph Neural Networks For rgbd Semantic Segmentation”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5199-5208, 2017. [cited by applicant]
S. Ren, et al., “Faster r-cnn: Towards Real-Time Object Detection with Region Proposal Networks”, In Advances in neural information processing systems, pp. 91-99, 2015. [cited by applicant]
G. Riegler, et al., “Octnet: Learning Deep 3D Representations at High Resolutions”, In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017. [cited by applicant]
O. Ronneberger, et al., “U-net: Convolutional Networks For Biomedical Image Segmentation”, In Medical Image Computing and Computer-Assisted Intervention (MICCAI), vol. 9351 of LNCS, pp. 234-241. Springer, 2015. (availab… [cited by applicant]
K. Simonyan, et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, arXiv preprint arXiv:1409.1556, 2014. [cited by applicant]
R. Socher, et al., “Convolutional-Recursive Deep Learning For 3D Object Classification”, In Advances in Neural Information Processing Systems, pp. 656-664, 2012. [cited by applicant]
S. Song, et al., “Deep Sliding Shapes For Amodal 3D Object Detection In rgb-d Images”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 808-816, 2016. [cited by applicant]
H. Su, et al., “Multi-View Convolutional Neural Networks for 3D Shape Recognition”, In Proc. ICCV, 2015. [cited by applicant]
C. Szegedy, et al., “Going Deeper With Convolutions”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1-9, 2015. [cited by applicant]
L. P. Tchapmi, et al., “Segcloud: Semantic Segmentation Of 3D Point Clouds”, In International Conference on 3D Vision (3DV), 2017. [cited by applicant]
M. Velas, et al., “CNN for Very Fast Ground Segmentation in Velodyne Lidar Data”, CoRR, abs/1709.02128, 2017. [cited by applicant]
M. Weinmann, et al., “Feature Relevance Assessment for the Semantic Interpretation of 3D Point Cloud Data”, ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, 5:W2, 2013. [cited by applicant]
M. Weinmann, et al., “Distinctive 2D and 3D Features for Automated Largescale Scene Analysis in Urban Areas”, Computers & Graphics, 49:47-57, 2015. [cited by applicant]
H. Zhao, et al., “Pyramid Scene Parsing Network”, arXiv preprint arXiv:1612.01105, 2016. [cited by applicant]