IP Library › Granted Patent US 12,190,520
Granted Patent B2
US 12,190,520 · App. 17/857,529 · Granted Jan 7, 2025

Pyramid architecture for multi-scale processing in point cloud segmentation

Inventors: Dong Nie (Bellevue, WA); Xiaofeng Ren (Yarrow Point, WA)
Assignee: Alibaba (China) Co., Ltd.
G06T7/11G06T3/40G06T2207/10028G06T2207/20016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,520
App. No.
17/857,529
Granted
Jan 7, 2025
Kind
B2
Abstract

This application describes cross-scale point cloud segmentation network architecture for and exemplary systems that utilize such network architecture for semantic segmentation of a point cloud. An embodiment of the network architecture includes an encoding path comprising a plurality of sequentially connected encoding nodes, a decoding path following the encoding path and comprising a plurality of sequentially connected decoding nodes, and a plurality of data links respectively corresponding to a plurality of levels of feature resolution, in which each of the plurality of data links connects one of the plurality of encoding nodes and one of the plurality of decoding nodes that have a same level of feature resolution.

Claims (64)

1. A computer-implemented method for point cloud segmentation comprising:

feeding a plurality of features extracted from an input point cloud into a point cloud segmentation network, wherein the point cloud segmentation network comprises:

an encoding path comprising a plurality of sequentially connected encoding nodes, wherein each encoding node decreases a feature resolution of an input of the encoding node;

a decoding path following the encoding path and comprising a plurality of sequentially connected decoding nodes, wherein each decoding node increases a feature resolution of an input of the decoding node;

a plurality of data links respectively corresponding to a plurality of levels of feature resolution, in which each of the plurality of data links connects one of the plurality of encoding nodes and one of the plurality of decoding nodes that have a same level of feature resolution, wherein:

at least one of the plurality of data links connects a first encoding node and a first decoding node and comprises one or more intermediate nodes between the first encoding node and the first decoding node, and

the at least one data link corresponds to a baseline feature resolution and exchanges data with (1) a first neighboring data link corresponding to a lower level of feature resolution than the baseline feature resolution and (2) a second neighboring data link corresponding to a higher level of feature resolution than the baseline feature resolution through the one or more intermediate nodes; and

obtaining an output from a last decoding node from the decoding path of the point cloud segmentation network for object classification or part segmentation.

2. The computer-implemented method of claim 1 , wherein when the at least one of the one or more intermediate nodes is a first intermediate node on the at least one data link, the first intermediate node is configured to:

receive a first input from a first node on the at least one data link, wherein the first node is the first encoding node connected by the data link;

receive a second input from a second node on the first neighboring data link corresponding to the lower level of feature resolution;

receive a third input from a third node on the second neighboring data link corresponding to the higher level of feature resolution; and

generate the output based on the first input, the second input, and the third input.

3. The computer-implemented method of claim 2 , wherein the second node is an encoding node following the first encoding node on the encoding path.

4. The computer-implemented method of claim 2 , wherein the third node is an intermediate node on the second neighboring data link corresponding to the higher level of feature resolution.

5. The computer-implemented method of claim 1 , wherein at least one intermediate node on the at least one data link corresponding to the baseline feature resolution is further configured to:

feed data to an intermediate node on the first neighboring data link corresponding to the lower level of feature resolution than the baseline feature resolution.

6. The computer-implemented method of claim 1 , wherein at least one intermediate node on the at least one data link corresponding to the baseline feature resolution is further configured to:

feed data to an intermediate node on the second neighboring data link corresponding to the higher level of feature resolution than the baseline feature resolution.

7. The computer-implemented method of claim 1 , wherein each of the plurality of encoding nodes is configured to perform subsampling to decrease the feature resolution.

8. The computer-implemented method of claim 1 , wherein each of the plurality of decoding nodes is configured to perform upsampling to increase the feature resolution.

9. The computer-implemented method of claim 1 , wherein the first decoding node is configured to:

receive a fourth input from a preceding decoding node on the decoding path;

receive a fifth input from a last intermediate node on the at least one data link;

receive a sixth input from a last intermediate node on the second neighboring data link corresponding to the higher feature resolution than the at least one data link; and

perform a feature fusion based on the fourth input, the fifth input, and the sixth input and feed a fusion result into a second decoding node that is subsequent to the first decoding node on the decoding path.

10. The computer-implemented method of claim 1 , wherein a first data link comprises more intermediate nodes than a second data link when the encoding node connected by the first data link has a higher feature resolution than the encoding node connected by the second data link.

11. The computer-implemented method of claim 2 , wherein the first input has a base feature resolution, the second input has a lower feature resolution than the base feature resolution and richer semantic information, and the third input has a higher feature resolution than the base feature resolution and richer detail information.

12. The computer-implemented method of claim 2 , wherein to generate the output based on the first input, the second input, and the third input, the first intermediate node is further configured to:

compute a semantic mask by applying a vector product operation on the first input and the second input;

compute a resolution mask by applying a vector addition operation on the first input and the third input;

transform the second input by applying the semantic mask;

transform the third input by applying the resolution mask;

transform the first input by applying a local aggregation on the first input; and

aggregate the first transformed input, the second transformed input, and the third transformed input to obtain the output.

13. The computer-implemented method of claim 12 , wherein prior to computing the semantic mask and the resolution mask, the first intermediate node is further configured to:

compress the first input, the second input, and the third input into a single-channel format using a multi-layer perceptron (MLP).

14. The computer-implemented method of claim 12 , wherein the semantic mask is computed by applying a sigmoid activation on an output of the vector product operation.

15. The computer-implemented method of claim 12 , wherein the resolution mask is computed by applying a sigmoid activation on an output of the vector addition operation.

16. The computer-implemented method of claim 12 , wherein to aggregate the first transformed input, the second transformed input, and the third transformed input, the first intermediate node is further configured to:

stack the first transformed input, the second transformed input, and the third transformed input to obtain multi-scale features; and

apply a multi-layer perceptron (MLP) to reduce channels of the multi-scale feature.

17. A cross-scale point cloud segmentation network architecture, comprising:

an encoding path comprising a plurality of sequentially connected encoding nodes, wherein each encoding node decreases a feature resolution of an input of the encoding node;

a decoding path following the encoding path and comprising a plurality of sequentially connected decoding nodes, wherein each decoding node increases a feature resolution of an input of the decoding node;

a plurality of data links respectively corresponding to a plurality of levels of feature resolution, in which each of the plurality of data links connects one of the plurality of encoding nodes and one of the plurality of decoding nodes that have a same level of feature resolution, wherein:

at least one of the plurality of data links connects a first encoding node and a first decoding node and comprises one or more intermediate nodes between the first encoding node and the first decoding node, and

at least one of the one or more intermediate nodes aggregates inputs from (1) a preceding intermediate node on the at least one data link corresponding to a baseline feature resolution, (2) an intermediate node on a first neighboring data link corresponding to a lower level of feature resolution than the baseline feature resolution, and (3) an intermediate node on a second neighboring data link corresponding to a higher level of feature resolution than the baseline feature resolution, generates an output based on the aggregated inputs and feeds the output into a next intermediate node towards a direction to the first decoding node.

18. The cross-scale point cloud segmentation network of claim 17 , wherein the at least one intermediate node on the at least one data link corresponding to the baseline feature resolution is further configured to:

feed the output to an intermediate node on the first neighboring data link corresponding to the lower level of feature resolution than the baseline feature resolution; and

feed the output to an intermediate node on the second neighboring data link corresponding to the higher level of feature resolution than the baseline feature resolution.

19. The cross-scale point cloud segmentation network of claim 17 , wherein when the at least one of the one or more intermediate nodes is a first intermediate node on the at least one data link, the first intermediate node is configured to:

receive a first input from the first encoding node;

receive a second input from a second encoding node on the first neighboring data link corresponding to the lower level of feature resolution;

receive a third input from a third encoding node on the second neighboring data link corresponding to the higher level of feature resolution; and

generate the output based on the first input, the second input, and the third input.

20. A non-transitory computer-readable storage medium for point cloud segmentation, configured with instructions executable by one or more processors to cause the one or more processors to perform operations comprising:

feeding a plurality of features extracted from an input point cloud into a point cloud segmentation network, wherein the point cloud segmentation network comprises:

an encoding path comprising a plurality of sequentially connected encoding nodes, wherein each encoding node decreases a feature resolution of an input of the encoding node;

a decoding path following the encoding path and comprising a plurality of sequentially connected decoding nodes, wherein each decoding node increases a feature resolution of an input of the decoding node;

a plurality of data links respectively corresponding to a plurality of levels of feature resolution, in which each of the plurality of data links connects one of the plurality of encoding nodes and one of the plurality of decoding nodes that have a same level of feature resolution, wherein:

at least one of the plurality of data links connects a first encoding node and a first decoding node and comprises one or more intermediate nodes between the first encoding node and the first decoding node, and

the at least one data link corresponds to a baseline feature resolution and exchanges data with (1) a first neighboring data link corresponding to a lower level of feature resolution than the baseline feature resolution and (2) a second neighboring data link corresponding to a higher level of feature resolution than the baseline feature resolution through the one or more intermediate nodes; and

obtaining an output from a last decoding node from the decoding path of the point cloud segmentation network for object classification or part segmentation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2022
From: NIE, DONG; REN, XIAOFENG
To: ALIBABA (CHINA) CO., LTD.
Reel/Frame 060399/0595 →
Continuity (1)
Related Publication 20240013399A1 · Jan 11, 2024
References Cited (116)
US 8494285B2 · Zhang · 2013 [cited by examiner]
US 8503730B2 · Kotaba · 2013 [cited by applicant]
US 8521418B2 · Ma et al. · 2013 [cited by applicant]
US 8699787B2 · Van Den et al. · 2014 [cited by applicant]
US 8860712B2 · Lowe et al. · 2014 [cited by applicant]
US 8948498B1 · Hickman et al. · 2015 [cited by applicant]
US 9082224B2 · Birtwistle et al. · 2015 [cited by applicant]
US 9153061B2 · Vaddadi et al. · 2015 [cited by applicant]
US 9412040B2 · Feng et al. · 2016 [cited by applicant]
US 9786062B2 · Sorkine-hornung et al. · 2017 [cited by applicant]
US 9865042B2 · Dai · 2018 [cited by examiner]
US 10178366B2 · Lucas · 2019 [cited by applicant]
US 10319146B2 · Steinbach et al. · 2019 [cited by applicant]
US 10321116B2 · Metzler et al. · 2019 [cited by applicant]
US 10504282B2 · Levinson et al. · 2019 [cited by applicant]
US 10679351B2 · El-Khamy · 2020 [cited by examiner]
US 10699477B2 · Levinson et al. · 2020 [cited by applicant]
US 11073619B2 · Ho · 2021 [cited by applicant]
US 20140037198A1 · Larlus-Larrondo · 2014 [cited by examiner]
US 20140328535A1 · Sorkine-hornung · 2014 [cited by applicant]
US 20150178988A1 · Montserrat Mora et al. · 2015 [cited by applicant]
US 20180276875A1 · Pylvaenaeinen et al. · 2018 [cited by applicant]
US 20190108639A1 · Tchapmi et al. · 2019 [cited by applicant]
US 20190114774A1 · Zhang · 2019 [cited by examiner]
US 20190139267A1 · Mokrushin · 2019 [cited by examiner]
US 20200099954A1 · Hemmer · 2020 [cited by examiner]
US 20200107033A1 · Joshi · 2020 [cited by examiner]
US 20200134375A1 · Zhan · 2020 [cited by examiner]
US 20200364570A1 · Kitamura · 2020 [cited by examiner]
US 20200364870A1 · Lee · 2020 [cited by examiner]
US 20200372676A1 · Tzur · 2020 [cited by applicant]
US 20210056730A1 · Ricard · 2021 [cited by examiner]
US 20210118163A1 · Guizilini · 2021 [cited by examiner]
US 20210192797A1 · Lasserre · 2021 [cited by examiner]
US 20210272327A1 · Wang · 2021 [cited by examiner]
US 20210350583A1 · Lasserre · 2021 [cited by examiner]
US 20210407147A1 · Flynn · 2021 [cited by examiner]
US 20220044358A1 · Wang · 2022 [cited by examiner]
US 20220058805A1 · Lee · 2022 [cited by examiner]
US 20220101489A1 · Nie · 2022 [cited by examiner]
US 20220121361A1 · Kamran · 2022 [cited by examiner]
US 20220164993A1 · Llach · 2022 [cited by examiner]
US 20220262002A1 · Wang · 2022 [cited by examiner]
US 20220292728A1 · Huang · 2022 [cited by examiner]
US 20220366612A1 · Taquet · 2022 [cited by examiner]
US 20220394283A1 · Cao · 2022 [cited by examiner]
US 20230048381A1 · Taquet · 2023 [cited by examiner]
US 20230072293A1 · Koh · 2023 [cited by examiner]
US 20230078840A1 · Salvi · 2023 [cited by examiner]
US 20230260197A1 · Chou · 2023 [cited by examiner]
US 20230410254A1 · Wang · 2023 [cited by examiner]
US 20230410377A1 · Zhang · 2023 [cited by examiner]
US 20240037840A1 · Ramesh Babu · 2024 [cited by examiner]
US 20240062515A1 · Oh · 2024 [cited by examiner]
US 20240070810A1 · Oosake · 2024 [cited by examiner]
US 20240127491A1 · Mammou · 2024 [cited by examiner]
US 20240212164A1 · Li · 2024 [cited by examiner]
Arief et al., Density-adaptive Sampling for Heterogeneous Point Cloud Object Segmentation in Autonomous Vehicle Applications, CVPR Workshop paper, 2019. [cited by applicant]
Armeni et al., “3D Semantic Parsing of Large-Scale Indoor Spaces,” CVPR paper, 2016. [cited by applicant]
Boulch, “ConvPoint: Continuous Convolutions for Point Cloud Processing,” Feb. 19, 2020. [cited by applicant]
Boulch et al., “FKAConv: Feature-Kernel Alignment for Point Cloud Convolution,” ACCV, 2020. [cited by applicant]
Chen et al., “DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs,” May 12, 2017. [cited by applicant]
Chen et al., “LSANet: Feature Learning on Point Sets by Local Spatial Aware Layer,” Jun. 20, 2019. [cited by applicant]
Dai et al., “Attentional Feature Fusion,” Nov. 9, 2020. [cited by applicant]
Dovrat et al., “Learning to Sample,” Apr. 1, 2019. [cited by applicant]
Engelmann et al., “Exploring Spatial Context for 3D Semantic Segmentation of Point Clouds,” Dec. 18, 2019. [cited by applicant]
Fan et al., “SCF-Net: Learning Spatial Contextual Features for Large-Scale Point Cloud Segmentation,” CVPR, 2021. [cited by applicant]
Hackel et al., “Semantic3D.net: A new Large-scale Point Cloud Classification Benchmark,” Apr. 12, 2017. [cited by applicant]
Hu et al., “RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds,” May 1, 2020. [cited by applicant]
Hu et al., “RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds,” CVPR 2020, May 1, 2020. [cited by applicant]
Hua et al., “Pointwise Convolutional Neural Networks,” Mar. 29, 2018. [cited by applicant]
Huang et al., “Recurrent Slice Networks for 3D Segmentation of Point Clouds,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018. [cited by applicant]
Jiang et al., “PointSIFT: A SIFT-like Network Module for 3D Point Cloud Semantic Segmentation,” Nov. 24, 2018. [cited by applicant]
Kamnitsas et al., “Efficient multi-scale 3D CNN with fully connected CRF for accurate brain lesion segmentation,” Feb. 2017. [cited by applicant]
Komarichev et al., “A-CNN: Annularly Convolutional Neural Networks on Point Clouds,” Apr. 16, 2019. [cited by applicant]
Lan et al., “Modeling Local Geometric Structure of 3D Point Clouds using Geo-CNN,” Nov. 19, 2018. [cited by applicant]
Landrieu et al., “Point Cloud Oversegmentation with Graph-Structured Deep Metric Learning,” Apr. 3, 2019. [cited by applicant]
Landrieu et al., “Large-scale Point Cloud Semantic Segmentation with Superpoint Graphs,” Mar. 28, 2018. [cited by applicant]
Lang et al., “SampleNet: Differentiable Point Cloud Sampling,” Apr. 4, 2020. [cited by applicant]
Li et al., “PointCNN: Convolution on X-Transformed Points,” Nov. 5, 2018. [cited by applicant]
Liu et al., “Path Aggregation Network for Instance Segmentation,” Sep. 18, 2018. [cited by applicant]
Liu et al., “A Closer Look at Local Aggregation Operators in Point Cloud Analysis,” 2020. [cited by applicant]
Lu et al., “CGA-Net: Category Guided Aggregation for Point Cloud Semantic Segmentation,” CVPR, 2021. [cited by applicant]
Ma et al. “Multi-Scale Point-Wise Convolutional Neural Networks for 3D Object Segmentation From LiDAR Point Clouds in Large-Scale Environments,” IEEE Transactions on Intelligent Transportation Systems, 2019. [cited by applicant]
Mao et al., “Interpolated Convolutional Networks for 3D Point Cloud Understanding,” Aug. 13, 2019. [cited by applicant]
Maturana et al., “VoxNet: A 3D Convolutional Neural Network for Real-Time Object Recognition,” 2015. [cited by applicant]
Nie et al., “Bidirectional Pyramid Networks for Semantic Segmentation,” ACCV, 2020. [cited by applicant]
Qi et al., “PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation,” Apr. 10, 2017. [cited by applicant]
Qi et al., “PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space,” Jun. 7, 2017. [cited by applicant]
Qiu et al., “Semantic Segmentation for Real Point Cloud Scenes via Bilateral Augmentation and Adaptive Fusion,” Apr. 13, 2021. [cited by applicant]
Riegler et al., “OctNet: Learning Deep 3D Representations at High Resolutions,” Apr. 10, 2017. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation,” May 18, 2015. [cited by applicant]
Roynard et al., “Classification of Point Cloud Scenes with Multiscale Voxel Deep Network,” Apr. 10, 2018. [cited by applicant]
Roynard et al., “Paris-Lille-3D: a large and high-quality ground truth urban point cloud dataset for automatic segmentation and classification,” Apr. 10, 2018. [cited by applicant]
Su et al., “SPLATNet: Sparse Lattice Networks for Point Cloud Processing,” May 9, 2018. [cited by applicant]
Sun et al., “Oriented Point Sampling for Plane Detection in Unorganized Point Clouds,” May 4, 2019. [cited by applicant]
Sun et al., “High-Resolution Representations for Labeling Pixels and Regions,” Apr. 9, 2019. [cited by applicant]
Tan et al., “EfficientDet: Scalable and Efficient Object Detection,” Jul. 27, 2020. [cited by applicant]
Tao et al, “Hierarchical multi-scale attention for semantic segmentation,” May 21, 2020. [cited by applicant]
Tatarchenko et al., “Tangent Convolutions for Dense Prediction in 3D,” Jul. 6, 2018. [cited by applicant]
Thomas et al., “Semantic Classification of 3D Point Clouds with Multiscale Spherical Neighborhoods,” Aug. 1, 2018. [cited by applicant]
Thomas et al., “KPConv: Flexible and Deformable Convolution for Point Clouds,” ICCV, 2019. [cited by applicant]
Wang et al., “Deep High-Resolution Representation Learning for Visual Recognition,” Mar. 13, 2020. [cited by applicant]
Wang et al., “Graph Attention Convolution for Point Cloud Semantic Segmentation,” 2019. [cited by applicant]
Wang et al., “Non-local Neural Networks,” 2018. [cited by applicant]
Woo et al., “CBAM: Convolutional Block Attention Module,” Jul. 18, 2018. [cited by applicant]
Wu et al., “PointConv: Deep Convolutional Networks on 3D Point Clouds,” Nov. 9, 2020. [cited by applicant]
Xu et al., “PAConv: Position Adaptive Convolution with Dynamic Kernel Assembling on Point Clouds,” Apr. 26, 2021. [cited by applicant]
Yang et al., “Modeling Point Clouds with Self-Attention and Gumbel Subset Sampling,” Apr. 6, 2019. [cited by applicant]
Ye et al., “Learning with Noisy Labels for Robust Point Cloud Segmentation,” 2021. [cited by applicant]
Ye et al., “3D Recurrent Neural Networks with Context Fusion for Point Cloud Semantic Segmentation,” ECCV 2018 paper, 2018. [cited by applicant]
Yu et al., “Deep Layer Aggregation,” Jan. 4, 2019. [cited by applicant]
Zhang et al., “ShellNet: Efficient Point Cloud Convolutional Neural Networks using Concentric Shells Statistics,” Aug. 17, 2019. [cited by applicant]
Zhao et al., “Pyramid Scene Parsing Network,” Apr. 27, 2017. [cited by applicant]
Zhao et al., “PointWeb: Enhancing Local Neighborhood Features for Point Cloud Processing,” 2019. [cited by applicant]
Kang et al., “PyramNet: Point Cloud Pyramid Attention Network and Graph Embedding Module for Classification and Segmentation,” Sep. 30, 2019. [cited by applicant]
Cited By (1)
US 12,450,690