IP Library › Granted Patent US 12,205,292
Granted Patent B2
US 12,205,292 · App. 17/378,155 · Granted Jan 21, 2025

Methods and systems for semantic segmentation of a point cloud

Inventors: Ran Cheng (Markham, CA); Ryan Razani (North York, CA); Bingbing Liu (Markham, CA)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06T7/11G01S17/89G01S17/931G06F18/24G06F18/253G06T2207/10028G06T2207/20016G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,292
App. No.
17/378,155
Granted
Jan 21, 2025
Kind
B2
Abstract

Systems, methods and apparatus for sematic segmentation of 3D point clouds using deep neural networks. The deep neural network generally has two primary subsystems: a multi-branch cascaded subnetwork that includes an encoder and a decoder, and is configured to receive a sparse 3D point cloud, and capture and fuse spatial feature information in the sparse 3D point cloud at multiple scales and multi hierarchical levels; and a spatial feature transformer subnetwork that is configured to transform the cascaded features generated by the multi-branch cascaded subnetwork and fuse these scaled features using a shared decoder attention framework to assist in the prediction of sematic classes for the sparse 3D point cloud.

Claims (60)

1. A method for semantic segmentation of a 3D point cloud, the method comprising:

processing a 3D point cloud to produce a sparse tensor;

feeding the sparse tensor as an input to each of a plurality of branches of an encoder of a neural network to produce a plurality of branch feature maps, N being a number of the plurality of branches, N being equal to or greater than 3, each ith branch respectively comprising i sequentially chained different encoder blocks to produce an ith branch feature map, i being an integer between 1 and N;

feeding the plurality of branch feature maps to a plurality of hierarchical attention blocks to generate a plurality of emphasized feature maps, wherein, for each pth branch of a 3rd to Nth branches, a pth branch feature map and a (p−2) th emphasized feature map are fed to a corresponding (p−1) th hierarchical attention block, the (p−2) th emphasized feature map is output by a (p−2) th hierarchical attention block, and wherein a first branch feature map and a second branch feature map are fed to a first hierarchical attention block;

feeding each emphasized feature map output by the plurality of hierarchical attention blocks to a spatial feature transformer to fuse each emphasized feature map of the plurality of hierarchical attention blocks and generate a fused feature map; and

processing the fused feature map and a final decoder block of a decoder to predict a class label for a plurality of points in the 3D point cloud.

2. The method of claim 1 , wherein processing the 3D point cloud to produce the sparse tensor is obtained by pre-processing the 3D point cloud to generate a voxel representation of the 3D point cloud.

3. The method of claim 2 , wherein the sparse tensor comprises for each point in the 3D point cloud, a set of coordinates and one or more associated features corresponding to the set of coordinates.

4. The method of claim 3 , wherein each set of coordinates is contained within a coordinate matrix, wherein the one or more associated features are contained within a feature matrix.

5. The method of claim 1 , further comprising, feeding an (N−1) th emphasized feature map output by an (N−1) th hierarchical attention block to a first decoder block.

6. The method of claim 5 , wherein the first decoder block is first of N decoder blocks.

7. The method of claim 6 , further comprising, feeding (N−1) encoder-decoder skip connection outputs from a first through (N−1) th encoder blocks of N encoder blocks to the N decoder blocks, wherein encoder-decoder skip connection outputs are fed to the N decoder blocks by reverse order of respective depth.

8. The method of claim 7 , wherein processing the fused feature map comprises feeding the fused feature map to an nth decoder block.

9. The method of claim 8 , further comprising fusing the fused feature map, an output of an (N−1) th decoder block and the output of first encoder blocks, wherein the fusing comprises concatenation followed by a convolution operation.

10. The method of claim 1 , further comprising scaling each emphasized feature map output by the plurality of hierarchical attention blocks to a common scale, prior to obtaining the fused feature map.

11. The method of claim 1 , further comprising assigning a weight to each of a plurality of channels, the plurality of channels corresponding to each output of the plurality of hierarchical attention blocks, prior to obtaining the fused feature map.

12. The method of claim 11 , wherein a kernel size of each encoder block is given according to:

K

=

⌊

N

+

2

-

p

2

M

⌋

+

3

wherein K is the kernel size, and M is block depth, and is a floor operation that rounds a value of

N

+

2

-

p

2

M

to a nearest integer value.

13. The method of claim 1 , wherein, for the first hierarchical attention block of the plurality of hierarchical attention blocks, the first hierarchical attention block comprises a first convolutional operation and a second convolutional operation.

14. The method of claim 13 , wherein, when a (p−1) th branch feature map and the pth branch feature map are fed to the corresponding (p−1) th hierarchical attention block, the pth branch feature map is fed to the second convolutional operation.

15. The method of claim 14 , wherein, when the (p−1) th branch feature map and the pth branch feature map are fed to the corresponding (p−1) th hierarchical attention block, the (p−1) th branch feature map is fed to the first convolutional operation.

16. The method of claim 15 , wherein, when the (p−1) th branch feature map and the pth branch feature map are fed to the corresponding (p−1) th hierarchical attention block, the pth branch feature map is upsampled and fed to the first convolutional operation.

17. The method of claim 16 , wherein, when the (p−1) th branch feature map and the pth branch feature map are fed to the corresponding (p−1) th hierarchical attention block, the (p−1) th branch feature map is downsampled and fed to the second convolutional operation.

18. The method of any one of claim 17 , further comprising:

adding a first output and a second output from the first convolutional operation and the second convolutional operation, respectively, to obtain an emphasized feature map from a hierarchical attention block.

19. An apparatus for semantic segmentation of a 3D point cloud, the apparatus comprising:

a memory storing executable instructions for implementing a neural network; and

at least one processor configured to execute the executable instructions to:

process a 3D point cloud to produce a sparse tensor;

feed the sparse tensor as an input to each of a plurality of branches of an encoder of the neural network to produce a plurality of branch feature maps, N being a number of the plurality of branches, N being equal to or greater than 3, each ith branch respectively comprising i sequentially chained different encoder blocks to produce an ith branch feature map, i being an integer between 1 and N;

feed the plurality of branch feature maps to a plurality of hierarchical attention blocks to generate a plurality of emphasized feature maps, wherein, for each pth branch of a 3rd to Nth branches, a pth branch feature map and the a (p−2) th emphasized feature map are fed to a corresponding (p−1) th hierarchical attention block, the (p−2) th emphasized feature map is output by a (p−2) th hierarchical attention block, and wherein a first branch feature map and a second branch feature map are fed to a first hierarchical attention block;

feed each emphasized feature map output by the plurality of hierarchical attention blocks to a spatial feature transformer to fuse each emphasized feature map of the plurality of hierarchical attention blocks and generate a fused feature map; and

process the fused feature map and a final decoder block to predict a label for a plurality of points in the 3D point cloud.

20. A non-transitory computer readable medium storing executable instructions which, when executed by a computer, cause at least one processor of the computer to:

process a 3D point cloud to produce a first sparse tensor;

feed the first sparse tensor as an input to each of a plurality of branches of an encoder of a neural network to produce a plurality of branch feature maps, N being a number of the plurality of branches, N being equal to or greater than 3, each ith branch respectively comprising i sequentially chained different encoder blocks to an ith branch feature map, i being an integer between 1 and N;

feed the plurality of branch feature maps to a plurality of hierarchical attention blocks to generate a plurality of emphasized feature maps, wherein, for each pth branch of a 3rd to Nth branches, a pth branch feature map and a (p−2) th emphasized feature map are fed to a corresponding (p−1) th hierarchical attention block, the (p−2) th emphasized feature map is output by a (p−2) th hierarchical attention block, and wherein a first branch feature map and a second branch feature map are fed to a first hierarchical attention block;

feed each emphasized feature map output by the plurality of hierarchical attention blocks to a spatial feature transformer to fuse each emphasized feature map of the plurality of hierarchical attention blocks and generate a fused feature map; and

process the fused feature map and a final decoder block to predict a label for a plurality of points in the 3D point cloud.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2021
From: CHENG, RAN; RAZANI, RYAN; LIU, BINGBING
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 057603/0939 →
Continuity (1)
Related Publication 20230035475A1 · Feb 2, 2023
References Cited (26)
US 10332001B2 · Rippel · 2019 [cited by examiner]
US 10970518B1 · Zhou et al. · 2021 [cited by applicant]
US 12079970B2 · Cheng · 2024 [cited by examiner]
US 20190096125A1 · Schulter · 2019 [cited by examiner]
US 20200349697A1 · Gao · 2020 [cited by examiner]
US 20210042557A1 · Roy · 2021 [cited by examiner]
US 20210082181A1 · Shi et al. · 2021 [cited by applicant]
US 20220164923A1 · Lee · 2022 [cited by examiner]
US 20220196798A1 · Chen · 2022 [cited by examiner]
US 20220318557A1 · Mohseni · 2022 [cited by examiner]
US 20240296605A1 · Kozlov · 2024 [cited by examiner]
CN 108665496A · 2018 [cited by applicant]
CN 111311611A · 2020 [cited by applicant]
EP 4012649A1 · 2022 [cited by examiner]
EP 4033402A1 · 2022 [cited by examiner]
GB 2602255A · 2022 [cited by examiner]
WO WO2022045495A1 · 2022 [cited by examiner]
Ivan David, “Method for Determining the Encoder Architecture of a Neural Network”, date filed: Jan. 26, 2021, date published: Jul. 27, 2022, English translation of EP 4033402 A1 (Year: 2021). [cited by examiner]
Wang Robin, “Optical Method”, date filed:Dec. 11, 2020, date published: Jun. 15, 2022, English translation of EP 4012649 A1 (Year: 2020). [cited by examiner]
Mehdi Bahri , “Method Of Generating A Latent Vector”, date filed Dec. 15, 2020, date published: 2022-06-39, English translation of GB 2602255 A (Year: 2020). [cited by examiner]
Feng, Guang et al. “Encoder Fusion Network with Co-Attention Embedding for Referring Image Segmentation.” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021): 15501-15510 (Year: 2021). [cited by examiner]
Choy et al., “4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 3075-3084. [cited by applicant]
Poudel et al., “Fast-SCNN: Fast Semantic Segmentation Network,” arXiv:1902.04502v1 [cs.CV], Feb. 12, 2019, retrieved online: https://arxiv.org/pdf/1902.04502.pdf. [cited by applicant]
Zhu et al., “Two-branch encoding and iterative attention decoding network for semantic segmentation,” Neural Computing and Applications, Sep. 1, 2020. [cited by applicant]
Yu et al., “BiSeNet V2: Bilateral Network with Guided Aggregation for Real-time Semantic Segmentation,” arXiv:2004.02147 [cs.CV], Apr. 5, 2020, retrieved online: https://arxiv.org/abs/2004.02147. [cited by applicant]
Choy, “High-dimensional convolutional neural networks for 3D perception,” Stanford University, Mar. 2020, retrieved online: http://purl.stanford.edu/fg022dx0979. [cited by applicant]