IP Library Patent Application 18434501
Patent Application
App. No. 18/434,501

DETECTION OF OBJECTS IN LIDAR POINT CLOUDS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/434,501
Abstract

A method of processing point cloud information includes converting points in a point cloud obtained from a lidar sensor into a voxel grid, generating, from the voxel grid, sparse voxel features by applying a multi-layer perceptron and one or more max pooling layers that reduce dimension of input data; applying a cascade of an encoder that performs a N-stage sparse-to-dense feature operation, a global context pooling (GCP) module, and an M-stage decoder that performs a dense-to-sparse feature generation operation. The GCP module bridges an output of a last stage of the N-stages with an input of a first stage of the M-stages, where N and M are positive integers. The GCP module comprises a multi-scale feature extractor; and performing one or more perception operations on an output of the M-stage decoder and/or an output of the GCP module.

Claims (55)

1 . A method of processing point cloud information, comprising:

converting points in a point cloud obtained from a lidar sensor into a voxel grid;

generating, from the voxel grid, sparse voxel features by applying a multi-layer perceptron and one or more max pooling layers that reduce dimension of input data;

applying a cascade of an encoder that performs a N-stage sparse-to-dense feature operation, a global context pooling (GCP) module, and an M-stage decoder that performs a dense-to-sparse feature generation operation,

wherein the GCP module bridges an output of a last stage of the N-stages with an input of a first stage of the M-stages, where N and M are positive integers; and

wherein the GCP module comprises a multi-scale feature extractor; and

performing one or more perception operations on an output of the M-stage decoder and/or an output of the GCP module.

2 . The method of claim 1 , wherein the one or more perception operations comprise obtaining a three-dimensional (3D) point-wise segmentation by attaching a 3D segmentation head to the output of a last stage of the M-stage decoder; obtaining voxel-level predictions at an output of the 3D segmentation head, and performing a de-voxelization on the voxel-level predictions to obtain point-wise segmentation results.

3 . The method of claim 2 , wherein the one or more perception operations comprise obtaining a panoptic segmentation by applying a second stage refinement to the point-wise segmentation results.

4 . The method of claim 3 , wherein the one or more perception operations comprise generating a panoptic segmentation result from the second stage refinement.

5 . The method of claim 1 , wherein N=4, and wherein stages of the N-stages have increasing channel width.

6 . The method of claim 1 , wherein each stage of the N-stage comprises a sparse convolution layer followed by two submanifold sparse convolution blocks.

7 . The method of claim 6 , wherein the sparse convolution layer of each of N-stages, except for first stage, has a stride of 2 such that spatial resolution is downsampled by a factor of 8 in the encoder.

8 . The method of claim 1 , wherein M=4 such that the decoder comprises symmetrical stages of 3D sparse deconvolution blocks with decreasing channel width except for a last stage.

9 . The method of claim 3 , wherein the second stage refinement is obtained by:

fusing detected boxes, voxel-wise features and bird's eye view (BEV) features;

applying a local coordinate transformation to points within each detected box;

calculating box classification scores and point-wise mask scores; and

determining output of the second stage refinement by fusing the box classification scores and point-wise mask scores and an output of a previous refinement stage.

10 . The method of claim 1 , wherein the GCP module is configured to operate as:

sparse voxel features into a dense feature map;

generating a 2D BEV feature map by concatenating features in different heights;

extracting long term contextual information by using a 2D convolutional neural network;

reshaping encoded BEV feature representation to a dense voxel map; and

transforming the dense voxel map to a sparse voxel by applying a dense-to-sparse conversion.

11 . A method of performing perception of a point cloud data obtained from a lidar, comprising:

generating a three-dimensional perception output from the point cloud data by processing the point cloud data through a cascade of three stages, wherein the cascade includes:

a first stage in which the point cloud data is encoded from a sparse representation to a dense representation;

a second stage in which features are extracted from the dense representation using a long-range contextual information to identify the features; and

a third stage in which the dense representation is transformed into a sparse representation from which the three-dimensional perception output is generated.

12 . The method of claim 11 , wherein the cascade of three stages comprises one or more perception operations comprising obtaining a three-dimensional (3D) point-wise segmentation by attaching a 3D segmentation head to the output of a last stage of an M-stage decoder;

obtaining voxel-level predictions at an output of the 3D segmentation head, and performing a de-voxelization on the voxel-level predictions to obtain point-wise segmentation results.

13 . The method of claim 12 , wherein the one or more perception operations comprise obtaining a panoptic segmentation by applying a second stage refinement to the point-wise segmentation results.

14 . The method of claim 13 , wherein the one or more perception operations comprise generating a panoptic segmentation result from the second stage refinement.

15 . The method of claim 11 , wherein the second stage comprises N stages having an increasing channel width.

16 . The method of claim 15 , wherein each stage of the N stages comprises a sparse convolution layer followed by two submanifold sparse convolution blocks.

17 . The method of claim 16 , wherein the sparse convolution layer of each of the N stages, except for first stage, has a stride of 2 such that spatial resolution is downsampled by a factor of 8 in the encoder.

18 . The method of claim 14 , wherein the second stage refinement is obtained by:

fusing detected boxes, voxel-wise features and bird's eye view (BEV) features;

applying a local coordinate transformation to points within each detected box;

calculating box classification scores and point-wise mask scores; and

determining output of the second stage refinement by fusing the box classification scores and point-wise mask scores and an output of a previous refinement stage.

19 . The method of claim 11 , wherein the second stage is configured to operate as:

sparse voxel features into a dense feature map;

generating a 2D BEV feature map by concatenating features in different heights;

extracting long term contextual information by using a 2D convolutional neural network;

reshaping encoded BEV feature representation to a dense voxel map; and

transforming the dense voxel map to a sparse voxel by applying a dense-to-sparse conversion.

20 . An image processing apparatus comprising one or more processors configured to implement a method, comprising:

converting points in a point cloud obtained from a lidar sensor into a voxel grid;

generating, from the voxel grid, sparse voxel features by applying a multi-layer perceptron and one or more max pooling layers that reduce dimension of input data;

applying a cascade of an encoder that performs a N-stage sparse-to-dense feature operation, a global context pooling (GCP) module, and an M-stage decoder that performs a dense-to-sparse feature generation operation,

wherein the GCP module bridges an output of a last stage of the N-stages with an input of a first stage of the M-stages, where N and M are positive integers; and

wherein the GCP module comprises a multi-scale feature extractor; and

performing one or more perception operations on an output of the M-stage decoder and/or an output of the GCP module.

Assignments (2)
CHANGE OF NAME Recorded Dec 3, 2025
From: TUSIMPLE, INC.
To: CREATEAI, INC.
Reel/Frame 073832/0553 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2024
From: YE, DONGQIANGZI; ZHOU, ZIXIANG; CHEN, WEIJIA; XIE, YUFEI; WANG, YU; WANG, PANQU; GE, LINGTING
To: TUSIMPLE, INC.
Reel/Frame 066397/0983 →