IP Library › Granted Patent US 12,583,464
Granted Patent B2
US 12,583,464 · App. 17/710,895 · Granted Mar 24, 2026

Streaming object detection and segmentation with polar pillars

Inventors: Sourabh Vora (Marina Del Rey, CA); Qi Chen (Baltimore, MD)
Assignee: Motional AD LLC
B60W50/0098B60W60/0015G01C21/3807
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,583,464
App. No.
17/710,895
Granted
Mar 24, 2026
Kind
B2
Abstract

Among other things, techniques for detecting objects in the environment surrounding a vehicle are described. A computer system is configured to receive a set of measurements from a sensor of a vehicle. The set of measurements includes a plurality of data points that represent a plurality of objects in a 3D space surrounding the vehicle. The system divides the 3D space into a plurality of pillars. The system then assigns each data point of the plurality of data points to a pillar in the plurality of pillars. The system generates a pseudo-image based on the plurality of pillars. The pseudo-image includes, for each pillar of the plurality of pillars, a corresponding feature representation of data points assigned to the pillar. The system detects the plurality of objects based on an analysis of the pseudo-image. The system then operates the vehicle based upon the detecting of the objects.

Claims (42)

1 . A system comprising:

one or more computer processors; and

one or more non-transitory storage media storing instructions which, when executed by the one or more computer processors, cause performance of operations comprising:

dividing streamed sectors of data points representing a three-dimensional (3D) space surrounding a vehicle into a plurality of polar pillars, wherein each polar pillar of the plurality of polar pillars comprises a slice of the 3D space that extends from a two-dimensional (2D) polar grid on a ground plane comprising wedged-shaped regions corresponding to the streamed sectors in the 3D space, wherein each data point of the sectors is assigned to a polar pillar in the plurality of polar pillars;

encoding the streamed sectors into a wedge-shaped region in a bird's eye view using polar pillars to obtain pillar-wise features on the polar grid, wherein 3D stacked polar pillar tensors are generated for non-empty polar pillars and convolution is iteratively applied to the 3D stacked polar pillar tensors to generate the pillar wise features;

generating a feature map based on the pillar-wise features, wherein each polar pillar of the plurality of polar pillars corresponds to a polar feature representation of data points assigned to the polar pillar;

inputting the feature map into a segmentation head, an object detection head, and a bounding box head of a network simultaneously;

outputting per-pixel segments, object classes, and bounding boxes from the segmentation head, object detection head, and bounding box head respectively, wherein the object detection head transforms the feature representation to a Cartesian representation for object detection and the bounding box head applies kernels to the data points of the feature representation based on a range for bounding box generation; and

operating the vehicle in the 3D space according to the per-pixel segments, object classes, and bounding boxes, wherein the streamed sectors are iteratively processed.

2 . The system of claim 1 , wherein transforming the feature representation to the Cartesian representation for object detection comprises transforming the feature representation of the data points assigned to a respective polar pillar from a polar representation in a 2D polar grid to a canonical Cartesian representation.

3 . The system of claim 1 , wherein applying kernels to data points of the feature representation based on a range for bounding box generation comprises applying kernels and normalization to data points assigned to respective polar pillars at different ranges.

4 . The system of claim 1 , wherein applying kernels to data points of the feature representation based on a range for bounding box generation comprises applying kernels to data points of the feature representation based on a range for at shared convolution layers of the segmentation head, the object detection head, or the bounding box head.

5 . The system of claim 1 , wherein the operations comprise performing panoptic fusion to identify different instances of a same object class and operating the vehicle in the 3D space according to panoptic segmentation.

6 . The system of claim 1 , wherein a backbone upsamples the feature map prior to inputting the feature map into the segmentation head of the network.

7 . The system of claim 1 , wherein the feature map is padded via multi-scale context padding prior to inputting the feature map into the segmentation head, object detection head, and bounding box head of the network.

8 . The system of claim 1 , wherein the 2D polar grid has substantially wedge shaped cells with varying cell sizes dependent upon a density of objects in a corresponding region of the 3D space surrounding the vehicle.

9 . The system of claim 1 , wherein the feature map is undistorted by interpolating features at Cartesian pillar locations using original pillar locations of the pillar-wise features.

10 . A method, comprising:

dividing, with at least one processor, streamed sectors of data points representing a three-dimensional (3D) space surrounding a vehicle into a plurality of polar pillars, wherein each polar pillar of the plurality of polar pillars comprises a slice of the 3D space that extends from a two-dimensional (2D) polar grid on a ground plane comprising wedged-shaped regions corresponding to the streamed sectors in the 3D space, wherein each data point of the sectors is assigned to a polar pillar in the plurality of polar pillars;

encoding, with the at least one processor, the streamed sectors into a wedge-shaped region in a bird's eye view using polar pillars to obtain pillar-wise features on the polar grid, wherein 3D stacked polar pillar tensors are generated for non-empty polar pillars and convolution is iteratively applied to the 3D stacked polar pillar tensors to generate the pillar wise features;

generating, with the at least one processor, a feature map based on the pillar-wise features, wherein each polar pillar of the plurality of polar pillars corresponds to a polar feature representation of data points assigned to the polar pillar;

inputting, with the at least one processor, the feature map into a segmentation head, an object detection head, and a bounding box head of a network simultaneously;

outputting, with the at least one processor, per-pixel segments, object classes, and bounding boxes from the segmentation head, object detection head, and bounding box head respectively, wherein the object detection head transforms the feature representation to a Cartesian representation for object detection and the bounding box head applies kernels to the data points of the feature representation based on a range for bounding box generation; and

operating, with the at least one processor, the vehicle in the 3D space according to the per-pixel segments, object classes, and bounding boxes, wherein the streamed sectors are iteratively processed.

11 . The method of claim 10 , wherein transforming the feature representation to the Cartesian representation for object detection comprises transforming the feature representation of the data points assigned to a respective polar pillar from a polar representation in a 2D polar grid to a canonical Cartesian representation.

12 . The method of claim 10 , wherein applying kernels to data points of the feature representation based on a range for bounding box generation comprises applying kernels and normalization to data points assigned to respective polar pillars at different ranges.

13 . The method of claim 10 , wherein applying kernels to data points of the feature representation based on a range for bounding box generation comprises applying kernels to data points of the feature representation based on a range for at shared convolution layers of the segmentation head, the object detection head, or the bounding box head.

14 . The method of claim 10 , comprising performing panoptic fusion to identify different instances of a same object class and operating the vehicle in the 3D space according to panoptic segmentation.

15 . The method of claim 10 , wherein a backbone upsamples the feature map prior to inputting the feature map into the segmentation head of the network.

16 . The method of claim 10 , wherein the feature map is padded via multi-scale context padding prior to inputting the feature map into the segmentation head, object detection head, and bounding box head of the network.

17 . At least one non-transitory storage media storing instructions that, when executed by at least one processor, cause the at least one processor to:

divide streamed sectors of data points representing a three-dimensional (3D) space surrounding a vehicle into a plurality of polar pillars, wherein each polar pillar of the plurality of polar pillars comprises a slice of the 3D space that extends from a two-dimensional (2D) polar grid on a ground plane comprising wedged-shaped regions corresponding to the streamed sectors in the 3D space, wherein each data point of the sectors is assigned to a polar pillar in the plurality of polar pillars;

encode the streamed sectors into a wedge-shaped region in a bird's eye view using polar pillars to obtain pillar-wise features on the polar grid, wherein 3D stacked polar pillar tensors are generated for non-empty polar pillars and convolution is iteratively applied to the 3D stacked polar pillar tensors to generate the pillar wise features;

generate a feature map based on the pillar-wise features, wherein each polar pillar of the plurality of polar pillars corresponds to a polar feature representation of data points assigned to the polar pillar;

input the feature map into a segmentation head, an object detection head, and a bounding box head of a network simultaneously;

output per-pixel segments, object classes, and bounding boxes from the segmentation head, object detection head, and bounding box head respectively, wherein the object detection head transforms the feature representation to a Cartesian representation for object detection and the bounding box head applies kernels to the data points of the feature representation based on a range for bounding box generation; and

operate the vehicle in the 3D space according to the per-pixel segments, object classes, and bounding boxes, wherein the streamed sectors are iteratively processed.

18 . The at least one non-transitory storage media of claim 17 , wherein transforming the feature representation to the Cartesian representation for object detection comprises transforming the feature representation of the data points assigned to a respective polar pillar from a polar representation in a 2D polar grid to a canonical Cartesian representation.

19 . The at least one non-transitory storage media of claim 17 , wherein applying kernels to data points of the feature representation based on a range for bounding box generation comprises applying kernels and normalization to data points assigned to respective polar pillars at different ranges.

20 . The at least one non-transitory storage media of claim 17 , wherein applying kernels to data points of the feature representation based on a range for bounding box generation comprises applying kernels to data points of the feature representation based on a range for at shared convolution layers of the segmentation head, the object detection head, or the bounding box head.

21 . The at least one non-transitory storage media of claim 17 , wherein the instructions comprise performing panoptic fusion to identify different instances of a same object class and operating the vehicle in the 3D space according to panoptic segmentation.

22 . The at least one non-transitory storage media of claim 17 , wherein a backbone upsamples the feature map prior to inputting the feature map into the segmentation head of the network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2022
From: VORA, SOURABH; CHEN, QI
To: MOTIONAL AD LLC
Reel/Frame 059468/0223 →
Continuity (2)
Provisional Application 63191887 · May 21, 2021
Related Publication 20220371606A1 · Nov 24, 2022
References Cited (28)
US 11798289B2 · Vora et al. · 2023 [cited by applicant]
US 20050004448A1 · Gurr et al. · 2005 [cited by applicant]
US 20180203124A1 · Izzat · 2018 [cited by examiner]
US 20180349746A1 · Vallespi-Gonzalez · 2018 [cited by examiner]
US 20190065824A1 · Gaudet · 2019 [cited by examiner]
US 20190311499A1 · Mammou · 2019 [cited by examiner]
US 20200093464A1 · Martins et al. · 2020 [cited by applicant]
US 20200111358A1 · Parchami · 2020 [cited by examiner]
US 20200150235A1 · Beijbom · 2020 [cited by examiner]
US 20200311569A1 · Ghosh · 2020 [cited by examiner]
US 20210150752A1 · Ayvaci et al. · 2021 [cited by applicant]
US 20210158043A1 · Hou · 2021 [cited by examiner]
US 20210312227A1 · Moradiannejad et al. · 2021 [cited by applicant]
US 20220032452A1 · Casas · 2022 [cited by examiner]
US 20220104463A1 · Spears et al. · 2022 [cited by applicant]
US 20220383640A1 · Vora et al. · 2022 [cited by applicant]
CN 112907685A · 2021 [cited by examiner]
JP 2021189917A · 2021 [cited by examiner]
JP-202189917-A English Translation (Year: 2024). [cited by examiner]
CN112907685A English Translation (Year: 2025). [cited by examiner]
[No Author Listed], “SAE International Standard J3016: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” SAE International, dated Sep. 2016, 30 pages. [cited by applicant]
Liu et al., “SSD: Single Shot Multibox Detector,” Presented at The 14th European Conference on Computer Vision—ECCV 2016, Amsterdam, The Netherlands, Oct. 8-16, 2016; Lecture Notes in Computer Science, 9905:21-37, avail… [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2022/030367, dated Aug. 23, 2022, 7 pages. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2022/030367, dated Nov. 30, 2023, 6 pages. [cited by applicant]
Chen et al., “PolarStream: Streaming Lidar Object Detection and Segmentation 0with Polar Pillars”, CoRR, Submitted on Mar. 24, 2022, arXiv:2106.07545v2, 13 pages. [cited by applicant]
Extended European Search Report in European Appln. No. 22805629.7, mailed on Feb. 17, 2025, 14 pages. [cited by applicant]
Liong et al., “AMVNet: Assertion-based Multi-View Fusion Network for LiDAR Semantic Segmentation”, CoRR, Submitted on Dec. 9, 2020, arXiv:2012.04934v1, 10 pages. [cited by applicant]
Stanisz et al., “Optimisation of the PointPillars network for 3D object detection in point clouds”, CoRR, Submitted on Jul. 1, 2020, arXiv:2007.00493v1, 7 pages. [cited by applicant]