IP Library › Granted Patent US 12,299,910
Granted Patent B2
US 12,299,910 · App. 17/885,949 · Granted May 13, 2025

Point cloud alignment systems for generating high definition maps for vehicle navigation

Inventors: Sherif Ahmed Morsy Nekkah (Singapore, SG); Nicole Alexandra Camous (Singapore, SG); Sergi Adipraja Widjaja (Singapore, SG); Venice Erin Baylon Liong (Singapore, SG); Xiaogang Wang (Singapore, SG)
Assignee: Motional AD LLC
G06T7/337G01S13/89G01S17/89G06T3/06G06T3/4046G06T3/60G06T2207/10028G06T2207/20081G06T2207/20084G06T2207/20221G06T2207/30241G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,910
App. No.
17/885,949
Granted
May 13, 2025
Kind
B2
Abstract

A method for determining a trajectory of a vehicle within a physical space based at least on an aligned point cloud may include generating a fused feature map by concatenating a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud. A machine learning model may be applied to determine, based at least on the fused feature map, a relative transform aligning the target point cloud to the source point cloud. An aligned target point cloud may be generated by transforming the target point cloud in accordance with the relative transform. Furthermore, a trajectory of a vehicle within the physical space may be determined based on at least the first relative transform. Related systems and computer program products are also provided.

Claims (56)

1. A method, comprising:

generating, using at least one data processor, a fused feature map by at least concatenating a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, the target point cloud corresponding to at least a portion of a same physical space as the source point cloud;

applying, using the at least one data processor, a machine learning model trained to determine, based at least on the fused feature map, a first relative transform aligning the target point cloud to the source point cloud, wherein the machine learning model determines the first relative transform by at least performing a weighted subsampling to extract, from the fused feature map, corresponding features from the source point cloud and the target point cloud such that the first feature map and the second feature map are downsampled into a combined feature map;

generating, using the at least one data processor, an aligned target point cloud by at least transforming the target point cloud in accordance with the first relative transform; and

determining, using the at least one data processor, a trajectory of a vehicle within the physical space based on at least the first relative transform.

2. The method of claim 1 , wherein the first feature map corresponds to an encoded representation of the source point cloud, and wherein the second feature map corresponds to an encoded representation of the target point cloud.

3. The method of claim 2 , further comprising:

encoding, using the at least one data processor, the source point cloud and the target point cloud such that one or more objects present in each point cloud are represented as one or more pillars, each of the one or more pillars bounding a plurality of points associated with a corresponding object.

4. The method of claim 1 , further comprising:

pre-training, using the at least one data processor, the machine learning model prior to training the machine learning model to determine the first relative transform, the pre-training includes reconstructing, using the at least one data processor, a point cloud by at least decoding an encoding of the point cloud deformed by one or more offsets determined by the machine learning model, and adjusting the machine learning model to minimize a difference between the point cloud and the decoded point cloud.

5. The method of claim 4 , further comprising:

pre-training, based at least on the difference between the point cloud and the decoded point cloud, an encoder generating the encoding of the point cloud.

6. The method of claim 1 , wherein the machine learning model includes a 1×1 convolution layer configured to perform the weighted subsampling.

7. The method of claim 1 , wherein the machine learning model further determines the first relative transform by at least flattening the combined feature map into a one-dimensional vector for processing by a fully connected layer of the machine learning model.

8. The method of claim 1 , further comprising:

applying, using the at least one data processor, the first relative transform to perform a coarse point registration of the target point cloud;

applying, using the at least one data processor, the machine learning model trained to determine a second relative transform to further align the aligned target point cloud to the source point cloud; and

applying, using the at least one data processor, the second relative transform to perform a fine point registration of the target point cloud.

9. The method of claim 1 , wherein the first relative transform includes a translation along at least one of an x-axis, a y-axis, and a z-axis.

10. The method of claim 1 , wherein the first relative transform includes a rotation around a fixed point of θ radians about a unit axis (X, Y, Z).

11. The method of claim 1 , wherein the source point cloud and the target point cloud comprise three dimensional point clouds.

12. A system, comprising:

at least one data processor, and

at least one non-transitory storage media storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

generating, using the at least one data processor, a fused feature map by at least concatenating a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, the target point cloud corresponding to at least a portion of a same physical space as the source point cloud;

applying, using the at least one data processor, a machine learning model trained to determine, based at least on the fused feature map, a first relative transform aligning the target point cloud to the source point cloud, wherein the machine learning model determines the first relative transform by at least performing a weighted subsampling to extract, from the fused feature map, corresponding features from the source point cloud and the target point cloud such that the first feature map and the second feature map are downsampled into a combined feature map;

generating, using the at least one data processor, an aligned target point cloud by at least transforming the target point cloud in accordance with the first relative transform; and

determining, using the at least one data processor, a trajectory of a vehicle within the physical space based on at least the first relative transform.

13. The system of claim 12 , wherein the first feature map corresponds to an encoded representation of the source point cloud, and wherein the second feature map corresponds to an encoded representation of the target point cloud.

14. The system of claim 13 , wherein the operations further comprise:

encoding, using the at least one data processor, the source point cloud and the target point cloud such that one or more objects present in each point cloud are represented as one or more pillars, each of the one or more pillars bounding a plurality of points associated with a corresponding object.

15. The system of claim 12 , wherein the operations further comprise:

pre-training, using the at least one data processor, the machine learning model prior to training the machine learning model to determine the first relative transform, the pre-training includes reconstructing, using the at least one data processor, a point cloud by at least decoding an encoding of the point cloud deformed by one or more offsets determined by the machine learning model, and adjusting the machine learning model to minimize a difference between the point cloud and the decoded point cloud.

16. The system of claim 15 , wherein the operations further comprise:

pre-training, based at least on the difference between the point cloud and the decoded point cloud, an encoder generating the encoding of the point cloud.

17. The system of claim 12 , wherein the machine learning model includes a 1×1 convolution layer configured to perform the weighted subsampling.

18. The system of claim 12 , wherein the machine learning model further determines the first relative transform by at least flattening the combined feature map into a one-dimensional vector for processing by a fully connected layer of the machine learning model.

19. The system of claim 12 , wherein the operations further comprise:

applying, using the at least one data processor, the first relative transform to perform a coarse point registration of the target point cloud;

applying, using the at least one data processor, the machine learning model trained to determine a second relative transform to further align the aligned target point cloud to the source point cloud; and

applying, using the at least one data processor, the second relative transform to perform a fine point registration of the target point cloud.

20. The system of claim 12 , wherein the first relative transform includes a translation along at least one of an x-axis, a y-axis, and a z-axis.

21. The system of claim 12 , wherein the first relative transform includes a rotation around a fixed point of θ radians about a unit axis (X, Y, Z).

22. The system of claim 12 , wherein the source point cloud and the target point cloud comprise three dimensional point clouds.

23. At least one non-transitory storage media storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

generating, using the at least one data processor, a fused feature map by at least concatenating a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, the target point cloud corresponding to at least a portion of a same physical space as the source point cloud;

applying, using the at least one data processor, a machine learning model trained to determine, based at least on the fused feature map, a first relative transform aligning the target point cloud to the source point cloud, wherein the machine learning model determines the first relative transform by at least performing a weighted subsampling to extract, from the fused feature map, corresponding features from the source point cloud and the target point cloud such that the first feature map and the second feature map are downsampled into a combined feature map;

generating, using the at least one data processor, an aligned target point cloud by at least transforming the target point cloud in accordance with the first relative transform; and

determining, using the at least one data processor, a trajectory of a vehicle within the physical space based on at least the first relative transform.

24. The at least one non-transitory storage media of claim 23 , wherein the first feature map corresponds to an encoded representation of the source point cloud, and wherein the second feature map corresponds to an encoded representation of the target point cloud.

25. The at least one non-transitory storage media of claim 24 , wherein the operations further comprise:

encoding, using the at least one data processor, the source point cloud and the target point cloud such that one or more objects present in each point cloud are represented as one or more pillars, each of the one or more pillars bounding a plurality of points associated with a corresponding object.

26. The at least one non-transitory storage media of claim 23 , wherein the operations further comprise:

pre-training, using the at least one data processor, the machine learning model prior to training the machine learning model to determine the first relative transform, the pre-training including reconstructing, using the at least one data processor, a point cloud by at least decoding an encoding of the point cloud deformed by one or more offsets determined by the machine learning model, and adjusting the machine learning model to minimize a difference between the point cloud and the decoded point cloud; and

pre-training, based at least on the difference between the point cloud and the decoded point cloud, an encoder generating the encoding of the point cloud.

27. The at least one non-transitory storage media of claim 23 , wherein the machine learning model includes a 1×1 convolution layer configured to perform the weighted subsampling.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2022
From: NEKKAH, SHERIF AHMED MORSY; CAMOUS, NICOLE ALEXANDRA; WIDJAJA, SERGI ADIPRAJA; LIONG, VENICE ERIN BAYLON; WANG, XIAOGANG
To: MOTIONAL AD LLC
Reel/Frame 060936/0851 →
Continuity (1)
Related Publication 20240054660A1 · Feb 15, 2024
References Cited (23)
US 11397242B1 · Zhang et al. · 2022 [cited by applicant]
US 11521394B2 · Beijbom et al. · 2022 [cited by applicant]
US 11698437B2 · Gerardo Castro et al. · 2023 [cited by applicant]
US 12033338B2 · Nekkah · 2024 [cited by examiner]
US 20170046840A1 · Chen et al. · 2017 [cited by applicant]
US 20180158235A1 · Wu · 2018 [cited by examiner]
US 20200043186A1 · Selviah et al. · 2020 [cited by applicant]
US 20200082560A1 · Nezhadarya et al. · 2020 [cited by applicant]
US 20210358137A1 · Lee et al. · 2021 [cited by applicant]
US 20210405638A1 · Boyraz · 2021 [cited by examiner]
US 20220164566A1 · Ye et al. · 2022 [cited by applicant]
US 20220414821A1 · Zhu · 2022 [cited by examiner]
US 20230074860A1 · Camous et al. · 2023 [cited by applicant]
US 20230177719A1 · Babin et al. · 2023 [cited by applicant]
US 20230182774A1 · Wang et al. · 2023 [cited by applicant]
Zhu, Minghan, Maani Ghaffari, and Huei Peng. “Correspondence-free point cloud registration with so (3)-equivariant implicit shape representations.” Conference on robot learning. PMLR, 2021. (Year: 2021). [cited by examiner]
[No. Author Listed], “Surface Vehicle Recommended Practice: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” SAE International, Standard J3016, Sep. 30, 2016, 30 page… [cited by applicant]
cvlibs.net [online], “Visual Odometry / SLAM Evaluation 2012,” 2012, retrieved Oct. 12, 2023, retrieved from URL <https://www.cvlibs.net/datasets/kitti/eval_odometry.php>, 10 pages. [cited by applicant]
Dai et al., “Deformable Convolutional Networks,” CoRR, revised Jun. 5, 2017, arXiv:1703.06211, 12 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2023/029721, mailed on Sep. 15, 2023, 8 pages. [cited by applicant]
Lang et al., “PointPillars: Fast Encoders for Object Detection from Point Clouds,” CoRR, revised May 7, 2019, arXiv:1812.05784, 9 pages. [cited by applicant]
nuscenes.org [online], “nuPlan,” 2020, retrieved on Oct. 12, 2023, retrieved from URL <https://www.nuscenes.org/nuplan>, 10 pages. [cited by applicant]
Segal et al., “Generalized-ICP,” Robotics: Science and Systems, 2009, 2(4):435, 8 pages. [cited by applicant]