IP Library › Granted Patent US 12,327,414
Granted Patent B2
US 12,327,414 · App. 17/827,963 · Granted Jun 10, 2025

Systems and methods for enhancement of 3D object detection using point cloud semantic segmentation and attentive anchor generation

Inventors: Ehsan Taghavi (Markham, CA); Ryan Razani (Toronto, CA); Bingbing Liu (Markham, CA)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06V20/58B60W30/00G01S17/89G06V10/26G06V10/762G06V10/764B60W2420/408
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,414
App. No.
17/827,963
Granted
Jun 10, 2025
Kind
B2
Abstract

Devices, systems, methods, and media are disclosed for performing an object detection task comprising: obtaining a semantic segmentation map representing a real-world space, the semantic segmentation map including an array of elements that each represent a respective location in the real-world space and are assigned a respective element classification label; clustering groups of the elements based on the assigned respective element classification labels to identify at least a first cluster of elements that have each been assigned the same respective element classification label; generating, based on a location of the first cluster within the semantic segmentation map, at least one anchor that defines a respective probable object location of a first dynamic object; and generating, based on the semantic segmentation map and the at least one anchor, a respective bounding box and object instance classification label for the first dynamic object.

Claims (54)

1. A method of performing an object detection task comprising:

obtaining a semantic segmentation map representing a real-world space, the semantic segmentation map including an array of elements that each represent a respective location in the real-world space, the array of elements including elements that are each assigned a respective element classification label selected from a set of possible classification labels that correspond to different classifications of dynamic objects;

clustering groups of the elements based on the assigned respective element classification labels to identify at least a first cluster of elements that have each been assigned the same respective element classification label;

generating, based on a location of the first cluster within the semantic segmentation map, a plurality of anchors, each of the plurality of anchors defining a different respective probable object location of a first dynamic object, wherein generating the plurality of anchors comprises:

computing an approximate location for the first dynamic object in the semantic segmentation map based on the locations of the elements of the first cluster;

generating a lower resolution map corresponding to the semantic segmentation map, and mapping the approximate location for the first dynamic object to a corresponding coarse element location in the lower resolution map:

generating a plurality of candidate anchors each indicating a different respective probable location of the first dynamic object relative to the coarse element location; and

mapping at least some of the plurality of candidate anchors to respective element locations of a higher resolution map to provide the plurality of anchors; and

generating, based on the semantic segmentation map and the at least one plurality of anchors, a respective bounding box and object instance classification label for the first dynamic object.

2. The method of claim 1 wherein computing the approximate location for the first dynamic object comprises determining a mean element location for the first cluster of elements based on the respective locations of the elements of the first cluster within the semantic segmentation map.

3. The method of claim 1 comprising sampling the plurality of candidate anchors to select only a subset of the plurality of candidate anchors to include in the mapping to the respective element locations of the higher resolution map.

4. The method of claim 1 wherein generating the plurality of candidate anchors comprises selecting, for each candidate anchor: an anchor geometry, an anchor orientation, and an anchor offset relative to the coarse element location.

5. The method of claim 1 , wherein clustering groups of the elements is performed to identify, in addition to the first cluster of elements, a plurality of further clusters that include elements that have each been assigned the same respective element classification label,

the method comprising, for each of the plurality of further clusters:

computing an approximate location in the semantic segmentation map for a respective dynamic object corresponding to the further cluster based on the location of the further cluster within the semantic segmentation map;

mapping the approximate location for the respective dynamic object to a corresponding coarse element location in the lower resolution map;

generating a respective plurality of candidate anchors each indicating a different respective probable location of the respective dynamic object; and

mapping at least some of the respective plurality of candidate anchors to respective element locations in the higher resolution map to provide a respective plurality of anchors for the further cluster, each anchor of the respective plurality of anchors defining a respective probable object location of the respective dynamic object in the higher resolution map,

the method further comprising:

generating a respective bounding box and object instance classification label for each of the respective dynamic objects represented in the plurality of further clusters based on the plurality of anchors provided for each of the plurality of further clusters.

6. The method of claim 5 comprising, prior to generating the respective bounding boxes and object instance classification labels for the first dynamic object and the respective dynamic objects represented in the plurality of further clusters, generating additional anchors according to a defined set of ad-hoc rules, each of the additional anchors defining a respective probable object location in the higher resolution map, wherein the generating the respective bounding boxes and object instance classification labels is also based on the additional anchors.

7. The method of claim 1 wherein obtaining the semantic segmentation map comprises obtaining a Light Detection and Ranging (LIDAR) frame of the real-world space using a LIDAR sensor and using a semantic segmentation model to assign the element classification labels used for the elements of the semantic segmentation map.

8. The method of claim 7 comprising applying a 3D to 2D conversion operation on an output of semantic segmentation model to generate the semantic segmentation map, wherein the semantic segmentation map represents a birds-eye-view (BEV) of the real-world space, and wherein the at least one anchor defines the respective probable object location of the first dynamic object with respect to the semantic segmentation map.

9. The method of claim 7 wherein the semantic segmentation map represents a 3D volume of the real-world space, and wherein the at least one anchor defines the respective probable object location of the first dynamic object with respect to the semantic segmentation map.

10. The method of claim 1 comprising controlling one or more of a steering and a speed of an autonomous vehicle based on the respective bounding box and object instance classification label for the first dynamic object.

11. A system comprising a processor device coupled to a memory, the memory storing executable instructions that when executed by the processor device configure the system to perform an object detection task comprising:

obtaining a semantic segmentation map representing a real-world space, the semantic segmentation map including an array of elements that each represent a respective location in the real-world space, the array of elements including elements that are each assigned a respective element classification label selected from a set of possible classification labels that correspond to different classifications of dynamic objects;

clustering groups of the elements based on the assigned respective element classification labels to identify at least a first cluster of elements that have each been assigned the same respective element classification label;

generating, based on a location of the first cluster within the semantic segmentation map, a plurality of anchors, each of the plurality of anchors defining a respective probable object location of a first dynamic object, wherein generating the plurality of anchors comprises:

computing an approximate location for the first dynamic object in the semantic segmentation map based on the locations of the elements of the first cluster;

generating a lower resolution map corresponding to the semantic segmentation map, and mapping the approximate location for the first dynamic object to a corresponding coarse element location in the lower resolution map;

generating a plurality of candidate anchors each indicating a different respective probable location of the first dynamic object relative to the coarse element location; and

mapping at least some of the plurality of candidate anchors to respective element locations of a higher resolution map to provide the plurality of anchors; and

generating, based on the semantic segmentation map and the at least one plurality of anchors, a respective bounding box and object instance classification label for the first dynamic object.

12. The system of claim 11 wherein computing the approximate location for the first dynamic object comprises determining a mean element location for the first cluster of elements based on the respective locations of the elements of the first cluster within the semantic segmentation map.

13. The system of claim 11 , the object detection task comprising sampling the plurality of candidate anchors to select only a subset of the plurality of candidate anchors to include in the mapping to the respective element locations of the higher resolution map.

14. The system of claim 11 wherein clustering groups of the elements is performed to identify, in addition to the first cluster of elements, a plurality of further clusters that include elements that have each been assigned the same respective element classification label,

the object detection task comprising, for each of the plurality of further clusters:

computing an approximate location in the semantic segmentation map for a respective dynamic object corresponding to the further cluster based on the location of the further cluster within the semantic segmentation map;

mapping the approximate location for the respective dynamic object to a corresponding coarse element location in the lower resolution map;

generating a respective plurality of candidate anchors each indicating a different respective probable location of the respective dynamic object; and

mapping at least some of the respective plurality of candidate anchors to respective element locations in the higher resolution map to provide a respective plurality of anchors for the further cluster, each anchor of the respective plurality of anchors defining a respective probable object location of the respective dynamic object in the higher resolution map,

the object detection task further comprising:

generating a respective bounding box and object instance classification label for each of the respective dynamic objects represented in the plurality of further clusters based on the plurality of anchors provided for each of the plurality of further clusters.

15. The system of claim 14 , the object detection task comprising, prior to generating the respective bounding boxes and object instance classification labels for the first dynamic object and the respective dynamic objects represented in the plurality of further clusters, generating additional anchors according to a defined set of ad-hoc rules, each of the additional anchors defining a respective probable object location in the higher resolution map, wherein the generating the respective bounding boxes and object instance classification labels is also based on the additional anchors.

16. A non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by a processor device of a computing system, cause the computing system to perform a method comprising:

obtaining a semantic segmentation map representing a real-world space, the semantic segmentation map including an array of elements that each represent a respective location in the real-world space, the array of elements including elements that are each assigned a respective element classification label selected from a set of possible classification labels that correspond to different classifications of dynamic objects;

clustering groups of the elements based on the assigned respective element classification labels to identify at least a first cluster of elements that have each been assigned the same respective element classification label;

generating, based on a location of the first cluster within the semantic segmentation map, a plurality of anchors, each of the plurality of anchors defining a respective probable object location of a first dynamic object, wherein generating the plurality of anchors comprises:

computing an approximate location for the first dynamic object in the semantic segmentation map based on the locations of the elements of the first cluster;

generating a lower resolution map corresponding to the semantic segmentation map, and mapping the approximate location for the first dynamic object to a corresponding coarse element location in the lower resolution map;

generating a plurality of candidate anchors each indicating a different respective probable location of the first dynamic object relative to the coarse element location; and

mapping at least some of the plurality of candidate anchors to respective element locations of a higher resolution map to provide the plurality of anchors; and

generating, based on the semantic segmentation map and the plurality of anchors, a respective bounding box and object instance classification label for the first dynamic object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2022
From: TAGHAVI, EHSAN; RAZANI, RYAN; LIU, BINGBING
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 062195/0693 →
Continuity (1)
Related Publication 20230410530A1 · Dec 21, 2023
References Cited (17)
US 11704806B2 · Choudhary · 2023 [cited by examiner]
US 20220120858A1 · Cennamo · 2022 [cited by examiner]
US 20220261593A1 · Yu · 2022 [cited by examiner]
US 20230281961A1 · Fazlali · 2023 [cited by examiner]
Mottaghi, Roozbeh, et al. “The role of context for object detection and semantic segmentation in the wild.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2014. [cited by examiner]
Felzenszwalb, Pedro F., et al. “Object detection with discriminatively trained part-based models.” IEEE transactions on pattern analysis and machine intelligence 32.9 (2009): 1627-1645. [cited by examiner]
Feng, Di, et al. “Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges.” IEEE Transactions on Intelligent Transportation Systems 22.3 (2020): 1341-1360. [cited by examiner]
Zuo, C., et al. “Det2Seg: A Two-Stage Approach for Road Object Segmentation from 3D Point Clouds,” 2019 IEEE Visual Communications and Image Processing (VCIP), Sydney, Australia, 2019. [cited by applicant]
NVidia. “Laser Focused: How Multi-View LidarNet Presents Rich Perspective for Self-Driving Cars”, Web blog, https://blogs.nvidia.com/blog/2020/03/11/drive-labs-multi-view-lidarnet-self-driving-cars/, Mar. 2020. [cited by applicant]
Ku, J., et al. “Joint 3D Proposal Generation and Object Detection from View Aggregation,” 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2018. [cited by applicant]
Erdal Aksoy, Eren, Saimir Baci, and Selcuk Cavdar. “SalsaNet: Fast Road and Vehicle Segmentation in LiDAR Point Clouds for Autonomous Driving.” arXiv preprint arXiv:1909.08291 (2019). [cited by applicant]
Dewan, Ayush, and Wolfram Burgard. “DeepTemporalSeg: Temporally Consistent Semantic Segmentation of 3D LiDAR Scans.” arXiv preprint arXiv:1906.06962 (2019). [cited by applicant]
Zhang, Feihu, et al. “Instance segmentation of lidar point clouds.” 2020 International Conference on Robotics and Automation (ICRA). IEEE, 2020 (to appear). [cited by applicant]
Lang, Alex H., et al. “Pointpillars: Fast encoders for object detection from point clouds.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2019. [cited by applicant]
Cortinhal, T., et al. “SalsaNext: Fast Semantic Segmentation of LiDAR Point Clouds for Autonomous Driving.” arXiv preprint arXiv:2003.03653 (2020). [cited by applicant]
Berman, M., et al. “The lovasz-softmax loss: a tractable surrogate for the optimization of the intersection-over-union measure in neural networks.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Re… [cited by applicant]
Ronneberger, Olaf, Philipp Fischer, and Thomas Brox. “U-net: Convolutional networks for biomedical image segmentation.” International Conference on Medical image computing and computer-assisted intervention. Springer, C… [cited by applicant]