IP Library › Granted Patent US 12,462,438
Granted Patent B2
US 12,462,438 · App. 17/553,403 · Granted Nov 4, 2025

Machine-learning for 3D object detection

Inventors: Asma Rejeb Sfar (Vélizy-Villacoublay, FR); Tom Durand (Vélizy-Villacoublay, FR); Ashad Hosenbocus (Vélizy-Villacoublay, FR)
Assignee: Dassault Systemes
G06T9/002G06N3/088G06V10/761G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,438
App. No.
17/553,403
Granted
Nov 4, 2025
Kind
B2
Abstract

A computer-implemented method of machine-learning for learning a neural network that encodes a super-point of a 3D point cloud into a latent vector. The method including obtaining a dataset of super-points. Each super-point is a set of points of a 3D point cloud. The set of points represents at least a part of an object. The method further includes learning the neural network based on the dataset of super-points. The learning includes minimizing a loss. The loss penalizes a disparity between two super-points. This constitutes improved machine-learning for 3D object detection.

Claims (54)

1 . A computer-implemented method of applying a neural network learnt according to machine-learning for learning a neural network configured to encode a super-point of a 3D point cloud into a latent vector, the method of machine-learning including obtaining a dataset of super-points, each super-point being a set of points of a 3D point cloud, the set of points representing at least a part of an object, and learning the neural network based on the dataset of super-points, the learning comprising minimizing a loss penalizing a disparity between two super-points,

the method comprising:

obtaining one or more first super-points of a first 3D point cloud representing a 3D scene; and

obtaining one or more second super-points of a second 3D point cloud representing a 3D object;

encoding, by applying the neural network, the one or more first super-points each into a respective first latent vector and the one or more second super-points each into a respective second latent vector; and

determining a similarity between each first super-point of one or more first super-points and each second super-point of one or more second super-points, by, for said each first super-point and said each second super-point, computing a similarity measure between the respective first latent vector encoding the first super-point and the respective second latent vector encoding the second super-point,

wherein the obtaining the one or more first super-points further includes:

obtaining one or more initial super-points of the first 3D point cloud; and

filtering the one or more initial super-points by selecting, among the initial super-points, each initial super-point for which:

a disparity between dimensions of the respective initial super-point and dimensions of at least one second super-point of the one or more second super-points is smaller than a predefined threshold; and

a disparity between a position of the respective initial super-point and a position of at least one second super-point of the one or more second super-points is smaller than a predefined threshold,

the selected initial super-points being the one or more first super-points, and

wherein the filtering further includes selecting, among the initial super-points, each initial super-point for which:

a distance between each dimension of the super-point and a corresponding dimension of at least one second super-point of the one or more second super-points is smaller than a predefined threshold;

a ratio between each dimension of the super-point and a corresponding dimension of at least one second super-point of the one or more second super-points is smaller than a maximal ratio and larger than a minimal ratio; and

a difference between a relative height from a closest support surface of the super-point and a relative height from a closest support surface of at least one second super-point of the one or more second super-points is smaller than a predefined threshold.

2 . The method of use of claim 1 , wherein the method further comprises:

determining, among first super-points each having a determined similarity with at least one second super-point that is larger than a predefined threshold, one or more groups of first super-points, each group of first super-points having a similar shape than the second 3D point cloud.

3 . The method of claim 2 , wherein the determining of the one or more groups further comprises:

obtaining a graph of second super-points, the graph of second super-points having nodes each representing a second super-point and edges each representing a geometrical relation between the two super-points represented by the nodes connected by the edges and having one or more geometrical attributes of the geometrical relation; and

forming one or more graphs each having nodes representing each a first super-point, by constructing edges each connecting two nodes and each having one or more geometrical attributes similar to geometrical attributes of an edge of the graph of second super-points, each formed graph corresponding to a respective group.

4 . The method of claim 3 , further comprising, for each group, determining a similarity score by measuring a similarity of the first super-points in the group with the one or more second super-points.

5 . A device comprising:

a processor; and

a computer readable data storage medium having recorded thereon

a computer program comprising instructions

for performing machine-learning for learning a neural network configured for encoding a super-point of a 3D point cloud into a latent vector by the processor being configured to obtain a dataset of super-points, each super-point being a set of points of a 3D point cloud, the set of points representing at least a part of an object, and learn the neural network based on the dataset of super-points, the learning comprising minimizing a loss penalizing a disparity between two super-points, and/or

for applying a neural network learnable according to the machine-learning by the processor being configured to:

obtain one or more first super-points of a first 3D point cloud representing a 3D scene, and

obtain one or more second super-points of a second 3D point cloud representing a 3D object,

encode, by applying the neural network, the one or more first super-points each into a respective first latent vector and the one or more second super-points each into a respective second latent vector; and

determine a similarity between each first super-point of one or more first super-points and each second super-point of one or more second super-points, by, for said each first super-point and said each second super-point, computing a similarity measure between the respective first latent vector encoding the first super-point and the respective second latent vector encoding the second super-point,

wherein the processor being configured to obtain the one or more first super-points further by being further configured to:

obtain one or more initial super-points of the first 3D point cloud; and

filter the one or more initial super-points by selecting, among the initial super-points, each initial super-point for which:

a disparity between dimensions of the respective initial super-point and dimensions of at least one second super-point of the one or more second super-points is smaller than a predefined threshold; and

a disparity between a position of the respective initial super-point and a position of at least one second super-point of the one or more second super-points is smaller than a predefined threshold,

the selected initial super-points being the one or more first super-points, and

wherein the processor is further configured to filter by being configured to select, among the initial super-points, each initial super-point for which:

a distance between each dimension of the super-point and a corresponding dimension of at least one second super-point of the one or more second super-points is smaller than a predefined threshold;

a ratio between each dimension of the super-point and a corresponding dimension of at least one second super-point of the one or more second super-points is smaller than a maximal ratio and larger than a minimal ratio; and

a difference between a relative height from a closest support surface of the super-point and a relative height from a closest support surface of at least one second super-point of the one or more second super-points is smaller than a predefined threshold.

6 . The device of claim 5 , wherein in the machine-learning, the loss is a reconstruction loss and the loss penalizes a disparity between a super-point and a reconstruction of the super-point.

7 . The device of claim 6 , wherein the disparity is a distance between the super-point and the reconstruction of the super-point.

8 . The device of claim 5 , wherein the processor is further configured to determine, among first super-points each having a determined similarity with at least one second super-point that is larger than a predefined threshold, one or more groups of first super-points, each group of first super-points having a similar shape than the second 3D point cloud.

9 . The device of claim 8 , wherein the processor is further configured to determine of the one or more groups by being configured to:

obtain a graph of second super-points, the graph of second super-points having nodes each representing a second super-point and edges each representing a geometrical relation between the two super-points represented by the nodes connected by the edges and having one or more geometrical attributes of the geometrical relation, and

form one or more graphs each having nodes representing each a first super-point, by constructing edges each connecting two nodes and each having one or more geometrical attributes similar to geometrical attributes of an edge of the graph of second super-points, each formed graph corresponding to a respective group.

10 . The device of claim 9 , wherein, for each group, the processor further configured to determine a similarity score by measuring a similarity of the first super-points in the group with the one or more second super-points.

11 . A non-transitory computer readable medium having stored thereon a program that when executed by the computer causes the computer to implement the method according to claim 1 .

12 . The method of claim 1 , wherein in the method of machine-learning the loss is a reconstruction loss and the loss penalizes a disparity between a super-point and a reconstruction of the super-point.

13 . The method of claim 12 , wherein in the method of machine-learning the disparity is a distance between the super-point and the reconstruction of the super-point.

14 . The method of claim 13 , wherein in the method of machine-learning the distance is a Chamfer distance or an Earth-Mover distance.

15 . The method of claim 1 , wherein in the method of machine-learning the learning is an unsupervised learning.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2022
From: REJEB SFAR, ASMA; DURAND, TOM; HOSENBOCUS, ASHAD
To: DASSAULT SYSTEMES
Reel/Frame 059114/0442 →
Priority Claims (1)
EP 20306588 · Dec 16, 2020 · regional
Continuity (1)
Related Publication 20220189070A1 · Jun 16, 2022
References Cited (33)
US 20200027247A1 · Minnen et al. · 2020 [cited by applicant]
US 20200099954A1 · Hemmer et al. · 2020 [cited by applicant]
JP 2020191077A · 2020 [cited by applicant]
JP 2024538685A · 2024 [cited by applicant]
Matsuzaki_2019 (Binary Representation for 3D Point Cloud Compression Based on Deep Auto-encoder, IEEE 8th Global Conference on Consumer Electronics (GCCE)) (Year: 2019). [cited by examiner]
Giancola_2019 (Leveraging Shape Completion for 3D Siamese Tracking, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2019) (Year: 2019). [cited by examiner]
Landrieu_2018 (Large-scale Point Cloud Semantic Segmentation with Superpoint Graphs, 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018), Jun. 2018, Salt Lake City, United States). (Year: 2018). [cited by examiner]
Ishtayeh_2014 (Similarity Threshold Determination for Test Document Clustering, 2014). (Year: 2014). [cited by examiner]
Landrieu_2017 (Weakly Supervised Segmentation-Aided Classification of Urban Scenes from 3D LIDAR point clouds, The international Achieves of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. XLII… [cited by examiner]
Matsuzaki et al.—“Binary Representation for 3D Point Cloud Compression Based on Deep Auto-encoder”, 2019 IEEE 8 [cited by applicant]
Giancola et al.—“Leveraging Shape Completion for 3D Siamese Tracking”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (10 pages). [cited by applicant]
Elbaz et al—“3D Point Cloud Registration for Localization using a Deep Neural Network Auto-Encoder”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (10 pages). [cited by applicant]
Yue et al.—“A LiDAR Point Cloud Generator: from a Virtual World to Autonomous Driving”, Mar. 31, 2018 (7 pages). [cited by applicant]
Yang et al.—“Toward the Repeatability and Robustness of the Local Reference Frame for 3D Shape Matching: An Evaluation”, IEEE Transactions on Image Processing, vol. 27, No. 8, Aug. 2018 (16 pages). [cited by applicant]
Xie et al.—“MLCVNet: Multi-Level Context VoteNet for 3D Object Detection”, Apr. 12, 2020 (10 pages). [cited by applicant]
Velzhev et al.—“Implicit Shape Models for Object Detection in 3D Point Clouds”, ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. 1-3, 2012 XXII ISPRS Congress, Aug. 25-Sep. 1, 20… [cited by applicant]
Tombari et al.—“Hough Voting for 3D Object Recognition under Occlusion and Clutter”, IPSJ Transactions on Computer Vision and Applications, vol. 4, Mar. 20-29, 2012, Information Processing Society of Japan (10 pages). [cited by applicant]
Spezialetti et al.—“Learning an Effective Equivariant 3D Descriptor Without Supervision”, ICCV, Sep. 15, 2019 (10 pages). [cited by applicant]
Rumelhart et al—“Learning Internal Representations by Error Propagation, Parallel distributed processing: explorations in the microstructure of cognition”, vol. 1, foundations, MIT Press, Cambridge, MA 1986, Basic Mecha… [cited by applicant]
Qi et al.—“Deep Hough Voting for 3D Object Dection in Point Clouds”, Aug. 22, 2019 (14 pages). [cited by applicant]
Mian et al.—“On the Repeatability and Quality of Keypoints for Local Feature-based 3D Object Retrieval from Cluttered Scenes”, Int. J. Comput Vis. (2010), 89, (14 pages). [cited by applicant]
Landrieu et al.—“Large-scale Point Cloud Semantic Segmentation with Superpoint Graphs”, Mar. 28, 2018 (12 pages). [cited by applicant]
Landrieu et al.—“Cut Pursuit: fast algorithms to learn piecewise constant functions on general weighted graphs”, SIAM Journal on Imaging Sciences, HAL Id:hal-01306779 https://halarchives-ouvertesfr/hal-01306779v4, submi… [cited by applicant]
Qi et al.—“P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds”, May 28, 2020 (10 pages). [cited by applicant]
Hackel et al.—“semantic3d.net: a New Large-Scale Point Cloud Classification Benchmark”, ISPRS, Apr. 12, 2017 (9 pages). [cited by applicant]
Guinard et al.—“Weakly Supervised Segmentation-Aided Classification of Urban Scenes From 3D Lidar Point Clouds”, The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. XL… [cited by applicant]
Geiger et al.—“Are we ready for Autonomous Driving the KITTI Vision Benchmark Suite”, CVPR 2012, (8 pages). [cited by applicant]
Elberink et al—“User-Assisted Object Detection by Segment Based Similarity Measures in Mobile Laser Scanner Data”, The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. … [cited by applicant]
Demantké et al.—“Dimensionality Based Scale Selection in 3D Lidar Point Clouds” International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, 2011,(6 pages). [cited by applicant]
Armeni et al.—“3D Semantic Parsing of Large-Scale Indoor Spaces”, CVPR 2016, http://buildingparser.stanford.edu (10 pages). [cited by applicant]
Bellekens et al.—“A Benchmark Survey of Rigid 3D Point Cloud Registration Algorithms” International Journal on Advances in Intelligent Systems, vol. 8, No. 1&2, 2015, http://wwwiariajournals.org/intelligent_systems/ (10… [cited by applicant]
Search Report dated Jun. 7, 2021 issued in corresponding European patent application No. 20306588.3 (10 pages). [cited by applicant]
Office Action dated Aug. 19, 2025, issued in counterpart JP Application No. 2021-203671, w/English Translation, (14 pages). [cited by applicant]