IP Library › Granted Patent US 12,397,817
Granted Patent B2
US 12,397,817 · App. 17/859,945 · Granted Aug 26, 2025

Representation learning for object detection from unlabeled point cloud sequences

Inventors: Xiangru Huang (Quincy, MA); Yue Wang (Cambridge, MA); Vitor Guizilini (Santa Clara, CA); Rares Andrei Ambrus (San Francisco, CA); Adrien David Gaidon (San Jose, CA); Justin Solomon (Somerville, MA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA; MASSACHUSETTS INSTITUTE OF TECHNOLOGY
B60W60/001G06V20/58B60W2420/403B60W2420/408B60W2554/4049
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,397,817
App. No.
17/859,945
Granted
Aug 26, 2025
Kind
B2
Abstract

A method of representation learning for object detection from unlabeled point cloud sequences is described. The method includes detecting moving object traces from temporally-ordered, unlabeled point cloud sequences. The method also includes extracting a set of moving objects based on the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences. The method further includes classifying the set of moving objects extracted from on the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences. The method also includes estimating 3D bounding boxes for the set of moving objects based on the classifying of the set of moving objects.

Claims (59)

1. A method of representation learning for object detection from unlabeled point cloud sequences, comprising:

detecting moving object traces from temporally-ordered, unlabeled point cloud sequences;

extracting a set of moving objects based on the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences;

classifying the set of moving objects extracted from on the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences;

estimating 3D bounding boxes for the set of moving objects based on the classifying of the set of moving objects;

labeling the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences as moving vehicles; and

planning a trajectory of an ego vehicle according to the labeled moving vehicles in a scene surrounding the ego vehicle.

2. The method of claim 1 , in which the detecting moving object traces comprises:

visualizing the temporally-ordered, unlabeled point cloud sequences in a world coordinate system according to an ego-motion;

removing ground points from the visualizing the temporally-ordered, unlabeled point cloud sequences to form a ground removed point cloud visualization;

estimating object cluster proposals according to a point cloud segmentation of the ground removed point cloud visualization; and

identifying the moving object traces from the object cluster proposal according to multi-object tracing.

3. The method of claim 1 , in which extracting the set of moving objects comprises:

training a single-frame semantic instance segmentation model to differentiate between feature vectors representing the set of moving objects and the feature vectors representing a background of the temporally-ordered, unlabeled point cloud sequences; and

identifying each of the feature vectors as a moving feature vector or a non-moving feature vector.

4. The method of claim 1 , further comprises training a feature extraction module to extract the set of moving objects based on the moving object traces detected from the sequence of temporally-ordered, unlabeled point clouds via self-supervised tasks.

5. The method of claim 1 , in which the set of moving objects are represented as a sequence of point clusters that correspond to corresponding one of the set of moving objects.

6. The method of claim 1 , further comprising:

training a first model to identify moving objects in a point cloud according to the labeled moving object traces;

inferring attributes of the bounding boxes from the labeled moving object traces; and

training a second model to detect objects in the point cloud according to the attributes of the bounding boxes inferred from the labeled moving object traces.

7. The method of claim 1 , in which estimating the 3D bounding boxes comprises:

registering the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences according to object class labels; and

assigning the 3D bounding boxes according to the object class labels.

8. The method of claim 1 , further comprising controlling the ego vehicle along the planned trajectory of the ego vehicle according to the labeled moving vehicles detected in the scene surrounding the ego vehicle.

9. A non-transitory computer-readable medium having program code recorded thereon for representation learning and object detection from unlabeled point cloud sequences, the program code being executed by a processor and comprising:

program code to detect moving object traces from temporally-ordered, unlabeled point cloud sequences;

program code to extract a set of moving objects based on the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences;

program code to classify the set of moving objects extracted from on the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences;

program code to estimate 3D bounding boxes for the set of moving objects based on the classifying of the set of moving objects;

program code to label the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences as moving vehicles; and

program code to plan a trajectory of an ego vehicle according to the labeled moving vehicles in a scene surrounding the ego vehicle.

10. The non-transitory computer-readable medium of claim 9 , in which the program code to detect moving object traces comprises:

program code to visualize the temporally-ordered, unlabeled point cloud sequences in a world coordinate system according to an ego-motion;

program code to remove ground points from the visualizing the temporally-ordered, unlabeled point cloud sequences to form a ground removed point cloud visualization;

program code to estimate object cluster proposals according to a point cloud segmentation of the ground removed point cloud visualization; and

program code to identify the moving object traces from the estimated object cluster proposal according to multi-object tracing.

11. The non-transitory computer-readable medium of claim 9 , in which the program code to extract the set of moving objects comprises:

program code to train a single-frame semantic instance segmentation model to differentiate between feature vectors representing the set of moving objects and the feature vectors representing a background of the temporally-ordered, unlabeled point cloud sequences; and

program code to identify each of the feature vectors as a moving feature vector or a non-moving feature vector.

12. The non-transitory computer-readable medium of claim 9 , further comprises program code to train a feature extraction module to extract the set of moving objects based on the moving object traces detected from the sequence of temporally-ordered, unlabeled point clouds via self-supervised tasks.

13. The non-transitory computer-readable medium of claim 9 , in which the set of moving objects are represented as a sequence of point clusters that correspond to corresponding one of the set of moving objects.

14. The non-transitory computer-readable medium of claim 9 , further comprising:

program code to train a first model to identify moving objects in a point cloud according to the labeled moving object traces;

program code to infer attributes of the bounding boxes from the labeled moving object traces; and

program code to train a second model to detect objects in the point cloud according to the attributes of the bounding boxes inferred from the labeled moving object traces.

15. The non-transitory computer-readable medium of claim 9 , in which the program code to estimate the 3D bounding boxes comprises:

program code to register the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences according to object class labels; and

program code to assign the 3D bounding boxes according to the object class labels.

16. The non-transitory computer-readable medium of claim 9 , further comprising program code to control the ego vehicle along the planned trajectory of the ego vehicle according to the labeled moving vehicles detected in the scene surrounding the ego vehicle.

17. A system of representation learning for object detection from unlabeled point cloud sequences, the system comprising:

a moving object trace detection module to detect moving object traces from temporally-ordered, unlabeled point cloud sequences;

a moving object extraction module to extract a set of moving objects based on the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences;

an object classification and labeling module to classify the set of moving objects extracted from on the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences;

a bounding box estimation module to estimate 3D bounding boxes for the set of moving objects based on the classifying of the set of moving objects, and to label the moving object traces detected from the sequence of temporally-ordered, unlabeled point cloud sequences as moving vehicles; and

a planner to plan a trajectory of an ego vehicle according to the labeled moving vehicles in a scene surrounding the ego vehicle.

18. The system of claim 17 , further comprising a single-frame semantic instance segmentation model trained to differentiate between feature vectors representing the set of moving objects and the feature vectors representing a background of the temporally-ordered, unlabeled point cloud sequences and to identify each of the feature vectors as a moving feature vector or a non-moving feature vector.

19. The system of claim 17 , in which the set of moving objects are represented as a sequence of point clusters that correspond to corresponding one of the set of moving objects.

20. The system of claim 17 , further comprising a controller to control the ego vehicle along the planner trajectory of the ego vehicle according to the 3D bounding boxes detected in the scene surrounding the ego vehicle.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2025
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 072646/0894 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2022
From: HUANG, XIANGRU; WANG, YUE; SOLOMON, JUSTIN
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 060456/0168 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2022
From: GUIZILINI, VITOR; AMRUS, RARES ANDREI; GAIDON, ADRIEN DAVID
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 060456/0361 →
Continuity (1)
Related Publication 20240010225A1 · Jan 11, 2024
References Cited (12)
US 10809361B2 · Vallespi-Gonzalez et al. · 2020 [cited by applicant]
US 20200320867A1 · Lewis · 2020 [cited by examiner]
US 20220153297A1 · Chen · 2022 [cited by examiner]
US 20230271607A1 · Kobashi · 2023 [cited by examiner]
US 20230271616A1 · Kobashi · 2023 [cited by examiner]
RU 2016145126A · 2018 [cited by applicant]
WO 2016170333A1 · 2016 [cited by applicant]
Luo, et al., “Self-Supervised Pillar Motion Learning for Autonomous Driving,” https://arxiv.org/abs/2104.08683, submitted on Apr. 18, 2021. [cited by applicant]
Lee, et al., “PillarFlow: End-to-end Birds-eye-view Flow Estimation for Autonomous Driving,” IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 2007-2013, Oct. 24, 2020-Jan. 24, 2021. [cited by applicant]
Qi, et al., “Offboard 3D Object Detection from Point Cloud Sequences,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6130-6140, 2021. [cited by applicant]
Chen, et al., “Moving Object Segmentation in 3D LiDAR Data: A Learning-based Approach Exploiting Sequential Data,” IEEE Robotics and Automation Letters, vol. 6, No. 4, pp. 6529-6536, Oct. 2021. [cited by applicant]
Yang, et al., “IPOD: Intensive Point-based Object Detector for Point Cloud,” https://arxiv.org/abs/1812.05276, submitted on Dec. 13, 2018. [cited by applicant]