IP Library › Granted Patent US 12,469,152
Granted Patent B2
US 12,469,152 · App. 18/195,466 · Granted Nov 11, 2025

System and method for three-dimensional multi-object tracking

Inventors: Jie Li (Los Altos, CA); Rares A. Ambrus (San Francisco, CA); Taraneh Sadjadpour (Stanford, CA); Christin Jeannette Bohg (Palo Alto, CA)
Assignees: Toyota Research Institute, Inc.; Toyota Jidosha Kabushiki Kaisha; The Board of Trustees of the Leland Stanford Junior University
G06T7/248G06T19/00G06T2207/10028G06T2207/30252G06T2210/12G06T2210/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,152
App. No.
18/195,466
Granted
Nov 11, 2025
Kind
B2
Abstract

Systems and methods for performing three-dimensional multi-object tracking are disclosed herein. In one example, a method includes the steps of determining a residual based on augmented current frame detection bounding boxes, augmented previous frame detection bounding boxes, augmented current frame shape descriptors, and augmented previous frame shape descriptors and predicting an affinity matrix using the residual. The residual indicates a spatiotemporal and shape similarity between current detections in a current frame point cloud data and previous detections in a previous frame point cloud data. The affinity matrix indicates associations between the previous detections and the current detections, as well as the augmented anchors.

Claims (59)

1 . A system for three-dimensional multi-object tracking comprising:

a processor; and

a memory having instructions that, when executed by the processor, cause the processor to:

determine a residual based on augmented current frame detection bounding boxes, augmented previous frame detection bounding boxes, augmented current frame shape descriptors, and augmented previous frame shape descriptors, the residual indicating a spatiotemporal and shape similarity between current detections in a current frame point cloud data and previous detections in a previous frame point cloud data, and

predict an affinity matrix using the residual, the affinity matrix indicating associations between the previous detections and the current detections, wherein the affinity matrix includes a first dimension that is correlated with the previous detections and includes information regarding a newborn track anchor and a false positive anchor and a second dimension that is correlated with the current detections and includes information regarding a false negative anchor and a dead track anchor,

wherein the affinity matrix identifies matches between tracks and detections, and

control a movement of a vehicle using the affinity matrix.

2 . The system of claim 1 , wherein the memory further includes instructions that, when executed by the processor, cause the processor to:

generate the augmented current frame detection bounding boxes by appending a bounding box of an object that was not detected in the current frame and a last known or previous frame bounding box that is not active to the current frame detection bounding boxes based on the current detections,

generate the augmented previous frame detection bounding boxes by appending a predicted bounding box that does not match an object and a first bounding box of a newly tracked object to previous frame detection bounding boxes based on the previous detections,

generate the augmented current frame shape descriptors by appending a false negative anchor's shape descriptor and a dead track anchor's shape descriptor to the current frame shape descriptors based on the current shape descriptors, and

generate the augmented previous frame shape descriptors by appending a false positive anchor's shape descriptor and newborn track anchor's shape descriptor to a previous frame shape descriptors based on the current shape descriptors.

3 . The system of claim 2 , wherein:

the current frame detection bounding boxes describe 3D locations, dimensions, and rotations of the current detections;

the previous frame detection bounding boxes describe 3D locations, dimensions, and rotation of the previous detections;

the previous frame shape descriptors describe shapes of the previous detections; and

the current frame shape descriptors describe shapes of the current detections.

4 . The system of claim 1 , wherein the memory further includes instructions for performing sequential track confidence refinement that, when executed by the processor, cause the processor to:

when the affinity matrix indicates that a current detection has a high probability of being a true positive, take a weighted average between a confidence of the current detection and an existing confidence of a track matched to the current detection, and

when the affinity matrix indicates that the current detection has a lower probability of being a true positive, downscale the existing confidence of the track matched to the current detection.

5 . The system of claim 1 , wherein the previous frame point cloud data and the current frame point cloud data was generated by a light detection and ranging sensor.

6 . The system of claim 5 , wherein the light detection and ranging sensor is mounted to the vehicle.

7 . A method for three-dimensional multi-object tracking comprising steps of:

determining a residual based on augmented current frame detection bounding boxes, augmented previous frame detection bounding boxes, augmented current frame shape descriptors, and augmented previous frame shape descriptors, the residual indicating a spatiotemporal and shape similarity between current detections in a current frame point cloud data and previous detections in a previous frame point cloud data;

predicting an affinity matrix using the residual, the affinity matrix indicating associations between the previous detection and the current detection, wherein the affinity matrix includes a first dimension that is correlated with the previous detections and includes information regarding a newborn track anchor and a false positive anchor and a second dimension that is correlated with the current detections and includes information regarding a false negative anchor and a dead track anchor; and

controlling a movement of a vehicle using the affinity matrix.

8 . The method of claim 7 , further comprising steps of:

generating the augmented current frame detection bounding boxes by appending a bounding box of an object that was not detected in the current frame and a last known or previous frame bounding box that is not active to current frame detection bounding boxes based on the current detections;

generating the augmented previous frame detection bounding boxes by appending a predicted bounding box that does not match an object and a first bounding box of a newly tracked object to previous frame detection bounding boxes based on the previous detections;

generating the augmented current frame shape descriptors by appending a false negative anchor's shape descriptor and a dead track anchor's shape descriptor to the current frame shape descriptors based on the current shape descriptors; and

generating the augmented previous frame shape descriptors by appending a false positive anchor's shape descriptor and newborn track anchor's shape descriptor to a previous frame shape descriptors based on the current shape descriptors.

9 . The method of claim 8 , wherein:

the current frame detection bounding boxes describe 3D locations, dimensions, and rotations of the current detections;

the previous frame detection bounding boxes describe 3D locations, dimensions, and rotation of the previous detections;

the previous frame shape descriptors describe shapes of the previous detections; and

the current frame shape descriptors describe shapes of the current detections.

10 . The method of claim 7 , further comprising performing sequential track confidence refinement, the steps of sequential track confidence refinement include:

when the affinity matrix indicates that a current detection has a high probability of being a true positive, taking a weighted average between a confidence of the current detection and an existing confidence of a track matched to the current detection; and

when the affinity matrix indicates that the current detection has a lower probability of being a true positive, downscaling the existing confidence of the track matched to the current detection.

11 . The method of claim 7 , wherein the previous frame point cloud data and the current frame point cloud data was generated by a light detection and ranging sensor.

12 . The method of claim 11 , wherein the light detection and ranging sensor is mounted to the vehicle.

13 . A non-transitory computer-readable medium storing instructions for performing three-dimensional multi-object tracking, the instructions, when executed by a processor, cause the processor to:

determine a residual based on augmented current frame detection bounding boxes, augmented previous frame detection bounding boxes, augmented current frame shape descriptors, and augmented previous frame shape descriptors, the residual indicating a spatiotemporal and shape similarity between current detections in a current frame point cloud data and previous detections in a previous frame point cloud data;

predict an affinity matrix using the residual, the affinity matrix indicating associations between the previous detection and the current detection, wherein the affinity matrix includes a first dimension that is correlated with the previous detections and includes information regarding a newborn track anchor and a false positive anchor and a second dimension that is correlated with the current detections and includes information regarding a false negative anchor and a dead track anchor; and

control a movement of a vehicle using the affinity matrix.

14 . The non-transitory computer-readable medium of claim 13 , further including instructions that, when executed by the processor, cause the processor to:

generate the augmented current frame detection bounding boxes by appending a bounding box of an object that was not detected in the current frame and a last known or previous frame bounding box that is not active to the current frame detection bounding boxes based on the current detections;

generate the augmented previous frame detection bounding boxes by appending a predicted bounding box that does not match an object and a first bounding box of a newly tracked object to a previous frame detection bounding boxes based on the previous detections;

generate the augmented current frame shape descriptors by appending a false negative anchor's shape descriptor and a dead track anchor's shape descriptor to the current frame shape descriptors based on the current shape descriptors; and

generate the augmented previous frame shape descriptors by appending a false positive anchor's shape descriptor and newborn track anchor's shape descriptor to a previous frame shape descriptors based on the current shape descriptors.

15 . The non-transitory computer-readable medium of claim 14 , wherein:

the current frame detection bounding boxes describe 3D locations, dimensions, and rotations of the current detections;

the previous frame detection bounding boxes describe 3D locations, dimensions, and rotation of the previous detections;

the previous frame shape descriptors describe shapes of the previous detections; and

the current frame shape descriptors describe shapes of the current detections.

16 . The non-transitory computer-readable medium of claim 13 , further including instructions for performing sequential track confidence refinement that, when executed by the processor, cause the processor to:

when the affinity matrix indicates that a current detection has a high probability of being a true positive, take a weighted average between a confidence of the current detection and an existing confidence of a track matched to the current detection; and

when the affinity matrix indicates that the current detection has a lower probability of being a true positive, downscale the existing confidence of the track matched to the current detection.

17 . The non-transitory computer-readable medium of claim 13 , wherein the previous frame point cloud data and the current frame point cloud data was generated by a light detection and ranging sensor.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2025
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 072957/0199 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2025
From: SADJADPOUR, TARANEH; BOHG, CHRISTIN JEANNETTE
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
Reel/Frame 072545/0173 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2023
From: LI, JIE; AMBRUS, RARES A.
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 063749/0585 →
Continuity (2)
Provisional Application 63421346 · Nov 1, 2022
Related Publication 20240153107A1 · May 9, 2024
References Cited (15)
US 10713491B2 · Zhu et al. · 2020 [cited by applicant]
US 20210072391A1 · Li · 2021 [cited by examiner]
Chiu, Hsu-kuang, et al. (“Probabilistic 3D multi-modal, multi-object tracking for autonomous driving.” 2021 IEEE international conference on robotics and automation (ICRA). IEEE, 2021.). [cited by examiner]
Sun, ShiJie, et al. “Deep affinity network for multiple object tracking.” IEEE transactions on pattern analysis and machine intelligence 43.1 (2019): 104-119. [cited by examiner]
Chiu et al. “Probabilistic 3d multi-modal, multi-object tracking for autonomous driving.” 2021 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2021. [cited by applicant]
Ross Girshick, “Fast r-cnn.” Proceedings of the IEEE international conference on computer vision. 2015. [cited by applicant]
Yin et al. “Center-based 3d object detection and tracking.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021. [cited by applicant]
Meyer et al. “Message passing algorithms for scalable multitarget tracking.” Proceedings of the IEEE 106.2 (2018): 221-259. [cited by applicant]
Wang et al. “DeepFusionMOT: A 3D Multi-Object Tracking Framework Based on Camera-LiDAR Fusion With Deep Association.” IEEE Robotics and Automation Letters 7.3 (2022): 8260-8267. [cited by applicant]
Weng et al. “3d multi-object tracking: A baseline and new evaluation metrics.” 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020. [cited by applicant]
Liang et al. “Pnpnet: End-to-end perception and prediction with tracking in the loop.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020. [cited by applicant]
Kim et al. “Eagermot: 3d multi-object tracking via sensor fusion.” 2021 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2021. [cited by applicant]
Sun et al. “Deep affinity network for multiple object tracking.” IEEE transactions on pattern analysis and machine Intelligence 43.1 (2019): 104-119. [cited by applicant]
Weng et al. “A baseline for 3d multi-object tracking.” arXiv preprint arXiv:1907.03961 1.2 (2019): 6. [cited by applicant]
Stearns et al. “SpOT: Spatiotemporal Modeling for 3D Object Tracking.” Computer Vision-ECCV 2022: 17th European Conference, Tel Aviv, Israel, Oct. 23-27, 2022, Proceedings, Part XXXVIII. Cham: Springer Nature Switzerlan… [cited by applicant]