IP Library › Granted Patent US 12,190,588
Granted Patent B2
US 12,190,588 · App. 17/339,413 · Granted Jan 7, 2025

Occlusion-aware multi-object tracking

Inventors: Dongdong Chen (Bellevue, WA); Qiankun Liu (Hefei, CN); Lu Yuan (Redmond, WA); Lei Zhang (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G06V20/52G06F18/2155G06N3/045G06T7/20G06V10/25G06V10/44G06T2207/30241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,588
App. No.
17/339,413
Granted
Jan 7, 2025
Kind
B2
Abstract

A system for tracking a target object across a plurality of image frames. The system comprises a logic machine and a storage machine. The storage machine holds instructions executable by the logic machine to calculate a trajectory for the target object over one or more previous frames occurring before a target frame. Responsive to assessing no detection of the target object in the target frame, the instructions are executable to predict an estimated region for the target object based on the trajectory, predict an occlusion center based on a set of candidate occluding locations for a set of other objects within a threshold distance of the estimated region, each location of the set of candidate occluding locations overlapping with the estimated region, and automatically estimate a bounding box for the target object in the target frame based on the occlusion center.

Claims (30)

1. A system for tracking a target object across a plurality of image frames, comprising:

a logic machine; and

a storage machine holding instructions executable by the logic machine to:

calculate a trajectory for the target object over one or more previous frames occurring before a target frame, wherein the target object is detected by tracking, in a similarity matrix, comparison values indicating similarity between object feature data for a first set of objects detected in a first previous frame, the first set of objects including the target object, and object feature data for a second set of objects in a second previous frame, wherein the similarity matrix includes:

a row for each object in a union of both of the first set of objects and the second set of objects; and

a column for each object in the union,

wherein each matrix element of the similarity matrix represents one comparison value between a pair of objects drawn from the union;

responsive to assessing no detection of the target object in the target frame:

upon determining that the target object is not detected in the target frame due to being occluded by a set of one or more other objects, predict an estimated region for the target object based on the trajectory;

predict an occlusion center based on a set of candidate occluding locations for the set of other objects within a threshold distance of the estimated region, each location of the set of candidate occluding locations overlapping with the estimated region; and

automatically estimate a bounding box for the target object in the target frame based on the occlusion center, wherein the bounding box is estimated via a trained machine learning system trained via supervised learning with image data and ground-truth bounding boxes.

2. The system of claim 1 , wherein the instructions are further executable to calculate a heatmap for the occlusion center by operating one or more convolutional neural network units.

3. The system of claim 1 , wherein estimating the bounding box includes operating a state machine configured to record state information including tracking status and tracked motion of a plurality of objects, and to estimate the bounding box based on such recorded state information.

4. The system of claim 1 , wherein estimating the bounding box includes operating a Kalman filter.

5. The system of claim 1 , further comprising a training storage device holding instructions executable to train a machine learning system to track objects based on one or more unsupervised labels.

6. The system of claim 5 , wherein the similarity matrix further includes

a placeholder column for the placeholder unsupervised labels.

7. The system of claim 5 , wherein the object feature data and the comparison values are generated by one or more differentiable functions, and wherein training the machine learning system to track objects includes configuring the one or more differentiable functions based on the unsupervised labels and one or more automatic supervision signals.

8. A method of tracking a target object across a plurality of image frames, the method comprising:

calculating a trajectory for the target object over one or more previous frames occurring before a target frame, wherein the target object is detected by tracking, in a similarity matrix, comparison values indicating similarity between object feature data for a first set of objects detected in a first previous frame, the first set of objects including the target object, and object feature data for a second set of objects in a second previous frame, wherein the similarity matrix includes:

a row for each object in a union of both of the first set of objects and the second set of objects; and

a column for each object in the union,

wherein each matrix element of the similarity matrix represents one comparison value between a pair of objects drawn from the union;

assessing no detection of the target object in the target frame due to the target object being occluded by a set of one or more other objects;

predicting an estimated region for the target object based on the trajectory;

predicting an occlusion center based on a set of candidate occluding locations for the set of other objects within a threshold distance of the estimated region, each location of the set of candidate occluding locations overlapping with the estimated region; and

automatically estimating a bounding box for the target object in the target frame based on the occlusion center, wherein the bounding box is estimated via a trained machine learning system trained via supervised learning with image data and ground-truth bounding boxes.

9. The method of claim 8 , further comprising calculating a heatmap for the occlusion center by operating one or more convolutional neural network units.

10. The method of claim 8 , wherein estimating the bounding box includes operating a state machine configured to record state information including tracking status and tracked motion of a plurality of objects, and to estimate the bounding box based on such recorded state information.

11. The method of claim 8 , wherein estimating the bounding box includes operating a Kalman filter.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2021
From: CHEN, DONGDONG; LIU, QIANKUN; YUAN, LU; ZHANG, LEI
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 056443/0899 →
Continuity (1)
Related Publication 20220391621A1 · Dec 8, 2022
References Cited (57)
US 6424370B1 · Courtney · 2002 [cited by examiner]
US 20050105764A1 · Han · 2005 [cited by examiner]
US 20190333233A1 · Hu · 2019 [cited by examiner]
US 20200265591A1 · Yang · 2020 [cited by examiner]
US 20220185625A1 · One · 2022 [cited by examiner]
US 20220300748A1 · Tokmakov · 2022 [cited by examiner]
US 20220301275A1 · Khadloya · 2022 [cited by examiner]
US 20220377242A1 · Camacho · 2022 [cited by examiner]
US 20230042004A1 · Kumar · 2023 [cited by examiner]
Babaee, et al., “A Dual CNN-RNN for Multiple People Tracking”, In Journal of Neurocomputing, vol. 368, Nov. 27, 2019, pp. 69-83. [cited by applicant]
Bergmann, et al., “Tracking Without Bells and Whistles”, In Proceedings of IEEE/CVF International Conference on Computer Vision, Oct. 27, 2019, 15 Pages. [cited by applicant]
Bernardin, et al., “Evaluating Multiple Object Tracking Performance: The CLEAR MOT Metrics”, In Journal on Image and Video Processing, vol. 2008, Dec. 2008, pp. 1-10. [cited by applicant]
Bewley, et al., “Simple Online and Realtime Tracking”, In Proceedings of the IEEE International Conference on Image Processing, Sep. 25, 2016, 5 Pages. [cited by applicant]
Braso, et al., “Learning a Neural Solver for Multiple Object Tracking”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 14, 2020, pp. 6247-6257. [cited by applicant]
Chi, et al., “PedHunter: Occlusion Robust Pedestrian Detector in Crowded Scenes”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, No. 07, Apr. 3, 2020, pp. 10639-10646. [cited by applicant]
Chu, et al., “Detection in Crowded Scenes: One Proposal, Multiple Predictions”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 12214-12223. [cited by applicant]
Chu, et al., “Online Multi-Object Tracking Using CNN-based Single Object Tracker with Spatial-Temporal Attention Mechanism”, In Proceedings of IEEE International Conference on Computer Vision, Oct. 22, 2017, pp. 4836-48… [cited by applicant]
Dendorfer, et al., “MOT20: A benchmark for multi object tracking in crowded scenes”, In Repository of arXiv:2003.09003v1, Mar. 19, 2020, pp. 1-7. [cited by applicant]
Dollar, et al., “Pedestrian Detection: A Benchmark”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 20, 2009, pp. 304-311. [cited by applicant]
Ess, et al., “A Mobile Vision System for Robust Multi-Person Tracking”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 23, 2008, 8 Pages. [cited by applicant]
Felzenszwalb, et al., “Object Detection with Discriminatively Trained Part Based Models”, In Journal of IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 32, No. 9, Sep. 22, 2009, pp. 1-20. [cited by applicant]
He, et al., “Mask R-CNN”, In Proceedings of the IEEE International Conference on Computer Vision, Oct. 22, 2017, pp. 1-12. [cited by applicant]
Hornakova, et al., “Lifted Disjoint Paths with Application in Multiple Object Tracking”, In Proceedings of International Conference on Machine Learning, Nov. 21, 2020, 12 Pages. [cited by applicant]
Karthik, et al., “Simple Unsupervised Multi-Object Tracking”, In Repository of arXiv:2006.02609v1, Jun. 4, 2020, op. 1-14. [cited by applicant]
Kingma, et al., “Adam: A Method for Stochastic Optimization”, In Repository of arXiv:1412.6980, Dec. 22, 2014, 9 Pages. [cited by applicant]
Law, et al., “CornerNet: Detecting Objects as Paired Keypoints”, In Proceedings of the European Conference on Computer Vision, Sep. 8, 2018, pp. 1-17. [cited by applicant]
Li, et al., “Learning to Associate: Hybrid-Boosted Multi-Target Tracker for Crowded Scene”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 20, 2009, pp. 2953-2960. [cited by applicant]
Liu, et al., “GSM: Graph Similarity Model for Multi-Object Tracking”, In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, Jul. 2020, pp. 530-536. [cited by applicant]
Liu, et al., “Real-Time Online Multi-Object Tracking in Compressed Domain”, In Journal of IEEE Access, vol. 7, Jun. 2019, pp. 76489-76499. [cited by applicant]
Maaten, et al., “Visualizing Data using t-SNE”, In Journal of Machine Learning Research, vol. 9, Nov. 2008, pp. 2579-2605. [cited by applicant]
Milan, et al., “MOT16: A Benchmark for Multi-Object Tracking”, In Repository of arXiv:1603.00831v2, May 3, 2016, pp. 1-12. [cited by applicant]
Pang, et al., “TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training Model”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 6308-6318. [cited by applicant]
Peng, et al., “Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking”, In Proceedings of European Conference on Computer Vision, Aug. 23, 2020, pp. 1-2… [cited by applicant]
Porzi, et al., “Learning Multi-Object Tracking and Segmentation from Automatic Annotations”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 6846-6855. [cited by applicant]
Redmon, et al., “You Only Look Once: Unified, Real-Time Object Detection”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 27, 2016, pp. 779-788. [cited by applicant]
Ren, et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, In Journal of IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, No. 6, Jun. 2017, pp. 1137-1149. [cited by applicant]
Sadeghian, et al., “Tracking The Untrackable: Learning to Track Multiple Cues with Long-Term Dependencies”, In Proceedings of the IEEE International Conference on Computer Vision, Oct. 22, 2017, pp. 300-311. [cited by applicant]
Shao, et al., “CrowdHuman: A Benchmark for Detecting Human in a Crowd”, In Repository of arXiv:1805.00123v1, Apr. 30, 2018, pp. 1-9. [cited by applicant]
Tang, et al., “Multiple People Tracking by Lifted Multicut and Person Re-identification”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp. 3539-3548. [cited by applicant]
Voigtlaender, et al., “MOTS: Multi-Object Tracking and Segmentation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 16, 2019, pp. 7942-7951. [cited by applicant]
Wang, et al., “CycAs: Self-supervised Cycle Association for Learning Re-Identifiable Descriptions”, In Repository of arXiv:2007.07577v1, Jul. 15, 2020, pp. 1-16. [cited by applicant]
Wang, et al., “Towards Real-Time Multi-Object Tracking”, In Repository of arXiv:1909.12605v2, Jul. 14, 2020, 17 Pages. [cited by applicant]
Wojke, et al., “Simple Online and Realtime Tracking with a Deep Association Metric”, In Proceedings of IEEE International Conference on Image Processing, Sep. 17, 2017, 5 Pages. [cited by applicant]
Xiao, et al., “Joint Detection and Identification Feature Learning for Person Search”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp. 3415-3424. [cited by applicant]
Yang, et al., “Exploit All the Layers: Fast and Accurate CNN Object Detector with Scale Dependent Pooling and Cascaded Rejection Classifiers”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni… [cited by applicant]
Yu, et al., “POI: Multiple Object Tracking with High Performance Detection and Appearance Feature”, In Proceedings of European Conference on Computer Vision, Oct. 8, 2016, pp. 1-7. [cited by applicant]
Zhang, et al., “CityPersons: A Diverse Dataset for Pedestrian Detection”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp. 3213-3221. [cited by applicant]
Zhang, et al., “FairMOT: On the Fairness of Detection and Re-Identification in Multiple Object Tracking”, In Repository of arXiv:2004.01888v5, Sep. 9, 2020, pp. 1-13. [cited by applicant]
Zhang, et al., “Multiplex Labeling Graph for Near-Online Tracking in Crowded Scenes”, In Journal of IEEE Internet of Things Journal, vol. 7, No. 9, Sep. 2020, pp. 7892-7902. [cited by applicant]
Zheng, et al., “Person Re-Identification in the Wild”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, 10 Pages. [cited by applicant]
Zhou, et al., “Objects as Points”, In Repository of arXiv:1904.07850v1, Apr. 16, 2019, pp. 1-12. [cited by applicant]
Zhou, et al., “Tracking Objects as Points”, In Proceedings of European Conference on Computer Vision, Aug. 23, 2020, pp. 1-22. [cited by applicant]
Zhu, et al., “Crowded Human Detection via an Anchor-pair Network”, In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Mar. 1, 2020, pp. 1391-1399. [cited by applicant]
Zhu, et al., “Online Multi-Object Tracking with Dual Matching Attention Networks”, In Proceedings of the European Conference on Computer Vision, Sep. 8, 2018, pp. 1-17. [cited by applicant]
Lu, et al., “An Occlusion Tolerent Method for Multi-Object Tracking”, In Proceedings of the 7th World Congress on Intelligent Control and Automation, Jun. 25, 2008, pp. 5105-5110. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/028647”, Mailed Date: Sep. 21, 2022, 10 Pages. [cited by applicant]
Xu, et al., “Partial Observation vs. Blind Tracking through Occlusion”, In Proceedings of 13th British Machine Vision Conference, Jan. 1, 2002, pp. 777-786. [cited by applicant]
Cited By (1)
US 12,548,170