IP Library › Granted Patent US 12,333,793
Granted Patent B2
US 12,333,793 · App. 17/296,560 · Granted Jun 17, 2025

Method for common detecting, tracking and classifying of objects

Inventors: Sikandar Amin (Neufahrn B. Freising, DE); Bharti Munjal (Munich, DE); Meltem Demirkus Brandlmaier (Munich, DE); Abdul Rafey Aftab (Munich, DE); Fabio Galasso (Rome, IT)
Assignee: OSRAM GMBH
G06V10/82G06F18/211G06F18/2415G06F18/2431G06N3/08G06V10/25G06V10/764G06V20/00G06V20/41G06V20/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,793
App. No.
17/296,560
Granted
Jun 17, 2025
Kind
B2
Abstract

A method for machine-based training of a computer-implemented network for common detecting, tracking, and classifying of at least one object in a video image sequence having a plurality of successive individual images. A combined error may be determined during the training, which error results from the errors of the determining of the class identification vector, determining of the at least one identification vector, the determining of the specific bounding box regression, and the determining of the inter-frame regression.

Claims (26)

1. A method for common detecting, tracking, and classifying of at least one object in a video image sequence having a multiplicity of successive frames by means of a trained computer-implemented network; wherein the method comprises:

receiving a first frame and a succeeding second frame;

detecting at least one object in the first frame and the succeeding second frame;

selecting an object from the first and second frames;

ascertaining at least one classification vector and a position for the object from the first and second frames;

ascertaining an association value on the basis of the ascertained classification vector and the position comprising providing a relative weighting between the at least one ascertained classification vector and the position by an already trained network during its execution in detecting, tracking and classifying, said relative weighting being dependent on the time between the received first and succeeding second frames;

generating a temporarily consistent and unique identification vector of the at least one object for each frame in response to the ascertained association value.

2. The method as claimed in claim 1 , wherein the detecting at least one object comprises:

generating a bounding box surrounding the at least one object;

generating a prediction for the bounding box from the first frame to the second frame;

generating a velocity vector for the bounding box of the first frame.

3. The method as claimed in claim 2 , wherein a bounding box is provided for each of the at least one object.

4. The method as claimed in claim 2 , wherein the selecting comprises at least one of:

selecting the bounding box of the first frame and selecting the bounding box of the second frame;

selecting the prediction and selecting the bounding box of the second frame; and

selecting the velocity vector and selecting the bounding box of the second frame.

5. The method as claimed in claim 1 , wherein the ascertaining at least one classification vector comprises:

acquiring features of the object;

calculating a unique feature vector from the acquired features; and

classifying the object from a group of predefined classes on the basis of the acquired features or on the basis of calculated feature vector.

6. The method as claimed in claim 1 , wherein the relative weighting between the ascertained classification vector and the position rises with increasing time or a falling frame rate between the first and second frames.

7. The method as claimed in claim 1 , wherein the generating a temporarily consistent and unique identification vector comprises a Hungarian combinatorial optimization method.

8. The method as claimed in claim 1 , wherein the unique identification of an object of a second frame that is not assignable to any object of a first frame is compared with the identification of an object of a third frame temporally proceeding the first frame.

9. A system for classifying objects on a computer which comprises:

a memory and one or more processors configured to perform the method as claimed in claim 1 .

10. A computer program product stored on a non-transitory computer readable medium and having instructions which, when executed on one or more processors, carry out the method as claimed in claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2022
From: AMIN, SIKANDAR; MUNJAL, BHARTI; BRANDLMAIER, MELTEM DEMIRKUS; AFTAB, ABDUL RAFEY; GALASSO, FABIO
To: OSRAM GMBH
Reel/Frame 059110/0956 →
Priority Claims (2)
DE 10 2018 220 274.5 · Nov 26, 2018 · national
DE 10 2018 220 276.1 · Nov 26, 2018 · national
Continuity (1)
Related Publication 20220027664A1 · Jan 27, 2022
References Cited (15)
US 20170046560A1 · Tsur · 2017 [cited by examiner]
US 20170053167A1 · Ren · 2017 [cited by examiner]
US 20180137892A1 · Ding · 2018 [cited by examiner]
US 20180341813A1 · Chen · 2018 [cited by examiner]
US 20200051254A1 · Habibian · 2020 [cited by examiner]
Feichtenhofer, Christoph et al., “Detect to Track and Track to Detect”, Oct. 22-29, 2017, pp. 3038-3046, J017 IEEE International Conference on Computer Vision {ICCV), IEEE,Venice, Italy (Year: 2017). [cited by examiner]
International Search Report issued for the corresponding PCT application No. PCT/EP2019/081317, mailed Feb. 20, 2020, 4 pages (for informational purposes only). [cited by applicant]
Feichtenhofer, Christoph et al., “Detect to Track and Track to Detect”, Oct. 22-29, 2017, pp. 3038-3046, 2017 IEEE International Conference on Computer Vision (ICCV), IEEE,Venice, Italy. [cited by applicant]
Xiao, Tong et al., “Joint Detection and Identification Feature Learning for Person Search”, Jul. 21-26, 2017, pp. 3376-3385, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Honolulu, HI, US… [cited by applicant]
Luo, Wenjie et al., “Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net”, IEEE, Jun. 18-23, 2018, pp. 3569-3577, 2018 IEEE/CVF Conference on Computer Vis… [cited by applicant]
Search Report issued for the German patent application No. 10 2018 220 274.5, issued Nov. 18, 2019, 13 pages (for informational purposes only). [cited by applicant]
Arun, Aditya et al., “Dissimilarity Coefficient based Weakly Supervised Object Detection”, Jun. 15-20, 2019, 14 pages, Conference: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Long B… [cited by applicant]
Bewley, Alex, “Vision based Detection and Tracking in Dynamic Environments with Minimal Supervision”, 2017, 194 pages, PhD Thesis, Queensland University of Technology. [cited by applicant]
Huang, Jonathan et al., “Speed/accuracy trade-offs for modern convolutional object detectors”, Jul. 21-26, 2017, pp. 7310-7319, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Honolulu, IH … [cited by applicant]
Wang, Hanyu et al., “MPNET: An End-to-End Deep Neural Network for Object Detection in Surveillance Video”, IEEE, May 24, 2018, pp. 30296-30308, IEEE Access, vol. 6, IEEE. [cited by applicant]