IP Library Granted Patent US 12,561,815
Granted Patent B2
US 12,561,815 · App. 18/310,216 · Granted Feb 24, 2026

Method of detecting object in video and video analysis terminal

Inventors: Kichang Yang (Seoul, KR); Youngki Lee (Seoul, KR); Juheon Yi (Seoul, KR); Kyungjin Lee (Seoul, KR)
Assignee: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
G06T7/20G06V10/25G06V10/44G06V10/762G06V10/764G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,815
App. No.
18/310,216
Granted
Feb 24, 2026
Kind
B2
Abstract

Provided is a video analysis terminal including a patch recommendation unit configured to recommend tracking-failure patches and new-object patches in a current frame of a video image, and a patch aggregation unit configured to generate a first patch cluster by collecting the tracking-failure patches recommended in the current frame, and generate a second patch cluster by collecting the new-object patches recommended in the current frame.

Claims (29)

1 . A video analysis terminal comprising:

at least one processor, and

memory storing instructions,

wherein the instructions, when executed by the at least one processor, cause the video analysis terminal to:

recommend tracking-failure patches and new-object patches in a current frame of a video image; and

generate a first patch cluster by collecting the tracking-failure patches recommended in the current frame, or generate a second patch cluster by collecting the new-object patches recommended in the current frame,

wherein the tracking-failure patches indicate regions for which tracking has failed in the current frame,

wherein the new-object patches indicate regions in which a new object is likely to be present but has not been detected and thus tracking has not been performed,

wherein the tracking-failure patches are extracted based on features, which are extracted from the current frame and imply a tracking failure, and machine learning for predicting a tracking failure degree based on the extracted features, and

wherein the extracted features comprise normalized cross correlation (NCC) between a bounding box in a frame before tracking and a bounding box after the tracking, a velocity of a bounding box, an acceleration of the bounding box, a gradient of a region around the bounding box, and confidence of detection.

2 . The video analysis terminal of claim 1 , wherein the instructions, when executed by the at least one processor, cause the vide analysis terminal to receive the first patch cluster or the second patch cluster, and detect an object, so as to improve an object detection speed.

3 . The video analysis terminal of claim 1 , wherein the first patch cluster or the second patch cluster has a rectangular shape.

4 . The video analysis terminal of claim 1 , wherein a size of the first patch cluster or the second patch cluster is adjusted according to a size and number of tracking-failure patches or new-object patches included in each of the first patch cluster or the second patch cluster.

5 . The video analysis terminal of claim 1 , wherein the instructions, when executed by the at least one processor, cause the video analysis terminal to collect the new-object patches and the tracking-failure patches in every t frame of the video image, before performing object detection.

6 . The video analysis terminal of claim 1 , wherein the instructions, when executed by the at least one processor, cause the video analysis terminal to recommend the new-object patches by using an edge intensity and a refresh interval.

7 . The video analysis terminal of claim 1 , wherein the instructions, when executed by the at least one processor, cause the video analysis terminal to generate the first patch cluster by classifying and arranging the collected tracking-failure patches according to error values.

8 . The video analysis terminal of claim 1 , wherein machine learning is performed by using a decision tree classification model, based on the extracted features, and then a degree of tracking failure is predicted by identifying an intersection over union (IoU), which is a degree of overlap between a tracked bounding box and a real object.

9 . A method, performed by a terminal, of performing video object detection, the method comprising:

recommending tracking-failure patches and new-object patches in a current frame of a video image; and

generating a first patch cluster by collecting the tracking-failure patches recommended in the current frame, or generating a second patch cluster by collecting the new-object patches recommended in the current frame,

wherein the tracking-failure patches indicate regions for which tracking has failed in the current frame, and

wherein the new-object patches indicate regions in which a new object is likely to be present but has not been detected and thus tracking has not been performed,

wherein the tracking-failure patches are extracted based on features, which are extracted from the current frame and imply a tracking failure, and machine learning for predicting a tracking failure degree extracted features, and

wherein the extracted features comprise normalized cross correlation (NCC) between a bounding box in a frame before tracking and a bounding box after the tracking, velocity of a bounding box, an acceleration of the bounding box, a gradient of a region around the bounding box, and confidence of detection.

10 . The method of claim 9 , further comprising receiving the first patch cluster or the second patch cluster, and detecting an object, so as to improve an object detection speed.

11 . The method of claim 9 , wherein a size of the first patch cluster or the second patch cluster is adjusted according to a size and number of tracking-failure patches or new-object patches included in each of the first patch cluster or the second patch cluster.

12 . The method of claim 9 , wherein the recommending comprises recommending the new-object patches by using an edge intensity and a refresh interval.

13 . The method of claim 9 , wherein the recommending comprises recommending the tracking-failure patches based on features, which are extracted from the current frame and imply a tracking failure, and machine learning for predicting a tracking failure degree based on the extracted features.

14 . A computer program stored in a non-transitory computer-readable recording medium, for executing, on the terminal, the method of performing the video object detection of claim 9 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2023
From: YANG, KICHANG; LEE, YOUNGKI; YI, JUHEON; LEE, KYUNGJIN
To: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Reel/Frame 063496/0596 →
Priority Claims (1)
KR 10-2022-0054394 · May 2, 2022 · national
Continuity (1)
Related Publication 20230351613A1 · Nov 2, 2023
References Cited (13)
US 7130446B2 · Rui et al. · 2006 [cited by applicant]
US 11076589B1 · Sibley · 2021 [cited by examiner]
US 20030235327A1 · Srinivasa · 2003 [cited by examiner]
US 20120093368A1 · Choi et al. · 2012 [cited by applicant]
US 20170256057A1 · Bak et al. · 2017 [cited by applicant]
US 20180082428A1 · Leung · 2018 [cited by examiner]
US 20190050994A1 · Fukagai · 2019 [cited by examiner]
JP 2002236929A · 2002 [cited by applicant]
KR 1020030045624A · 2003 [cited by applicant]
KR 100970119B1 · 2010 [cited by applicant]
KR 1020150033047A · 2015 [cited by applicant]
KR 1020190048208A · 2019 [cited by applicant]
Office Action issued Feb. 27, 2025 in Korean Patent Application No. 10-2022-0054394. [cited by applicant]