IP Library Granted Patent US 11,935,296
Granted Patent B2
US 11,935,296 · App. 17/411,728 · Granted Mar 19, 2024

Apparatus and method for online action detection

Inventors: Jin Young Moon (Daejeon, KR); Hyung Il Kim (Daejeon, KR); Jong Youl Park (Daejeon, KR); Kang Min Bae (Daejeon, KR); Ki Min Yun (Daejeon, KR)
Assignee: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
G06V20/41G06V20/46G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,935,296
App. No.
17/411,728
Granted
Mar 19, 2024
Kind
B2
Abstract

Provided is an apparatus for online action detection, the apparatus including a feature extraction unit configured to extract a chunk-level feature of a video chunk sequence of a streaming video, a filtering unit configured to perform filtering on the chunk-level feature, and an action classification unit configured to classify an action class using the filtered chunk-level feature.

Claims (32)

1. An apparatus for online action detection, the apparatus comprising:

a feature extraction unit configured to extract a chunk-level feature of a video chunk sequence of a streaming video;

a filtering unit configured to perform filtering on the chunk-level features; and

an action classification unit configured to classify an action class using the filtered chunk-level features, wherein

the filtering unit configured to receive a chunk-level feature sequence and infer a relation of an action instance represented by a current chunk and other chunks to generate a filtered chunk-level feature sequence to be used for action classification, and

the filtering unit configured to predict an action relevance of each chunk with the current point in time so as to generate filtered features in which a feature of a chunk related to the current point in time is emphasized and a feature of a chunk unrelated to the current point in time is filtered out.

2. The apparatus of claim 1 , wherein the feature extraction unit includes:

a video frame extraction unit configured to extract frames from a video segment;

a single chunk feature generation unit configured to generate a single chunk feature for each chunk; and

a chunk-level feature extraction unit configured to generate the chunk-level feature sequence using the single chunk features.

3. The apparatus of claim 2 , wherein the video frame extraction unit divides a video segment from a previous time to a current time to be processed into video chunks of the same length and extracts frames from the video segment corresponding to a preset number of the chunks.

4. The apparatus of claim 3 , wherein the video frame extraction unit extracts each frame from the video segment or extracts a red-green-blue (RGB) frame or a flow frame through sampling.

5. The apparatus of claim 1 , wherein the action classification unit acquires an action class probability for a current action included in the input video segment.

6. The apparatus of claim 5 , wherein the action classification unit receives the filtered chunk-level feature sequence and outputs the action class probability of a current action for each class including action classes and a background.

7. A method of online action detection, the method comprising the steps of:

(a) extracting a chunk-level feature of a video chunk sequence of a streaming video;

(b) performing filtering on the chunk-level features of the input video segment; and

(c) classifying an action class and outputting an action class probability using the chunk-level features filtered in the step (b), wherein

the step (b) includes performing the filtering using relevance between the chunk-level feature and an action instance,

the step (b) includes receiving the chunk-level feature sequence and inferring a relation of an action instance represented by a current chunk and other chunks to generate a filtered chunk-level feature sequence to be used for action classification, and

the step (b) includes predicting an action relevance of each chunk with the current point in time so as to generate filtered features in which a feature of a chunk related to the current point in time is emphasized and a feature of a chunk unrelated to the current point in time is filtered out.

8. The method of claim 7 , wherein the step (a) includes extracting frames from a video segment, generating a single chunk feature for each chunk, and generating a chunk-level feature sequence using the single chunk features.

9. The method of claim 8 , wherein the step (a) includes dividing a video segment from a previous time to a current time into video chunks of the same length and extracting each frame from the video segment corresponding to a preset number of the chunks or extracting frames through sampling.

10. The method of claim 7 , wherein the step (c) includes receiving the filtered chunk-level feature sequence and outputting the action class probability of a current point in time for each class including the action class and a background.

11. An apparatus for online action detection, comprising:

an input unit configured to receive a video chunk sequence of a streaming video;

a memory in which a program for detecting an action using the video chunk sequence is stored; and

a processor configured to execute the program,

wherein the processor extracts a chunk-level feature of the video chunk sequence, performs filtering on the chunk-level feature to generate a chunk-level feature sequence to be used for action classification, and classifies an action class and outputs an action class probability using the chunk-level feature sequence, wherein

the processor configured to infer a relation of an action instance represented by a current chunk and other chunks to generate a filtered chunk-level feature sequence to be used for action classification, and

the processor configured to predict an action relevance of each chunk with the current point in time so as to generate filtered features in which a feature of a chunk related to the current point in time is emphasized and a feature of a chunk unrelated to the current point in time is filtered out.

12. The apparatus of claim 11 , wherein the processor outputs the action class probability of a current point in time for each class including the action class and a background using the chunk-level feature sequence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2021
From: MOON, JIN YOUNG; KIM, HYUNG IL; PARK, JONG YOUL; BAE, KANG MIN; YUN, KI MIN
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 057289/0068 →
Priority Claims (1)
KR 10-2020-0106794 · Aug 25, 2020 · national
Continuity (1)
Related Publication 20220067382A1 · Mar 3, 2022
Cited By (1)
US 12,555,375