IP Library Granted Patent US 12,639,948
Granted Patent B2
US 12,639,948 · App. 18/688,343 · Granted May 26, 2026

Method and system for video temporal action proposal generation

Inventors: Ping Luo (Hong Kong, HK); Jiannan Wu (Beijing, CN); Jiajun Shen (Hong Kong, HK); Lan Ma (Hong Kong, HK)
Assignees: University of Hong Kong Versitech Limited; TCL Technology Group Corporation
G06V20/49G06V10/44G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,948
App. No.
18/688,343
Granted
May 26, 2026
Kind
B2
Abstract

A system and method for video temporal action proposal generation are provided. It processes video features extracted from the input video through an encoder to obtain video encoding features with global information, extracts corresponding interest segment features from the video encoding features using pre-trained proposal segments, and provides them to the decoder. The decoder generates segment features based on the interest segment features corresponding to each proposal segment and the pre-trained proposal features. These are then provided to the prediction module, generating temporal action proposal results based on the decoder's segment features. The solution in embodiments of the present invention can effectively capture global context information of the video, obtaining video encoding features with stronger representational capabilities. By introducing the several learnable proposal segments to extract the corresponding position-based feature sequences from the video encoding features, the training convergence speed is enhanced, and computational burden is reduced.

Claims (17)

1 . A system for video temporal action proposal generation comprising a feature extraction module, a feature processing module, and a prediction module, wherein:

the feature extraction module is used to extract, from an input video, video features related to the input video;

the feature processing module comprises a pre-trained encoder and decoder, wherein the encoder, based on the video features from the feature extraction module, obtains video encoding features with global information, extracts interest segment features corresponding to each proposal segment from the video encoding features through pre-trained proposal segments, and provides them to the decoder, wherein the decoder, based on the interest segment features corresponding to each proposal segment and pre-trained proposal features corresponding to the proposal segments, generates segment features and provides them to the prediction module;

the prediction module generates temporal action proposal results based on the segment features from the decoder, comprising proposal boundaries and confidence scores.

2 . The system according to claim 1 , wherein the encoder comprises a graph attention layer, a multi-head self-attention layer, and a feed-forward layer, the encoder adds results of the video features and position coding and uses them as a value vector input for the multi-head self-attention layer, and, simultaneously, the encoder provides the results as input to be processed by the graph attention layer, wherein an output thereof undergoes a linear transformation to obtain a query vector and a key vector for the multi-head self-attention layer.

3 . The system according to claim 1 , wherein the decoder includes a multi-head self-attention layer, a sparse interaction module, and a feed-forward layer, and wherein the decoder processes the proposal features corresponding to the proposal segment through the multi-head self-attention layer and then provides them to the sparse interaction module, for performing sparse interaction with the interest segment features corresponding to the proposal segment; wherein an output of the sparse interaction module is processed through the feed-forward layer to obtain the segment features.

4 . The system according to claim 1 , wherein the feature processing module is constructed based on a transformer model.

5 . The system according to claim 1 , wherein the prediction module performs boundary regression and binary classification prediction based on the segment features from the decoder.

6 . A method for generating temporal action proposal generation using the system according to claim 1 , comprising:

step S 1 ) extracting video features from an input video through a feature extraction module;

step S 2 ) processing the extracted video features using an encoder to obtain video encoding features with global context information of the input video;

step S 3 ) utilizing each of pre-trained multiple proposal segments to extract corresponding interest segment features from the video encoding features;

step S 4 ) through the decoder, generating segment features based on the interest segment features corresponding to each proposal segment and the pre-trained proposal features corresponding to the proposal segments;

step S 5 ) employing a prediction module to perform boundary regression and binary classification prediction based on the segment features from the decoder, so as to output corresponding temporal action proposal results.

7 . The method according to claim 6 , wherein the encoder comprises a graph attention layer, a multi-head self-attention layer, and a feed-forward layer, wherein the step S 2 ) comprises taking results of adding the video features and position coding as a value vector input for the multi-head self-attention layer, and, simultaneously, taking the results as input to be processed by the graph attention layer, wherein an output thereof undergoes a linear transformation to obtain a query vector and a key vector for the multi-head self-attention layer.

8 . The method according to claim 6 , wherein the decoder includes a multi-head self-attention layer, a sparse interaction module, and a feed-forward layer, wherein the step S 4 ) comprises processing the proposal features corresponding to each proposal segment through the multi-head self-attention layer and then feeding them into the sparse interaction module for performing sparse interaction with the interest segment features corresponding to the proposal segment; wherein an output of the sparse interaction module is processed through the feed-forward layer to obtain the segment features.

9 . A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, wherein the program, when executed, implements the method according to claim 6 .

Assignments (2)
CHANGE OF NAME Recorded Mar 6, 2026
From: VERSITECH LIMITED
To: UNIVERSITY OF HONG KONG VERSITECH LIMITED
Reel/Frame 075020/0740 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2024
From: LUO, PING; WU, JIANNAN; SHEN, JIAJUN; MA, LAN
To: VERSITECH LIMITED; TCL TECHNOLOGY GROUP CORPORATION
Reel/Frame 066658/0834 →
Priority Claims (1)
CN 202111049034.6 · Sep 8, 2021 · national
Continuity (1)
Related Publication 20250069397A1 · Feb 27, 2025
References Cited (13)
US 20120219213A1 · Wang et al. · 2012 [cited by applicant]
US 20220058396A1 · Chen · 2022 [cited by examiner]
CN 110163129A · 2019 [cited by applicant]
CN 110852256A · 2020 [cited by applicant]
CN 111327949A · 2020 [cited by applicant]
CN 111372123A · 2020 [cited by applicant]
CN 112183588A · 2021 [cited by applicant]
CN 112906586A · 2021 [cited by applicant]
Tianwei Lin et al., “Temporal Convolution Based Action Proposal: Submission to ActivityNet 2017,” Computer Vision and Pattern Recognition, 2018, 1707.06750v3, p. 1-4. [cited by applicant]
Jialin Gao et al., “Accurate Temporal Action Proposal Generation with Relation-Aware Pyramid Network,” The Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020, vol. 34 (7), p. 10810-10817. [cited by applicant]
Tianwei Lin et al., “BMN: Boundary-Matching Network for Temporal Action Proposal Generation,” IEEE/CVF International Conference on Computer Vision, 2019, p. 3889-3898. [cited by applicant]
Chuming Lin et al., “Fast Learning of Temporal Action Proposal via Dense Boundary Generator,” The Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020, vol. 34 (7), p. 11499-11506. [cited by applicant]
Haisheng Su et al., “BSN++: Complementary Boundary Regressor with Scale-Balanced Relation Modeling for Temporal Action Proposal Generation,” The Thirty-Fifth AAAI Conference on Artificial Intelligence, 2021, p. 2602-261… [cited by applicant]