IP Library › Granted Patent US 12,749,326
Granted Patent B2
US 12,749,326 · App. 18/667,244 · Granted Sep 29, 2026

Active sparse labeling of video frames

Inventors: Yogesh Singh Rawat (Orlando, FL); Aayush Jung Bahadur Rana (Orlando, FL)
Assignee: University of Central Florida Research Foundation, Inc.
G06V20/70G06V10/774
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,749,326
App. No.
18/667,244
Granted
Sep 29, 2026
Kind
B2
Abstract

An active sparse labeling system that provides high performance and low annotation costs by performing partial instance annotation (i.e., sparse labeling) by frame level selection to annotate the most informative frames, thereby improving action detection task efficiencies. The active sparse labeling system utilizes a frame level cost estimation to determine the utility of each frame in a video based on the frame's impact on action detection. The system includes an adaptive proximity-aware uncertainty model, which is an uncertainty-based frame scoring mechanism. The adaptive proximity-aware uncertainty model estimates a frame's utility using the uncertainty of detections of the frame's proximity to existing annotations, thereby determining a diverse set of frames in a video which are effective for learning the task of dense video understanding (such as action detection). In addition, the active sparse labeling system includes a loss formulation training model (max-Gaussian weighted loss) that uses weighted pseudo-labeling.

Claims (65)

1 . A system for active sparse labeling of multimedia input to facilitate dense video understanding, comprising at least one computer processor configured to:

a) receive a multimedia input comprising a plurality of frames;

b) apply an adaptive proximity-aware uncertainty selection model to iteratively select a subset of frames from the multimedia input, wherein the selection model is configured to:

i. calculate frame-level uncertainty for each frame based on pixel-wise confidence scores of localization;

ii. determine an adaptive distance metric based on the proximity of a new frame to previously selected frames;

iii. compute a selection score for each frame based on the frame-level uncertainty and the adaptive distance metric;

iv. select frames for annotation when their selection scores exceed a predefined threshold, thereby ensuring a diverse and informative set of frames is chosen across the temporal domain of the multimedia input;

c) annotate the selected subset of frames to create a labeled dataset;

d) use the labeled dataset to train an action detection model to achieve or exceed a predetermined precision benchmark; and

e) update the action detection model iteratively by reapplying the adaptive proximity-aware uncertainty selection model to further refine frame selection based on updated model insights and annotations until the precision benchmark is met or an annotation cost budget is exhausted.

2 . The system of claim 1 , wherein the precision benchmark algorithm is selected from the group consisting of video-metric average precision (v-mAP), frame-metric average precision (f-mAP) and mean average precision (mAP).

3 . The system of claim 1 , wherein the max-Gaussian weighted loss model is further configured to assign weights to each frame based on the proximity of their localization to ground-truth annotations, whereby frames closer to the ground-truth data are given higher weight.

4 . The system of claim 3 , wherein the max-Gaussian weighted loss model adjusts the localization loss for each frame based on the assigned weights.

5 . The system of claim 4 , wherein the localization loss adjustments are applied iteratively, with each iteration of the training cycle using the updated frame weights to continuously refine the action detection model's accuracy.

6 . The system of claim 3 , wherein the weights assigned by the max-Gaussian weighted loss model follow a Gaussian distribution centered on the frame's distance to the nearest ground-truth annotation, with a variance that adjusts adaptively based on the overall performance of the action detection model in previous training iterations.

7 . The system of claim 6 , wherein the variance of the Gaussian distribution used to assign weights to frames is reduced as the difference between the action detection model's performance and the predetermined precision benchmark decreases, thus focusing training more on frames that significantly deviate from desired outcomes.

8 . The system of claim 1 , wherein the adaptive proximity-aware uncertainty selection model incorporates an intra-sample approach whereby the subset of frames selected represents different temporal segments of the multimedia input.

9 . The system of claim 1 , wherein the adaptive proximity-aware uncertainty selection model employs an uncertainty-based scoring mechanism that assigns priority to frames to those with higher uncertainty scores for subsequent annotation.

10 . The system of claim 1 , further comprising a pseudo-label generation function configured to:

a. generate pseudo-labels for non-selected frames utilizing interpolation and a spatio-temporal superpixel method; and

b. refine pseudo-labels over time responsive to newly acquired annotations and prior predictions.

11 . The system of claim 10 , wherein the spatio-temporal superpixel method aggregates visually similar pixels across frames to extend sparse annotations.

12 . The system of claim 1 , further comprising a scoring mechanism configured to compute a selection score for each frame by applying proximity and uncertainty metrics for targeted frame selection.

13 . The system of claim 1 , wherein the system applies a mix of annotation types including bounding boxes, pixel-wise masks, and scribbles.

14 . The system of claim 1 , wherein the computer processor is configured to apply the adaptive proximity-aware uncertainty selection model iteratively across the multimedia input.

15 . A system for active sparse labeling of multimedia input to facilitate dense video understanding, the system comprising at least one computer processor configured to:

a. receive a multimedia input comprising a plurality of frames;

b. apply an adaptive proximity-aware uncertainty selection model to iteratively select a subset of frames from the multimedia input, the model configured to:

i. calculate frame-level uncertainty for each frame based on pixel-wise confidence scores of localization;

ii. determine an adaptive distance metric based on the proximity of a new frame to previously selected frames;

iii. compute a selection score for each frame based on the frame-level uncertainty and the adaptive distance metric;

iv. select frames for annotation when their selection scores exceed a predefined threshold, ensuring a diverse and informative set of frames is chosen across the temporal domain of the multimedia input;

c. annotate the selected subset of frames to create a labeled dataset;

d. use the labeled dataset to train an action detection model to achieve or exceed a predetermined precision benchmark;

e. update the action detection model iteratively by reapplying the adaptive proximity-aware uncertainty selection model to further refine frame selection based on updated model insights and annotations until the precision benchmark is met or an annotation cost budget is exhausted;

f. employ an uncertainty-based scoring mechanism that assigns priority to frames with higher uncertainty scores for subsequent annotation;

g. generate pseudo-labels for non-selected frames utilizing interpolation and a spatio-temporal superpixel method and refine pseudo-labels over time responsive to newly acquired annotations and prior predictions;

h. assign weights to each frame based on the proximity of their localization to ground-truth annotations, whereby frames closer to the ground-truth data are given higher weight;

i. adjust the localization loss for each frame based on the assigned weights; and

j. apply the adaptive proximity-aware uncertainty selection model iteratively across the multimedia input.

16 . A method for active sparse labeling of multimedia input to facilitate dense video understanding, implemented by at least one computer processor, the method comprising:

a. receiving a multimedia input comprising a plurality of frames;

b. applying an adaptive proximity-aware uncertainty selection model to iteratively select a subset of frames from the multimedia input, the model configured to:

i. calculate frame-level uncertainty for each frame based on pixel-wise confidence scores of localization;

ii. determine an adaptive distance metric based on the proximity of a new frame to previously selected frames;

iii. compute a selection score for each frame based on the frame-level uncertainty and the adaptive distance metric;

iv. select frames for annotation when their selection scores exceed a predefined threshold, ensuring a diverse and informative set of frames is chosen across the temporal domain of the multimedia input;

c. annotating the selected subset of frames to create a labeled dataset;

d. using the labeled dataset to train an action detection model to achieve or exceed a predetermined precision benchmark; and

e. updating the action detection model iteratively by reapplying the adaptive proximity-aware uncertainty selection model to further refine frame selection based on updated model insights and annotations until the precision benchmark is met or an annotation cost budget is exhausted.

17 . The method of claim 16 , wherein the precision benchmark algorithm is selected from the group consisting of video-metric average precision (v-mAP), frame-metric average precision (f-mAP) and mean average precision (mAP).

18 . The method of claim 16 , further comprising assigning weights to each frame based on the proximity of their localization to ground-truth annotations, whereby frames closer to the ground-truth data are given higher weight.

19 . The method of claim 18 , further comprising adjusting the localization loss for each frame based on the assigned weights.

20 . The method of claim 19 , wherein the localization loss adjustments are applied iteratively, with each iteration of the training cycle using the updated frame weights to continuously refine the action detection model's accuracy.

21 . The method of claim 18 , wherein the weights assigned by a max-Gaussian weighted loss model follow a Gaussian distribution centered on the frame's distance to the nearest ground-truth annotation, with a variance that adjusts adaptively based on the overall performance of the action detection model in previous training iterations.

22 . The method of claim 21 , wherein the variance of the Gaussian distribution used to assign weights to frames is reduced as the difference between the action detection model's performance and the predetermined precision benchmark decreases, thus focusing training more on frames that significantly deviate from desired outcomes.

23 . The method of claim 16 , wherein the adaptive proximity-aware uncertainty selection model incorporates an intra-sample approach whereby the subset of frames selected represents different temporal segments of the multimedia input.

24 . The method of claim 16 , wherein the adaptive proximity-aware uncertainty selection model employs an uncertainty-based scoring mechanism that assigns priority to frames with higher uncertainty scores for subsequent annotation.

25 . The method of claim 16 , further comprising:

a. generating pseudo-labels for non-selected frames utilizing interpolation and a spatio-temporal superpixel method; and

b. refining pseudo-labels over time responsive to newly acquired annotations and prior predictions.

26 . The method of claim 25 , wherein the spatio-temporal superpixel method aggregates visually similar pixels across frames to extend sparse annotations.

27 . The method of claim 16 , further comprising computing a selection score for each frame by applying proximity and uncertainty metrics for targeted frame selection.

28 . The method of claim 16 , wherein a mix of annotation types including bounding boxes, pixel-wise masks, and scribbles is applied.

29 . The method of claim 16 , wherein the adaptive proximity-aware uncertainty selection model is applied iteratively across the multimedia input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2024
From: RAWAT, YOGESH SINGH; RANA, AAYUSH JUNG BAHADUR
To: UNIVERSITY OF CENTRAL FLORIDA RESEARCH FOUNDATION, INC.
Reel/Frame 068030/0850 →
Continuity (2)
Provisional Application 63514482 · Jul 19, 2023
Related Publication 20250029410A1 · Jan 23, 2025
References Cited (54)
US 9111146B2 · Dunlop · 2015 [cited by examiner]
US 10140508B2 · Zhang · 2018 [cited by examiner]
US 10614310B2 · Polak · 2020 [cited by examiner]
US 10924800B2 · Song · 2021 [cited by examiner]
US 11335093B2 · Shrivastava · 2022 [cited by examiner]
US 11756210B2 · Wang · 2023 [cited by examiner]
US 11895343B2 · Mittal · 2024 [cited by examiner]
US 12307756B2 · Goldin · 2025 [cited by examiner]
US 12437519B2 · Tsai · 2025 [cited by examiner]
US 20170262996A1 · Jain · 2017 [cited by examiner]
US 20240153272A1 · Li · 2024 [cited by examiner]
US 20250062040A1 · Brunner · 2025 [cited by examiner]
CN 116189058A · 2023 [cited by examiner]
Joshua Gleason, Carlos D Castillo, and Rama Chellappa. Real-time detection of activities in untrimmed videos. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, pp. 117-125, 2… [cited by applicant]
Mamshad Nayeem Rizve, Ugur Demir, Praveen Tirupattur, Aayush Jung Rana, Kevin Duarte, Ishan R Dave, Yogesh S Rawat, and Mubarak Shah. Gabriella: An online system for real-time activity detection in untrimmed security vi… [cited by applicant]
Rui Hou, Chen Chen, and Mubarak Shah. Tube convolutional neural network (t-cnn) for action detection in videos. In IEEE International Conference on Computer Vision, 2017. [cited by applicant]
Kevin Duarte, Yogesh Rawat, and Mubarak Shah. Videocapsulenet: A simplified network for action detection. In Advances in Neural Information Processing Systems, pp. 7610-7619, 2018. [cited by applicant]
Xitong Yang, Xiaodong Yang, Ming-Yu Liu, Fanyi Xiao, Larry S Davis, and Jan Kautz. Step: Spatio-temporal progressive learning for video action detection. In Proceedings of the IEEE Conference on Computer Vision and Patt… [cited by applicant]
Dong Li, Zhaofan Qiu, Qi Dai, Ting Yao, and Tao Mei. Recurrent tubelet proposal and recognition networks for action detection. In Proceedings of the European conference on computer vision (ECCV), pp. 303-318, 2018. [cited by applicant]
Kirill Gavrilyuk, Amir Ghodrati, Zhenyang Li, and Cees GM Snoek. Actor and action video segmentation from a sentence. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5958-5966, 2018. [cited by applicant]
Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6299-6308, 2017. [cited by applicant]
Pascal Mettes, Cees GM Snoek, and Shih-Fu Chang. Localizing actions from video labels and pseudo-annotations. arXiv preprint arXiv: 1707.09143, 2017. [cited by applicant]
Pascal Mettes and Cees GM Snoek. Pointly-supervised action localization. International Journal of Computer Vision, 127(3):263-281, 2019. [cited by applicant]
Philippe Weinzaepfel, Zaid Harchaoui, and Cordelia Schmid. Learning to track for spatiotemporal action localization. In Proceedings of the IEEE international conference on computer vision, pp. 3164-3172, 2015. [cited by applicant]
Guilhem Cheron, Jean-Baptiste Alayrac, Ivan Laptev, and Cordelia Schmid. A flexible model for training action localization with varying levels of supervision. In Proceedings of the 32nd International Conference on Neura… [cited by applicant]
Victor Escorcia, Cuong D Dao, Mihir Jain, Bernard Ghanem, and Cees Snoek. Guess where? actor-supervision for spatiotemporal action localization. Computer Vision and Image Understanding, 192:102886, 2020. [cited by applicant]
Shiwei Zhang, Lin Song, Changxin Gao, and Nong Sang. GLnet: Global local network for weakly supervised action localization. IEEE Transactions on Multimedia, 22(10):2610-2622, 2019. [cited by applicant]
Philippe Weinzaepfel, Xavier Martin, and Cordelia Schmid. Human action localization with sparse spatial supervision. arXiv preprint arXiv:1605.05197, 2016. [cited by applicant]
Aayush J. Rana and Yogesh S. Rawat. We don't need thousand proposals: Single shot actor-action detection in videos. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 2960-29… [cited by applicant]
Anurag Arnab, Chen Sun, Arsha Nagrani, and Cordelia Schmid. Uncertainty-aware weakly supervised action detection from untrimmed videos. In European Conference on Computer Vision, pp. 751-768. Springer, 2020. [cited by applicant]
Burr Settles. Active learning literature survey. 2009. Computer Sciences Technical Report 1648. University of Wisconsin—Madison. [cited by applicant]
Xin Li and Yuhong Guo. Adaptive active learning for image classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 859-866, 2013. [cited by applicant]
Keze Wang, Dongyu Zhang, Ya Li, Ruimao Zhang, and Liang Lin. Cost-effective active learning for deep image classification. IEEE Transactions on Circuits and Systems for Video Technology, 27(12):2591-2600, 2016. [cited by applicant]
Ajay J Joshi, Fatih Porikli, and Nikolaos Papanikolopoulos. Multi-class active learning for image classification. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 2372-2379. IEEE, 2009. [cited by applicant]
Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Brij B Gupta, Xiaojiang Chen, and Xin Wang. A survey of deep active learning. ACM Computing Surveys (CSUR), 54(9):1-40, 2021. [cited by applicant]
Yarin Gal, Riashat Islam, and Zoubin Ghahramani. Deep bayesian active learning with image data. In International Conference on Machine Learning, pp. 1183-1192. PMLR, 2017. [cited by applicant]
Ashish Kapoor, Kristen Grauman, Raquel Urtasun, and Trevor Darrell. Active learning with gaussian processes for object categorization. In 2007 IEEE 11th international conference on computer vision, pp. 1-8. IEEE, 2007. [cited by applicant]
Hamed H Aghdam, Abel Gonzalez-Garcia, Joost van de Weijer, and Antonio M López. Active learning for deep detection neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3672-36… [cited by applicant]
Alex Holub, Pietro Perona, and Michael C Burl. Entropy-based active learning for object recognition. In 2008 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, pp. 1-8. IEEE, 2008. [cited by applicant]
Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel. Bayesian active learning for classification and preference learning. arXiv preprint arXiv: 1112.5745, 2011. [cited by applicant]
Carl Vondrick and Deva Ramanan. Video annotation and tracking with active learning. Advances in Neural Information Processing Systems, 24:28-36, 2011. [cited by applicant]
Javad Zolfaghari Bengar, Abel Gonzalez-Garcia, Gabriel Villalonga, Bogdan Raducanu, Hamed Habibi Aghdam, Mikhail Mozerov, Antonio M Lopez, and Joost van de Weijer. Temporal coherence for active learning in videos. In Pr… [cited by applicant]
Fabian Caba Heilbron, Joon-Young Lee, Hailin Jin, and Bernard Ghanem. What do i annotate next? An empirical study of active learning for action localization. In Proceedings of the European Conference on Computer Vision … [cited by applicant]
Soumya Roy, Asim Unmesh, and Vinay p. Namboodiri. Deep active learning for object detection. In BMVC, p. 91, 2018. [cited by applicant]
Jiwoong Choi, Ismail Elezi, Hyuk-Jae Lee, Clement Farabet, and Jose M Alvarez. Active learning for deep object detection via probabilistic modeling. In Proceedings of the IEEE/CVF International Conference on Computer Vi… [cited by applicant]
Tianning Yuan, Fang Wan, Mengying Fu, Jianzhuang Liu, Songcen Xu, Xiangyang Ji, and Qixiang Ye. Multiple instance active learning for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pa… [cited by applicant]
Alireza Fathi, Maria Florina Balcan, Xiaofeng Ren, and James M Rehg. Combining self training and active learning for video segmentation. In Proceedings of the British Machine Vision Conference (BMVC), 2011. [cited by applicant]
Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning, pp. 1050-1059. PMLR, 2016. [cited by applicant]
Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah. In defense of pseudo-labeling: An uncertainty-aware pseudo-label selection framework for semi-supervised learning. In International Conference on Lea… [cited by applicant]
Rajat Modi, Aayush Jung Rana, Akash Kumar, Praveen Tirupattur, Shruti Vyas, Yogesh Rawat, and Mubarak Shah. Video action detection: Analysing limitations and challenges. In Proceedings of the IEEE/CVF Conference on Comp… [cited by applicant]
Vicky Kalogeiton, Philippe Weinzaepfel, Vittorio Ferrari, and Cordelia Schmid. Action tubelet detector for spatio- temporal action localization. In Proceedings of the IEEE International Conference on Computer Vision, pp… [cited by applicant]
Akash Kumar and Yogesh Singh Rawat. End-to-end semi-supervised learning for video action detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14700-14710, 2022. [cited by applicant]
Rana, A., & Rawat, Y. S. Are all Frames Equal? Active Sparse Labeling for Video Action Detection. In Advances in Neural Information Processing Systems, 36, 2022. [cited by applicant]
Rana, A., & Rawat, Y. S. (2023). Hybrid Active Learning via Deep Clustering for Video Action Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023. [cited by applicant]