IP Library › Granted Patent US 12,299,982
Granted Patent B2
US 12,299,982 · App. 16/931,228 · Granted May 13, 2025

Systems and methods for partially supervised online action detection in untrimmed videos

Inventors: Mingfei Gao (Sunnyvale, CA); Yingbo Zhou (Mountain View, CA); Ran Xu (Mountain View, CA); Caiming Xiong (Menlo Park, CA)
Assignee: Salesforce, Inc.
G06V20/44G06F17/18G06F18/2113G06F18/214G06F18/2431G06N3/084G06V10/764G06V10/82G06V20/20G06V20/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,982
App. No.
16/931,228
Granted
May 13, 2025
Kind
B2
Abstract

Embodiments described herein provide systems and methods for a partially supervised training model for online action detection. Specifically, the online action detection framework may include two modules that are trained jointly—a Temporal Proposal Generator (TPG) and an Online Action Recognizer (OAR). In the training phase, OAR performs both online per-frame action recognition and start point detection. At the same time, TPG generates class-wise temporal action proposals serving as noisy supervisions for OAR. TPG is then optimized with the video-level annotations. In this way, the online action detection framework can be trained with video-category labels only without pre-annotated segment-level boundary labels.

Claims (81)

1. A method of training an online action detection (OAD) neural network model using a training dataset of untrimmed videos having video-level labels without annotated labels indicating whether a specific video frame contains an action start of a specific action class, the method comprising:

receiving, by a communication interface, an input of the training dataset of untrimmed videos including a set of video-level labels indicating one or more action classes that emerge in the untrimmed videos, wherein the untrimmed videos are used for training the OAD neural network model with no annotated label indicating action starts of the one or more action classes;

generating, by a feature extractor neural network model implemented on one or more hardware processors, feature representations from the training dataset of the untrimmed videos;

generating, by a temporal proposal generator (TPG) neural network model implemented on the one or more hardware processors and receiving the feature representations from the feature extractor neural network, class-wise temporal proposals indicating a respective estimated action start for each action class of the one or more action classes based on the feature representations and the set of video-level labels;

generating, by an online action recognizer (OAR) neural network model implemented on the one or more hardware processors and receiving the feature representations from the feature extractor neural network and the class-wise temporal proposals from the TPG neural network model, per-frame action scores over action classes indicating whether each frame of the untrimmed video contains each specific action class, and a class-agnostic start score indicating whether the respective frame contains a start of any action based on the feature representations and the class-wise temporal proposals;

training the OAD neural network model comprising the TPG neural network model and the OAR neural network model according to a loss metric computed based on the per-frame action scores, the class-agnostic start scores and based on the class-wise temporal proposals as pseudo ground-truth labels; and

generating, by the trained OAD neural network model, a predicted action start for an input real-time video stream.

2. The method of claim 1 , wherein the class-wise temporal proposals indicating the estimated action start for each action class are generated by:

generating, for each untrimmed video with a respective video-level annotation, per-frame action scores for each action class using supervised offline action localization problem; and

computing a video-level score for each action class based on the generated per-frame action scores.

3. The method of claim 2 , further comprising:

selecting a first set of action classes corresponding to the video-level scores higher than a first threshold;

applying a second threshold on a subset of the per-frame scores that corresponds to the first set of action classes along a temporal axis; and

obtaining the class-wise temporal proposals based temporal constraints of the subset of the per-frame scores.

4. The method of claim 2 , further comprising:

computing a predicted video-level probability based on the video-level score for each action class; and

computing a multiple instance learning loss by comparing the respective video-level annotation and the predicted video-level probability.

5. The method of claim 2 , further comprising:

computing a first region feature representation corresponding to a first region of an untrimmed video having a first activity level and a second region feature representation corresponding to a second region of the untrimmed video having a second activity level;

identifying, for the untrimmed video, another untrimmed video that shares a common video-level label for a specific action class with the untrimmed video; and

computing a pair-wise co-activity similarity loss for the untrimmed video and the another untrimmed video based on the first region feature representation and the second region feature representation.

6. The method of claim 5 , wherein the first region feature representation is computed using a temporal attention vector and a vector of the feature representations,

wherein the temporal attention vector is computed by applying temporal softmax over the per-frame action scores.

7. The method of claim 1 , wherein the generating the per-frame action scores and the class-agnostic start score comprises:

updating, at each timestep, a hidden state and a cell state of a long short-term memory based on an input of the feature representations;

applying max pooling on a set of hidden states between a current timestep and a past timestep to obtain an average hidden state;

generating the per-frame action scores based on the hidden state and a first vector of classifier parameters; and

generating the class-agnostic score based on the average hidden state and a second vector of classifier parameters.

8. The method of claim 7 , further comprising:

converting the class-wise temporal proposals to per-frame action labels and binary start labels;

computing a frame loss using a cross entropy loss between the per-frame action labels and the generated per-frame action scores; and

computing a start loss using a focal loss between the binary start labels and the generated class-agnostic score.

9. The method of claim 8 , wherein the loss metric is computed by a weighted sum of an online action recognizer loss and a temporal proposal generator loss,

wherein the online action recognizer loss includes a sum of the frame loss and the start loss, and

wherein the temporal proposal generator loss includes a multiple instance learning loss and a pair-wise co-activity similarity loss.

10. The method of claim 9 , wherein the input of untrimmed videos comprises a first untrimmed video having video-level annotations only, and a second untrimmed video having frame-level annotations, and wherein the OAD neural network model is trained by a combination of the first untrimmed video, and the second untrimmed video supervised by the frame-level annotations.

11. A system for training an online action detection (OAD) neural network model using a training dataset of untrimmed videos having video-level labels without annotated labels indicating whether a specific video frame contains an action start of a specific action class, the system comprising:

a communication interface that receives an input of the training dataset of untrimmed videos including a set of video-level labels indicating one or more action classes that emerge in the untrimmed videos, wherein the untrimmed videos are used for training the OAD neural network model with no annotated label indicating action starts of the one or more action classes;

a memory storing

a feature extractor neural network model, a temporal proposal generator (TPG) neural network model and the OAD neural network model, and a plurality of processor-executable instructions; and

a processor reading from the memory and executing the plurality of processor-executable instructions to:

generate, by a feature extractor neural network model implemented on one or more hardware processors, feature representations from the input of the untrimmed videos;

generate, by a temporal proposal generator (TPG) neural network model implemented on the one or more hardware processors and receiving the feature representations from the feature extractor neural network, class-wise temporal proposals indicating a respective estimated action start for each action class of the one or more action classes based on the feature representations and the set of video-level labels;

generate, by an online action recognizer (OAR) neural network model implemented on the one or more hardware processors and receiving the feature representations from the feature extractor neural network and the class-wise temporal proposals from the TPG neural network model, per-frame action scores over action classes indicating whether each frame of the untrimmed video contains each specific action class, and a class-agnostic start score indicating whether the respective frame contains a start of any action based on the feature representations and the class-wise temporal proposals;

train the OAD neural network model comprising the TPG neural network model and the OAR neural network model according to a loss metric computed based on the per-frame action scores, the class-agnostic start scores and based on the class-wise temporal proposals as pseudo ground-truth labels; and

generate, by the trained OAD neural network model, a predicted action start for an input real-time video stream.

12. The system of claim 11 , wherein the class-wise temporal proposals indicating an estimated action start label for each action class are generated by:

generating, for each untrimmed video with a respective video-level annotation, per-frame action scores for each action class using supervised offline action localization problem; and

computing a video-level score for each action class based on the generated per-frame action scores.

13. The system of claim 12 , wherein the processor is further configured to:

select a first set of action classes corresponding to the video-level scores higher than a first threshold;

apply a second threshold on a subset of the per-frame scores that corresponds to the first set of action classes along a temporal axis; and

obtain the class-wise temporal proposals based temporal constraints of the subset of the per-frame scores.

14. The system of claim 12 , wherein the processor is further configured to:

compute a predicted video-level probability based on the video-level score for each action class; and

compute a multiple instance learning loss by comparing the respective video-level annotation and the predicted video-level probability.

15. The system of claim 12 , wherein the processor is further configured to:

compute a first region feature representation corresponding to a first region of an untrimmed video having a first activity level and a second region feature representation corresponding to a second region of the untrimmed video having a second activity level;

identify, for the untrimmed video, another untrimmed video that shares a common video-level label for a specific action class with the untrimmed video; and

compute a pair-wise co-activity similarity loss for the untrimmed video and the another untrimmed video based on the first region feature representation and the second region feature representation.

16. The system of claim 15 , wherein the first region feature representation is computed using a temporal attention vector and a vector of the feature representations,

wherein the temporal attention vector is computed by applying temporal softmax over the per-frame action scores.

17. The system of claim 11 , wherein the generating the per-frame action scores and the class-agnostic start score comprises:

updating, at each timestep, a hidden state and a cell state of a long short-term memory based on an input of the feature representations;

applying max pooling on a set of hidden states between a current timestep and a past timestep to obtain an average hidden state;

generating the per-frame action scores based on the hidden state and a first vector of classifier parameters; and

generating the class-agnostic score based on the average hidden state and a second vector of classifier parameters.

18. The system of claim 17 , wherein the processor is further configured to:

convert the class-wise temporal proposals to per-frame action labels and binary start labels;

compute a frame loss using a cross entropy loss between the per-frame action labels and the generated per-frame action scores; and

compute a start loss using a focal loss between the binary start labels and the generated class-agnostic score.

19. The system of claim 18 , wherein the loss metric is computed by a weighted sum of an online action recognizer loss and a temporal proposal generator loss,

wherein the online action recognizer loss includes a sum of the frame loss and the start loss, and

wherein the temporal proposal generator loss includes a multiple instance learning loss and a pair-wise co-activity similarity loss.

20. A non-transitory processor-readable storage medium storing a feature extractor neural network model, a temporal proposal generator (TPG) neural network model and the OAD neural network model and processor-executable instructions for training the OAD neural network model using a training dataset of untrimmed videos having video-level labels without annotated labels indicating whether a specific video frame contains an action start of a specific action class, the processor-executable instructions executable by a processor to perform:

receiving, by a communication interface, an input of the training dataset of untrimmed videos including a set of video-level labels indicating one or more action classes that emerge in the untrimmed videos, wherein the untrimmed videos are used for training the OAD neural network model with no annotated label indicating action starts of the one or more action classes;

generating, by a feature extractor neural network model implemented on one or more hardware processors, feature representations from the training dataset of the untrimmed videos;

generating, by a temporal proposal generator (TPG) neural network model implemented on the one or more hardware processors and receiving the feature representations from the feature extractor neural network, class-wise temporal proposals indicating a respective estimated action start for each action class of the one or more action classes based on the feature representations and the set of video-level labels;

generating, by an online action recognizer (OAR) neural network model implemented on the one or more hardware processors and receiving the feature representations from the feature extractor neural network and the class-wise temporal proposals from the TPG neural network model, per-frame action scores over action classes indicating whether each frame of the untrimmed video contains each specific action class, and a class-agnostic start score indicating whether the respective frame contains a start of any action based on the feature representations and the class-wise temporal proposals;

training the OAD neural network model comprising the TPG neural network model and the OAR neural network model according to a loss metric computed based on the per-frame action scores, the class-agnostic start scores and based on the class-wise temporal proposals as pseudo ground-truth labels; and

generating, by the trained OAD neural network model, a predicted action start for an input real-time video stream.

Assignments (2)
CHANGE OF NAME Recorded Aug 4, 2026
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 076118/0548 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 16, 2020
From: GAO, MINGFEI; ZHOU, YINGBO; XU, RAN; XIONG, CAIMING
To: SALESFORCE.COM, INC.
Reel/Frame 053233/0440 →
Continuity (2)
Provisional Application 63023402 · May 12, 2020
Related Publication 20210357687A1 · Nov 18, 2021
References Cited (111)
US 9432671B2 · Campanelli et al. · 2016 [cited by applicant]
US 10282663B2 · Socher et al. · 2019 [cited by applicant]
US 10474709B2 · Paulus · 2019 [cited by applicant]
US 10521465B2 · Paulus · 2019 [cited by applicant]
US 10542270B2 · Zhou et al. · 2020 [cited by applicant]
US 10546217B2 · Albright et al. · 2020 [cited by applicant]
US 10558750B2 · Lu et al. · 2020 [cited by applicant]
US 10565305B2 · Lu et al. · 2020 [cited by applicant]
US 10565306B2 · Lu et al. · 2020 [cited by applicant]
US 10565318B2 · Bradbury · 2020 [cited by applicant]
US 10565493B2 · Merity et al. · 2020 [cited by applicant]
US 10573295B2 · Zhou et al. · 2020 [cited by applicant]
US 10592767B2 · Trott et al. · 2020 [cited by applicant]
US 10699060B2 · McCann · 2020 [cited by applicant]
US 10747761B2 · Zhong et al. · 2020 [cited by applicant]
US 10776581B2 · McCann et al. · 2020 [cited by applicant]
US 10783875B2 · Hosseini-Asl et al. · 2020 [cited by applicant]
US 20160234464A1 · Loce · 2016 [cited by examiner]
US 20160350653A1 · Socher et al. · 2016 [cited by applicant]
US 20170024645A1 · Socher et al. · 2017 [cited by applicant]
US 20170032280A1 · Socher · 2017 [cited by applicant]
US 20170140240A1 · Socher et al. · 2017 [cited by applicant]
US 20180096219A1 · Socher · 2018 [cited by applicant]
US 20180121787A1 · Hashimoto et al. · 2018 [cited by applicant]
US 20180121788A1 · Hashimoto et al. · 2018 [cited by applicant]
US 20180121799A1 · Hashimoto et al. · 2018 [cited by applicant]
US 20180129931A1 · Bradbury et al. · 2018 [cited by applicant]
US 20180129937A1 · Bradbury et al. · 2018 [cited by applicant]
US 20180129938A1 · Xiong et al. · 2018 [cited by applicant]
US 20180268287A1 · Johansen et al. · 2018 [cited by applicant]
US 20180268298A1 · Johansen et al. · 2018 [cited by applicant]
US 20180336453A1 · Merity et al. · 2018 [cited by applicant]
US 20180373682A1 · McCann et al. · 2018 [cited by applicant]
US 20180373987A1 · Zhang et al. · 2018 [cited by applicant]
US 20190108400A1 · Escorcia · 2019 [cited by examiner]
US 20190130248A1 · Zhong et al. · 2019 [cited by applicant]
US 20190130249A1 · Bradbury et al. · 2019 [cited by applicant]
US 20190130273A1 · Keskar et al. · 2019 [cited by applicant]
US 20190130312A1 · Xiong et al. · 2019 [cited by applicant]
US 20190130896A1 · Zhou et al. · 2019 [cited by applicant]
US 20190188568A1 · Keskar et al. · 2019 [cited by applicant]
US 20190213482A1 · Socher et al. · 2019 [cited by applicant]
US 20190251431A1 · Keskar et al. · 2019 [cited by applicant]
US 20190258714A1 · Zhong et al. · 2019 [cited by applicant]
US 20190258939A1 · Min et al. · 2019 [cited by applicant]
US 20190286073A1 · Asl et al. · 2019 [cited by applicant]
US 20190295530A1 · Hosseini-Asl et al. · 2019 [cited by applicant]
US 20190355270A1 · McCann et al. · 2019 [cited by applicant]
US 20190362020A1 · Paulus et al. · 2019 [cited by applicant]
US 20200005765A1 · Zhou et al. · 2020 [cited by applicant]
US 20200057805A1 · Lu et al. · 2020 [cited by applicant]
US 20200065651A1 · Merity et al. · 2020 [cited by applicant]
US 20200084465A1 · Zhou et al. · 2020 [cited by applicant]
US 20200089757A1 · Machado et al. · 2020 [cited by applicant]
US 20200090033A1 · Ramachandran et al. · 2020 [cited by applicant]
US 20200090034A1 · Ramachandran et al. · 2020 [cited by applicant]
US 20200103911A1 · Ma et al. · 2020 [cited by applicant]
US 20200104643A1 · Hu et al. · 2020 [cited by applicant]
US 20200104699A1 · Zhou et al. · 2020 [cited by applicant]
US 20200105272A1 · Wu et al. · 2020 [cited by applicant]
US 20200117854A1 · Lu et al. · 2020 [cited by applicant]
US 20200117861A1 · Bradbury · 2020 [cited by applicant]
US 20200142917A1 · Paulus · 2020 [cited by applicant]
US 20200175305A1 · Trott et al. · 2020 [cited by applicant]
US 20200184020A1 · Hashimoto et al. · 2020 [cited by applicant]
US 20200234113A1 · Liu · 2020 [cited by applicant]
US 20200272940A1 · Sun et al. · 2020 [cited by applicant]
US 20200285704A1 · Rajani et al. · 2020 [cited by applicant]
US 20200285705A1 · Zheng et al. · 2020 [cited by applicant]
US 20200285706A1 · Singh et al. · 2020 [cited by applicant]
US 20200285993A1 · Liu et al. · 2020 [cited by applicant]
US 20200302178A1 · Gao et al. · 2020 [cited by applicant]
US 20200302236A1 · Gao et al. · 2020 [cited by applicant]
Shou, Zheng, et al. “Online action detection in untrimmed, streaming videos-modeling and evaluation.” ECCV. vol. 1. No. 2. 2018. (Year: 2018). [cited by examiner]
Xu, Mingze, et al. “Temporal recurrent networks for online action detection.” Proceedings of the IEEE/CVF international conference on computer vision. 2019. (Year: 2019). [cited by examiner]
Song, Xiaolin, et al. “Temporal-spatial mapping for action recognition.” IEEE Transactions on Circuits and Systems for Video Technology 30.3 (2019): 748-759. (Year: 2019). [cited by examiner]
Yoshikawa, Taizo, Viktor Losing, and Emel Demircan. “Machine learning for human movement understanding.” Advanced Robotics 34.13 (2020): 828-844. (Year: 2020). [cited by examiner]
Parsa, Behnoosh, et al. “Predicting ergonomic risks during indoor object manipulation using spatiotemporal convolutional networks.” CORR. 2019. (Year: 2019). [cited by examiner]
Tang, Fengxiao, et al. “On extracting the spatial-temporal features of network traffic patterns: A tensor based deep learning model.” 2018 International Conference on Network Infrastructure and Digital Content (IC-NIDC)… [cited by examiner]
Eliáš, RNDr Petr. “Action Recognition, Annotation, and Searching in Motion Data.” (Year: 2019). [cited by examiner]
Liu, Jiaying, et al. “Multi-modality multi-task recurrent neural network for online action detection.” IEEE Transactions on Circuits and Systems for Video Technology 29.9 (2018): 2667-2682. (Year: 2018). [cited by examiner]
Mahmud, Tahmida, Mahmudul Hasan, and Amit K. Roy-Chowdhury. “Joint prediction of activity labels and starting times in untrimmed videos.” Proceedings of the IEEE International conference on Computer Vision. 2017. (Year:… [cited by examiner]
P. Bojanowski, et al. Weakly supervised action labeling in videos under ordering constraints. In European Conference on Computer Vision, pp. 628-643. Springer, 2014. [cited by applicant]
S. Buch, V. Escorcia, C. Shen, B. Ghanem, and J. C. Niebles. SST: Single-stream temporal action proposals. In CVPR, 2017. [cited by applicant]
J. Carreira and A. Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, 2017. [cited by applicant]
Y.-W. Chao, S. Vijayanarasimhan, B. Seybold, D. A. Ross, J. Deng, and R. Sukthankar. Rethinking the faster r-cnn architecture for temporal action localization. In CVPR, 2018. [cited by applicant]
X. Dai, B. Singh, G. Zhang, L. S. Davis, and Y. Q. Chen. Temporal context network for activity localization in videos. In ICCV, 2017. [cited by applicant]
R. De Geest, E. Gavves, A. Ghodrati, Z. Li, C. Snoek, and T. Tuytelaars. Online action detection. In ECCV, 2016.; arxiv:1604.06506v2 [cs.CV] Aug. 30, 2016. [cited by applicant]
O. Duchenne, I. Laptev, J. Sivic, F. R. Bach, and J. Ponce. Automatic annotation of human actions in video. In ICCV, vol. 1, pp. 3-2, 2009. [cited by applicant]
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR, 2015. [cited by applicant]
J. Gao, Z. Yang, and R. Nevatia. RED: Reinforced encoder-decoder networks for action anticipation. In BMVC, 2017. [cited by applicant]
J. Gao, Z. Yang, C. Sun, K. Chen, and R. Nevatia. TURN TAP: Temporal unit regression network for temporal action proposals. ICCV, 2017. [cited by applicant]
M. Gao, M. Xu, L. S. Davis, R. Socher, and C. Xiong. Startnet: Online detection of action start in untrimmed videos. In IEEE International Conference on Computer Vision (ICCV), 2019. [cited by applicant]
D.-A. Huang, L. Fei-Fei, and J. C. Niebles. Connectionist temporal modeling for weakly supervised action labeling. In European Conference on Computer Vision, pp. 137-153. Springer, 2016. [cited by applicant]
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar. THUMOS challenge: Action recognition with a large No. of classes. http://crcv.ucf.edu/THUMOS14/, 2014. [cited by applicant]
D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv:1412.6980, 2014. [cited by applicant]
T. N. Kipf and M. Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016. [cited by applicant]
P. Lee, Y. Uh, and H. Byun. Background suppression network for weakly-supervised temporal action localization. In AAAI, 2020. [cited by applicant]
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, pp. 2980-2988, 2017. [cited by applicant]
D. Liu, T. Jiang, and Y. Wang. Completeness modeling and context separation for weakly supervised temporal action localization. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019. [cited by applicant]
S. Narayan, H. Cholakkal, F. Shahbaz Khan, and L. Shao. 3c-net: Category count and center loss for weakly-supervised action localization. arXiv preprint arXiv:1908.08216, 2019. [cited by applicant]
S. Paul, S. Roy, and A. K. Roy-Chowdhury. W-talc: Weakly-supervised temporal activity localization and classification. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 563-579, 2018. [cited by applicant]
Z. Shou, H. Gao, L. Zhang, K. Miyazawa, and S.-F. Chang. Autoloc: Weakly-supervised temporal action localization in untrimmed videos. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 154-171, 201… [cited by applicant]
Z. Shou, J. Pan, J. Chan, K. Miyazawa, H. Mansour, A. Vetro, X. Giro-i Nieto, and S.-F. Chang. Online action detection in untrimmed, streaming videos-modeling and evaluation. In ECCV, 2018. [cited by applicant]
Z. Shou, D. Wang, and S.-F. Chang. Temporal action localization in untrimmed videos via multi-stage cnns. In CVPR, 2016. [cited by applicant]
L. Wang, Y. Xiong, D. Lin, and L. Van Gool. Untrimmednets for weakly supervised action recognition and detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4325-4334, 2017. [cited by applicant]
H. Xu, A. Das, and K. Saenko. R-C3D: Region convolutional 3d network for temporal activity detection. In ICCV, 2017. [cited by applicant]
M. Xu, M. Gao, Y.-T. Chen, L. S. Davis, and D. J. Crandall. Temporal recurrent networks for online action detection. In IEEE International Conference on Computer Vision (ICCV), 2019. [cited by applicant]
Y. Yuan, Y. Lyu, X. Shen, I. W. Tsang, and D.-Y. Yeung. Marginalized average attentional network for weakly-supervised learning. In International Conference on Learning Representations (ICLR), 2019. [cited by applicant]
R. Zeng, W. Huang, M. Tan, Y. Rong, P. Zhao, J. Huang, and C. Gan. Graph convolutional networks for temporal action localization. In Proceedings of the IEEE International Conference on Computer Vision, pp. 7094-7103, 20… [cited by applicant]
Y. Zhao, Y. Xiong, L. Wang, Z. Wu, X. Tang, and D. Lin. Temporal action detection with structured segment networks. In ICCV, 2017. [cited by applicant]
Cited By (1)
US 12,548,335