IP Library › Granted Patent US 12,548,335
Granted Patent B2
US 12,548,335 · App. 18/308,542 · Granted Feb 10, 2026

Weakly supervised action segmentation

Inventors: Reza Ghoddoosian (San Jose, CA); Isht Dwivedi (Mountain View, CA); Nakul Agarwal (San Francisco, CA); Behzad Dariush (San Ramon, CA)
Assignee: Honda Motor Co., Ltd.
G06V20/49G06V10/774G06V10/82G06V20/41G06V20/44G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,335
App. No.
18/308,542
Granted
Feb 10, 2026
Kind
B2
Abstract

According to one aspect, weakly-supervised action segmentation may include performing feature extraction to extract one or more features associated with a current frame of a video including a series of one or more actions, feeding one or more of the features to a recognition network to generate a predicted action score for the current frame of the video, feeding one or more of the features and the predicted action score to an action transition model to generate a potential subsequent action, feeding the potential subsequent action and the predicted action score to a hybrid segmentation model to generate a predicted sequence of actions from a first frame of the video to the current frame of the video, and segmenting or labeling one or more frames of the video based on the predicted sequence of actions from the first frame of the video to the current frame of the video.

Claims (33)

1 . A system for weakly-supervised action segmentation, comprising:

a memory storing one or more instructions; and

a processor executing one or more of the instructions stored on the memory to perform:

performing feature extraction to extract one or more features associated with a current frame and a previous frame of a video including a series of one or more actions;

feeding the one or more of features to a recognition network to generate a predicted action score for the current frame of the video;

feeding the one or more of features and the predicted action score to an action transition model to generate a potential subsequent action; and

feeding a segmentation result from the previous frame, the potential subsequent action, and the predicted action score to a hybrid segmentation model to generate a predicted sequence of actions from a first frame of the video to the current frame of the video and a segmentation result for the current frame.

2 . The system for weakly-supervised action segmentation of claim 1 , wherein the hybrid segmentation model generates the predicted sequence of actions based on a predicted action length for a predicted action associated with the predicted action score.

3 . The system for weakly-supervised action segmentation of claim 1 , wherein the action transition model generates the potential subsequent action based on a transcript of one or more known sequences of actions, the one or more of features, and the predicted action score.

4 . The system for weakly-supervised action segmentation of claim 1 , wherein the hybrid segmentation model generates a predicted sequence of action lengths corresponding to the predicted sequence of actions.

5 . The system for weakly-supervised action segmentation of claim 4 , wherein the processor detects one or more errors associated with the predicted sequence of action lengths and the predicted sequence of actions based on an error function.

6 . The system for weakly-supervised action segmentation of claim 1 , wherein the hybrid segmentation model is based on an unconstrained Viterbi algorithm.

7 . A computer-implemented method for weakly-supervised action segmentation, comprising:

performing feature extraction to extract one or more features associated with a current frame and a previous frame of a video including a series of one or more actions;

feeding the one or more of features to a recognition network to generate a predicted action score for the current frame of the video;

feeding the one or more of features and the predicted action score to an action transition model to generate a potential subsequent action; and

feeding a segmentation result from the previous frame, the potential subsequent action, and the predicted action score to a hybrid segmentation model to generate a predicted sequence of actions from a first frame of the video to the current frame of the video and a segmentation result for the current frame.

8 . The computer-implemented method for weakly-supervised action segmentation of claim 7 , wherein the hybrid segmentation model generates the predicted sequence of actions based on a predicted action length for a predicted action associated with the predicted action score.

9 . The computer-implemented method for weakly-supervised action segmentation of claim 7 , wherein the action transition model generates the potential subsequent action based on a transcript of one or more known sequences of actions, the one or more of features, and the predicted action score.

10 . The computer-implemented method for weakly-supervised action segmentation of claim 7 , wherein the hybrid segmentation model generates a predicted sequence of action lengths corresponding to the predicted sequence of actions.

11 . The computer-implemented method for weakly-supervised action segmentation of claim 10 , comprising detecting one or more errors associated with the predicted sequence of action lengths and the predicted sequence of actions based on an error function.

12 . The computer-implemented method for weakly-supervised action segmentation of claim 7 , wherein the hybrid segmentation model is based on an unconstrained Viterbi algorithm.

13 . A system for weakly-supervised action segmentation, comprising:

a memory storing one or more instructions; and

a processor executing one or more of the instructions stored on the memory to perform:

performing feature extraction to extract one or more features associated with a current frame and a previous frame of a video including a series of one or more actions;

feeding the one or more of features to a recognition network to generate a predicted action score for the current frame of the video;

feeding the one or more of features and the predicted action score to an action transition model to generate a potential subsequent action;

feeding a segmentation result from the previous frame, the potential subsequent action, and the predicted action score to a hybrid segmentation model to generate a predicted sequence of actions from a first frame of the video to the current frame of the video and a segmentation result for the current frame; and

segmenting or labeling one or more frames of the video based on the predicted sequence of actions from the first frame of the video to the current frame of the video.

14 . The system for weakly-supervised action segmentation of claim 13 , wherein the hybrid segmentation model generates the predicted sequence of actions based on a predicted action length for a predicted action associated with the predicted action score.

15 . The system for weakly-supervised action segmentation of claim 13 , wherein the action transition model generates the potential subsequent action based on a transcript of one or more known sequences of actions, the one or more of features, and the predicted action score.

16 . The system for weakly-supervised action segmentation of claim 13 , wherein the hybrid segmentation model generates a predicted sequence of action lengths corresponding to the predicted sequence of actions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2023
From: GHODDOOSIAN, REZA; DWIVEDI, ISHT; AGARWAL, NAKUL; DARIUSH, BEHZAD
To: HONDA MOTOR CO., LTD.
Reel/Frame 063468/0783 →
Continuity (1)
Related Publication 20240371166A1 · Nov 7, 2024
References Cited (147)
US 6965861B1 · Dailey · 2005 [cited by examiner]
US 11017556B2 · Yang · 2021 [cited by examiner]
US 11055516B2 · Zhu · 2021 [cited by examiner]
US 11260872B2 · Chen · 2022 [cited by examiner]
US 11538564B2 · Yao et al. · 2022 [cited by applicant]
US 11600069B2 · Lin · 2023 [cited by examiner]
US 11626195B2 · Lyman et al. · 2023 [cited by applicant]
US 11636681B2 · Wang · 2023 [cited by examiner]
US 11669792B2 · Lyman et al. · 2023 [cited by applicant]
US 11679500B2 · Chen · 2023 [cited by examiner]
US 11748677B2 · Prosky et al. · 2023 [cited by applicant]
US 11790297B2 · Lyman et al. · 2023 [cited by applicant]
US 11790655B2 · Yuan · 2023 [cited by examiner]
US 11823106B2 · Lyman et al. · 2023 [cited by applicant]
US 11887354B2 · Zhang et al. · 2024 [cited by applicant]
US 11971884B2 · Aggarwal · 2024 [cited by examiner]
US 12182974B2 · Nakamura · 2024 [cited by examiner]
US 12299982B2 · Gao · 2025 [cited by examiner]
US 20060161814A1 · Wocke et al. · 2006 [cited by applicant]
US 20170140285A1 · Dotan-Cohen · 2017 [cited by examiner]
US 20190102908A1 · Yang · 2019 [cited by examiner]
US 20190205629A1 · Zhu · 2019 [cited by examiner]
US 20200114924A1 · Chen · 2020 [cited by examiner]
US 20200160064A1 · Wang · 2020 [cited by examiner]
US 20200160176A1 · Mehrasa · 2020 [cited by examiner]
US 20200160974A1 · Yao et al. · 2020 [cited by applicant]
US 20200272823A1 · Liu · 2020 [cited by examiner]
US 20210201043A1 · Yuan · 2021 [cited by examiner]
US 20210216782A1 · Lin · 2021 [cited by examiner]
US 20210357687A1 · Gao · 2021 [cited by examiner]
US 20220180622A1 · Zhang et al. · 2022 [cited by applicant]
US 20220184806A1 · Chen · 2022 [cited by examiner]
US 20220215915A1 · Lyman et al. · 2022 [cited by applicant]
US 20220245141A1 · Aggarwal · 2022 [cited by examiner]
US 20220261967A1 · Nakamura · 2022 [cited by examiner]
US 20220318555A1 · Ben-Ari · 2022 [cited by examiner]
US 20220327834A1 · Liu · 2022 [cited by examiner]
US 20230274580A1 · Yao · 2023 [cited by examiner]
US 20230386203A1 · Girdhar · 2023 [cited by examiner]
US 20240292073A1 · Khalil · 2024 [cited by examiner]
CN 110852256A · 2020 [cited by examiner]
CN 111259775A · 2020 [cited by examiner]
CN 115471771A · 2022 [cited by examiner]
CN 115588230A · 2023 [cited by examiner]
Ng et al., “Forecasting future action sequences with attention: a new approach to weakly supervised action forecasting.” IEEE Transactions on Image Processing 29 (2020): 8880-8891. (Year: 2020). [cited by examiner]
Elmi et al., “Deep-Cogn: skeleton-based human action recognition for cognitive behavior assessment.” In 2022 IEEE 34th International Conference on Tools with Artificial Intelligence (ICTAI), pp. 692-699. IEEE, 2022. (Ye… [cited by examiner]
Ren et al., “Diffnet: Discriminative feature fusion network of multisurface skeleton project images for action recognition.” In Proceedings of the 2021 13th International Conference on Machine Learning and Computing, pp… [cited by examiner]
Ren et al., “CAA: Candidate-Aware Aggregation for Temporal Action Detection.” In Proceedings of the 29th ACM International Conference on Multimedia, pp. 4930-4938. 2021. (Year: 2021). [cited by examiner]
Burges et al., “Shortest path segmentation: A method for training a neural network to recognize character strings.” In Proc. Int. Joint Conf. Neural Networks, vol. 3, pp. 165-172. 1992. (Year: 1992). [cited by examiner]
Kuehne et al., “A Hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation.” IEEE Transactions on Pattern Analysis & Machine Intelligence 42, No. 04 (2020): 765-779. (Year: 2020). [cited by examiner]
Shen et al., “Semi-Weakly-Supervised Learning of Complex Actions from Instructional Task Videos,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 2022, pp. 3334-3344 (Ye… [cited by examiner]
Zou et al., “A Temporal Convolutional Network for Weakly Supervised Action Segmentation,” 2021 7th IEEE International Conference on Network Intelligence and Digital Content (IC-NIDC), Beijing, China, 2021, pp. 359-363 (… [cited by examiner]
Kuehne et al., “An end-to-end generative framework for video segmentation and recognition.” arXiv preprint arXiv:1509.01947 (2016). (Year: 2016). [cited by examiner]
CN-110852256-A (machine translation) (Year: 2020). [cited by examiner]
CN-111259775-A (machine translation) (Year: 2020). [cited by examiner]
CN-115471771-A (machine translation) (Year: 2022). [cited by examiner]
CN-115588230-A (machine translation) (Year: 2023). [cited by examiner]
Ghoddoosian et al., “Weakly-Supervised Action Segmentation and Unseen Error Detection in Anomalous Instructional Videos,” 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2023, pp. 10094-… [cited by examiner]
Notice of Allowance of U.S. Appl. No. 17/590,379 dated Aug. 7, 2024, 33 pages. [cited by applicant]
Andra Acsintoae, Andrei Florescu, Mariana-Iuliana Georgescu, Tudor Mare, Paul Sumedrea, Radu Tudor Ionescu, Fahad Shahbaz Khan, and Mubarak Shah. Ubnormal: New benchmark for supervised open-set video anomaly detection. … [cited by applicant]
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien. Unsupervised learning from narrated instruction videos. In Proceedings of the IEEE Conference on Computer Vis… [cited by applicant]
Marcella Astrid, Muhammad Zaigham Zaheer, Jae-Yeong Lee, and Seung-Ik Lee. Learning not to reconstruct anomalies. arXiv preprint arXiv:2110.09742, 2021. [cited by applicant]
Nadine Behrmann, S Alireza Golestaneh, Zico Kolter, Jürgen Gall, and Mehdi Noroozi. Unified fully and timestamp supervised temporal action segmentation via sequence to sequence translation. In Computer Vision—ECCV 2022:… [cited by applicant]
Yizhak Ben-Shabat, Xin Yu, Fatemeh Saleh, Dylan Campbell, Cristian Rodriguez-Opazo, Hongdong Li, and Stephen Gould. The ikea asm dataset: Understanding people assembling furniture through actions, objects and pose. In P… [cited by applicant]
Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6299-6308, 2017. [cited by applicant]
Chien-Yi Chang, De-An Huang, Yanan Sui, Li Fei-Fei, and Juan Carlos Niebles. D3tw: Discriminative differentiable dynamic time warping for weakly supervised action alignment and segmentation. In Proceedings of the IEEE/C… [cited by applicant]
Xiaobin Chang, Frederick Tung, and Greg Mori. Learning discriminative prototypes with dynamic time warping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8395-8404, 2021. [cited by applicant]
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision: Collection, pipeline an… [cited by applicant]
Hanqiu Deng, Zhaoxiang Zhang, Shihao Zou, and Xingyu Li. Bi-directional frame interpolation for unsupervised video anomaly detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, … [cited by applicant]
Ehsan Elhamifar and Zwe Naing. Unsupervised procedure learning via joint dynamic summarization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6341-6350, 2019. [cited by applicant]
Alireza Fathi, Xiaofeng Ren, and James M Rehg. Learning to recognize objects in egocentric activities. In CVPR 2011, pp. 3281-3288. IEEE, 2011. [cited by applicant]
Mingfei Gao, Yingbo Zhou, Ran Xu, Richard Socher, and Caiming Xiong. Woad: Weakly supervised online action detection in untrimmed videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit… [cited by applicant]
Reza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Chiho Choi, and Behzad Dariush. Weakly-supervised online action segmentation in multi-view instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Visio… [cited by applicant]
Reza Ghoddoosian, Saif Sayed, and Vassilis Athitsos. Hierarchical modeling for task recognition and action segmentation in weakly-labeled instructional videos. In Proceedings of the IEEE/CVF Winter Conference on Applica… [cited by applicant]
Hilde Kuehne, Ali Arslan, and Thomas Serre. The language of actions: Recovering the syntax and semantics of goaldirected human activities. In Proceedings of the IEEE conference on computer vision and pattern recognition… [cited by applicant]
Hilde Kuehne, Alexander Richard, and Juergen Gall. Weakly supervised learning of actions from transcripts. Computer Vision and Image Understanding, 163:78-89, 2017. [cited by applicant]
Anna Kukleva, Hilde Kuehne, Fadime Sener, and Jurgen Gall. Unsupervised learning of action classes with continuous temporal embedding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition… [cited by applicant]
Dongha Lee, Sehun Yu, Hyunjun Ju, and Hwanjo Yu. Weakly supervised temporal anomaly segmentation with dynamic time warping. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7355-7364, 2021. [cited by applicant]
Sangmin Lee, Hak Gu Kim, and Yong Man Ro. Bman: Bidirectional multi-scale aggregation networks for abnormal event detection. IEEE Transactions on Image Processing, 29:2395-2408, 2019. [cited by applicant]
Jun Li, Peng Lei, and Sinisa Todorovic. Weakly supervised energy-based learning for action segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6243-6251, 2019. [cited by applicant]
Yin Li, Miao Liu, and JamesMRehg. In the eye of beholder: Joint learning of gaze and actions in first person video. In Proceedings of the European conference on computer vision (ECCV), pp. 619-635, 2018. [cited by applicant]
Zijia Lu and Ehsan Elhamifar. Weakly-supervised action segmentation and alignment via transcript-aware union-ofsubspaces learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8085-809… [cited by applicant]
Didik Purwanto, Yie-Tarng Chen, and Wen-Hsien Fang. Dance with self-attention: A new look of conditional random fields on anomaly detection in videos. In Proceedings of the IEEE/CVF International Conference on Computer … [cited by applicant]
Yicheng Qian, Weixin Luo, Dongze Lian, Xu Tang, Peilin Zhao, and Shenghua Gao. Svip: Sequence verification for procedures in videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, … [cited by applicant]
Alexander Richard, Hilde Kuehne, and Juergen Gall. Weakly supervised action learning with mn based fine-to-coarse modeling. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 754-763, … [cited by applicant]
Alexander Richard, Hilde Kuehne, Ahsan Iqbal, and Juergen Gall. Neuralnetwork-viterbi: A framework for weakly supervised video learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, … [cited by applicant]
Nicolae-Cǎtǎlin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B Moeslund, and Mubarak Shah. Self-supervised predictive convolutional attentive block for anomaly detection. In Proc… [cited by applicant]
Marcus Rohrbach, Anna Rohrbach, Michaela Regneri, Sikandar Amin, Mykhaylo Andriluka, Manfred Pinkal, and Bernt Schiele. Recognizing fine-grained and composite activities using hand-centric features and script data. Inte… [cited by applicant]
Yaser Souri, Yazan Abu Farha, Fabien Despinoy, Gianpiero Francesca, and Juergen Gall. Fifa: Fast inference approximation for action segmentation. In DAGM German Conference on Pattern Recognition, pp. 282-296. Springer, … [cited by applicant]
Yaser Souri, Mohsen Fayyaz, Luca Minciullo, Gianpiero Francesca, and Juergen Gall. Fast weakly supervised action segmentation using mutual consistency. IEEE Transactions on Pattern Analysis and Machine Intelligence, 202… [cited by applicant]
Sebastian Stein and Stephen J McKenna. Combining embedded accelerometers with computer vision for recognizing food preparation activities. In Proceedings of the 2013 ACM international joint conference on Pervasive and u… [cited by applicant]
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE Conference on Co… [cited by applicant]
Kamalakar Vijay Thakare, Yash Raghuwanshi, Debi Prosad Dogra, Heeseung Choi, and Ig-Jae Kim. Dyannet: A scene dynamicity guided self-trained video anomaly detection network. In Proceedings of the IEEE/CVF Winter Confere… [cited by applicant]
Xiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao, Zhengrong Zuo, Changxin Gao, and Nong Sang. Oadtr: Online action detection with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Visio… [cited by applicant]
Jhih-Ciangwu, He-Yen Hsieh, Ding-Jie Chen, Chiou-Shann Fuh, and Tyng-Luh Liu. Self-supervised sparse representation for video anomaly detection. In Computer Vision—ECCV 2022: 17th European Conference, Tel Aviv, Israel, … [cited by applicant]
Guang Yu, Siqi Wang, Zhiping Cai, Xinwang Liu, Chuanfu Xu, and Chengkun Wu. Deep anomaly discovery from unlabeled videos via normality advantage and self-paced refinement. In Proceedings of the IEEE/CVF Conference on Co… [cited by applicant]
Christopher Zach, Thomas Pock, and Horst Bischof. A duality based approach for realtime tv-I 1 optical flow. In Joint pattern recognition symposium, pp. 214-223. Springer, 2007. [cited by applicant]
Astrid, and Seung-Ik Lee. Claws: Clustering assisted weakly supervised learning with normalcy suppression for anomalous event detection. In Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2… [cited by applicant]
M Zaigham Zaheer, Arif Mahmood, M Haris Khan, Mattia Segu, Fisher Yu, and Seung-Ik Lee. Generative cooperative learning for unsupervised video anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vis… [cited by applicant]
Yuansheng Zhu, Wentao Bao, and Qi Yu. Towards open set video anomaly detection. In Computer Vision—ECCV 2022: 17th European Conference, Tel Aviv, Israel, Oct. 23-27, 2022, Proceedings, Part XXXIV, pp. 395-412. Springer,… [cited by applicant]
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. Crosstask weakly supervised learning from instructional videos. In Proceedings of the IEEE/CVF Conference on Com… [cited by applicant]
Yazan Abu Farha and Juergen Gall. Uncertainty-aware anticipation of activities. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pp. 0-0, 2019. [cited by applicant]
Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli, Li Fei-Fei, and Juan Carlos Niebles. Procedure planning in Instructional videos. arXiv preprint arXiv:1907.01172, 2019. [cited by applicant]
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555, 2014. [cited by applicant]
K Deepak, G Srivathsan, S Roshan, and S Chandrakala. Deep multi-view representation learning for video anomaly detection using spatiotemporal autoencoders. Circuits, Systems, and Signal Processing, 40(3):1333-1349, 2021. [cited by applicant]
Li Ding and Chenliang Xu. Weakly-supervised action segmentation with iterative soft boundary assignment. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6508-6516, 2018. [cited by applicant]
Hyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung, and Changick Kim. Learning to discriminate information for online action detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Chenyou Fan, Jangwon Lee, Mingze Xu, Krishna Kumar Singh, Yong Jae Lee, David J Crandall, and Michael S Ryoo. Identifying first-person camera wearers in thirdperson videos. In Proceedings of the IEEE Conference on Compu… [cited by applicant]
Yazan Abu Farha and Jurgen Gall. Ms-tcn: Multi-stage temporal convolutional network for action segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3575-3584, 2019. [cited by applicant]
Antonino Furnari and Giovanni Maria Farinella. What would you expect? anticipating egocentric actions with rollingunrolling Istms and modality attention. In Proceedings of the IEEE/CVF International Conference on Comput… [cited by applicant]
Jiyang Gao, Zhenheng Yang, and Ram Nevatia. Red: Reinforced encoder-decoder networks for action anticipation. arXiv preprint arXiv:1707.04818, 2017. [cited by applicant]
Mingfei Gao, Mingze Xu, Larry S Davis, Richard Socher, and Caiming Xiong. Startnet: Online detection of action start in untrimmed videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5… [cited by applicant]
Shang-Hua Gao, Qi Han, Zhong-Yu Li, Pai Peng, Liang Wang, and Ming-Ming Cheng. Global2local: Efficient structure search for video action segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat… [cited by applicant]
Reza Ghoddoosian, Saif Sayed, and Vassilis Athitsos. Action duration prediction for segment-level alignment of weaklylabeled videos. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p… [cited by applicant]
Sanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram N Syed, Andrey Konin, Zeeshan Zia, and Quoc-Huy Tran. Learning by aligning videos in time. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R… [cited by applicant]
Hsuan-I Ho, Wei-Chen Chiu, and Yu-Chiang Frank Wang. Summarizing first-person videos from third persons' points of view. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 70-85, 2018. [cited by applicant]
Yuchi Ishikawa, Seito Kasai, Yoshimitsu Aoki, and Hirokatsu Kataoka. Alleviating over-segmentation errors by detecting action boundaries. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Visi… [cited by applicant]
Qiuhong Ke, Mario Fritz, and Bernt Schiele. Timeconditioned action anticipation in one shot. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9925-9934, 2019. [cited by applicant]
Sateesh Kumar, Sanjay Haresh, Awais Ahmed, Andrey Konin, M Zeeshan Zia, and Quoc-Huy Tran. Unsupervised activity segmentation by joint representation learning and online clustering. 2021. [cited by applicant]
Yaman Kumar, Mayank Aggarwal, Pratham Nawal, Shin'ichi Satoh, Rajiv Ratn Shah, and Roger Zimmermann. Harnessing ai for speech reconstruction using multi-view silent video feed. In Proceedings of the 26th ACM internation… [cited by applicant]
Zhe Li, Yazan Abu Farha, and Jurgen Gall. Temporal action segmentation from timestamp supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8365-8374, 2021. [cited by applicant]
Yunyu Liu, Lichen Wang, Yue Bai, Can Qin, Zhengming Ding, and Yun Fu. Generative view-correlation adaptation for semi-supervised multi-view learning. In European Conference on Computer Vision, pp. 318-334. Springer, 202… [cited by applicant]
Tahmida Mahmud, Mahmudul Hasan, and Amit K Roy-Chowdhury. Joint prediction of activity labels and starting times in untrimmed videos. In Proceedings of the IEEE International conference on Computer Vision, pp. 5773-5782… [cited by applicant]
Jingjing Meng, Suchen Wang, Hongxing Wang, Junsong Yuan, and Yap-Peng Tan. Video summarization via Multiview representative selection. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pp… [cited by applicant]
Rameswar Panda and Amit K Roy-Chowdhury. Multi-view surveillance video summarization via joint embedding and sparse optimization. IEEE Transactions on Multimedia, 19(9):2010-2021, 2017. [cited by applicant]
Paritosh Parmar and Brendan Tran Morris. What and how well you performed? a multitask learning approach to action quality assessment. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.… [cited by applicant]
Florent Perronnin and Christopher Dance. Fisher kernels on visual vocabularies for image categorization. In 2007 IEEE conference on computer vision and pattern recognition, pp. 1-8. IEEE, 2007. [cited by applicant]
AJ Piergiovanni and Michael S Ryoo. Recognizing actions in videos from unseen viewpoints. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4124-4132, 2021. [cited by applicant]
Sanqing Qu, Guang Chen, Dan Xu, Jinhu Dong, Fan Lu, and Alois Knoll. Lap-net: Adaptive features sampling via learning action progression for online action detection. arXiv preprint arXiv:2011.07915, 2020. [cited by applicant]
Charles Ringer and Mihalis A Nicolaou. Deep unsupervised multi-view detection of video game stream highlights. In Proceedings of the 13th International Conference on the Foundations of Digital Games, pp. 1-6, 2018. [cited by applicant]
Saquib Sarfraz, Naila Murray, Vivek Sharma, Ali Diba, Luc Van Gool, and Rainer Stiefelhagen. Temporally-weighted hierarchical clustering for unsupervised action segmentation. In Proceedings of the IEEE/CVF Conference on… [cited by applicant]
Fadime Sener, Dipika Singhania, and Angela Yao. Temporal aggregate representations for long-range video understanding. In European Conference on Computer Vision, pp. 154-171. Springer, 2020. [cited by applicant]
Fadime Sener and Angela Yao. Unsupervised learning and segmentation of complex activities from video. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8368-8376, 2018. [cited by applicant]
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, Sergey Levine, and Google Brain. Time-contrastive networks: Self-supervised learning from video. In 2018 IEEE international conferenc… [cited by applicant]
Zheng Shou, Junting Pan, Jonathan Chan, Kazuyuki Miyazawa, Hassan Mansour, Anthony Vetro, Xavier Giro-I Nieto, and Shih-Fu Chang. Online detection of action start in untrimmed, streaming videos. In Proceedings of the Eu… [cited by applicant]
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. Actor and observer: Joint modeling of first and third-person videos. In Proceedings of the IEEE Conference on Computer Vision and Pa… [cited by applicant]
Andrew Viterbi. Error bounds for convolutional codes and an asymptotically optimum decoding algorithm. IEEE transactions on Information Theory, 13(2):260-269, 1967. [cited by applicant]
Shruti Vyas, Yogesh S Rawat, and Mubarak Shah. Multiview action recognition using cross-view video prediction. In ECCV, pp. 427-444. Springer, 2020. [cited by applicant]
Dongang Wang, Wanli Ouyang, Wen Li, and Dong Xu. Dividing and aggregating network for multi-view action recognition. In ECCV, pp. 451-467, 2018. [cited by applicant]
Heng Wang and Cordelia Schmid. Action recognition with improved trajectories. In Proceedings of the IEEE international conference on computer vision, pp. 3551-3558, 2013. [cited by applicant]
Lichen Wang, Zhengming Ding, Zhiqiang Tao, Yunyu Liu, and Yun Fu. Generative multi-view human action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6212-6221, 2019. [cited by applicant]
Zhenzhi Wang, Ziteng Gao, Limin Wang, Zhifeng Li, and Gangshan Wu. Boundary-aware cascade networks for temporal action segmentation. In European Conference on Computer Vision, pp. 34-51. Springer, 2020. [cited by applicant]
Bo Xiong, Haoqi Fan, Kristen Grauman, and Christoph Feichtenhofer. Multiview pseudo-labeling for semi-supervised learning from video. arXiv preprint arXiv:2104.00682, 2021. [cited by applicant]
Mingze Xu, Mingfei Gao, Yi-Ting Chen, Larry S Davis, and David J Crandall. Temporal recurrent networks for online action detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5532-55… [cited by applicant]
Mingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li, Wei Xia, Zhuowen Tu, and Stefano Soatto. Long short-term transformer for online action detection. arXiv preprint arXiv:2107.03377, 2021. [cited by applicant]
Bowen Zhang, Hao Chen, Meng Wang, and Yuanjun Xiong. Online action detection in streaming videos with time buffers. arXiv preprint arXiv:2010.03016, 2020. [cited by applicant]
Peisen Zhao, Lingxi Xie, Ya Zhang, Yanfeng Wang, and Qi Tian. Privileged knowledge distillation for online action detection. arXiv preprint arXiv:2011.09158, 2020. [cited by applicant]