Girdhar R., et al., “Anticipative Video Transformer,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 10, 2021, pp. 13485-13495. (Year: 2021).
[cited by examiner]
Zhao S., et al., “Point transformer,” 2021, In ICCV, 11 pages.
[cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2023/023490, mailed Dec. 5, 2024, 6 pages.
[cited by applicant]
Girdhar R., et al., “Anticipative Video Transformer,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 10, 2021, pp. 13485-13495.
[cited by applicant]
Lample G., et al., “Cross-Lingual Language Model Pretraining,” Neural Information Processing Systems 32 (NeurIPS), 2019, pp. 1-11.
[cited by applicant]
Lewis M., et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,” arXiv preprint arXiv:1910.13461, 2019, 10 pages.
[cited by applicant]
Li X., et al., “Directional Temporal Modeling for Action Recognition,” European Conference on Computer Vision (ECCV), 2020, 17 pages.
[cited by applicant]
Li Y., et al., “In the Eye of Beholder: Joint Learning of Gaze and Actions in First Person Video”, The European Conference on Computer Vision (ECCV), 2018, pp. 619-635.
[cited by applicant]
Liu D., et al., “Non-Local Recurrent Network for Image Restoration,” Advances in Neural Information Processing Systems (NeurIPS), 2019, 10 pages.
[cited by applicant]
Liu M., et al., “Forecasting Human Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video,” European Conference on Computer Vision (ECCV), 2020, pp. 704-721.
[cited by applicant]
Liu W., et al., “Future Frame Prediction for Anomaly Detection—A New Baseline,” Computer Vision and Pattern Recognition (CVPR), 2018, pp. 6536-6545.
[cited by applicant]
Long X., et al., “Attention Clusters: Purely Attention based Local Feature Integration for Video Classification,” Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7834-7843.
[cited by applicant]
Luc P., et al., “Predicting Future Instance Segmentation by Forecasting Convolutional Features,” European Conference on Computer Vision (ECCV), 2018, 16 pages.
[cited by applicant]
Ma S., et al., “Learning Activity Progression in LSTMs for Activity Detection and Early Detection,” Computer Vision and Pattern Recognition (CVPR), 2016, pp. 1942-1950.
[cited by applicant]
Miech A., et al., “Learnable Pooling with Context Gating for Video Classification,” arXiv:1706.06905v2 [cs.CV], Mar. 5, 2018, 8 pages.
[cited by applicant]
Miech A., et al., “Leveraging the Present to Anticipate the Future in Videos,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2019, 8 pages.
[cited by applicant]
Nagarajan T., et al., “Ego-Topo: Environment Affordances from Egocentric Video,” The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 163-172.
[cited by applicant]
Neimark D., et al., “Video Transformer Network,” IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, 2021, pp. 3163-3172.
[cited by applicant]
Oord A., et al., “Representation Learning with Contrastive Predictive Coding,” Machine Learning, 2018, pp. 1-13.
[cited by applicant]
Peters M.E., et al., “Deep Contextualized Word Representations,” In Proceedings of ACL, Jun. 1-6, 2018, pp. 2227-2237.
[cited by applicant]
Piaget J., “La naissance de l'intelligence chez l'enfant. [The Origins of Intelligence in Children],” 1935, 442 pages. (English Translation).
[cited by applicant]
Radford A., et al., “Language Models are Unsupervised Multitask Learners,” 2019, 24 pages.
[cited by applicant]
Radford A., et al., “Improving Language Understanding by Generative Pre-Training,” 2018, 12 pages.
[cited by applicant]
Raffel C., et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” arXiv preprint arXiv: 1910.10683, 2019, 67 pages.
[cited by applicant]
Ren S., et al., “Faster R-CNN: Towards Real-Time Object Detection With Region Proposal Networks,” Advances in Neural Information Processing Systems, 2015, pp. 91-99.
[cited by applicant]
Rhinehart N., et al., “First-Person Activity Forecasting with Online Inverse Reinforcement Learning,” The IEEE International Conference on Computer Vision (ICCV), 2017, pp. 3696-3705.
[cited by applicant]
Richard A., et al., “Weakly Supervised Action Learning with RNN Based Fine-to-Coarse Modeling,” In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 754-763.
[cited by applicant]
Rodriguez C., et al., “Action Anticipation by Predicting Future Dynamic Images,” European Conference on Computer Vision (ECCV) Workshop, 2018, pp. 1-16.
[cited by applicant]
Sener F., et al., “Technical Report: Temporal Aggregate Representations,” arXiv:2106.03152, 2021, 4 pages.
[cited by applicant]
Sener F., et al., “Temporal Aggregate Representations for Long-Range Video Understanding,” In European Conference on Computer Vision (ECCV), Jul. 30, 2020, 23 pages.
[cited by applicant]
Shi Y., et al., “Action Anticipation with RBF Kernelized Feature Mapping RNN,” European Conference on Computer Vision (ECCV), 2018, pp. 1-17.
[cited by applicant]
Shou M.Z., et al., “Generic Event Boundary Detection: A Benchmark for Event Segmentation,” arXiv: 2101.10511, 2021, 15 pages.
[cited by applicant]
Simonyan K., et al., “Two-Stream Convolutional Networks for Action Recognition in Videos, ” Advances in Neural Information Processing Systems, 2014, vol. 27, pp. 568-576.
[cited by applicant]
Soomro K., et al., “UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild,” arXiv:1212.0402v1 [cs.CV], Dec. 3, 2012, 7 pages.
[cited by applicant]
Stein S., et al., “Combining Embedded Accelerometers with Computer Vision for Recognizing Food Preparation Activities,” UbiComp, Sep. 2013, pp. 729-738.
[cited by applicant]
Sun C., et al., “Contrastive Bidirectional Transformer for Temporal Representation Learning,” arXiv preprint arXiv: 1906.05743, 2019.
[cited by applicant]
Sun C., et al., “VideoBERT: A Joint Model for Video and Language Representation Learning,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 7464-7473.
[cited by applicant]
Touvron H., et al., “Training Data-Efficient Image Transformers and Distillation Through Attention,” In International Conference on Machine Learning (ICML), 2021, 11 pages.
[cited by applicant]
Tran D., et al., “A Closer Look at Spatiotemporal Convolutions for Action Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 6450-6459.
[cited by applicant]
Tran D., et al., “Learning Spatiotemporal Features with 3D Convolutional Networks,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015, pp. 4489-4497.
[cited by applicant]
Tran D., et al., “Video Classification with Channel-Separated Convolutional Networks,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 5551-5560.
[cited by applicant]
Vaswani A., et al., “Attention Is All You Need,” Advances in Neural Information Processing Systems, December (NIPS), 2017, pp. 5998-6008.
[cited by applicant]
Vondrick C., et al., “Anticipating Visual Representations from Unlabeled Video,” Computer Vision and Pattern Recognition (CVPR), 2016, pp. 98-106.
[cited by applicant]
Wang L., et al., “Temporal Segment Networks: Towards Good Practices for Deep Action Recognition,” In European Conference on Computer Vision (ECCV), Sep. 17, 2016, 17 pages.
[cited by applicant]
Wang Q., et al., “Learning Deep Transformer Models for Machine Translation,” 2019, In ACL, arXiv:1906.01787v1 [cs.CL], 13 pages.
[cited by applicant]
Wang X., et al., “Non-Local Neural Networks, ” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7794-7803.
[cited by applicant]
Wei D., et al., “Learning and Using the Arrow of Time,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 8052-8060.
[cited by applicant]
Wu C., et al., “Long-Term Feature Banks for Detailed Video Understanding,” In Proceedings of the IEEE/CVF Conference onComputer Vision and Pattern Recognition, 2019, pp. 284-293.
[cited by applicant]
Xie S., et al., “Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification,” Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 318-335, DOI: 10.1007/978-3-03…
[cited by applicant]
Yamamuro Y., et al., Submission to Epic-Kitchens Action Anticipation Challenge 2021, In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021, 4 pages.
[cited by applicant]
Yang C., et al., “Video Representation Learning With Visual Tempo Consistency,” 2020, In arXiv preprint arXiv:2006.15489, 11 pages.
[cited by applicant]
Yu Y., et al., “Learning to Anticipate Egocentric Actions by Imagination,” IEEE Transactions on Image Processing, 2021, vol. 30, 10 pages.
[cited by applicant]
Zatsarynna O., et al., “Multi-Modal Temporal Convolutional Network for Anticipating Actions in Egocentric Videos,” In CVPR Workshop, 2021, 10 pages.
[cited by applicant]
Zhang Y., et al., “VidTr: Video Transformer without Convolutions,” arXiv:2104.11746, 2021, 11 pages.
[cited by applicant]
Abnar S., et al., “Quantifying Attention Flow in Transformers,” Artificial Intelligence Computation and Language (ACL), May 31, 2020, 8pages.
[cited by applicant]
Arandjelovic R., et al., “Look, Listen and Learn,” International Conference on Computer Vision (ICCV), 2017, pp. 609-617.
[cited by applicant]
Arnab A., et al., “ViVIT: A Video Vision Transformer,” IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 6836-6846.
[cited by applicant]
Bertasius G., et al., “Classifying, Segmenting, and Tracking Object Instances in Video with Mask Propagation,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 9739-9748.
[cited by applicant]
Bertasius G., et al., “Cobe: Contextualized Object Embeddings from Narrated Instructional Video,” In Neural Information Processing Systems 33 (NeurIPS 2020), 2020, 13 pages.
[cited by applicant]
Bertasius G, et al., “Is Space-Time Attention All You Need for Video Understanding?,” Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 2021, 11 pages.
[cited by applicant]
Brown T.B., et al., “Language Models are Few-Shot Learners,” arXiv:2005.14165, 2020, 75 pages.
[cited by applicant]
Buades A., et al., “A Non-Local Algorithm for Image Denoising,” IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), 2005, 6 pages.
[cited by applicant]
Cao Y., et al., “GcNet: Non-Local Networks Meet Squeeze-Excitation Networks and Beyond,” In IEEE/CVF International Conference on Computer Vision (ICCV), 2019, 10 pages.
[cited by applicant]
Carion N., et al., “End-to-End Object Detection with Transformers,” In European Conference Computer Vision (ECCV), 2020, 26 pages.
[cited by applicant]
Carreira J., et al., “Quo Vadis, Action Recognition? A New Model and The Kinetics Dataset,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 6299-6308.
[cited by applicant]
Damen D., et al., “Rescaling Egocentric Vision: Collection Pipeline and Challenges for Epic-Kitchens-100,” arXiv preprint arXiv:2006.13256, Sep. 17, 2021, 20 pages.
[cited by applicant]
Damen D., et al., “Scaling Egocentric Vision: The EPIC-KITCHENS Dataset”, The European Conference on Computer Vision (ECCV), 2018, pp. 720-736.
[cited by applicant]
De G.R., et al., “Modeling Temporal Structure with LSTM for Online Action Detection,” IEEE Winter Conference on Applications of Computer Vision (WACV), 2018, 9 pages.
[cited by applicant]
Dessalene E., et al., “Forecasting Action through Contact Representations from First Person Video,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2021, 12 pages.
[cited by applicant]
Devlin J., et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” In Proceedings of the 2019Confer-ence of the North American Chapter of the Association for ComputationalLinguistics:…
[cited by applicant]
Dosovitskiy A., et al., “An Image is Worth 16x16 words: Transformers for Image Recognition at Scale,” International Conference on Learning (ICLR), 2021, 21 pages.
[cited by applicant]
Fan H., et al., “Multiscale Vision Transformers,” In IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 6824-6835.
[cited by applicant]
Farha Y.A., et al., “When will you do what?—Anticipating Temporal Occurrences of Activities,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 5343-5352.
[cited by applicant]
Feichtenhofer C., et al., “SlowFast Networks for Video Recognition,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 6202-6211.
[cited by applicant]
Fernando B., et al., “Self-Supervised Video Representation Learning With Odd-One-Out Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 3636-3645.
[cited by applicant]
Furnari A., et al., “Leveraging Uncertainty to Rethink Loss Functions and Evaluation Measures for Egocentric Action Anticipation,” European Conference on Computer Vision (ECCV), 2018, 17 pages.
[cited by applicant]
Furnari A., et al., “Rolling-Unrolling LSTMs for Action Anticipation from First-Person Video,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2020, 16 pages.
[cited by applicant]
Furnari A., et al., “What Would you Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality Attention”, The IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 6252-6261.
[cited by applicant]
Gammulle H., et al., “Predicting the Future: A Jointly Learnt Model for Action Anticipation,” International Conference on Computer Vision (ICCV), 2019, pp. 5562-5571.
[cited by applicant]
Gao J., et al., “Red: Reinforced Encoder-Decoder Networks for Action Anticipation,” The British Machine Vision Conference (BMVC), 2017, 11 pages.
[cited by applicant]
Ghadiyaram D., et al., “Largescale Weakly-Supervised Pre-Training for Video Action Recognition,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 12046-12055.
[cited by applicant]
Girdhar R., et al., “ActionVLAD: Learning Spatio-Temporal Aggregation for Action Classification,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 971-980.
[cited by applicant]
Girdhar R., et al., “Anticipative Video Transformer @ EPIC-Kitchens Action Anticipation Challenge 2021,” Conference on Computer Vision and Pattern (CVPR), 2021, 5 pages.
[cited by applicant]
Girdhar R., et al., “Attentional Pooling for Action Recognition,” Conference on Neural Information Processing Systems (NeurIPS), 2017, 12 pages.
[cited by applicant]
Girdhar R., et al., “CATER: A Diagnostic Dataset for Compositional Actions and TEmporal Reasoning,” International Conference on Learning Representations (ICLR), Apr. 5, 2020, 16 pages.
[cited by applicant]
Girdhar R., et al., “Video Action Transformer Network,” In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2019, pp. 244-253.
[cited by applicant]
Goyal P., et al., “Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour,” arXiv preprint, arXiv: 1706.02677, Apr. 30, 2018, 12 pages.
[cited by applicant]
Goyal R., et al., “The “Something Something” Video Database for Learning and Evaluating Visual Common Sense,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 5842-5850.
[cited by applicant]
Gu X., et al., “Transaction: ICL-SJTU Submission to Epic-Kitchens Action Anticipation Challenge 2021,” Computer Vision and Pattern Recognition (CVPR), Jul. 28, 2021, 4 pages.
[cited by applicant]
Han T., et al., “Memory-Augmented Dense Predictive Coding for Video Representation Learning,” In European Conference on Computer Vision (ECCV), Aug. 3, 2020, 23 pages.
[cited by applicant]
Han T., et al., “Video Representation Learning by Dense Predictive Coding,” IEEE/CVF International Conference on Computer Vision (ICCV), 2019, 10 pages.
[cited by applicant]
Huang D.A., et al., “Action-Reaction: Fore-Casting the Dynamics of Human Interaction,” European Conference on Computer Vision (ECCV), 2014, pp. 489-504.
[cited by applicant]
Jain A., et al., “Recurrent Neural Networks for Driver Activity Anticipation via Sensory-Fusion Architecture,” IEEE International Conference on Robotics and Automation (ICRA), Sep. 16, 2015, 8 pages.
[cited by applicant]
Jayaraman D., et al., “Learning Image Representations Tied to Ego-Motion,” In IEEE International Conference on Computer Vision (ICCV), 2015, pp. 1413-1421.
[cited by applicant]
Jayaraman D., et al., “Slow and Steady Feature Analysis: Higher Order Temporal Coherence in Video,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 3852-3861.
[cited by applicant]
Karpathy A., et al., “Large-scale Video Classification with Convolutional Neural Networks,” In Proceedings of 2014 EEE Conference on Computer Vision and Pattern Recognition, Computer Science Department, Stanford Univers…
[cited by applicant]
Kay W., et al., “The Kinetics Human Action Video Dataset,” arXiv preprint, arXiv :1705.06950, May 19, 2017, 22 pages.
[cited by applicant]
Kim D., et al., “Self- Supervised Video Representation Learning with Space-Time Cubic Puzzles,” In Proceedings of the AAAI Conference on Artificial Intelligence, 2019, vol. 33, pp. 8545-8552.
[cited by applicant]
Kitani K.M., et al., “Activity forecasting,” European Conference on Computer Vision (ECCV), 2012, 14 pages.
[cited by applicant]
Kong S., et al., “Low-Rank Bilinear Pooling for Fine-Grained Classification,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 365-374.
[cited by applicant]
Koppula H.S., et al., “Anticipating Human Activities using Object Affordances for Reactive Robotic Response,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 38, Issue No. 1, May 2015, pp. 1…
[cited by applicant]
Korbar B., et al., “Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization,” Advances in Neural Information Processing Systems (NeurIPS), 2018, 12 pages.
[cited by applicant]
Kuehne H., et al., “HMDB: A Large Video Database for Human Motion Recognition,” 2011 International Conference on Computer Vision, DOI: 10.1109/ICCV.2011.6126543, 2011, 9 Pages.
[cited by applicant]
Kuehne H., et al., “The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities,” Computer Vision and Pattern Recognition (CVPR), 2014, pp. 780-787.
[cited by applicant]
International Search report and Written Opinion for International Application No. PCT/US2023/023490, mailed Jul. 25, 2023, 6 pages.
[cited by applicant]