IP Library › Granted Patent US 12,406,496
Granted Patent B2
US 12,406,496 · App. 17/824,402 · Granted Sep 2, 2025

Anticipative video transformer model for future action anticipation

Inventors: Rohit Girdhar (Jersey City, NJ); Kristen Lorraine Grauman (Austin, TX)
G06V20/41G06V10/62G06V10/776G06V10/82G06V20/10G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,496
App. No.
17/824,402
Granted
Sep 2, 2025
Kind
B2
Abstract

In particular embodiments, a computing system may receive a video comprising a plurality of image frames. The system may generate, for each image frame in the plurality of image frames and using a spatial-attention encoder, an image-frame feature corresponding to the image frame. For each image-frame feature, the system may generate, using a temporal-attention decoder, a predicted future feature based on one or more of the image-frame features corresponding to one or more of the plurality of image frames that precede a time associated with the predicted future feature. The system may generate a future action anticipation based on the predicted future feature. The future action anticipation corresponds to an anticipation of a future action occurring after a sequence of actions observed in the plurality of images frames in the video.

Claims (48)

1. A method, implemented by a computing system, comprising:

receiving a video comprising a plurality of image frames;

generating, for the plurality of image frames and using a spatial-attention encoder, one or more image-frame features corresponding to one or more image frames of the plurality of image frames;

for the one or more image-frame features, generating, using a temporal-attention decoder, a predicted future feature based on the one or more image-frame features corresponding to the one or more image frames that precede a time associated with the predicted future feature; and

generating a video representation of a future action anticipation based on the predicted future feature, wherein the future action anticipation corresponds to an anticipation of a future action occurring after a sequence of actions observed in the plurality of image frames in the video, and wherein the video representation is configured for display on a user interface.

2. The method of claim 1 , further comprising:

generating a predicted action class label for the predicted future feature using a linear classifier.

3. The method of claim 2 , further comprising:

determining a first loss function by comparing the future action anticipation with ground-truth future action;

determining a second loss function by comparing predicted future features, generated by the temporal-attention decoder, with ground-truth future features;

determining a third loss function by comparing predicted action class labels corresponding to the predicted future features with ground-truth action class labels; and

training a machine-learning model based on the first loss fuction, the second loss fuction, and the third loss function, wherein the machine-learning model is used to generate the future action anticipation.

4. The method of claim 3 , wherein the machine-learning model comprises the spatial-attention encoder and the temporal-attention decoder.

5. The method of claim 3 , wherein the machine-learning model is an end-to-end attention-based model that attends to the plurality of image frames in the video to anticipate one or more future actions.

6. The method of claim 1 , further comprising:

generating a long-term anticipation comprising a sequence of future actions occurring after the sequence of actions observed in the plurality of images frames in the video.

7. The method of claim 1 , further comprising:

displaying the video representation of the future action anticipation on the user interface of a device worn by a user.

8. The method of claim 7 , wherein the device worn by the user is an augmented-reality device.

9. The method of claim 7 , wherein the video is an egocentric video captured using one or more cameras of the device worn by the user.

10. The method of claim 1 , wherein generating the future action anticipation comprises:

decoding the predicted future feature of the video into a distribution over semantic action classes using a linear classifier.

11. The method of claim 1 , wherein generating, using the spatial-attention encoder, the one or more image-frame features corresponding to the one or more image frames comprises:

splitting the one or more image frames into a plurality of patches;

determining, for the plurality of patches, patch features corresponding to the plurality of patches;

adding, to the patch features, a classification token and a spatial position embedding; and

generating a feature corresponding to the one or more image frames based on the patch features and the spatial position embeddings.

12. The method of claim 1 , wherein the spatial-attention encoder is a transformer encoder comprising one or more of a multi-head self-attention component and a feed forward network.

13. The method of claim 1 , wherein the generating, using the temporal-attention decoder, the predicted future feature comprises:

adding a temporal position encoding to the image-frame features generated using the spatial-attention encoder; and

generating the predicted future feature at a current frame image in the video after attending to an image-frame feature corresponding to the current image frame and one or more image-frame features corresponding to one or more image frames preceding the current image frame in the video.

14. The method of claim 13 , wherein the temporal-attention decoder is a transformer decoder with casual masked attention.

15. The method of claim 14 , wherein the casual masked attention ensures that the temporal-attention decoder attends to specific portions or image frames of the video for the future action anticipation.

16. The method of claim 1 , wherein the sequence of actions observed in the plurality of image frames in the video is associated with one or more action tasks.

17. The method of claim 16 , wherein the future action anticipation comprises a subsequent action task that is likely to be performed or followed after the one or more action tasks associated with the sequence of actions observed in the plurality of image frames in the video.

18. The method of claim 1 , wherein the plurality of image frames is sequential.

19. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

receive a video comprising a plurality of image frames;

generate, for the plurality of image frames and using a spatial-attention encoder, one or more image-frame features corresponding to the one or more image frames of the plurality of image frames;

for the one or more image-frame features, generate, using a temporal-attention decoder, a predicted future feature based on the one or more image-frame features corresponding to the one or more image frames that precede a time associated with the predicted future feature; and

generate a video representation of a future action anticipation based on the predicted future feature, wherein the future action anticipation corresponds to an anticipation of a future action occurring after a sequence of actions observed in the plurality of image frames in the video, and wherein the video representation is configured for display on a user interface.

20. A system comprising:

one or more processors; and

one or more computer-readable non-transitory storage media coupled to one or more of the processors and comprising instructions operable when executed by one or more of the processors to cause the system to:

receive a video comprising a plurality of image frames;

generate, for the plurality of image frames and using a spatial-attention encoder, an one or more image-frame features corresponding to one or more image frames of the plurality of image frames;

for the one or more image-frame features, generate, using a temporal-attention decoder, a predicted future feature based on the one or more image-frame features corresponding to the one or more image frames that precede a time associated with the predicted future feature; and

generate a video representation of a future action anticipation based on the predicted future feature, wherein the future action anticipation corresponds to an anticipation of a future action occurring after a sequence of actions observed in the plurality of image frames in the video, and wherein the video representation is configured for display on a user interface.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2022
From: GIRDHAR, ROHIT; GRAUMAN, KRISTEN LORRAINE
To: META PLATFORMS, INC.
Reel/Frame 060724/0700 →
Continuity (1)
Related Publication 20230386203A1 · Nov 30, 2023
References Cited (107)
US 20200160064A1 · Wang · 2020 [cited by examiner]
US 20230351759A1 · Girase · 2023 [cited by examiner]
Girdhar R., et al., “Anticipative Video Transformer,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 10, 2021, pp. 13485-13495. (Year: 2021). [cited by examiner]
Zhao S., et al., “Point transformer,” 2021, In ICCV, 11 pages. [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2023/023490, mailed Dec. 5, 2024, 6 pages. [cited by applicant]
Girdhar R., et al., “Anticipative Video Transformer,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 10, 2021, pp. 13485-13495. [cited by applicant]
Lample G., et al., “Cross-Lingual Language Model Pretraining,” Neural Information Processing Systems 32 (NeurIPS), 2019, pp. 1-11. [cited by applicant]
Lewis M., et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,” arXiv preprint arXiv:1910.13461, 2019, 10 pages. [cited by applicant]
Li X., et al., “Directional Temporal Modeling for Action Recognition,” European Conference on Computer Vision (ECCV), 2020, 17 pages. [cited by applicant]
Li Y., et al., “In the Eye of Beholder: Joint Learning of Gaze and Actions in First Person Video”, The European Conference on Computer Vision (ECCV), 2018, pp. 619-635. [cited by applicant]
Liu D., et al., “Non-Local Recurrent Network for Image Restoration,” Advances in Neural Information Processing Systems (NeurIPS), 2019, 10 pages. [cited by applicant]
Liu M., et al., “Forecasting Human Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video,” European Conference on Computer Vision (ECCV), 2020, pp. 704-721. [cited by applicant]
Liu W., et al., “Future Frame Prediction for Anomaly Detection—A New Baseline,” Computer Vision and Pattern Recognition (CVPR), 2018, pp. 6536-6545. [cited by applicant]
Long X., et al., “Attention Clusters: Purely Attention based Local Feature Integration for Video Classification,” Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7834-7843. [cited by applicant]
Luc P., et al., “Predicting Future Instance Segmentation by Forecasting Convolutional Features,” European Conference on Computer Vision (ECCV), 2018, 16 pages. [cited by applicant]
Ma S., et al., “Learning Activity Progression in LSTMs for Activity Detection and Early Detection,” Computer Vision and Pattern Recognition (CVPR), 2016, pp. 1942-1950. [cited by applicant]
Miech A., et al., “Learnable Pooling with Context Gating for Video Classification,” arXiv:1706.06905v2 [cs.CV], Mar. 5, 2018, 8 pages. [cited by applicant]
Miech A., et al., “Leveraging the Present to Anticipate the Future in Videos,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2019, 8 pages. [cited by applicant]
Nagarajan T., et al., “Ego-Topo: Environment Affordances from Egocentric Video,” The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 163-172. [cited by applicant]
Neimark D., et al., “Video Transformer Network,” IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, 2021, pp. 3163-3172. [cited by applicant]
Oord A., et al., “Representation Learning with Contrastive Predictive Coding,” Machine Learning, 2018, pp. 1-13. [cited by applicant]
Peters M.E., et al., “Deep Contextualized Word Representations,” In Proceedings of ACL, Jun. 1-6, 2018, pp. 2227-2237. [cited by applicant]
Piaget J., “La naissance de l'intelligence chez l'enfant. [The Origins of Intelligence in Children],” 1935, 442 pages. (English Translation). [cited by applicant]
Radford A., et al., “Language Models are Unsupervised Multitask Learners,” 2019, 24 pages. [cited by applicant]
Radford A., et al., “Improving Language Understanding by Generative Pre-Training,” 2018, 12 pages. [cited by applicant]
Raffel C., et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” arXiv preprint arXiv: 1910.10683, 2019, 67 pages. [cited by applicant]
Ren S., et al., “Faster R-CNN: Towards Real-Time Object Detection With Region Proposal Networks,” Advances in Neural Information Processing Systems, 2015, pp. 91-99. [cited by applicant]
Rhinehart N., et al., “First-Person Activity Forecasting with Online Inverse Reinforcement Learning,” The IEEE International Conference on Computer Vision (ICCV), 2017, pp. 3696-3705. [cited by applicant]
Richard A., et al., “Weakly Supervised Action Learning with RNN Based Fine-to-Coarse Modeling,” In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 754-763. [cited by applicant]
Rodriguez C., et al., “Action Anticipation by Predicting Future Dynamic Images,” European Conference on Computer Vision (ECCV) Workshop, 2018, pp. 1-16. [cited by applicant]
Sener F., et al., “Technical Report: Temporal Aggregate Representations,” arXiv:2106.03152, 2021, 4 pages. [cited by applicant]
Sener F., et al., “Temporal Aggregate Representations for Long-Range Video Understanding,” In European Conference on Computer Vision (ECCV), Jul. 30, 2020, 23 pages. [cited by applicant]
Shi Y., et al., “Action Anticipation with RBF Kernelized Feature Mapping RNN,” European Conference on Computer Vision (ECCV), 2018, pp. 1-17. [cited by applicant]
Shou M.Z., et al., “Generic Event Boundary Detection: A Benchmark for Event Segmentation,” arXiv: 2101.10511, 2021, 15 pages. [cited by applicant]
Simonyan K., et al., “Two-Stream Convolutional Networks for Action Recognition in Videos, ” Advances in Neural Information Processing Systems, 2014, vol. 27, pp. 568-576. [cited by applicant]
Soomro K., et al., “UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild,” arXiv:1212.0402v1 [cs.CV], Dec. 3, 2012, 7 pages. [cited by applicant]
Stein S., et al., “Combining Embedded Accelerometers with Computer Vision for Recognizing Food Preparation Activities,” UbiComp, Sep. 2013, pp. 729-738. [cited by applicant]
Sun C., et al., “Contrastive Bidirectional Transformer for Temporal Representation Learning,” arXiv preprint arXiv: 1906.05743, 2019. [cited by applicant]
Sun C., et al., “VideoBERT: A Joint Model for Video and Language Representation Learning,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 7464-7473. [cited by applicant]
Touvron H., et al., “Training Data-Efficient Image Transformers and Distillation Through Attention,” In International Conference on Machine Learning (ICML), 2021, 11 pages. [cited by applicant]
Tran D., et al., “A Closer Look at Spatiotemporal Convolutions for Action Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 6450-6459. [cited by applicant]
Tran D., et al., “Learning Spatiotemporal Features with 3D Convolutional Networks,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015, pp. 4489-4497. [cited by applicant]
Tran D., et al., “Video Classification with Channel-Separated Convolutional Networks,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 5551-5560. [cited by applicant]
Vaswani A., et al., “Attention Is All You Need,” Advances in Neural Information Processing Systems, December (NIPS), 2017, pp. 5998-6008. [cited by applicant]
Vondrick C., et al., “Anticipating Visual Representations from Unlabeled Video,” Computer Vision and Pattern Recognition (CVPR), 2016, pp. 98-106. [cited by applicant]
Wang L., et al., “Temporal Segment Networks: Towards Good Practices for Deep Action Recognition,” In European Conference on Computer Vision (ECCV), Sep. 17, 2016, 17 pages. [cited by applicant]
Wang Q., et al., “Learning Deep Transformer Models for Machine Translation,” 2019, In ACL, arXiv:1906.01787v1 [cs.CL], 13 pages. [cited by applicant]
Wang X., et al., “Non-Local Neural Networks, ” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7794-7803. [cited by applicant]
Wei D., et al., “Learning and Using the Arrow of Time,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 8052-8060. [cited by applicant]
Wu C., et al., “Long-Term Feature Banks for Detailed Video Understanding,” In Proceedings of the IEEE/CVF Conference onComputer Vision and Pattern Recognition, 2019, pp. 284-293. [cited by applicant]
Xie S., et al., “Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification,” Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 318-335, DOI: 10.1007/978-3-03… [cited by applicant]
Yamamuro Y., et al., Submission to Epic-Kitchens Action Anticipation Challenge 2021, In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021, 4 pages. [cited by applicant]
Yang C., et al., “Video Representation Learning With Visual Tempo Consistency,” 2020, In arXiv preprint arXiv:2006.15489, 11 pages. [cited by applicant]
Yu Y., et al., “Learning to Anticipate Egocentric Actions by Imagination,” IEEE Transactions on Image Processing, 2021, vol. 30, 10 pages. [cited by applicant]
Zatsarynna O., et al., “Multi-Modal Temporal Convolutional Network for Anticipating Actions in Egocentric Videos,” In CVPR Workshop, 2021, 10 pages. [cited by applicant]
Zhang Y., et al., “VidTr: Video Transformer without Convolutions,” arXiv:2104.11746, 2021, 11 pages. [cited by applicant]
Abnar S., et al., “Quantifying Attention Flow in Transformers,” Artificial Intelligence Computation and Language (ACL), May 31, 2020, 8pages. [cited by applicant]
Arandjelovic R., et al., “Look, Listen and Learn,” International Conference on Computer Vision (ICCV), 2017, pp. 609-617. [cited by applicant]
Arnab A., et al., “ViVIT: A Video Vision Transformer,” IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 6836-6846. [cited by applicant]
Bertasius G., et al., “Classifying, Segmenting, and Tracking Object Instances in Video with Mask Propagation,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 9739-9748. [cited by applicant]
Bertasius G., et al., “Cobe: Contextualized Object Embeddings from Narrated Instructional Video,” In Neural Information Processing Systems 33 (NeurIPS 2020), 2020, 13 pages. [cited by applicant]
Bertasius G, et al., “Is Space-Time Attention All You Need for Video Understanding?,” Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 2021, 11 pages. [cited by applicant]
Brown T.B., et al., “Language Models are Few-Shot Learners,” arXiv:2005.14165, 2020, 75 pages. [cited by applicant]
Buades A., et al., “A Non-Local Algorithm for Image Denoising,” IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), 2005, 6 pages. [cited by applicant]
Cao Y., et al., “GcNet: Non-Local Networks Meet Squeeze-Excitation Networks and Beyond,” In IEEE/CVF International Conference on Computer Vision (ICCV), 2019, 10 pages. [cited by applicant]
Carion N., et al., “End-to-End Object Detection with Transformers,” In European Conference Computer Vision (ECCV), 2020, 26 pages. [cited by applicant]
Carreira J., et al., “Quo Vadis, Action Recognition? A New Model and The Kinetics Dataset,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 6299-6308. [cited by applicant]
Damen D., et al., “Rescaling Egocentric Vision: Collection Pipeline and Challenges for Epic-Kitchens-100,” arXiv preprint arXiv:2006.13256, Sep. 17, 2021, 20 pages. [cited by applicant]
Damen D., et al., “Scaling Egocentric Vision: The EPIC-KITCHENS Dataset”, The European Conference on Computer Vision (ECCV), 2018, pp. 720-736. [cited by applicant]
De G.R., et al., “Modeling Temporal Structure with LSTM for Online Action Detection,” IEEE Winter Conference on Applications of Computer Vision (WACV), 2018, 9 pages. [cited by applicant]
Dessalene E., et al., “Forecasting Action through Contact Representations from First Person Video,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2021, 12 pages. [cited by applicant]
Devlin J., et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” In Proceedings of the 2019Confer-ence of the North American Chapter of the Association for ComputationalLinguistics:… [cited by applicant]
Dosovitskiy A., et al., “An Image is Worth 16x16 words: Transformers for Image Recognition at Scale,” International Conference on Learning (ICLR), 2021, 21 pages. [cited by applicant]
Fan H., et al., “Multiscale Vision Transformers,” In IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 6824-6835. [cited by applicant]
Farha Y.A., et al., “When will you do what?—Anticipating Temporal Occurrences of Activities,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 5343-5352. [cited by applicant]
Feichtenhofer C., et al., “SlowFast Networks for Video Recognition,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 6202-6211. [cited by applicant]
Fernando B., et al., “Self-Supervised Video Representation Learning With Odd-One-Out Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 3636-3645. [cited by applicant]
Furnari A., et al., “Leveraging Uncertainty to Rethink Loss Functions and Evaluation Measures for Egocentric Action Anticipation,” European Conference on Computer Vision (ECCV), 2018, 17 pages. [cited by applicant]
Furnari A., et al., “Rolling-Unrolling LSTMs for Action Anticipation from First-Person Video,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2020, 16 pages. [cited by applicant]
Furnari A., et al., “What Would you Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality Attention”, The IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 6252-6261. [cited by applicant]
Gammulle H., et al., “Predicting the Future: A Jointly Learnt Model for Action Anticipation,” International Conference on Computer Vision (ICCV), 2019, pp. 5562-5571. [cited by applicant]
Gao J., et al., “Red: Reinforced Encoder-Decoder Networks for Action Anticipation,” The British Machine Vision Conference (BMVC), 2017, 11 pages. [cited by applicant]
Ghadiyaram D., et al., “Largescale Weakly-Supervised Pre-Training for Video Action Recognition,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 12046-12055. [cited by applicant]
Girdhar R., et al., “ActionVLAD: Learning Spatio-Temporal Aggregation for Action Classification,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 971-980. [cited by applicant]
Girdhar R., et al., “Anticipative Video Transformer @ EPIC-Kitchens Action Anticipation Challenge 2021,” Conference on Computer Vision and Pattern (CVPR), 2021, 5 pages. [cited by applicant]
Girdhar R., et al., “Attentional Pooling for Action Recognition,” Conference on Neural Information Processing Systems (NeurIPS), 2017, 12 pages. [cited by applicant]
Girdhar R., et al., “CATER: A Diagnostic Dataset for Compositional Actions and TEmporal Reasoning,” International Conference on Learning Representations (ICLR), Apr. 5, 2020, 16 pages. [cited by applicant]
Girdhar R., et al., “Video Action Transformer Network,” In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2019, pp. 244-253. [cited by applicant]
Goyal P., et al., “Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour,” arXiv preprint, arXiv: 1706.02677, Apr. 30, 2018, 12 pages. [cited by applicant]
Goyal R., et al., “The “Something Something” Video Database for Learning and Evaluating Visual Common Sense,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 5842-5850. [cited by applicant]
Gu X., et al., “Transaction: ICL-SJTU Submission to Epic-Kitchens Action Anticipation Challenge 2021,” Computer Vision and Pattern Recognition (CVPR), Jul. 28, 2021, 4 pages. [cited by applicant]
Han T., et al., “Memory-Augmented Dense Predictive Coding for Video Representation Learning,” In European Conference on Computer Vision (ECCV), Aug. 3, 2020, 23 pages. [cited by applicant]
Han T., et al., “Video Representation Learning by Dense Predictive Coding,” IEEE/CVF International Conference on Computer Vision (ICCV), 2019, 10 pages. [cited by applicant]
Huang D.A., et al., “Action-Reaction: Fore-Casting the Dynamics of Human Interaction,” European Conference on Computer Vision (ECCV), 2014, pp. 489-504. [cited by applicant]
Jain A., et al., “Recurrent Neural Networks for Driver Activity Anticipation via Sensory-Fusion Architecture,” IEEE International Conference on Robotics and Automation (ICRA), Sep. 16, 2015, 8 pages. [cited by applicant]
Jayaraman D., et al., “Learning Image Representations Tied to Ego-Motion,” In IEEE International Conference on Computer Vision (ICCV), 2015, pp. 1413-1421. [cited by applicant]
Jayaraman D., et al., “Slow and Steady Feature Analysis: Higher Order Temporal Coherence in Video,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 3852-3861. [cited by applicant]
Karpathy A., et al., “Large-scale Video Classification with Convolutional Neural Networks,” In Proceedings of 2014 EEE Conference on Computer Vision and Pattern Recognition, Computer Science Department, Stanford Univers… [cited by applicant]
Kay W., et al., “The Kinetics Human Action Video Dataset,” arXiv preprint, arXiv :1705.06950, May 19, 2017, 22 pages. [cited by applicant]
Kim D., et al., “Self- Supervised Video Representation Learning with Space-Time Cubic Puzzles,” In Proceedings of the AAAI Conference on Artificial Intelligence, 2019, vol. 33, pp. 8545-8552. [cited by applicant]
Kitani K.M., et al., “Activity forecasting,” European Conference on Computer Vision (ECCV), 2012, 14 pages. [cited by applicant]
Kong S., et al., “Low-Rank Bilinear Pooling for Fine-Grained Classification,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 365-374. [cited by applicant]
Koppula H.S., et al., “Anticipating Human Activities using Object Affordances for Reactive Robotic Response,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 38, Issue No. 1, May 2015, pp. 1… [cited by applicant]
Korbar B., et al., “Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization,” Advances in Neural Information Processing Systems (NeurIPS), 2018, 12 pages. [cited by applicant]
Kuehne H., et al., “HMDB: A Large Video Database for Human Motion Recognition,” 2011 International Conference on Computer Vision, DOI: 10.1109/ICCV.2011.6126543, 2011, 9 Pages. [cited by applicant]
Kuehne H., et al., “The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities,” Computer Vision and Pattern Recognition (CVPR), 2014, pp. 780-787. [cited by applicant]
International Search report and Written Opinion for International Application No. PCT/US2023/023490, mailed Jul. 25, 2023, 6 pages. [cited by applicant]