IP Library › Granted Patent US 12,725,413
Granted Patent B2
US 12,725,413 · App. 18/411,928 · Granted Sep 1, 2026

System and method for modeling local and global spatio-temporal context in video for video recognition

Inventors: Syed Talal Wasim (Abu Dhabi, AE); Muhammad Uzair Khattak (Abu Dhabi, AE); Muzammal Naseer (Abu Dhabi, AE); Salman Khan (Abu Dhabi, AE); Fahad Shahbaz Khan (Abu Dhabi, AE)
Assignee: Mohamed bin Zayed University of Artificial Intelligence
G06V20/41G06V10/7715G06V20/50
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,413
App. No.
18/411,928
Granted
Sep 1, 2026
Kind
B2
Abstract

A system and a method for modeling local and global spatio-temporal context in a video for video recognition includes obtaining an input feature map and transforming the input feature map using linear functions to generate a spatial feature map and a temporal feature map corresponding to a video. The method further includes generating hierarchical contextual feature maps based on the spatial feature map and the temporal feature map that represent a context of the video at multiple levels of granularity. The method further includes aggregating the hierarchical contextual feature maps based on gating weights to obtain a spatial modulator and a temporal modulator that are representative of an aggregated context across the multiple levels. The method further includes obtaining an output spatio-temporal feature map based on the spatial modulator, the temporal modulator, and a query token associated with the video.

Claims (62)

1 . A method for modeling local and global spatio-temporal context in a video for video recognition, the method comprising:

obtaining an input feature map corresponding to a video;

transforming the input feature map using linear functions to generate a spatial feature map and a temporal feature map, wherein the spatial feature map is representative of intra-frame features in a frame of the video and the temporal feature map is representative of inter-frame features across frames of the video;

generating hierarchical contextual feature maps for the spatial feature map and the temporal feature map by applying:

depth-wise convolutions at multiple levels to the spatial feature map to generate a level-specific spatial feature map for each level, and

point-wise convolutions at multiple levels to the temporal feature map to generate a level-specific temporal feature map for each level;

aggregating the hierarchical contextual feature maps by:

aggregating the level-specific spatial feature map for the multiple levels using a first set of gating weights to obtain a spatial modulator, and

aggregating the level-specific temporal feature map for the multiple levels using a second set of gating weights to obtain a temporal modulator; and

obtaining an output spatio-temporal feature map based on the spatial modulator, the temporal modulator, and a query token associated with the video.

2 . The method of claim 1 , wherein the depth-wise convolutions and the point-wise convolutions are implemented using a GeLU activation function.

3 . The method of claim 1 , wherein generating the hierarchical contextual feature maps includes:

performing a global-pooling operation on the level-specific spatial feature map corresponding to the highest level of the multiple levels to obtain a top-level spatial feature map, and

performing the global-pooling operation on the level-specific temporal feature map corresponding to the highest level to obtain a top-level temporal feature map, wherein the top-level spatial feature map and the top-level temporal feature map are representative of a global context of the video.

4 . The method of claim 1 , wherein aggregating the hierarchical contextual feature maps includes:

obtaining a dot product of the first set of gating weights and the level-specific spatial feature map corresponding to each level to generate a first set of dot products, wherein the first set of dot products includes a dot product of an additional level-specific spatial feature map and the first set of gating weights corresponding to a level above the multiple levels,

aggregating the first set of dot products to obtain an aggregated spatial feature map, and

applying a first linear function to the aggregated spatial feature map to obtain the spatial modulator.

5 . The method of claim 1 , wherein aggregating the hierarchical contextual feature maps includes:

obtaining a dot product of the second set of gating weights and the level-specific temporal feature map corresponding to each level to generate a second set of dot products, wherein the second set of dot products includes a dot product of an additional level-specific temporal feature map and the second set of gating weights corresponding to a level above the multiple levels,

aggregating the second set of dot products to obtain an aggregated temporal feature map, and

applying a second linear function to the aggregated temporal feature map to obtain the temporal modulator.

6 . The method of claim 1 , wherein obtaining the output spatio-temporal feature map includes:

performing an element-wise multiplication between the query token, the spatial modulator and the temporal modulator.

7 . The method of claim 1 , wherein the query token is obtained by applying a linear function to the input feature map.

8 . The method of claim 1 further comprising:

executing a video recognition model using the output spatio-temporal feature map to classify the video into a specified classification.

9 . A method for modeling local and global spatio-temporal context in a video for video recognition, the method comprising:

obtaining a spatial feature map and a temporal feature map for a video;

generating hierarchical contextual feature maps based on the spatial feature map and the temporal feature map that represent a context of the video at multiple levels of granularity;

aggregating the hierarchical contextual feature maps based on gating weights to obtain a spatial modulator and a temporal modulator that are representative of an aggregated context across the multiple levels; and

obtaining an output spatio-temporal feature map based on the spatial modulator, the temporal modulator and a query token associated with the video.

10 . The method of claim 9 , wherein generating and aggregating the hierarchical contextual feature maps for the spatial feature map is performed independent of generating and aggregating the hierarchical contextual feature maps for the temporal feature map.

11 . The method of claim 9 , wherein generating the hierarchical contextual feature maps includes:

applying depth-wise convolutions at different levels to the spatial feature map to generate a level-specific spatial feature map for each level, and

applying point-wise convolutions at different levels to the temporal feature map to generate a level-specific temporal feature map for each level.

12 . The method of claim 11 further comprising:

performing a global-pooling operation on the level-specific spatial feature map corresponding to the highest level of the multiple levels to obtain a top-level spatial feature map, and

performing the global-pooling operation on the level-specific temporal feature map corresponding to the highest level to obtain a top-level temporal feature map, wherein the top-level spatial feature map and the top-level temporal feature map are representative of a global context of the video.

13 . The method of claim 11 , wherein aggregating the hierarchical contextual feature maps includes:

aggregating the level-specific spatial feature maps using a first set of gating weights to obtain the spatial modulator, and

aggregating the level-specific temporal feature maps using a second set of gating weights to obtain a temporal modulator.

14 . The method of claim 13 , wherein aggregating the hierarchical contextual feature maps includes:

obtaining a dot product of the first set of gating weights and the level-specific spatial feature map corresponding to each level to generate a first set of dot products,

aggregating the first set of dot products to obtain an aggregated spatial feature map, and

applying a first linear function to the aggregated spatial feature map to obtain the spatial modulator.

15 . The method of claim 14 , wherein the first set of dot products includes a dot product of a top-level spatial feature map and the first set of gating weights corresponding to a level above the multiple levels.

16 . The method of claim 13 , wherein aggregating the hierarchical contextual feature maps includes:

obtaining a dot product of the second set of gating weights and the level-specific temporal feature map corresponding to each level to generate a second set of dot products,

aggregating the second set of dot products to obtain an aggregated temporal feature map, and

applying a second linear function to the aggregated temporal feature map to obtain the temporal modulator.

17 . The method of claim 16 , wherein the second set of dot products includes a dot product of a top-level temporal feature map and the second set of gating weights corresponding to a level above the multiple levels.

18 . The method of claim 9 , wherein obtaining the output spatio-temporal feature map includes:

performing an element-wise multiplication between the query token, the spatial modulator and the temporal modulator.

19 . The method of claim 9 , wherein the spatial feature map is generated by transforming an input feature map corresponding to the video using a first linear function, and wherein the temporal feature map by transforming the input feature map using a second linear function.

20 . A system comprising:

a memory storing set of instructions; and

a processor configured to execute the set of instructions to cause the system to perform a method of:

obtaining a spatial feature map and a temporal feature map for a video;

generating hierarchical contextual feature maps based on the spatial feature map and the temporal feature map that represent a context of the video at multiple levels of granularity;

aggregating the hierarchical contextual feature maps based on gating weights to obtain a spatial modulator and a temporal modulator that are representative of an aggregated context across the multiple levels; and

obtaining an output spatio-temporal feature map based on the spatial modulator, the temporal modulator and a query token associated with the video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2024
From: WASIM, SYED TALAL; KHATTAK, MUHAMMAD UZAIR; NASEER, MUZAMMAL; KHAN, SALMAN; KHAN, FAHAD SHAHBAZ
To: MOHAMED BIN ZAYED UNIVERSITY OF ARTIFICIAL INTELLIGENCE
Reel/Frame 066129/0866 →
Continuity (1)
Related Publication 20250232583A1 · Jul 17, 2025
References Cited (10)
US 11361546B2 · Carreira · 2022 [cited by examiner]
US 20210248761A1 · Liu · 2021 [cited by examiner]
WO 2022111506A1 · 2022 [cited by applicant]
Li, L., Kong, T., Sun, F., Liu, H. (2019). Deep Point-Wise Prediction for Action Temporal Proposal. In: Gedeon, T., Wong, K., Lee, M. (eds) Neural Information Processing. ICONIP 2019. Lecture Notes in Computer Science( … [cited by examiner]
Z. Wei et al., “DMFormer: Closing the gap Between CNN and Vision Transformers,” ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1-5, d… [cited by examiner]
Jianwei Yang, Chunyuan Li, Xiyang Dai, Lu Yuan, Jianfeng Gao. Focal Modulation Networks. 36th Conference on Neural Information Processing Systems. arXiv:2203.11926. (Year: 2022). [cited by examiner]
Syed Talal Wasim, Muhammad Uzair Khattak, Muzammal Naseer, Salman Khan, Mubarak Shah, Fahad Shahbaz Khan. Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition. arXiv:2307.06947v4. (Year: 2023). [cited by examiner]
Fazry et al. ; Change Detection of High-Resolution Remote Sensing Images Through Adaptive Focal Modulation on Hierarchical Feature Maps ; IEEEAccess vol. 11 ; Jul. 5, 2023 ; 19 Pages. [cited by applicant]
Chen et al. ; Space-time video super-resolution using long-term temporal feature aggregation ; Autonomous Intelligent Systems ; 2023 ; 9 Pages. [cited by applicant]
Yang et al. ; A Spatio-Temporal Motion Network for Action Recognition Based on Spatial Attention ; MDPI Entropy, 24 ; Mar. 4, 2022 ; 19 Pages. [cited by applicant]