IP Library Granted Patent US 12,586,413
Granted Patent B2
US 12,586,413 · App. 17/754,685 · Granted Mar 24, 2026

Method for recognizing activities using separate spatial and temporal attention weights

Inventors: Gianpiero Francesca (Brussels, BE); Luca Minciullo (Brussels, BE); Lorenzo Garattoni (Brussels, BE); Srijan Das (Nice, FR); Rui Dai (Biot, FR); Francois Bremond (Villeneuve Loubet, FR)
Assignee: TOYOTA JIDOSHA KABUSHIKI KAISHA
G06V40/20G06N3/0442G06N3/045G06N3/0464G06V10/462G06V10/82G06V20/41G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,413
App. No.
17/754,685
Granted
Mar 24, 2026
Kind
B2
Abstract

A device and a method for recognizing person activity in a sequence of frames ( 100 ) comprising: obtaining a set of consecutives 3D poses ( 103 ), obtaining a feature map ( 102 ), obtaining a vector of spatiotemporal features, obtaining a matrix of spatial attention weights, obtaining a matrix of temporal attention weights ( 110 ), modulating ( 106 ) the feature map using the matrix of spatial attention weights to obtain a spatially modulated feature map, modulating ( 111 ) the feature map using the vector of temporal attention weights to obtain a temporally modulated feature map, performing a convolution ( 114 ) of the spatially modulated feature map and of the temporally modulated feature map to obtain a convoluted feature map, performing a classification ( 115 ) using the convoluted feature map so as to determine the activity of the person in the video.

Claims (34)

1 . A method for recognizing person activity in a video comprising a sequence of frames, each frame showing at least a portion of the person, the method comprising:

obtaining a set of consecutive 3D poses of the person using the sequence of frames, each of the consecutive 3D poses illustrating a posture of the person from a frame of the sequence of frames, and each of the consecutive 3D poses being associated with an instant in the sequence of frames,

obtaining a feature map elaborated using a first encoder neural network configured to receive the sequence of frames as input and to output the feature map having dimensions associated with time, space, and a number of channels,

obtaining a vector of spatiotemporal features using a second recurrent neural network configured to receive the set of consecutive 3D poses of the person as input,

a third neural network receiving the vector of spatiotemporal features as input and outputting a matrix of spatial attention weights, wherein each weight indicates an importance of a location in the matrix, wherein the third neural network comprises a first fully connected layer, a hyperbolic tangent layer, a second fully connected layer, and a sigmoid layer,

a fourth neural network, different from the third neural network, the fourth neural network receiving the vector of spatiotemporal features as input and outputting a matrix of temporal attention weights, wherein each weight indicates a saliency of an instant in the sequence of frames, wherein the fourth neural network comprises a first fully connected layer, a hyperbolic tangent layer, a second fully connected layer, and a Softmax layer,

obtaining a spatially-modulated feature map by modulating the feature map using the matrix of spatial attention weights,

obtaining a temporally-modulated feature map, different from the spatially-modulated feature map, by modulating the feature map using the matrix of temporal attention weights,

performing a convolution of the spatially modulated feature map and of the temporally modulated feature map to obtain a convoluted feature map,

performing a classification using the convoluted feature map so as to determine the activity of the person in the video.

2 . The method according to claim 1 , wherein the first encoder neural network includes a portion of an inflated 3D convolutional neural network.

3 . The method according to claim 1 , further comprising, prior to the performing the convolution:

performing a Global Average Pooling on the spatially-modulated feature map; and

performing a Global Average Pooling on the temporally modulated feature map.

4 . The method according to claim 1 , wherein the performing the convolution comprises performing a 1×1×1 convolution.

5 . The method according to claim 1 , wherein the performing the classification comprises using a Softmax function.

6 . The method according to claim 1 , wherein each of the consecutive 3D poses comprises a set of 3D coordinates (x_j) indicating positions of joints of a given skeleton.

7 . The method according to claim 1 , further comprising:

a preliminary training step of at least one of the first encoder neural network, the second recurrent neural network, the third neural network, and the fourth neural network.

8 . The method according to claim 1 , further comprising:

a preliminary training step comprising:

determining a loss using a cross-entropy loss, determining a loss based on the matrix of spatial attention weights, and determining a loss based on the matrix of temporal attention weights.

9 . A device for recognizing person activity in a video comprising a sequence of frames, each frame showing at least a portion of the person, the device comprising:

a module for obtaining a set of consecutive 3D poses of the person using the sequence of frames, each of the consecutive 3D poses illustrating a posture of the person from a frame of the sequence of frames, and each of the consecutive 3D poses being associated with an instant in the sequence of frames,

a first encoder neural network configured to receive the sequence of frames as input and to output the feature map having dimensions associated with time, space, and a number of channels,

a second neural network configured to receive the set of consecutive 3D poses of the person as input and to output a vector of spatiotemporal features,

a third neural network configured to receive the vector of spatiotemporal features as input and to output a matrix of spatial attention weights, wherein each weight indicates an importance of a location in the matrix, wherein the third neural network comprises a first fully connected layer, a hyperbolic tangent layer, a second fully connected layer, and a sigmoid layer,

a fourth neural network, different from the third neural network, configured to receive the vector of spatiotemporal features as input and to output a matrix of temporal attention weights, wherein each weight indicates a saliency of an instant in the sequence of frames, wherein the fourth neural network comprises a first fully connected layer, a hyperbolic tangent layer, a second fully connected layer, and a Softmax layer,

a module for obtaining a spatially-modulated feature map by modulating the feature map using the matrix of spatial attention weights,

a module for obtaining a temporally-modulated feature map, different from the spatially-modulated feature map, by modulating the feature map using the matrix of temporal attention weights,

a module for performing a convolution of the spatially modulated feature map and of the temporally modulated feature map to obtain a convoluted feature map,

a module for performing a classification using the convoluted feature map so as to determine the activity of the person in the video.

10 . A system comprising the device of claim 9 and comprising a video acquisition module configured to obtain the video.

11 . A non-transitory computer-readable medium comprising instructions stored thereon that when executed by a processor cause the processor to execute instructions for executing the steps of the method according to claim 1 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2024
From: TOYOTA MOTOR EUROPE
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 068672/0268 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2023
From: FRANCESCA, GIANPIERO; MINCIULLO, LUCA; GARA TTONI, LORENZO; DAS, SRIJAN; DAI, RUI; BREMOND, FRANCOIS
To: TOYOTA MOTOR EUROPE
Reel/Frame 062821/0087 →
Continuity (1)
Related Publication 20230134967A1 · May 4, 2023
References Cited (9)
US 20200074227A1 · Lan · 2020 [cited by examiner]
Yun (“Two-person interaction detection using body-pose features and multiple instance learning”, IEEE Computer Society Conference, Jun. 2012, pp. 1-8) (Year: 2012). [cited by examiner]
Hou Jingxuan et al: “Spatial-Temporal Attention Res-TCN for Skeleton-Based Dynamic Hand Gesture Recognition”, Jan. 23, 2019 (Jan. 23, 2019), Robocup 2008: Robot Soccer World Cup XII ; [Lecture Notes in Computer Science]… [cited by applicant]
Sijie Song et al: “An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 18, 20… [cited by applicant]
Chan Wensong et al: “Select and Focus: Action Recognition with Spatial-Temporal Attention”, Aug. 2, 2019 (Aug. 2, 2019), Robocup 2008: Robot Soccer World Cup XII; [Lecture Notes in Computer Science], Springer Internatio… [cited by applicant]
Fabien Baradel et al: “Human Action Recognition: Pose-based Attention draws focus to Hands”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Dec. 20, 2017 (Dec. 20, 2017), XP… [cited by applicant]
Joao Carreira et al: “Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset”, arXIV.1705.07750v3 [cs.CV] Feb. 12, 2018. [cited by applicant]
Christoph Feichtenhofer et al: “SlowFast Networks for Video Recognition”, arXIV.1812.03982v3 [cs.CV] Oct. 29, 2019. [cited by applicant]
Fabien Baradel et al: “Glimpse Clouds: Human Activity Recognition from Unstructured Feature Points”, arXIV.1802.07898v4 [cs.CV] Aug. 21, 2018. [cited by applicant]