Method for video recognition capable of encoding spatial and temporal relationships of concepts using contextual features
The proposed invention aims at encoding contextual information for video analysis and understanding, by encoding spatial and temporal relationships of objects and the main agent in a scene. The main target application of the invention is human activity recognition. The encoding of such spatial and temporal relationships may be crucial to distinguish different categories of human activities and may be important to help in the discrimination of different video categories, aiming at video classification, retrieval, categorization and other video analysis applications.
1. A method for video recognition using contextual features capable of encoding spatial and temporal relationships of concepts, the method comprising performing, by at least one processor, operations including:
acquiring input video data from a video;
processing the input video data to detect concepts in the video;
computing contextual features from the detected concepts, wherein the computing contextual features includes:
computing, by the Egocentric Pyramid, spatial relationships of detected concepts in relation to a main agent of the video as concept-agent pairings;
computing pairings between concepts as concept-concept pairings; and
making use of the computed pairings to determine temporal relationships of the concepts, using the Temporal Egocentric Relational Network, to generate prediction scores for the concepts; and
outputting the generated prediction scores from the Temporal Egocentric Relational Network.
2. The method according to claim 1 , wherein the acquiring input video data comprises splitting the video into t video segments of equal size T and then, from each video segment, sampling a random snippet S i with length |S i | such that |S i |≤T.
3. The method according to claim 1 , wherein the computing contextual features from the detected concepts includes attributing scores to captured context to determine the concepts and agents in the video.
4. The method according to claim 3 , wherein the Egocentric Pyramid considers as the main agent in the video to be the concept with the highest attributed score obtained by the detected concepts.
5. The method according to claim 1 , wherein when more than one agent is in the video, a number of agents is a same number of Egocentric Pyramids, where each Egocentric Pyramid is considered a separate concept as the agent in the video.
6. The method according to claim 1 , wherein the Temporal Egocentric Relational Network determines the temporal relationships from both egocentric pairings and concept pairings.
7. The method according to claim 1 , wherein the Temporal Egocentric Relational Network determines the temporal relationships from egocentric pairings.
8. The method according to claim 1 , wherein the Temporal Egocentric Relational Network determines the temporal relationships from concept pairings.
9. The method according to claim 1 , wherein the Temporal Egocentric Relational Network uses the computed pairings to determine features and a classifier in a unified way.
10. The method according to claim 1 , wherein the Temporal Egocentric Relational Network is configured to determine concept information over time.
11. The method according to claim 1 , wherein the Temporal Egocentric Relational Network is defined as:
TERN( S )= ( R Φ ( S 1 ), R Φ ( S 2 ), . . . , R Φ ( S t )),
where S t is the video, R Φ is a relational network with parameters Φ, is a pooling operation and the relational network R Φ , given parameters Φ=[ϕ 1 ,ϕ 2 ], is defined as R Φ (O)=ƒ ϕ 1 (1/n 2 Σ o i ,o j g ϕ 2 (o i , o j )), where O={o i } i=1 n represents an input set of n detected concepts (e.g., objects), where o i is the i-th concept such that o i ∈ ƒ ; and functions ƒ ϕ 1 and g ϕ 2 are stacked multi-layer perceptrons (MLP) parameterized by parameters ϕ 1 and ϕ 2 , respectively.