IP Library › Granted Patent US 11,416,774
Granted Patent B2
US 11,416,774 · App. 16/849,350 · Granted Aug 16, 2022

Method for video recognition capable of encoding spatial and temporal relationships of concepts using contextual features

Inventors: Jesimon Barreto Santos (Minas Gerais, BR); Victor Hugo Cunha de Melo (Minas Gerais, BR); William Robson Schwartz (Minas Gerais, BR); Otávio Augusto Bizetto Penatti (São Paulo, BR)
Assignees: SAMSUNG ELECTRONICA DA AMAZONIA LTDA.; UNIVERSIDADE FEDERAL DE MINAS GERAIS-UFMG
G06N20/00G06F16/583G06K9/6215
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,416,774
App. No.
16/849,350
Granted
Aug 16, 2022
Kind
B2
Abstract

The proposed invention aims at encoding contextual information for video analysis and understanding, by encoding spatial and temporal relationships of objects and the main agent in a scene. The main target application of the invention is human activity recognition. The encoding of such spatial and temporal relationships may be crucial to distinguish different categories of human activities and may be important to help in the discrimination of different video categories, aiming at video classification, retrieval, categorization and other video analysis applications.

Claims (20)

1. A method for video recognition using contextual features capable of encoding spatial and temporal relationships of concepts, the method comprising performing, by at least one processor, operations including:

acquiring input video data from a video;

processing the input video data to detect concepts in the video;

computing contextual features from the detected concepts, wherein the computing contextual features includes:

computing, by the Egocentric Pyramid, spatial relationships of detected concepts in relation to a main agent of the video as concept-agent pairings;

computing pairings between concepts as concept-concept pairings; and

making use of the computed pairings to determine temporal relationships of the concepts, using the Temporal Egocentric Relational Network, to generate prediction scores for the concepts; and

outputting the generated prediction scores from the Temporal Egocentric Relational Network.

2. The method according to claim 1 , wherein the acquiring input video data comprises splitting the video into t video segments of equal size T and then, from each video segment, sampling a random snippet S i with length |S i | such that |S i |≤T.

3. The method according to claim 1 , wherein the computing contextual features from the detected concepts includes attributing scores to captured context to determine the concepts and agents in the video.

4. The method according to claim 3 , wherein the Egocentric Pyramid considers as the main agent in the video to be the concept with the highest attributed score obtained by the detected concepts.

5. The method according to claim 1 , wherein when more than one agent is in the video, a number of agents is a same number of Egocentric Pyramids, where each Egocentric Pyramid is considered a separate concept as the agent in the video.

6. The method according to claim 1 , wherein the Temporal Egocentric Relational Network determines the temporal relationships from both egocentric pairings and concept pairings.

7. The method according to claim 1 , wherein the Temporal Egocentric Relational Network determines the temporal relationships from egocentric pairings.

8. The method according to claim 1 , wherein the Temporal Egocentric Relational Network determines the temporal relationships from concept pairings.

9. The method according to claim 1 , wherein the Temporal Egocentric Relational Network uses the computed pairings to determine features and a classifier in a unified way.

10. The method according to claim 1 , wherein the Temporal Egocentric Relational Network is configured to determine concept information over time.

11. The method according to claim 1 , wherein the Temporal Egocentric Relational Network is defined as:

TERN( S )= ( R Φ ( S 1 ), R Φ ( S 2 ), . . . , R Φ ( S t )),

where S t is the video, R Φ is a relational network with parameters Φ, is a pooling operation and the relational network R Φ , given parameters Φ=[ϕ 1 ,ϕ 2 ], is defined as R Φ (O)=ƒ ϕ 1 (1/n 2 Σ o i ,o j g ϕ 2 (o i , o j )), where O={o i } i=1 n represents an input set of n detected concepts (e.g., objects), where o i is the i-th concept such that o i ∈ ƒ ; and functions ƒ ϕ 1 and g ϕ 2 are stacked multi-layer perceptrons (MLP) parameterized by parameters ϕ 1 and ϕ 2 , respectively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2020
From: SANTOS, JESIMON BARRETO; CUNHA DE MELO, VICTOR HUGO; SCHWARTZ, WILLIAM ROBSON; PENATTI, OTÁVIO AUGUSTO BIZETTO
To: SAMSUNG ELETRÔNICA DA AMAZÔNIA LTDA.; UNIVERSIDADE FEDERAL DE MINAS GERAIS - UFMG
Reel/Frame 052822/0599 →
Priority Claims (2)
BR 10 2019 022207 7 · Oct 23, 2019 · national
BR 10 2019 024569 7 · Nov 21, 2019 · national
Continuity (1)
Related Publication 20210125100A1 · Apr 29, 2021