IP Library Granted Patent US 11,625,953
Granted Patent B2
US 11,625,953 · App. 16/903,527 · Granted Apr 11, 2023

Action recognition using implicit pose representations

Inventors: Philippe Weinzaepfel (Montbonnot-Saint-Martin, FR); Gregory Rogez (Gières, FR)
Assignee: NAVER CORPORATION
G06V40/20G06K9/628G06K9/6269
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,625,953
App. No.
16/903,527
Granted
Apr 11, 2023
Kind
B2
Abstract

A computer-implemented method of recognition of actions performed by individuals includes: by one or more processors, obtaining images including at least a portion of an individual; by the one or more processors, based on the images, generating implicit representations of poses of the individual in the images; and by the one or more processors, determining an action performed by the individual and captured in the images by classifying the implicit representations of the poses of the individual.

Claims (27)

1. A computer-implemented method of recognition of actions performed by individuals, the method comprising:

by one or more processors, obtaining images including at least a portion of an individual;

by the one or more processors, based on the images, generating implicit representations of poses of the individual in the images, the implicit representations including feature vectors and not including key points and not including 2 dimensional (2D) or three dimensional (3D) individual skeleton data; and

by the one or more processors, determining an action performed by the individual and captured in the images by classifying the implicit representations of the poses of the individual.

2. The method of claim 1 , further comprising deriving key points with three dimensional (3D) coordinates corresponding to a pose of a skeleton of the individual from the implicit representations.

3. The method of claim 1 , wherein the generating the implicit representations includes generating the implicit representations of the poses of the individual in the images using a first neural network.

4. The method of claim 3 , further comprising training the first neural network to determine at least one of two dimensional (2D) and three dimensional (3D) poses of individuals in images.

5. The method of claim 1 , wherein determining the action performed by the individual includes determining the action by classifying the implicit representations of the poses using a second neural network.

6. The method of claim 5 , further comprising training the second neural network to classify feature vectors according to actions by individuals.

7. The method of claim 1 , wherein the individual includes a human, the poses include human poses, and the action includes an action performed by a human.

8. The method of claim 1 , wherein the images include a sequence of images from a video.

9. The method of claim 1 , wherein:

generating the implicit representations includes generating the implicit representations based on the images, respectively;

the method further includes concatenating the implicit representations to produce a concatenated implicit representation; and

the determining the action includes performing a convolution on the concatenated implicit representation to produce a final implicit representation and determining the action performed by the individual by classifying the final implicit representation.

10. The method of claim 9 , wherein the convolution is a one dimensional (1D) convolution.

11. The method of claim 9 , wherein the convolution is a one dimensional (1D) temporal convolution.

12. The method of claim 1 further comprising, by the one or more processors, determining candidate boxes around individual captured in the images,

wherein generating the implicit representations includes generating the implicit representations based on the candidate boxes.

13. The method of claim 12 further comprising, by the one or more processors, extracting tubes from the images based on the candidate boxes.

14. The method of claim 12 , wherein determining the candidate boxes includes determining the candidate boxes using a regional proposal network (RPN) module.

15. A system, comprising:

one or more processors; and

memory including code that, when executed by the one or more processors, perform functions including:

obtaining images including at least a portion of an individual;

based on the images, generating implicit representations of poses of the individual in the images, the implicit representations including feature vectors and not including key points and not including 2 dimensional (2D) or three dimensional (3D) individual skeleton data; and

determining an action performed by the individual and captured in the images by classifying the implicit representations of the poses of the individual.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2020
From: WEINZAEPFEL, PHILIPPE; ROGEZ, GREGORY
To: NAVER CORPORATION
Reel/Frame 052960/0691 →
Priority Claims (1)
EP 19306098 · Sep 11, 2019 · regional
Continuity (1)
Related Publication 20210073525A1 · Mar 11, 2021