IP Library › Granted Patent US 11,790,682
Granted Patent B2
US 11,790,682 · App. 17/388,260 · Granted Oct 17, 2023

Image analysis using neural networks for pose and action identification

Inventors: Razwan Ghafoor (Sutton Coldfield, GB); Peter Rennert (Sutton Coldfield, GB); Hichame Moriceau (London, GB)
Assignee: STANDARD COGNITION, CORP.
G06V40/103
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,790,682
App. No.
17/388,260
Filed
Jul 29, 2021
Granted
Oct 17, 2023
Kind
B2
Examiner
HON, MING Y
Art Unit
2666
USPC
382/156
Abstract

An apparatus for performing image analysis to identify human actions represented in an image, comprising: a joint-determination module configured to analyse an image depicting one or more people using a first computational neural network to determine a set of joint candidates for the one or more people depicted in the image; a pose-estimation module configured to derive pose estimates from the set of joint candidates that estimate a body configuration for the one or more people depicted in the image; and an action-identification module configured to analyse a region of interest within the image identified from the derived pose estimates using a second computational neural network to identify an action performed by a person depicted in the image.

Claims (32)

1. An apparatus for performing image analysis to identify human actions represented in an image, comprising:

a joint-determination module configured to analyse an image depicting one or more people using a first computational neural network to determine a set of joint candidates for the one or more people depicted in the image,

wherein the first computational neural network processes, as input, the image on a pixel-by-pixel basis to generate, an output vector containing a number of elements, each element within the number of elements storing a probability that a respective pixel within the image depicts a joint, and wherein the output vector is evaluated to determine the set of joint candidates;

a pose-estimation module configured to derive pose estimates from the set of joint candidates that estimate a body configuration for the one or more people depicted in the image; and

an action-identification module configured to analyse a region of interest within the image identified from the derived pose estimates using a second computational neural network to identify an action performed by a person depicted in the image.

2. An apparatus as claimed in claim 1 , wherein the region of interest defines a sub-region of the image, and action-identification module is configured to analyse only the sub- region of the image using the second computational neural network.

3. An apparatus as claimed in claim 1 , wherein the action-identification module is configured to analyse the region of interest to identify objects of a specified object class, and to identify the action in response to detecting an object of the specified class in the region of interest.

4. An apparatus as claimed in claim 1 , wherein the apparatus further comprises an image-region module configured to identify the region of interest within the image from the derived pose estimates.

5. An apparatus as claimed in claim 4 , wherein the region of interest bounds a specified subset of joints of a derived pose estimate.

6. An apparatus as claimed in claim 4 , wherein the region of interest bounds one or more derived pose estimates.

7. An apparatus as claimed in claim 4 , wherein the image-region module is configured to identify a region of interest that bounds terminal ends of a derived pose estimate, and the action-identification module is configured to analyse the identified region of interest using the second computational neural network to identify whether the person depicted in the image is holding an object of a specified class or not.

8. An apparatus as claimed in claim 1 , wherein the action-identification module is configured to identify the action from a class of actions including: scanning an item at a point-of-sale; and selecting an item for purchase.

9. An apparatus as claimed in claim 1 , wherein the first network is a convolutional neural network.

10. An apparatus as claimed in claim 1 , wherein the second network is a convolutional neural network.

11. An apparatus for performing image analysis to identify human actions represented in a sequence of images, comprising:

a joint-determination module configured to analyse the sequence of images depicting one or more people using a first computational neural network to determine a set of joint candidates for the one or more people depicted in the sequence of images;

a pose-estimation module configured to derive pose estimates from the set of joint candidates that estimate a body configuration for the one or more people depicted in the sequence of images; and

an action-identification module configured to analyse the derived pose estimates using a second computational neural network to identify an action performed by a person depicted in the sequence of images, wherein the second computational neural network is a recursive neural network.

12. An apparatus for performing image analysis to identify human actions represented in an image, comprising:

a joint-determination module configured to analyse the image depicting one or more people using a first computational neural network to determine a set of joint candidates for the one or more people depicted in the image;

a pose-estimation module configured to derive pose estimates from the set of joint candidates that estimate a body configuration for the one or more people depicted in the image,

the pose-estimation module further comprising a joint classifier configured to calculate, for a particular joint candidate, a set of probabilities, wherein a respective probability, within the set of probabilities, corresponds to a likelihood that the particular joint candidate is a particular body part from a set of body part classes; and

an action-identification module configured to analyse the derived pose estimates using a second computational neural network to identify an action performed by a person depicted in the image.

13. An apparatus as claimed in claim 12 , the apparatus further comprising an extractor module configured to extract from each derived pose estimate values for a set of one or more parameters characterising the pose estimate, wherein the action-identification module is configured to use the second computational neural network to identify an action performed by a person depicted in the one or more images in dependence on the extracted parameter values.

14. An apparatus as claimed in claim 13 , wherein the set of one or more parameters relate to a specified subset of joints of the pose estimate.

15. An apparatus as claimed in claim 13 , wherein the set of one or more parameters comprises at least one of: joint position for specified joints; joint angles between specified connected joints; joint velocity for specified joints; and the distance between specified pairs of joints.

16. An apparatus as claimed in claim 13 , wherein the one or more images is a series of multiple images, and the action-identification module is configured to identify the action performed by the person depicted in the series of images from changes in their derived pose estimate over the series of images and wherein the action-identification module is configured to use the second computational neural network to identify an action performed by the person from the change in the extracted parameter values characterising their derived pose estimate over the series of images.

17. An apparatus as claimed in claim 12 , wherein the one or more images is a series of multiple images, and the action-identification module is configured to identify the action performed by the person depicted in the series of images from changes in their derived pose estimate over the series of images.

18. An apparatus as claimed in claim 12 , wherein each image of the one or more images depicts a plurality of people, and the action-identification module is configured to analyse the derived pose estimates for each of the plurality of people for the one or more images using the second computational neural network and to identify an action performed by each person depicted in the one or more images.

19. An apparatus as claimed in claim 12 , wherein the pose-estimate module is configured to derive the pose estimates from the set of joint candidates and further from imposed anatomical constraints on the joints.

20. An apparatus as claimed in claim 12 , wherein:

the joint classifier is configured to calculate a set of conditional probabilities, wherein a conditional probability, within the set of conditional probabilities, corresponds to a likelihood that the set of joint candidates belongs to a same person given that candidate joints, of the set of candidate joints, belong to corresponding body part classes.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2022
From: GHAFOOR, RAZWAN; RENNERT, PETER; MORICEAU, HICHAME
To: STANDARD COGNITION, CORP.
Reel/Frame 058732/0728 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2021
From: THIRDEYE LABS LIMITED
To: STANDARD COGNITION, CORP.
Reel/Frame 057920/0421 →
Priority Claims (1)
GB 1703914 · Mar 10, 2017 · national
Continuity (2)
Continuation 16492781
Related Publication 20220012478A1 · Jan 13, 2022
Cited By (2)
US 12,670,740 US 12,711,808