IP Library Granted Patent US 12,033,436
Granted Patent B2
US 12,033,436 · App. 17/944,418 · Granted Jul 9, 2024

Methods and apparatus for human pose estimation from images using dynamic multi-headed convolutional attention

Inventors: Alec Diaz-Arias (Columbia, MO); Dmitriy Shin (Columbia, MO); Jean E. Robillard (Iowa City, IA); Mitchell Messmore (Iowa City, IA); John Rachid (Iowa City, IA)
Assignee: INSEER Inc.
G06V40/23G06V10/806G06V40/103G16H20/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,033,436
App. No.
17/944,418
Granted
Jul 9, 2024
Kind
B2
Abstract

An apparatus for 3D human pose estimation using dynamic multi-headed convolutional attention mechanism is presented. The apparatus contains two dynamic multi-headed convolutional attention mechanism with spatial attention and another with temporal attention that leverages the spatial attention mechanism to extract frame-wise inter-joint dependencies by analyzing sections of limbs that are related. The temporal attention mechanism extracts global inter-frame relationships by analyzing correlations between the temporal profile of joints. The temporal profile mechanism leads to a more diverse temporal attention map while achieving substantial parameter reduction.

Claims (41)

1. An apparatus, comprising:

at least a processor; and

a memory operatively coupled to the processor, the memory storing instructions that when executed cause the processor to:

receive a plurality of image frames, each image frame from the plurality of image frames containing a measured temporal joint data of a subject;

receive a plurality of joint localization overlays from a spatial joint machine-learning model;

execute a trained limb segment machine-learning model to identify a plurality of frame interrelations using the plurality of image frames and the plurality of joint localization overlays as an input, the trained limb segment machine-learning model trained using an interrelation training set that contains limb segment image data correlated to a limb matrix in a motion sequence; and

generate a temporal joints profile based on the plurality of frame interrelation.

2. The apparatus of claim 1 , wherein the instruction to cause the processor to generate the temporal joints profile is based on a plurality of video inputs from a plurality of sensors pointing at the subject from different angles, the plurality of sensors including the sensor.

3. The apparatus of claim 1 , wherein the memory further stores instructions that when executed, cause the processor to extract a plurality of scaled frame interrelations based on a convolutional filter size prior to generating the temporal joints profile.

4. The apparatus of claim 1 , wherein the memory further stores instructions to cause the processor to independently generate a plurality of temporal joints profiles for a plurality of subjects captured by the sensor.

5. A non-transitory, processor-readable medium storing processor-executable instructions to cause the processor to:

receive, from a sensor operatively coupled to the processor, a plurality of image frames containing a measured temporal joint data of a subject across at least two frames of image data from the sensor;

receive the plurality of joint localization overlays from the spatial joint machine-learning model;

execute a trained limb segment machine-learning model to identify a plurality of frame interrelations using the plurality of image frames and the plurality of joint localization overlays as an input, the trained limb segment machine-learning model trained using an interrelation training set that contains limb segment image data correlated to a limb matrix in a motion sequence;

compute a plurality of averaged convoluted quantitative identifiers based on the plurality of image frames;

execute a trained multi-headed temporal profile machine-learning model to output a plurality of temporal joints interrelations using the plurality of averaged convoluted quantitative identifiers and the plurality of frame interrelations as an input, the trained multi-headed temporal profile machine-learning mode using a temporal joints training set; and

generate an aggregated temporal joints profile based on the temporal joints interrelations.

6. The non-transitory, processor-readable medium of claim 5 , the non-transitory, processor-readable medium further causing the processor to receive, from a plurality of sensors, a plurality of video inputs, each sensor from the plurality of sensors pointing to the at least a subject from a different angle.

7. The non-transitory, processor-readable medium of claim 5 , the non-transitory, processor-readable medium further stores instructions to cause the processor to extract a plurality of scaled frame interrelations based on a convolutional filter size prior.

8. The non-transitory, processor-readable medium of claim 5 , wherein the non-transitory, processor-readable medium further stores instructions to cause the processor to compute the plurality of averaged convoluted quantitative identifiers and generate the aggregated temporal joints profile simultaneously.

9. The non-transitory, processor-readable medium of claim 5 , wherein the non-transitory, processor-readable medium further stores instructions to cause the processor to:

generate the plurality of frame interrelations based on the plurality of averaged convoluted quantitative identifiers; and

reduce sparsity of a plurality of joint interrelations between the plurality of frame interrelations temporal joints model using the plurality of averaged convoluted quantitative identifiers.

10. The non-transitory, processor-readable medium of claim 5 , wherein the non-transitory, processor-readable medium further stores instructions to cause the processor to map the plurality of averaged convoluted quantitative identifiers to an output attention matrix of a plurality of output attention matrices, each output attention matrix representing a sequential pose prior to generating the aggregated temporal joints profile.

11. A non-transitory, processor-readable medium storing processor-executable instructions to cause the processor to:

receive a temporal joints training set that includes a concatenated temporal profile correlated to a concatenated temporal sequence;

train a multi-headed temporal profile machine-learning model of a plurality of machine-learning models using the temporal joints training set,

the multi-headed temporal profile machine-learning model, when executed, outputting a plurality of temporal joints interrelations using as an input a plurality of averaged convoluted quantitative identifiers computed based on a plurality of image frames containing a measured temporal joint data of a subject across at least two frames of image data from a sensor operatively coupled to the processor;

receive a plurality of joint localization overlays from a spatial joint machine-learning model; and

train a limb segment machine-learning model from the plurality of machine-learning models using an interrelation training set that contains a limb segment image data correlated to a limb matrix in a motion sequence,

the limb segment machine-learning model, when executed, uses the plurality of joint localization overlays as an input and outputs a plurality of frame interrelations.

12. The non-transitory, processor-readable medium of claim 11 , the non-transitory, processor-readable medium further causing the processor to:

train the spatial joint machine-learning model of the plurality of machine-learning models using a spatial joints training set containing a two-dimensional human pose correlated to a two-dimensional joints profile,

the spatial joint machine-learning model, when executed, using the plurality of image frames as an input, the spatial joint machine-learning model outputs the plurality of joint localization overlays.

13. The non-transitory, processor-readable medium of claim 11 , the non-transitory, processor-readable medium further causing the processor to:

receive the plurality of frame interrelations from the limb segment machine-learning model; and

train a temporal profile machine-learning model of the plurality of machine-learning models using a temporal sequence training set containing a temporal pose sequence correlated to a temporal joint sequence,

the temporal profile machine-learning model, when executed, using the plurality of frame interrelations as an input and outputting a temporal joints profile.

14. The non-transitory, processor-readable medium of claim 11 , wherein the plurality of machine-learning models includes deep learning.

15. The non-transitory, processor-readable medium of claim 11 , wherein each image frame from the plurality of image frames contains a measured temporal joint data of the subject.

16. The non-transitory, processor-readable medium of claim 11 , wherein the plurality of image frames are received from a plurality of sensors that point at the subject from a plurality of different angles.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: DIAZ-ARIAS, ALEC; SHIN, DMITRIY; ROBILLARD, JEAN E.; MESSMORE, MITCHELL; RACHID, JOHN
To: INSEER INC.
Reel/Frame 062084/0854 →
Continuity (2)
Continuation 17740650 · May 10, 2022
Related Publication 20230368578A1 · Nov 16, 2023