Recovery of 3D human pose by jointly learning metrics and mixtures of experts
View Patent ↗Systems and methods are disclosed for determining 3D human pose by generating an Appearance and Position Context (APC) local descriptor that achieves selectivity and invariance while requiring no background subtraction; jointly learning visual words and pose regressors in a supervised manner; and estimating the 3D human pose.
1. A method to determine a 3D human pose, comprising:
a. learning visual words for human pose estimation through a supervised method;
b. deriving a separate metric for each visual word from labeled image-to-pose pairs through supervised learning;
c. representing a multi-modal distribution of the 3D human pose space conditioned on a feature space with a Bayesian mixture of experts (BME) model; and
d. jointly optimizing metric learning and the BME model by an iterative gradient ascent process.
2. The method of claim 1 , comprising obtaining visual words by an unsupervised clustering operation.
3. The method of claim 1 , comprising learning an individual distance metric for each visual word.