Stable pose estimation with analysis by synthesis
One embodiment of the present invention sets forth a technique for generating a pose estimation model. The technique includes generating one or more trained components included in the pose estimation model based on a first set of training images and a first set of labeled poses associated with the first set of training images, wherein each labeled pose includes a first set of positions on a left side of an object and a second set of positions on a right side of the object. The technique also includes training the pose estimation model based on a set of reconstructions of a second set of training images, wherein the set of reconstructions is generated by the pose estimation model from a set of predicted poses outputted by the one or more trained components.
1 . A computer-implemented method for generating a pose estimation model, the computer-implemented method comprising:
generating one or more trained components included in the pose estimation model based on one or more supervised losses computed from (i) a first set of predicted poses associated with a first set of training images and (ii) a first set of labeled poses associated with the first set of training images, wherein the first set of training images depict a first set of articulated objects against a first set of backgrounds;
generating, via execution of the one or more trained components, a set of predicted poses based on input that includes a second set of training images that depict a second set of articulated objects in a first set of poses against a second set of backgrounds;
generating, via execution of an image renderer included in the pose estimation model, a set of output images based on input that includes (i) the set of predicted poses and (ii) a set of reference images that depict the second set of articulated objects in a second set of poses against the second set of backgrounds; and
training the one or more trained components and the image renderer based on one or more unsupervised losses computed between (i) the set of output images and (ii) the second set of training images to generate a trained pose estimation model.
2 . The computer-implemented method of claim 1 , further comprising fine tuning the trained pose estimation model based on a third set of training images of a first object and the one or more unsupervised losses.
3 . The computer-implemented method of claim 1 , further comprising synthesizing the first set of training images and the first set of labeled poses prior to generating the one or more trained components.
4 . The computer-implemented method of claim 1 , further comprising further training the trained pose estimation model based on a third set of training images and a second set of labeled poses associated with the third set of training images.
5 . The computer-implemented method of claim 1 , further comprising applying the pose estimation model to a target image to estimate a first set of positions on a left side of a first object depicted within the target image and a second set of positions on a right side of the first object depicted within the target image.
6 . The computer-implemented method of claim 1 , wherein the one or more trained components comprise an image encoder that generates a skeleton image from an input image, and wherein the skeleton image comprises a plurality of channels that indicate different sets of pixel locations for different parts of an object within the input image.
7 . The computer-implemented method of claim 6 , wherein the one or more trained components further comprise a pose estimator that converts the skeleton image into a first set of pixel locations associated with a first set of joint positions and a second set of pixel locations associated with a second set of joint positions.
8 . The computer-implemented method of claim 7 , wherein the one or more trained components further comprise an uplift model that converts the first set of pixel locations and the second set of pixel locations into a set of three-dimensional (3D) coordinates.
9 . The computer-implemented method of claim 8 , wherein the image renderer generates an output image included in the set of output images based on input that includes (i) a projection of the set of 3D coordinates onto pixel locations in an analytic skeleton image and (ii) a reference image included in the set of reference images.
10 . The computer-implemented method of claim 1 , wherein the first set of training images comprises a set of synthetic images and the second set of training images comprises a set of non-rendered images that are captured by cameras.
11 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:
generating one or more trained components included in a pose estimation model based on one or more supervised losses computed from (i) a first set of predicted poses associated with a first set of training images that depict a first set of articulated objects against a first set of backgrounds and (ii) a first set of labeled poses associated with the first set of training images; and
training the one or more trained components and an image renderer included in the pose estimation model based on one or more unsupervised losses computed between (i) a second set of training images that depict a second set of articulated objects against a second set of backgrounds and (ii) a set of reconstructions of the second set of training images to generate a trained pose estimation model, wherein the set of reconstructions is generated by the image renderer from a set of predicted poses outputted by the one or more trained components.
12 . The one or more non-transitory computer-readable media of claim 11 , wherein the instructions further cause the one or more processors to perform the step of fine tuning the trained pose estimation model based on a third set of training images of a first object.
13 . The one or more non-transitory computer-readable media of claim 11 , wherein the instructions further cause the one or more processors to perform the step of synthesizing the first set of training images and the first set of labeled poses prior to generating the one or more trained components.
14 . The one or more non-transitory computer-readable media of claim 11 , wherein generating the one or more trained components comprises training an image encoder that generates a skeleton image from an input image based on an error between a set of limbs included in the skeleton image and a ground truth pose associated with the input image.
15 . The one or more non-transitory computer-readable media of claim 14 , wherein training the pose estimation model comprises further training the image encoder based on a discriminator loss associated with the input image and a set of unpaired poses.
16 . The one or more non-transitory computer-readable media of claim 11 , wherein generating the one or more trained components comprises training a pose estimator based on one or more errors between a predicted pose generated by the pose estimator from an input image and a ground truth pose for the input image.
17 . The one or more non-transitory computer-readable media of claim 11 , wherein training the pose estimation model comprises training the image renderer based on one or more losses associated with a reconstruction of a first image of a first object generated by the image renderer, wherein the reconstruction is generated by the image renderer based on a predicted pose associated with the first image and a second input image of the first object.
18 . The one or more non-transitory computer-readable media of claim 17 , wherein the one or more losses comprise at least one of a perceptual loss, a discriminator loss, or a discriminator feature matching loss.
19 . The one or more non-transitory computer-readable media of claim 11 , wherein the first set of labeled poses comprises a first set of joints on a left side of an object and a second set of joints on a right side of the object.
20 . A system, comprising:
one or more memories that store instructions, and
one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to:
execute a trained pose estimation model based on an input image, wherein the trained pose estimation model is generated by:
generating one or more trained components included in a pose estimation model based on one or more supervised losses computed from (i) a first set of predicted poses associated with a first set of training images that depict a first set of articulated objects against a first set of backgrounds and (ii) a first set of labeled poses associated with the first set of training images; and
training the one or more trained components and an image renderer included in the pose estimation model based on one or more unsupervised losses computed between (i) a second set of training images that depict a second set of articulated objects against a second set of backgrounds and (ii) a set of reconstructions of the second set of training images, wherein the set of reconstructions is generated by the image renderer from a set of predicted poses outputted by the one or more trained components; and
receive, as output of the one or more trained components, one or more poses associated with an object depicted in the input image.