IP Library › Granted Patent US 12,725,289
Granted Patent B2
US 12,725,289 · App. 17/748,398 · Granted Sep 1, 2026

Stable pose estimation with analysis by synthesis

Inventors: Martin Guay (Zurich, CH); Dominik Tobias Borer (Zurich, CH); Jakob Joachim Buhmann (Zurich, CH)
Assignees: Disney Enterprises, Inc.; ETH Zürich (Eidgenössische Technische Hochschule Zürich)
G06T7/70G06N20/00G06T9/002G06V10/774G06T2207/20081G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,289
App. No.
17/748,398
Granted
Sep 1, 2026
Kind
B2
Abstract

One embodiment of the present invention sets forth a technique for generating a pose estimation model. The technique includes generating one or more trained components included in the pose estimation model based on a first set of training images and a first set of labeled poses associated with the first set of training images, wherein each labeled pose includes a first set of positions on a left side of an object and a second set of positions on a right side of the object. The technique also includes training the pose estimation model based on a set of reconstructions of a second set of training images, wherein the set of reconstructions is generated by the pose estimation model from a set of predicted poses outputted by the one or more trained components.

Claims (32)

1 . A computer-implemented method for generating a pose estimation model, the computer-implemented method comprising:

generating one or more trained components included in the pose estimation model based on one or more supervised losses computed from (i) a first set of predicted poses associated with a first set of training images and (ii) a first set of labeled poses associated with the first set of training images, wherein the first set of training images depict a first set of articulated objects against a first set of backgrounds;

generating, via execution of the one or more trained components, a set of predicted poses based on input that includes a second set of training images that depict a second set of articulated objects in a first set of poses against a second set of backgrounds;

generating, via execution of an image renderer included in the pose estimation model, a set of output images based on input that includes (i) the set of predicted poses and (ii) a set of reference images that depict the second set of articulated objects in a second set of poses against the second set of backgrounds; and

training the one or more trained components and the image renderer based on one or more unsupervised losses computed between (i) the set of output images and (ii) the second set of training images to generate a trained pose estimation model.

2 . The computer-implemented method of claim 1 , further comprising fine tuning the trained pose estimation model based on a third set of training images of a first object and the one or more unsupervised losses.

3 . The computer-implemented method of claim 1 , further comprising synthesizing the first set of training images and the first set of labeled poses prior to generating the one or more trained components.

4 . The computer-implemented method of claim 1 , further comprising further training the trained pose estimation model based on a third set of training images and a second set of labeled poses associated with the third set of training images.

5 . The computer-implemented method of claim 1 , further comprising applying the pose estimation model to a target image to estimate a first set of positions on a left side of a first object depicted within the target image and a second set of positions on a right side of the first object depicted within the target image.

6 . The computer-implemented method of claim 1 , wherein the one or more trained components comprise an image encoder that generates a skeleton image from an input image, and wherein the skeleton image comprises a plurality of channels that indicate different sets of pixel locations for different parts of an object within the input image.

7 . The computer-implemented method of claim 6 , wherein the one or more trained components further comprise a pose estimator that converts the skeleton image into a first set of pixel locations associated with a first set of joint positions and a second set of pixel locations associated with a second set of joint positions.

8 . The computer-implemented method of claim 7 , wherein the one or more trained components further comprise an uplift model that converts the first set of pixel locations and the second set of pixel locations into a set of three-dimensional (3D) coordinates.

9 . The computer-implemented method of claim 8 , wherein the image renderer generates an output image included in the set of output images based on input that includes (i) a projection of the set of 3D coordinates onto pixel locations in an analytic skeleton image and (ii) a reference image included in the set of reference images.

10 . The computer-implemented method of claim 1 , wherein the first set of training images comprises a set of synthetic images and the second set of training images comprises a set of non-rendered images that are captured by cameras.

11 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:

generating one or more trained components included in a pose estimation model based on one or more supervised losses computed from (i) a first set of predicted poses associated with a first set of training images that depict a first set of articulated objects against a first set of backgrounds and (ii) a first set of labeled poses associated with the first set of training images; and

training the one or more trained components and an image renderer included in the pose estimation model based on one or more unsupervised losses computed between (i) a second set of training images that depict a second set of articulated objects against a second set of backgrounds and (ii) a set of reconstructions of the second set of training images to generate a trained pose estimation model, wherein the set of reconstructions is generated by the image renderer from a set of predicted poses outputted by the one or more trained components.

12 . The one or more non-transitory computer-readable media of claim 11 , wherein the instructions further cause the one or more processors to perform the step of fine tuning the trained pose estimation model based on a third set of training images of a first object.

13 . The one or more non-transitory computer-readable media of claim 11 , wherein the instructions further cause the one or more processors to perform the step of synthesizing the first set of training images and the first set of labeled poses prior to generating the one or more trained components.

14 . The one or more non-transitory computer-readable media of claim 11 , wherein generating the one or more trained components comprises training an image encoder that generates a skeleton image from an input image based on an error between a set of limbs included in the skeleton image and a ground truth pose associated with the input image.

15 . The one or more non-transitory computer-readable media of claim 14 , wherein training the pose estimation model comprises further training the image encoder based on a discriminator loss associated with the input image and a set of unpaired poses.

16 . The one or more non-transitory computer-readable media of claim 11 , wherein generating the one or more trained components comprises training a pose estimator based on one or more errors between a predicted pose generated by the pose estimator from an input image and a ground truth pose for the input image.

17 . The one or more non-transitory computer-readable media of claim 11 , wherein training the pose estimation model comprises training the image renderer based on one or more losses associated with a reconstruction of a first image of a first object generated by the image renderer, wherein the reconstruction is generated by the image renderer based on a predicted pose associated with the first image and a second input image of the first object.

18 . The one or more non-transitory computer-readable media of claim 17 , wherein the one or more losses comprise at least one of a perceptual loss, a discriminator loss, or a discriminator feature matching loss.

19 . The one or more non-transitory computer-readable media of claim 11 , wherein the first set of labeled poses comprises a first set of joints on a left side of an object and a second set of joints on a right side of the object.

20 . A system, comprising:

one or more memories that store instructions, and

one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to:

execute a trained pose estimation model based on an input image, wherein the trained pose estimation model is generated by:

generating one or more trained components included in a pose estimation model based on one or more supervised losses computed from (i) a first set of predicted poses associated with a first set of training images that depict a first set of articulated objects against a first set of backgrounds and (ii) a first set of labeled poses associated with the first set of training images; and

training the one or more trained components and an image renderer included in the pose estimation model based on one or more unsupervised losses computed between (i) a second set of training images that depict a second set of articulated objects against a second set of backgrounds and (ii) a set of reconstructions of the second set of training images, wherein the set of reconstructions is generated by the image renderer from a set of predicted poses outputted by the one or more trained components; and

receive, as output of the one or more trained components, one or more poses associated with an object depicted in the input image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2022
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 059982/0421 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2022
From: GUAY, MARTIN; BORER, DOMINIK TOBIAS; BUHMANN, JAKOB JOACHIM
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH; ETH ZÜRICH (EIDGENÖSSISCHE TECHNISCHE HOCHSCHULE ZÜRICH)
Reel/Frame 059964/0141 →
Continuity (2)
Provisional Application 63194566 · May 28, 2021
Related Publication 20220392099A1 · Dec 8, 2022
References Cited (57)
US 10679046B1 · Black · 2020 [cited by examiner]
US 20210064925A1 · Shih · 2021 [cited by examiner]
US 20220188559A1 · Baek · 2022 [cited by examiner]
US 20220301304A1 · Hampali · 2022 [cited by examiner]
US 20230298204A1 · Wang · 2023 [cited by examiner]
US 20230386072A1 · Yao · 2023 [cited by examiner]
US 20240193809A1 · Ostadabbas · 2024 [cited by examiner]
Huang et al. “Invariant Representation Learning for Infant Pose Estimation with Small Data” Computer Vision and Pattern Recognition https://arxiv.org/abs/2010.06100v3 v3 Dec. 2, 2020 (Year: 2020). [cited by examiner]
Sun, X., Xiao, B., Wei, F., Liang, S., & Wei, Y. (2018). Integral human pose regression. In Proceedings of the European conference on computer vision (ECCV) (pp. 529-545). (Year: 2018). [cited by examiner]
Ma, L., Jia, X., Sun, Q., Schiele, B., Tuytelaars, T., & Van Gool, L. (2017). Pose guided person image generation. Advances in neural information processing systems, 30. (Year: 2017). [cited by examiner]
Pavlakos, Georgios, et al. “Expressive body capture: 3d hands, face, and body from a single image.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019. (Year: 2019). [cited by examiner]
Tripathi, S., Chandra, S., Agrawal, A., Tyagi, A., Rehg, J. M., & Chari, V. (2019). Learning to generate synthetic data via compositing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti… [cited by examiner]
Hoffmann, D. T., Tzionas, D., Black, M. J., & Tang, S. (Sep. 2019). Learning to train with synthetic humans. In German conference on pattern recognition (pp. 609-623). Cham: Springer International Publishing. (Year: 201… [cited by examiner]
Aberman et al., “Deep Video-Based Performance Cloning”, Eurographics 2019, DOI: 10.1111/cgf.13632, vol. 38, No. 2, 2019, pp. 219-233. [cited by applicant]
Andriluka et al., “2D Human Pose Estimation: New Benchmark and State of the Art Analysis”, In 2014 IEEE Conference on Computer Vision and Pattern Recognition, 2014, 8 pages. [cited by applicant]
Bala et al., “Automated markerless pose estimation in freely moving macaques with OpenMonkeyStudio”, https://doi.org/10.1038/s41467-020-18441-5, Nature Communications, vol. 11, No. 1, Dec. 2020, 12 pages. [cited by applicant]
Borer et al., “Augmenting Cats and Dogs: Procedural Texturing for Generalized Pet Tracking”, DOI: 10.520/0010333/70122032, In Proceedings of the 16th International Joint Conference on Computer Vision, Imaging and Comput… [cited by applicant]
Borer et al., “Rig-Space Neural Rendering: Compressing the Rendering of Characters for Previs, Real-Time Animation And High-Quality Asset Re-Use”, In Proceedings of the 16th International Joint Conference on Computer Vi… [cited by applicant]
Cao et al., “OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”, arXiv:1812.08008, 2018, pp. 7291-7299. [cited by applicant]
Ionescu et al., “Latent Structured Models for Human Pose Estimation”, IEEE International Conference on Computer Vision (ICCV), Nov. 2011, 8 pages. [cited by applicant]
Chan et al., “Everybody Dance Now”, In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 5932-5941. [cited by applicant]
Chen et al., “Unsupervised 3D Pose Estimation with Geometric Self-Supervision”, In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 5707-5717. [cited by applicant]
Cmu, “Cmu Motion Capture Database”, Retrieved from http://mocap.cs.cmu.edu/, 2001, 2 pages. [cited by applicant]
Dosovitskiy et al., “Generating Images with Perceptual Similarity Metrics based on Deep Networks”, Advances in Neural Information Processing Systems, vol. 29, Curran Associates, 2016, 9 pages. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets”, In Proceedings of the 27th International Conference on Neural Information Processing Systems, vol. 2, 2014, 9 pages. [cited by applicant]
Habermann et al., “DeepCap: Monocular Human Performance Capture Using Weak Supervision”, DOI:10.1109/CVPR42600.2020.00510, In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 5051-5062. [cited by applicant]
Ionescu et al., “Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 36, No. 7, Jul. 2014, pp. 1325-… [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks”, DOI 10.1109/CVPR.2017.632, In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 5967-5976. [cited by applicant]
Jakab et al., “Unsupervised Learning of Object Landmarks Through Conditional Image Generation”, In Advances in Neural Information Processing Systems, vol. 31, 2018, 12 pages. [cited by applicant]
Jakab et al., “Self-Supervised Learning of Interpretable Keypoints from Unlabelled Videos”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 8784-8794. [cited by applicant]
Kanazawa et al., “End-to-End Recovery of Human Shape and Pose”, In Computer Vision and Pattern Recognition (CVPR), DOI 10.1109/CVPR.2018.00744, 2018, pp. 7122-7131. [cited by applicant]
Kanazawa et al., “WarpNet: Weakly Supervised Matching for Singleview Reconstruction”, DOI 10.1109/CVPR.2016.354, In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 3253-3261. [cited by applicant]
Kundu et al., “Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image Synthesis”, DOI 10.1109/CVPR.44600.2020.00619, In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 6151-6… [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context”, In ECCV, 2014, 16 pages. [cited by applicant]
Loper et al., “SMPL: A Skinned Multi-Person Linear Model”, ACM Transactions on Graphics, vol. 34, No. 6, Article 248, DOI: http://doi.acm.org/10.1145/2816795.2818013, Nov. 2015, pp. 248:1-248:16. [cited by applicant]
Lorenz et al., “Unsupervised Part-Based Disentangling of Object Shape and Appearance”, In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 10947-10956. [cited by applicant]
Mao et al., “Least Squares Generative Adversarial Networks”, DOI 10.1109/ICCV.2017.304, In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Oct. 2017, pp. 2813-2821. [cited by applicant]
Martinez et al., “A Simple yet Effective Baseline for 3d Human Pose Estimation”, DOI 10.1109/ICCV.2017.288, IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2659-2668. [cited by applicant]
Newell et al., “Stacked Hourglass Networks for Human Pose Estimation”, In ECCV, 2016, 17 pages. [cited by applicant]
Park et al., “Semantic Image Synthesis with Spatially-Adaptive Normalization”, DOI 10.1109/CVPR.2019.00244, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 2332-2341. [cited by applicant]
Pavlakos et al., “Learning To Estimate 3D Human Pose and Shape from a Single Color Image”, DOI 10.1109/CVPR.2018.00055, In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 459-468. [cited by applicant]
Rhodin et al., “Unsupervised Geometry-Aware Representation for 3D Human Pose Estimation”, In Computer Vision, ECCV, Apr. 3, 2018, 17 pages. [cited by applicant]
Sarkar et al., “HumanGAN: A Generative Model of Humans Images”, abs/2103.06902, Mar. 11, 2021, 21 pages. [cited by applicant]
Schmidtke et al., “Unsupervised Human Pose Estimation Through Transforming Shape Templates”, DOI 10.1109/CVPR46437.2021, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, p… [cited by applicant]
Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, In International Conference on Learning Representations, 2015, 14 pages. [cited by applicant]
Thewlis et al., “Unsupervised Learning of Object Landmarks by Factorized Spatial Embeddings”, DOI 10.1109/ICCV.2017.348, In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Oct. 2017, pp. 3229… [cited by applicant]
Tobin et al., “Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World”, 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Sep. 24-28, 2017, pp. 23-30. [cited by applicant]
Toshev et al., “DeepPose: Human Pose Estimation via Deep Neural Networks”, In 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 1653-1660. [cited by applicant]
Varol et al., “Learning from Synthetic Humans”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, pp. 4627-4635. [cited by applicant]
Villegas et al., “Learning to Generate Long-term Future via Hierarchical Prediction”, In Proceedings of the 34th International Conference on Machine Learning, vol. 70, ICML, 2017, pp. 3560-3569. [cited by applicant]
Wang et al., “High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs”, In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 8798-8807. [cited by applicant]
Wei et al., “Convolutional Pose Machines”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 4724-4732. [cited by applicant]
Yang et al., “3D Human Pose Estimation in the Wild by Adversarial Learning”, In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 5255-5264. [cited by applicant]
Zhang et al., “Unsupervised Discovery of Object Landmarks as Structural Representations”, In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2018, pp. 2694-2703. [cited by applicant]
Zhu et al., “Unpaired Image-To-Image Translation using Cycle-Consistent Adversarial Networks”, In Computer Vision (ICCV), 2017 IEEE International Conference on Computer Vision, 2017, pp. 2242-2251. [cited by applicant]
Zuffi et al., “Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture from Images “in the wild””, In International Conference on Computer Vision, Oct. 2019, pp. 5358-5367. [cited by applicant]
Zuffi et al., “3D Menagerie: Modeling the 3D Shape and Pose of Animals”, In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 5524-5532. [cited by applicant]