IP Library Granted Patent US 12,394,089
Granted Patent B2
US 12,394,089 · App. 17/493,045 · Granted Aug 19, 2025

Pose estimation systems and methods trained using motion capture data

Inventors: Fabien Baradel (Grenoble, FR); Romain Bregier (Grenoble, FR); Thibault Groueix (Grenoble, FR); Ioannis Kalantidis (Grenoble, FR); Philippe Weinzaepfel (Montbonnot-Saint-Martin, FR); Gregory Rogez (Gières, FR)
Assignee: NAVER CORPORATION
G06T7/75G06F18/214G06N3/08G06T7/251G06T2200/08G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,089
App. No.
17/493,045
Granted
Aug 19, 2025
Kind
B2
Abstract

A training system includes: a neural network model configured to determine three-dimensional coordinates of joints, respectively, representing poses of animals in images, where the neural network model is trained using a first training dataset including: images including animals; and coordinates of joints of the animals in the images, respectively; and a training module configured to, after the training of the neural network model using the first training dataset, train the neural network model using a second training dataset including motion capture data, where the motion capture data does not include images of animals and includes measured coordinates at points, respectively, on animals.

Claims (91)

1. A training system comprising:

a neural network model configured to determine three-dimensional coordinates of joints, respectively, representing poses of animals in images,

wherein the neural network model is trained using a first training dataset including:

images including animals; and

coordinates of joints of the animals in the images, respectively; and

a training module configured to, after the training of the neural network model using the first training dataset, train the neural network model using a second training dataset including motion capture data,

wherein the motion capture data does not include images of animals and includes measured coordinates at points, respectively, on animals,

wherein the neural network model includes batch normalization layers, and

wherein the training module is configured to, after the training of the neural network model using the first training dataset, train the neural network model by adjusting only parameters of the batch normalization layers and maintaining all other parameters of all other portions of the neural network model constant after the training of the neural network model using the first training dataset,

wherein the training module is configured to, after the training of the neural network model using the first dataset, train the neural network model using texturized body models of animals with backgrounds applied.

2. The training system of claim 1 wherein the training module is configured to, after the training of the neural network model using the first training dataset, train the neural network model further using data from the first training dataset.

3. The training system of claim 2 wherein the training module is configured to generate the texturized body models of animals with backgrounds applied by:

rendering body models of animals based on the motion capture data;

texturizing the body models using textures from a texture dataset; and

apply backgrounds to the body models from a background dataset.

4. The training system of claim 3 wherein the training module is further configured to selectively:

render second models of the animals in the images using the coordinates of the joints of the animals, respectively;

texturize the second models using textures from the texture dataset;

apply second backgrounds to the models from the background dataset; and

train the neural network model using the texturized second models with the second backgrounds.

5. The training system of claim 1 wherein the model is the SPIN model.

6. The training system of claim 1 wherein the motion capture data in the second training dataset includes motion capture data from the AMASS dataset.

7. A system, comprising:

a neural network model configured to determine three-dimensional coordinates of joints, respectively, representing a pose of an animal in an image,

wherein the neural network model is trained using:

a first training dataset including:

images including animals; and

coordinates of joints of the animals in the images, respectively; and

a second training dataset including motion capture data that:

does not include images of animals; and

includes measured coordinates at points, respectively, on animals;

wherein the neural network model includes batch normalization layers, and

wherein the neural network model is trained by adjusting only parameters of the batch normalization layers and maintaining all other parameters of all other portions of the neural network model constant after the training of the neural network model using the first training dataset; and

a camera configured to capture images; and

a control module configured to selectively actuate an actuator based on three-dimensional coordinates of joints representing a pose of an animal determined by the neural network model based on an image captured by the camera,

wherein the neural network model is further trained based on texturized body models of animals with generated backgrounds.

8. A training system, comprising:

a neural network model configured to determine three-dimensional coordinates of joints, respectively, representing poses of animals in sequences of images;

a training module configured to train the neural network model using a training dataset including sets of time series of motion capture data,

wherein the motion capture data does not include images of animals and includes measured coordinates at points, respectively, on animals,

wherein the training module is configured to render a time series of body models of animals based on one of the sets of time series of motion capture data; and

a masking module configured to replace a first predetermined percentage of the body models in the time series of body models with different body models of animals,

wherein the training module is configured to train the neural network model based on outputs of the neural network model generated in response to the time series of body models.

9. The training system of claim 8 wherein the training module is configured to:

texturize the body models using textures from a texture dataset;

apply backgrounds to the body models from a background dataset; and

train the neural network model using the texturized body models with the backgrounds.

10. The training system of claim 9 wherein the masking module is further configured to replace at least one of the body models in the time series of models with a predetermined masking token.

11. The training system of claim 10 wherein the masking module is configured to replace a second predetermined percentage of the body models in the time series of body models with a predetermined masking token.

12. The training system of claim 11 wherein the second predetermined percentage is approximately 12.5 percent.

13. The training system of claim 9 further comprising a noise module configured to add noise to at least one of the body models in the time series of body models.

14. The training system of claim 13 wherein the noise includes Gaussian noise.

15. The training system of claim 9 further comprising a position encoding module configured to add positional encodings to the time series of body models.

16. The training system of claim 9 wherein the training module is configured to train the neural network model based on minimizing a pose loss determined based on differences between:

three-dimensional coordinates determined by the neural network model based on the time series of body models of animals; and

coordinates of the motion capture data stored in the training dataset.

17. The training system of claim 9 wherein the training module is configured to train the neural network model based on minimizing a three-dimensional keypoint loss determined based on differences between:

three-dimensional coordinates determined by the neural network model based on the time series of body models of animals; and

coordinates of the motion capture data stored in the training dataset.

18. A system, comprising:

a neural network model configured to determine three-dimensional coordinates of joints, respectively, representing poses of animals,

wherein the neural network model is trained using a training dataset including motion capture data that:

does not include images of animals; and

includes measured coordinates at points, respectively, on animals; and

wherein the neural network model includes batch normalization layers, and

wherein the neural network model is trained by adjusting only parameters of the batch normalization layers and maintaining all other parameters of all other portions of the neural network model constant after the training of the neural network model using the training dataset;

a camera configured to capture time series of images; and

a control module configured to selectively actuate an actuator based on three-dimensional coordinates of joints representing a pose of an animal determined by the neural network model based on a time series of images captured by the camera,

wherein the neural network model is further trained based on texturized body models of animals with generated backgrounds.

19. A training method comprising:

training, using a first training dataset, a neural network model to determine three-dimensional coordinates of joints, respectively, representing poses of animals in images, the first training dataset including:

images including animals; and

coordinates of joints of the animals in the images, respectively; and

after the training of the neural network model using the first training dataset, training the neural network model using a second training dataset including sets of time series of motion capture data,

wherein the motion capture data does not include images of animals and includes measured coordinates at points, respectively, on animals,

wherein the training the neural network using the second training dataset includes:

rendering a time series of body models of animals based on one of the sets of time series of motion capture data; replacing a first predetermined percentage of the body models in the time series of models with different body models of animals; and training the neural network model based on outputs of the neural network model generated in response to the time series of body models.

20. A method, comprising:

using a neural network model, determining three-dimensional coordinates of joints, respectively, representing a pose of an animal in an image,

wherein the neural network model is trained using:

a first training dataset including:

images including animals; and

coordinates of joints of the animals in the images, respectively; and

a second training dataset including motion capture data that:

does not include images of animals; and

includes measured coordinates at points, respectively, on animals,

wherein the neural network model includes batch normalization layers, and

the neural network model is trained by adjusting only parameters of the batch normalization layers and maintaining all other parameters of all other portions of the neural network model constant after the training of the neural network model using the first training dataset; and

capturing images using a camera; and

selectively actuating an actuator based on three-dimensional coordinates of joints representing a pose of an animal determined by the neural network model based on an image captured by the camera,

wherein the neural network model is further trained based on texturized body models of animals with generated backgrounds.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2024
From: NAVER LABS CORPORATION
To: NAVER CORPORATION
Reel/Frame 068820/0495 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2021
From: BARADEL, FABIEN; BRÉGIER, ROMAIN; GROUEIX, THIBAULT; KALANTIDIS, IOANNIS; WEINZAEPFEL, PHILIPPE; ROGEZ, GREGORY
To: NAVER CORPORATION; NAVER LABS CORPORATION
Reel/Frame 057690/0457 →
Continuity (1)
Related Publication 20230119559A1 · Apr 20, 2023
References Cited (50)
US 10679046B1 · Black · 2020 [cited by examiner]
US 20210200190A1 · Guo · 2021 [cited by examiner]
US 20210326601A1 · Tang · 2021 [cited by examiner]
US 20220101113A1 · Tam · 2022 [cited by examiner]
US 20230048497A1 · Choi · 2023 [cited by examiner]
US 20230067081A1 · Serackis · 2023 [cited by examiner]
L. Müller, A. A. A. Osman, S. Tang, C.-H. P. Huang and M. J. Black, “On Self-Contact and Human Pose,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, Jun. 1, 2021, pp. 998… [cited by examiner]
Adrien Gaidon, Antonio M. Lopez, and Florent Perronnin. The reasonable effectiveness of synthetic visual data. IJCV, 2018. [cited by applicant]
Angjoo Kanazawa, Jason Y Zhang, Panna Felsen, and Jitendra Malik. Learning 3D human dynamics from video. In CVPR, 2019. [cited by applicant]
Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In CVPR, 2018. [cited by applicant]
Anurag Arnab, Carl Doersch, and Andrew Zisserman. Exploiting temporal context for 3D human pose estimation in the wild. In CVPR, 2019. [cited by applicant]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. In NeurIPS, 2017. [cited by applicant]
Carl Doersch and Andrew Zisserman. Sim2real transfer learning for 3D human pose estimation: motion to the rescue. In NeurIPS, 2019. [cited by applicant]
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments. IEEE Trans. PAMI, 2014. [cited by applicant]
Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J Black, and Peter V Gehler. Unite the people: Closing the loop between 3D and 2D human representations. In CVPR, 2017. [cited by applicant]
Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3d human pose estimation in the wild using improved cnn supervision. In 3DV, 2017. [cited by applicant]
Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu, Srinath Sridhar, Gerard Pons-Moll, and Christian Theobalt. Single-shot multi-person 3d pose estimation from monocular rgb. In 3DV, 2018. [cited by applicant]
Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter V. Gehler, Javier Romero, and Michael J. Black. Keep it SMPL: automatic estimation of 3D human pose and shape from a single image. In ECCV, 2016. [cited by applicant]
Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao. LSUN: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015. [cited by applicant]
Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa. Human mesh recovery from multiple shots. arXiv preprint arXiv:2012.09843, 2020. [cited by applicant]
Georgios Pavlakos, Nikos Kolotouros, and Kostas Daniilidis. Texturepose: Supervising human mesh estimation with texture consistency. In ICCV, 2019. [cited by applicant]
Gregory Rogez and Cordelia Schmid. Mocap-guided data augmentation for 3D pose estimation in the wild. In NIPS, 2016. [cited by applicant]
Gul Varol, Duygu Ceylan, Bryan Russell, Jimei Yang, Ersin Yumer, Ivan Laptev, and Cordelia Schmid. Bodynet: Volumetric inference of 3d human body shapes. In ECCV, 2018. [cited by applicant]
Gul Varol, Javier Romero, Xavier Martin, Naureen Mahmood, Michael J. Black, Ivan Laptev, and Cordelia Schmid. Learning from synthetic humans. In CVPR, 2017. [cited by applicant]
Gyeongsik Moon and Kyoung Mu Lee. I2I-meshnet: Imageto-Iixel prediction network for accurate 3D human pose and mesh estimation from a single rgb image. arXiv preprint arXiv:2008.03713, 2020. [cited by applicant]
Hongsuk Choi, Gyeongsik Moon, and Kyoung Mu Lee. Beyond static features for temporally consistent 3D human pose and shape from a video. arXiv preprint arXiv:2011.08627, 2020. [cited by applicant]
Hongsuk Choi, Gyeongsik Moon, and Kyoung Mu Lee. Pose2mesh: Graph convolutional network for 3D human pose and mesh recovery from a 2d human pose. In ECCV, 2020. [cited by applicant]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. [cited by applicant]
Jogendra Nath Kundu, Ambareesh Revanur, Govind Vitthal Waghmare, Rahul Mysore Venkatesh, and R Venkatesh Babu. Unsupervised cross-modal alignment for multi-person 3d pose estimation. In ECCV, 2020. [cited by applicant]
Jonathan Frankle, David J Schwab, and Ari S Morcos. Training batchnorm and only batchnorm: On the expressive power of random features in cnns. In ICLR, 2021. [cited by applicant]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: a skinned multiperson linear model. ACM Transactions on Graphics, 2015. [cited by applicant]
Mohamed Omran, Christoph Lassner, Gerard Pons-Moll, Peter Gehler, and Bernt Schiele. Neural body fitting: Unifying deep learning and model based human pose and shape estimation. In 3DV, 2018. [cited by applicant]
Muhammed Kocabas, Nikos Athanasiou, and Michael J Black. Vibe: Video inference for human body pose and shape estimation. In CVPR, 2020. [cited by applicant]
Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele. 2D human pose estimation: New benchmark and state of the art analysis. In CVPR, 2014. [cited by applicant]
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael Black. AMASS: Archive of Motion Capture as Surface Shapes. In ICCV, 2019. [cited by applicant]
Nikos Kolotouros, Georgios Pavlakos, and Kostas Daniilidis. Convolutional mesh regression for single-image human shape reconstruction. In CVPR, 2019. [cited by applicant]
Nikos Kolotouros, Georgios Pavlakos, Michael J Black, and Kostas Daniilidis. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In ICCV, 2019. [cited by applicant]
Sam Johnson and Mark Everingham. Clustered pose and nonlinear appearance models for human pose estimation. In BMVC, 2010. [cited by applicant]
Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. [cited by applicant]
Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto. Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing. In ISMIR, 2019. [cited by applicant]
Timo von Marcard, Roberto Henschel, Michael J. Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering Accurate 3D Human Pose in the Wild Using IMUs and a Moving Camera. In ECCV, 2018. [cited by applicant]
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. [cited by applicant]
Tyler Zhu, Per Karlsson, and Christoph Bregler. Simpose: Effectively learning densepose and surface normals of people from simulated data. In ECCV, 2020. [cited by applicant]
Vincent Leroy, Philippe Weinzaepfel, Romain Bregier, Hadrien Combaluzier, and Gregory Rogez. SMPLy bench-marking 3D human pose estimation in the wild. In 3DV, 2020. [cited by applicant]
Wenzheng Chen, Huan Wang, Yangyan Li, Hao Su, Zhenhua Wang, Changhe Tu, Dani Lischinski, Daniel Cohen-Or, and Baoquan Chen. Synthesizing training images for boosting human 3D pose estimation. In 3DV, 2016. [cited by applicant]
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neural networks. In CVPR, 2019. [cited by applicant]
Yinghao Huang, Federica Bogo, Christoph Lassner, Angjoo Kanazawa, Peter V. Gehler, Ijaz Akhter, and Michael J. Black. Towards accurate markerless human shape and pose estimation over time. In 3DV, 2017. [cited by applicant]
Yu Rong, Ziwei Liu, Cheng Li, Kaidi Cao, and Chen Change Loy. Delving deep into hybrid annotations for 3D human recovery in the wild. In ICCV, 2019. [cited by applicant]
Yu Sun, Yun Ye, Wu Liu, Wenpeng Gao, Yili Fu, and Tao Mei. Human mesh recovery from monocular images via a skeleton-disentangled representation. In ICCV, 2019. [cited by applicant]
Zhengyi Luo, S Alireza Golestaneh, and Kris M Kitani. 3d human motion estimation via motion compression and refinement. In ACCV, 2020. [cited by applicant]