IP Library › Granted Patent US 12,205,311
Granted Patent B1
US 12,205,311 · App. 18/382,838 · Granted Jan 21, 2025

System for training neural networks that predict the parameters of a human mesh model using dense depth and part-based UV map annotations

Inventors: Batuhan Karagoz (Ankara, TR); Emre Akbas (Ankara, TR); Bedirhan Uguz (Pittsburgh, PA); Ozhan Suat (Ankara, TR); Necip Berme (Worthington, OH); Mohan Chandra Baro (Columbus, OH)
Assignee: Bertec Corporation
G06T7/50G06V10/751G06V20/46G06T2207/20081G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,311
App. No.
18/382,838
Granted
Jan 21, 2025
Kind
B1
Abstract

A system for training neural networks that predict the parameters of a human mesh model is disclosed herein. The system includes at least one camera and a data processor configured to execute computer executable instructions for: receiving a first frame and a second frame of a video from the at least one camera; extracting first and second image data from the first and second frames of the video; inputting the sequence of frames of the video into a human mesh estimator module, the human mesh estimator module estimating mesh parameters from the sequence of frames of the video so as to determine a predicted mesh; and generating a training signal for input into the human mesh estimator module by using a two-dimensional keypoint loss module that compares a first set of two-dimensional image-based keypoints to a second set of two-dimensional model-based keypoints.

Claims (17)

1. A system for training neural networks that predict the parameters of a human mesh model, the system comprising:

at least one camera, the at least one camera configured to capture a sequence of frames of a video, the sequence of frames of the video including at least a first frame and a second frame; and

a data processor including at least one hardware component, the data processor configured to execute computer executable instructions, the computer executable instructions comprising instructions for:

receiving the first frame and the second frame of the video from the at least one camera;

extracting first image data from the first frame of the video;

extracting second image data from the second frame of the video;

inputting the sequence of frames of the video into a human mesh estimator module, the human mesh estimator module estimating mesh parameters from the sequence of frames of the video so as to determine a human mesh model;

generating a training signal for input into the human mesh estimator module by using a rigid transform loss module that calculates rigid transformations between the first and second frames of the sequence of frames for one or more body parts, warps a predicted mesh for the first frame from the human mesh model using the rigid transformations, and compares the warped mesh with a predicted mesh for the second frame from the human mesh model; and

generating another training signal for input into the human mesh estimator module by using a two-dimensional keypoint loss module that compares a first set of two-dimensional image-based keypoints to a second set of two-dimensional model-based keypoints.

2. The system according to claim 1 , wherein the data processor is further configured to estimate the first set of two-dimensional image-based keypoints from at least one of the first image data or the second image data using a two-dimensional keypoint estimator.

3. The system according to claim 1 , wherein the data processor is further configured to estimate the first set of two-dimensional image-based keypoints from a training dataset that comprises two-dimensional ground truth keypoints.

4. The system according to claim 1 , wherein the data processor is further configured to estimate the second set of two-dimensional model-based keypoints by projecting the second set of two-dimensional model-based keypoints from the human mesh model.

5. The system according to claim 1 , wherein, when generating the training signal for input into the human mesh estimator module using the two-dimensional keypoint loss module, the data processor is further configured to minimize numerical differences between corresponding keypoints of the first set of two-dimensional image-based keypoints and the second set of two-dimensional model-based keypoints so as to ensure a precise alignment between the human mesh model and real-world images in the video.

6. The system according to claim 1 , wherein the data processor is further configured to generate depth maps and IUV maps for the first frame and the second frame of the video.

7. The system according to claim 6 , wherein the data processor is configured to generate the depth maps using a human depth estimator module.

8. The system according to claim 6 , wherein the at least one camera comprises an RGB-D camera that outputs both image data and depth data, and the data processor is configured to generate the depth maps using the depth data from the RGB-D camera.

9. The system according to claim 1 , wherein the data processor is configured to estimate the second set of two-dimensional model-based keypoints by additionally considering a movement of the human mesh model over a period of time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: KARAGOZ, BATUHAN; AKBAS, EMRE; UGUZ, BEDIRHAN; SUAT, OZHAN; BERME, NECIP; BARO, MOHAN CHANDRA
To: BERTEC CORPORATION
Reel/Frame 065312/0130 →
Continuity (2)
Continuation In Part 18195777 · May 10, 2023
Provisional Application 63354497 · Jun 22, 2022
References Cited (139)
US 6038488A · Barnes et al. · 2000 [cited by applicant]
US 6113237A · Ober et al. · 2000 [cited by applicant]
US 6152564A · Ober et al. · 2000 [cited by applicant]
US 6295878B1 · Berme · 2001 [cited by applicant]
US 6354155B1 · Berme · 2002 [cited by applicant]
US 6389883B1 · Berme et al. · 2002 [cited by applicant]
US 6936016B2 · Berme et al. · 2005 [cited by applicant]
US 8181541B2 · Berme · 2012 [cited by applicant]
US 8315822B2 · Berme et al. · 2012 [cited by applicant]
US 8315823B2 · Berme et al. · 2012 [cited by applicant]
US D689388S · Berme · 2013 [cited by applicant]
US D689389S · Berme · 2013 [cited by applicant]
US 8543540B1 · Wilson et al. · 2013 [cited by applicant]
US 8544347B1 · Berme · 2013 [cited by applicant]
US 8643669B1 · Wilson et al. · 2014 [cited by applicant]
US 8700569B1 · Wilson et al. · 2014 [cited by applicant]
US 8704855B1 · Berme et al. · 2014 [cited by applicant]
US 8764532B1 · Berme · 2014 [cited by applicant]
US 8847989B1 · Berme et al. · 2014 [cited by applicant]
US D715669S · Berme · 2014 [cited by applicant]
US 8902249B1 · Wilson et al. · 2014 [cited by applicant]
US 8915149B1 · Berme · 2014 [cited by applicant]
US 9032817B2 · Berme et al. · 2015 [cited by applicant]
US 9043278B1 · Wilson et al. · 2015 [cited by applicant]
US 9066667B1 · Berme et al. · 2015 [cited by applicant]
US 9081436B1 · Berme et al. · 2015 [cited by applicant]
US 9168420B1 · Berme et al. · 2015 [cited by applicant]
US 9173596B1 · Berme et al. · 2015 [cited by applicant]
US 9200897B1 · Wilson et al. · 2015 [cited by applicant]
US 9277857B1 · Berme et al. · 2016 [cited by applicant]
US D755067S · Berme et al. · 2016 [cited by applicant]
US 9404823B1 · Berme et al. · 2016 [cited by applicant]
US 9414784B1 · Berme et al. · 2016 [cited by applicant]
US 9468370B1 · Shearer · 2016 [cited by applicant]
US 9517008B1 · Berme et al. · 2016 [cited by applicant]
US 9526443B1 · Berme et al. · 2016 [cited by applicant]
US 9526451B1 · Berme · 2016 [cited by applicant]
US 9558399B1 · Jeka et al. · 2017 [cited by applicant]
US 9568382B1 · Berme et al. · 2017 [cited by applicant]
US 9622686B1 · Berme et al. · 2017 [cited by applicant]
US 9763604B1 · Berme et al. · 2017 [cited by applicant]
US 9770203B1 · Berme et al. · 2017 [cited by applicant]
US 9778119B2 · Berme et al. · 2017 [cited by applicant]
US 9814430B1 · Berme et al. · 2017 [cited by applicant]
US 9829311B1 · Wilson et al. · 2017 [cited by applicant]
US 9854997B1 · Berme et al. · 2018 [cited by applicant]
US 9916011B1 · Berme et al. · 2018 [cited by applicant]
US 9927312B1 · Berme et al. · 2018 [cited by applicant]
US 10010248B1 · Shearer · 2018 [cited by applicant]
US 10010286B1 · Berme et al. · 2018 [cited by applicant]
US 10085676B1 · Berme et al. · 2018 [cited by applicant]
US 10117602B1 · Berme et al. · 2018 [cited by applicant]
US 10126186B2 · Berme et al. · 2018 [cited by applicant]
US 10216262B1 · Berme et al. · 2019 [cited by applicant]
US 10231662B1 · Berme et al. · 2019 [cited by applicant]
US 10264964B1 · Berme et al. · 2019 [cited by applicant]
US 10331324B1 · Wilson et al. · 2019 [cited by applicant]
US 10342473B1 · Berme et al. · 2019 [cited by applicant]
US 10390736B1 · Berme et al. · 2019 [cited by applicant]
US 10413230B1 · Berme et al. · 2019 [cited by applicant]
US 10463250B1 · Berme et al. · 2019 [cited by applicant]
US 10527508B2 · Berme et al. · 2020 [cited by applicant]
US 10555688B1 · Berme et al. · 2020 [cited by applicant]
US 10646153B1 · Berme et al. · 2020 [cited by applicant]
US 10722114B1 · Berme et al. · 2020 [cited by applicant]
US 10736545B1 · Berme et al. · 2020 [cited by applicant]
US 10765936B2 · Berme et al. · 2020 [cited by applicant]
US 10803990B1 · Wilson et al. · 2020 [cited by applicant]
US 10853970B1 · Akbas et al. · 2020 [cited by applicant]
US 10856796B1 · Berme et al. · 2020 [cited by applicant]
US 10860843B1 · Berme et al. · 2020 [cited by applicant]
US 10945599B1 · Berme et al. · 2021 [cited by applicant]
US 10966606B1 · Berme · 2021 [cited by applicant]
US 11033453B1 · Berme et al. · 2021 [cited by applicant]
US 11052288B1 · Berme et al. · 2021 [cited by applicant]
US 11054325B2 · Berme et al. · 2021 [cited by applicant]
US 11074711B1 · Akbas et al. · 2021 [cited by applicant]
US 11097154B1 · Berme et al. · 2021 [cited by applicant]
US 11158422B1 · Wilson et al. · 2021 [cited by applicant]
US 11182924B1 · Akbas et al. · 2021 [cited by applicant]
US 11262231B1 · Berme et al. · 2022 [cited by applicant]
US 11262258B2 · Berme et al. · 2022 [cited by applicant]
US 11301045B1 · Berme et al. · 2022 [cited by applicant]
US 11311209B1 · Berme et al. · 2022 [cited by applicant]
US 11321868B1 · Akbas et al. · 2022 [cited by applicant]
US 11337606B1 · Berme et al. · 2022 [cited by applicant]
US 11348279B1 · Akbas et al. · 2022 [cited by applicant]
US 11458362B1 · Berme et al. · 2022 [cited by applicant]
US 11521373B1 · Akbas et al. · 2022 [cited by applicant]
US 11540744B1 · Berme · 2023 [cited by applicant]
US 11604106B2 · Berme et al. · 2023 [cited by applicant]
US 11631193B1 · Akbas et al. · 2023 [cited by applicant]
US 11688139B1 · Karagoz et al. · 2023 [cited by applicant]
US 11705244B1 · Berme · 2023 [cited by applicant]
US 11712162B1 · Berme et al. · 2023 [cited by applicant]
US 11790536B1 · Berme et al. · 2023 [cited by applicant]
US 11798182B1 · Karagoz et al. · 2023 [cited by applicant]
US 11816258B1 · Berme et al. · 2023 [cited by applicant]
US 11826601B1 · Berme · 2023 [cited by applicant]
US 11850078B1 · Berme · 2023 [cited by applicant]
US 11857331B1 · Berme et al. · 2024 [cited by applicant]
US 11865407B1 · Berme et al. · 2024 [cited by applicant]
US 11911147B1 · Berme et al. · 2024 [cited by applicant]
US 20030216656A1 · Berme et al. · 2003 [cited by applicant]
US 20080228110A1 · Berme · 2008 [cited by applicant]
US 20110277562A1 · Berme · 2011 [cited by applicant]
US 20120266648A1 · Berme et al. · 2012 [cited by applicant]
US 20120271565A1 · Berme et al. · 2012 [cited by applicant]
US 20150096387A1 · Berme et al. · 2015 [cited by applicant]
US 20160245711A1 · Berme et al. · 2016 [cited by applicant]
US 20160334288A1 · Berme et al. · 2016 [cited by applicant]
US 20180024015A1 · Berme et al. · 2018 [cited by applicant]
US 20190078951A1 · Berme et al. · 2019 [cited by applicant]
US 20200139229A1 · Berme et al. · 2020 [cited by applicant]
US 20200408625A1 · Berme et al. · 2020 [cited by applicant]
US 20210333163A1 · Berme et al. · 2021 [cited by applicant]
US 20220178775A1 · Berme et al. · 2022 [cited by applicant]
Kocabas, Muhammed, Nikos Athanasiou, and Michael J. Black. “Vibe: Video inference for human body pose and shape estimation.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020. (Year… [cited by examiner]
Nguyen, Phong, et al. “Human View Synthesis using a Single Sparse RGB-D Input.” arXiv preprint arXiv:2112.13889 (2021). (Year: 2021). [cited by examiner]
Yang, Ji, et al. “3D pose estimation and future motion prediction from 2D images.” Pattern Recognition 124 (2021): 108439. (Year: 2021). [cited by examiner]
M. Kocabas, S. Karagoz, and E. Akbas. Multiposenet: Fast multi-person pose estimation using pose residual network. In European Conference on Computer Vision. (Jul. 2018) pp. 1-17. [cited by applicant]
R. Güler, N. Neverova, and I. Kokkinos. Densepose: Dense human pose estimation in the wild. In EEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). (Jun. 2018) pp. 7297-7306. [cited by applicant]
Y. Jafarian and H. S. Park. Learning high fidelity depths of dressed humans by watching social media dance videos. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). (Mar. 2021) pp. 12753-12762. [cited by applicant]
C. Mayer, M. Danelljan, D. P. Paudel, and L. Van Gool. Learning target candidate association to keep track of what not to track. arXiv preprint arXiv:2103.16556 (Mar. 2021) pp. 13444-13454. [cited by applicant]
Y. Raaj, H. Idrees, G. Hidalgo, and Y. Sheikh. Efficient online multi-person 2d pose tracking with recurrent spatio-temporal affinity fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco… [cited by applicant]
S. Tang, F. Tan, K. Cheng, Z. Li, S. Zhu, and P. Tan. A neural network for detailed human depth estimation from a single image. In Proceedings of the IEEE/CVF International Conference on Computer Vision. (Oct. 2019) pp.… [cited by applicant]
Guan, Shanyan, et al. “Out-of-domain human mesh reconstruction via dynamic bilevel online adaptation.” IEEE Transactions on Pattern Analysis and Machine Intelligence 45.4 (2022): 5070-5086. (Year: 2022). [cited by applicant]
Guler, Riza Alp, and Iasonas Kokkinos. “Holopose: Holistic 3d human reconstruction in-the-wild.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. (Year: 2019). [cited by applicant]
Hu, Tao, et al. “Hvtr: Hybrid volumetric-textural rendering for human avatars.” 2022 International Conference on 3D Vision (3DV). IEEE, 2022. (Year: 2022). [cited by applicant]
Kim, Hyomin, et al. “LaplacianFusion: Detailed 3D Clothed-Human Body Reconstruction.” ACM Transactions on Graphics (TOG) 41.6 (2022): 1-14. (Year: 2022). [cited by applicant]
Kundu, Jogendra Nath, et al. “Appearance consensus driven self-supervised human mesh recovery.” Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2020, Proceedings, Part I 16. Springer Intern… [cited by applicant]
Li, Kun, et al. “Image-guided human reconstruction via multi-scale graph transformation networks.” IEEE Transactions on Image Processing 30 (2021): 5239-5251. (Year: 2021). [cited by applicant]
Liu, Xun-Yu, et al. “3D Human Pose and Shape Estimation from Video.” 2022, CPSCom, pp. 617-624 (Year: 2022). [cited by applicant]
Ma, Liqian, et al. “Direct Dense Pose Estimation.” 2021 International Conference on 3D Vision (3DV). IEEE, 2021. (Year: 2021). [cited by applicant]
Qiao, Yi-Ling, Alexander Gao, and Ming Lin. “Neuphysics: Editable neural geometry and physics from monocular videos.” Advances in Neural Information Processing Systems 35 (2022): 12841-12854. (Year: 2022). [cited by applicant]
Wang, Zhe. Robust Estimation of 3D Human Body Pose with Geometric Priors. Diss. University of California, Irvine, 2021. (Year: 2021). [cited by applicant]
Wang, Zhe, Jimei Yang, and Charless Fowlkes. “The best of both worlds: combining model-based and nonparametric approaches for 3D human body estimation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Patt… [cited by applicant]
Xu, Yuanlu, Song-Chun Zhu, and Tony Tung. “Denserac: Joint 3d pose and shape estimation by dense render-and-compare.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. (Year: 2019). [cited by applicant]
Notice of Allowance in U.S. Appl. No. 18/195,777, mailed on Jun. 20, 2023. [cited by applicant]
Cited By (1)
US 12,518,418