IP Library › Granted Patent US 12,198,374
Granted Patent B2
US 12,198,374 · App. 17/231,952 · Granted Jan 14, 2025

Method for training SMPL parameter prediction model, computer device, and storage medium

Inventors: Shuang Sun (Shenzhen, CN); Chen Li (Shenzhen, CN); Yuwing Tai (Shenzhen, CN); Jiaya Jia (Shenzhen, CN); Xiaoyong Shen (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06T7/74G06F18/22G06N3/08G06T7/70G06T17/00G06V10/809G06V10/82G06V40/11G06T2207/20044
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,374
App. No.
17/231,952
Granted
Jan 14, 2025
Kind
B2
Abstract

A method for training an SMPL parameter prediction model, including: obtaining a sample picture; inputting the sample picture into a pose parameter prediction model to obtain a predicted pose parameter; inputting the sample picture into a shape parameter prediction model to obtain a predicted shape parameter; calculating model prediction losses according to an SMPL parameter prediction model and annotation information of the sample picture; and updating the pose parameter prediction model and the shape parameter prediction model according to the model prediction losses.

Claims (81)

1. A method for training a skinned multi-person linear model (SMPL) parameter prediction model, performed by a computer device, the method comprising:

obtaining a sample picture, the sample picture containing a human body image;

inputting the sample picture into a pose parameter prediction model to obtain a predicted pose parameter, the predicted pose parameter being a parameter used for indicating a human body pose in the SMPL parameter prediction model;

inputting the sample picture into a shape parameter prediction model to obtain a predicted shape parameter, the predicted shape parameter being a parameter used for indicating a human body shape in the SMPL parameter prediction model;

constructing a three-dimensional human body model according to the predicted pose parameter and the predicted shape parameter;

calculating model prediction losses according to the three-dimensional human body model based on the SMPL parameter prediction model and annotation information of the sample picture; and

reversely training the pose parameter prediction model and the shape parameter prediction model the according to the model prediction losses.

2. The method according to claim 1 , wherein the model prediction losses comprise a first model prediction loss; and the calculating model prediction losses according to the three-dimensional human body model based on the SMPL parameter prediction model and annotation information of the sample picture comprises:

calculating the first model prediction loss according to the three-dimensional human body model based on the SMPL parameter prediction model and annotated SMPL parameters in the annotation information, the annotated SMPL parameters comprising an annotated pose parameter and an annotated shape parameter.

3. The method according to claim 2 , wherein the calculating the first model prediction loss according to the three-dimensional human body model based on the SMPL parameter prediction model and annotated SMPL parameters in the annotation information comprises:

calculating a first Euclidean distance between the annotated pose parameter and the predicted pose parameter;

calculating a second Euclidean distance between the annotated shape parameter and the predicted shape parameter; and

determining the first model prediction loss according to the first Euclidean distance and the second Euclidean distance.

4. The method according to claim 1 , wherein the model prediction losses comprise a second model prediction loss; and the calculating the model prediction losses according to the three-dimensional human body model based on the SMPL parameter prediction model and the annotation information of the sample picture comprises:

calculating the second model prediction loss according to predicted joint coordinates of joints in the three-dimensional human body model and annotated joint coordinates of joints in the annotation information.

5. The method according to claim 4 , wherein the annotated joint coordinates comprise three-dimensional annotated joint coordinates and/or two-dimensional annotated joint coordinates; and

the calculating the second model prediction loss according to predicted joint coordinates of joints in the three-dimensional human body model and annotated joint coordinates of joints in the annotation information comprises:

calculating third Euclidean distances between three-dimensional predicted joint coordinates of the joints in the three-dimensional human body model and the three-dimensional annotated joint coordinates; and

calculating the second model prediction loss according to the third Euclidean distances.

6. The method according to claim 5 , wherein the calculating third Euclidean distances between three-dimensional predicted joint coordinates of the joints in the three-dimensional human body model and the three-dimensional annotated joint coordinates comprises:

determining the three-dimensional predicted joint coordinates of the joints in the three-dimensional human body model according to vertex coordinates of model vertices around the joints in the three-dimensional human body model; and

calculating the third Euclidean distances between the three-dimensional predicted joint coordinates and the three-dimensional annotated joint coordinates.

7. The method according to claim 5 , wherein the calculating the second model prediction loss according to the third Euclidean distances comprises:

calculating fourth Euclidean distances between two-dimensional predicted joint coordinates of the joints in the three-dimensional human body model and the two-dimensional annotated joint coordinates; and

calculating the second model prediction loss according to the fourth Euclidean distances.

8. The method according to claim 7 , wherein the pose parameter prediction model is further configured to output a projection parameter according to the inputted sample picture, the projection parameter being used for projecting points in a three-dimensional space into a two-dimensional space; and

the calculating fourth Euclidean distances between two-dimensional predicted joint coordinates of the joints in the three-dimensional human body model and the two-dimensional annotated joint coordinates comprises:

determining the three-dimensional predicted joint coordinates of the joints in the three-dimensional human body model according to vertex coordinates of model vertices around the joints in the three-dimensional human body model;

performing projection processing on the three-dimensional predicted joint coordinates according to the projection parameter to obtain the two-dimensional predicted joint coordinates; and

calculating the fourth Euclidean distances between the two-dimensional predicted joint coordinates and the two-dimensional annotated joint coordinates.

9. The method according to claim 1 , wherein the model prediction losses comprise a third model prediction loss; and the calculating the model prediction losses according to the three-dimensional human body model based on the SMPL parameter prediction model and the annotation information of the sample picture comprises:

calculating the third model prediction loss according to a predicted two-dimensional human body contour of the three-dimensional human body model and an annotated two-dimensional human body contour in the annotation information.

10. The method according to claim 9 , wherein the pose parameter prediction model is further configured to output a projection parameter according to the inputted sample picture, the projection parameter being used for projecting points in a three-dimensional space into a two-dimensional space; and

the calculating the third model prediction loss according to a predicted two-dimensional human body contour of the three-dimensional human body model and an annotated two-dimensional human body contour in the annotation information comprises:

projecting model vertices in the three-dimensional human body model into a two-dimensional space according to the projection parameter, and generating the predicted two-dimensional human body contour;

calculating a first contour loss and a second contour loss according to the predicted two-dimensional human body contour and the annotated two-dimensional human body contour; and

determining the third model prediction loss according to the first contour loss and the second contour loss.

11. The method according to claim 10 , wherein the calculating a first contour loss and a second contour loss according to the predicted two-dimensional human body contour and the annotated two-dimensional human body contour comprises:

calculating a first shortest distance from each contour point in the predicted two-dimensional human body contour to the annotated two-dimensional human body contour, and

calculating the first contour loss according to the first shortest distance corresponding to each contour point in the predicted two-dimensional human body contour; and

calculating a second shortest distance from each contour point in the annotated two-dimensional human body contour to the predicted two-dimensional human body contour, and calculating the second contour loss according to the second shortest distance corresponding to each contour point in the annotated two-dimensional human body contour.

12. The method according to claim 11 , wherein the calculating the first contour loss according to the first shortest distance corresponding to each contour point in the predicted two-dimensional human body contour comprises:

determining a first weight of the contour point corresponding to each model vertex according to visibility of the joint to which the each model vertex in the three-dimensional human body model belongs; and

calculating the first contour loss according to the first weight and the first shortest distance corresponding to each contour point in the predicted two-dimensional human body contour, wherein

in a case that the joint to which a model vertex belongs is visible, the first weight of the contour point corresponding to the model vertex is 1; and in a case that the joint to which a model vertex belongs is invisible, the first weight of the contour point corresponding to the model vertex is 0.

13. The method according to claim 11 , wherein the calculating the first contour loss according to the first shortest distance corresponding to each contour point in the predicted two-dimensional human body contour comprises:

determining the predicted joint coordinates of the joint to which each model vertex in the three-dimensional human body model belongs;

determining a second weight of the contour point corresponding to each model vertex according to a fifth Euclidean distance between the predicted joint coordinates and the annotated joint coordinates, the second weight and the fifth Euclidean distance having a negative correlation; and

calculating the first contour loss according to the second weight and the first shortest distance corresponding to each contour point in the predicted two-dimensional human body contour.

14. The method according to claim 1 , further comprising:

performing regularization on the predicted shape parameter to obtain a fourth model prediction loss; and

updating the pose parameter prediction model and the shape parameter prediction model according to the fourth model prediction loss.

15. The method according to claim 1 , comprising:

obtaining a target picture, the target picture containing a human body image;

inputting the target picture into the updated pose parameter prediction model to obtain a predicted pose parameter;

inputting the target picture into the updated shape parameter prediction model to obtain a predicted shape parameter; and

constructing a target three-dimensional human body model according to the predicted pose parameter and the predicted shape parameter.

16. A computer device, comprising a processor and a memory, the memory storing a plurality of instructions that, when executed by the processor, cause the computer device to perform a plurality of operations including:

obtaining a sample picture, the sample picture containing a human body image;

inputting the sample picture into a pose parameter prediction model to obtain a predicted pose parameter, the predicted pose parameter being a parameter used for indicating a human body pose in the SMPL parameter prediction model;

inputting the sample picture into a shape parameter prediction model to obtain a predicted shape parameter, the predicted shape parameter being a parameter used for indicating a human body shape in the SMPL parameter prediction model;

constructing a three-dimensional human body model according to the predicted pose parameter and the predicted shape parameter;

calculating model prediction losses according to the three-dimensional human body model based on the SMPL parameter prediction model and annotation information of the sample picture; and

reversely training the pose parameter prediction model and the shape parameter prediction model according to the model prediction losses.

17. The computer device according to claim 16 , wherein the plurality of operations further comprise:

obtaining a target picture, the target picture containing a human body image;

inputting the target picture into the updated pose parameter prediction model to obtain a predicted pose parameter;

inputting the target picture into the updated shape parameter prediction model to obtain a predicted shape parameter; and

constructing a target three-dimensional human body model according to the predicted pose parameter and the predicted shape parameter.

18. A non-transitory computer-readable storage medium, storing a plurality of instructions that, when executed by a processor of a computer device, cause the computer device to perform a plurality of operations including:

obtaining a sample picture, the sample picture containing a human body image;

inputting the sample picture into a pose parameter prediction model to obtain a predicted pose parameter, the predicted pose parameter being a parameter used for indicating a human body pose in the SMPL parameter prediction model;

inputting the sample picture into a shape parameter prediction model to obtain a predicted shape parameter, the predicted shape parameter being a parameter used for indicating a human body shape in the SMPL parameter prediction model;

constructing a three-dimensional human body model according to the predicted pose parameter and the predicted shape parameter;

calculating model prediction losses according to the three-dimensional human body model based on the SMPL parameter prediction model and annotation information of the sample picture; and

reversely training the pose parameter prediction model and the shape parameter prediction model according to the model prediction losses.

19. The non-transitory computer-readable storage medium according to claim 18 , wherein the plurality of operations further comprise:

obtaining a target picture, the target picture containing a human body image;

inputting the target picture into the updated pose parameter prediction model to obtain a predicted pose parameter;

inputting the target picture into the updated shape parameter prediction model to obtain a predicted shape parameter; and

constructing a target three-dimensional human body model according to the predicted pose parameter and the predicted shape parameter.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2021
From: SUN, SHUANG; LI, CHEN; TAI, YUWING; JIA, JIAYA; SHEN, XIAOYONG
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 057124/0375 →
Priority Claims (1)
CN 201910103414.X · Feb 1, 2019 · national
Continuity (2)
Continuation PCTCN2020072023 · Jan 14, 2020
Related Publication 20210232924A1 · Jul 29, 2021
References Cited (20)
US 10529137B1 · Black · 2020 [cited by examiner]
US 10679046B1 · Black · 2020 [cited by examiner]
US 11302064B2 · Li · 2022 [cited by examiner]
US 20180315230A1 · Black · 2018 [cited by examiner]
US 20190371080A1 · Sminchisescu · 2019 [cited by examiner]
US 20210241522A1 · Guler · 2021 [cited by examiner]
CN 108594997A · 2018 [cited by applicant]
CN 109285215A · 2019 [cited by applicant]
CN 109859296A · 2019 [cited by applicant]
Bogo et al., “Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image,” B. Leibe et al. (Eds.): ECCV 2016, Part V, LNCS 9909, pp. 561-578, 2016 (Year: 2016). [cited by examiner]
Pengfei Dou et al. “End-to-End 3D Face Reconstruction with Deep Neural Networks,” 2017 IEEE International Conference on Computer Vision, University of Houston Computational Biomedicine Lab, Nov. 9, 2017, 10 pgs. [cited by applicant]
Tencent Technology, ISR, PCT/CN2020/072023, Apr. 3, 2020, 2 pgs. [cited by applicant]
Extended European Search Report, EP20748016.1, Sep. 19, 2022, 8 pgs. [cited by applicant]
Meysam Madadi et al., “SMPLR: Deep SMPL Reverse for 3D Human Pose and Shape Recovery”, arxiv.org, Cornell University Library, Dec. 27, 2018, XP081457828, 11 pgs. [cited by applicant]
Christoph Lassner et al., “Unite the People: Closing the Loop Between 3D and 2D Human Representations”, arxiv.org, Cornell University Library, Jan. 10, 2017, XP080740529, 10 pgs. [cited by applicant]
Federica Bogo et al., “Keep It SMPL Automatic Estimation of 3D Human Pose and Shape from a Single Image”, Sep. 16, 2016, Advances in Biometrics: International Conference, ICB 2007, XP047355099, 2 pgs., Retrieved from th… [cited by applicant]
Ernesto Brau et al., “3D Human Pose Estimation via Deep Learning from 2D Annotations”, 2016 Fourth International Conference on 3D Vision (3DV), IEEE, Oct. 25, 2016, XP033027667, 3 pgs., Retrieved from the Internet: http… [cited by applicant]
Georgios Pavlakos et al., “Learning to Estimate 3D Human Pose and Shape from a Single Color Image”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Jun. 18, 2018, XP033476006, 2 pgs., Retrieve… [cited by applicant]
Tencent Technology, WO, PCT/CN2020/072023, Apr. 3, 2020, 5 pgs. [cited by applicant]
Tencent Technology, IPRP, PCT/CN2020/072023, Jul. 27, 2021, 6 pgs. [cited by applicant]
Cited By (1)
US 12,524,965