IP Library › Granted Patent US 12,633,055
Granted Patent B2
US 12,633,055 · App. 18/248,315 · Granted May 19, 2026

Method and apparatus for training parameter estimation models, device, and storage medium

Inventors: Xiaowei Zhang (Guangzhou, CN); Zhongyuan Hu (Guangzhou, CN); Gengdai Liu (Guangzhou, CN)
Assignee: BIGO TECHNOLOGY PTE. LTD.
G06T17/20G06T15/50G06T19/20G06V10/60G06V10/774G06V10/776G06V10/82G06V40/171
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,055
App. No.
18/248,315
Granted
May 19, 2026
Kind
B2
Abstract

Provided is a method for training parameter estimation models. The method includes: estimating reconstruction parameters specified for three-dimensional face reconstruction by inputting training samples in a face image training set to a pre-constructed neural network model, and reconstructing three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to a pre-constructed three-dimensional morphable model; calculating a plurality of loss functions of a plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, and adjusting weights corresponding to the plurality of loss functions; and generating fitting loss functions based on the plurality of loss functions and the weights corresponding to the plurality of loss functions, and acquiring a trained parameter estimation model by performing an inverse correction on the neural network model using the fitting loss functions.

Claims (74)

1 . A method for training parameter estimation models, comprising:

estimating reconstruction parameters specified for three-dimensional face reconstruction by inputting training samples in a face image training set to a pre-constructed neural network model, and reconstructing three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to a pre-constructed three-dimensional morphable model;

calculating a plurality of loss functions of a plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, and adjusting weights corresponding to the plurality of loss functions; and

generating fitting loss functions based on the plurality of loss functions and the weights corresponding to the plurality of loss functions, and acquiring a trained parameter estimation model by performing an inverse correction on the pre-constructed neural network model using the fitting loss functions;

wherein the three-dimensional morphable model comprises a dual principal component analysis model and a single principal component analysis model, wherein the dual principal component analysis model comprises a three-dimensional average face, a first principal component base indicating a change of face appearance, and a second principal component base indicating a change of face expression, and the single principal component analysis model comprises an average face albedo and a third principal component base indicating a change of a face albedo; and

reconstructing the three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model comprises:

inputting reconstruction parameters matched with the first principal component base and reconstruction parameters matched with the second principal component base to the dual principal component analysis model, and acquiring a three-dimensional morphable face by deforming the three-dimensional average face; and

inputting the three-dimensional morphable face and reconstruction parameters matched with the third principal component base to the single principal component analysis model, and acquiring the reconstructed three-dimensional faces by performing an albedo correction on the three-dimensional morphable face based on the average face albedo.

2 . The method according to claim 1 , wherein the plurality of loss functions of the plurality of pieces of two-dimensional supervision information comprise an image pixel loss function, a key point loss function, an identity feature loss function, an albedo penalty function, and a regular term corresponding to a target reconstruction parameter in the reconstruction parameters specified for the three-dimensional face reconstruction.

3 . The method according to claim 2 , wherein

in a case of calculating the image pixel loss function between the three-dimensional faces and the training samples, the method further comprises:

segmenting skin masks from the training samples; and

calculating the image pixel loss function between the three-dimensional faces and the training samples comprises:

acquiring the image pixel loss function by calculating, based on the skin masks, pixel errors of same pixel points in facial skin regions in the three-dimensional faces and the training samples.

4 . The method according to claim 2 , wherein

in a case of calculating the key point loss function between the three-dimensional faces and the training samples, the method further comprises:

extracting a plurality of key feature points at a predetermined position from the training samples, and determining visibility of the plurality of key feature points; and

determining a plurality of visible key feature points based on the visibility of the plurality of key feature points; and

calculating the key point loss function between the three-dimensional faces and the training samples comprises:

acquiring the key point loss function by calculating position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples.

5 . The method according to claim 4 , wherein prior to calculating the position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples, the method further comprises:

dynamically selecting, based on head gestures in the three-dimensional faces, three-dimensional mesh vertexes matched with the plurality of visible key feature points from the three-dimensional faces, and calculating the position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples by determining position information of the three-dimensional mesh vertexes in the three-dimensional faces as reconstruction positions of the plurality of visible key feature points.

6 . The method according to claim 2 , wherein

in a case of calculating the identity feature loss function between the three-dimensional faces and the training samples, the method further comprises:

acquiring first identity features corresponding to the training samples and second identity features corresponding to the three-dimensional faces by inputting the training samples and the reconstructed three-dimensional faces to a pre-constructed face recognition model; and

calculating the identity feature loss function between the three-dimensional faces and the training samples comprises:

calculating the identity feature loss function based on similarities between the first identity features and the second identity features.

7 . The method according to claim 6 , wherein the first identity features corresponding to the training samples are further calculated by:

collecting a plurality of face images, containing same faces as the training samples, shot from a plurality of angles, and extracting a plurality of identity sub-features corresponding to the plurality of face images by inputting the plurality of face images to the pre-constructed three-dimensional morphable model; and

acquiring the first identity features corresponding to the training samples by integrating the extracted plurality of identity sub-features.

8 . The method according to claim 2 , wherein

in a case of calculating the albedo penalty function of the three-dimensional faces, the method further comprises:

calculating albedos of a plurality of vertexes in the three-dimensional faces; and

calculating the albedo penalty function of the three-dimensional faces comprises:

calculating the albedo penalty function based on the albedos of the plurality of vertexes in the three-dimensional faces and a predetermined albedo interval.

9 . The method according to claim 1 , wherein prior to reconstructing the three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model, the method further comprises:

collecting three-dimensional face scan data with uniform illumination under a multi-dimensional data source, and acquiring the three-dimensional average face, the average face albedo, the first principal component base, the second principal component base, and the third principal component base by performing a deformation analysis, an expression change analysis, and an albedo analysis on the three-dimensional face scan data.

10 . The method according to claim 1 , wherein the three-dimensional morphable model further comprises an illumination parameter indicating a change of face illumination, a position parameter indicating face translation, and a rotation parameter indicating a head gesture.

11 . The method according to claim 1 , wherein prior to calculating the plurality of loss functions of the plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, the method further comprises:

rendering the three-dimensional faces by a differentiable renderer, and training the parameter estimation model based on the rendered three-dimensional faces.

12 . The method according to claim 1 , wherein upon acquiring the trained parameter estimation models by performing the inverse correction on the pre-constructed neural network model using the fitting loss functions, the method further comprises:

inputting to-be-reconstructed two-dimensional face images to the parameter estimation model, estimating the reconstruction parameters specified for the three-dimensional face reconstruction, and reconstructing three-dimensional faces corresponding to the two-dimensional face images by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model.

13 . A computer device for training parameter estimation models, comprising:

at least one processor; and

a memory, configured to store at least one program;

wherein the at least one processor, when running the at least one program, is caused to perform:

estimating reconstruction parameters specified for three-dimensional face reconstruction by inputting training samples in a face image training set to a pre-constructed neural network model, and reconstructing three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to a pre-constructed three-dimensional morphable model;

calculating a plurality of loss functions of a plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, and adjusting weights corresponding to the plurality of loss functions; and

generating fitting loss functions based on the plurality of loss functions and the weights corresponding to the plurality of loss functions, and acquiring a trained parameter estimation model by performing an inverse correction on the pre-constructed neural network model using the fitting loss functions;

wherein the three-dimensional morphable model comprises a dual principal component analysis model and a single principal component analysis model, wherein the dual principal component analysis model comprises a three-dimensional average face, a first principal component base indicating a change of face appearance, and a second principal component base indicating a change of face expression, and the single principal component analysis model comprises an average face albedo and a third principal component base indicating a change of a face albedo; and

reconstructing the three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model comprises:

inputting reconstruction parameters matched with the first principal component base and reconstruction parameters matched with the second principal component base to the dual principal component analysis model, and acquiring a three-dimensional morphable face by deforming the three-dimensional average face; and

inputting the three-dimensional morphable face and reconstruction parameters matched with the third principal component base to the single principal component analysis model, and acquiring the reconstructed three-dimensional faces by performing an albedo correction on the three-dimensional morphable face based on the average face albedo.

14 . The computer device for training parameter estimation models according to claim 13 , wherein the plurality of loss functions of the plurality of pieces of two-dimensional supervision information comprise an image pixel loss function, a key point loss function, an identity feature loss function, an albedo penalty function, and a regular term corresponding to a target reconstruction parameter in the reconstruction parameters specified for the three-dimensional face reconstruction.

15 . The computer device for training parameter estimation models according to claim 14 , wherein the at least one processor, when running the at least one program, is caused to perform:

segmenting skin masks from the training samples; and

acquiring the image pixel loss function by calculating, based on the skin masks, pixel errors of same pixel points in facial skin regions in the three-dimensional faces and the training samples.

16 . The computer device for training parameter estimation models according to claim 14 , wherein the at least one processor, when running the at least one program, is caused to perform:

extracting a plurality of key feature points at a predetermined position from the training samples, and determining visibility of the plurality of key feature points;

determining a plurality of visible key feature points based on the visibility of the plurality of key feature points; and

acquiring the key point loss function by calculating position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples.

17 . The computer device for training parameter estimation models according to claim 16 , wherein the at least one processor, when running the at least one program, is caused to perform:

dynamically selecting, based on head gestures in the three-dimensional faces, three-dimensional mesh vertexes matched with the plurality of visible key feature points from the three-dimensional faces, and calculating the position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples by determining position information of the three-dimensional mesh vertexes in the three-dimensional faces as reconstruction positions of the plurality of visible key feature points.

18 . The computer device for training parameter estimation models according to claim 14 , wherein the at least one processor, when running the at least one program, is caused to perform:

acquiring first identity features corresponding to the training samples and second identity features corresponding to the three-dimensional faces by inputting the training samples and the reconstructed three-dimensional faces to a pre-constructed face recognition model; and

calculating the identity feature loss function based on similarities between the first identity features and the second identity features.

19 . A non-transitory computer-readable storage medium, storing a computer program thereon, wherein the program, when run by a processor, causes the processor to perform:

estimating reconstruction parameters specified for three-dimensional face reconstruction by inputting training samples in a face image training set to a pre-constructed neural network model, and reconstructing three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to a pre-constructed three-dimensional morphable model;

calculating a plurality of loss functions of a plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, and adjusting weights corresponding to the plurality of loss functions; and

generating fitting loss functions based on the plurality of loss functions and the weights corresponding to the plurality of loss functions, and acquiring a trained parameter estimation model by performing an inverse correction on the pre-constructed neural network model using the fitting loss functions;

wherein the three-dimensional morphable model comprises a dual principal component analysis model and a single principal component analysis model, wherein the dual principal component analysis model comprises a three-dimensional average face, a first principal component base indicating a change of face appearance, and a second principal component base indicating a change of face expression, and the single principal component analysis model comprises an average face albedo and a third principal component base indicating a change of a face albedo; and

reconstructing the three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model comprises:

inputting reconstruction parameters matched with the first principal component base and reconstruction parameters matched with the second principal component base to the dual principal component analysis model, and acquiring a three-dimensional morphable face by deforming the three-dimensional average face; and

inputting the three-dimensional morphable face and reconstruction parameters matched with the third principal component base to the single principal component analysis model, and acquiring the reconstructed three-dimensional faces by performing an albedo correction on the three-dimensional morphable face based on the average face albedo.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2023
From: ZHANG, XIAOWEI; HU, ZHONGYUAN; LIU, GENGDAI
To: BIGO TECHNOLOGY PTE. LTD.
Reel/Frame 063260/0559 →
Priority Claims (1)
CN 202011211255.4 · Nov 3, 2020 · national
Continuity (1)
Related Publication 20240296624A1 · Sep 5, 2024
References Cited (21)
US 11443484B2 · Vesdapunt · 2022 [cited by examiner]
US 20160314619A1 · Luo et al. · 2016 [cited by applicant]
US 20210256752A1 · Chen · 2021 [cited by applicant]
US 20210286977A1 · Chen et al. · 2021 [cited by applicant]
CN 109191507A · 2019 [cited by applicant]
CN 109934300A · 2019 [cited by applicant]
CN 109978989A · 2019 [cited by applicant]
CN 110414370A · 2019 [cited by applicant]
CN 110619676A · 2019 [cited by applicant]
CN 111428667A · 2020 [cited by applicant]
CN 112529999A · 2021 [cited by applicant]
Extended European Search Report Communication Pursuant to Rule 62 EPC for European Application No. 21888418.7 dated Sep. 13, 2024, which is a foreign counterpart application to this application. [cited by applicant]
Lin, Jiangke, et al.; “Towards High-Fidelity 3D Face Reconstruction From In-the-wild Images Using Graph Convolutional Networks”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 13,… [cited by applicant]
International Search Report of the International Searching Authority for State Intellectual Property Office of the People's Republic of China in PCT application No. PCT/CN2021/125575 issued on Jan. 19, 2022, which is an… [cited by applicant]
Blanz, Volker, et al.; “A Morphable Model for the Synthesis of 3D Faces”, In ACM Transactions on Graphic (Proceedings of SIGGRAPH):187-194, 1999. [cited by applicant]
Chen, Ke; “3D Face Reconstruction from Single Facial Image Based on Convolutional Neural Network”, Chinese Selected Doctoral Dissertations and Master's Theses Full-Text Databases(Master) Information Science and Technolo… [cited by applicant]
Luo, Yao; “Research on the Method of 3D Face Reconstruction from a Single Image”, Chinese Selected Doctoral Dissertations and Master's Theses Full-Text Databases(Master) Information Science and Technology Series, No. 07… [cited by applicant]
Zhou, Jian, et al.; “3D Face Reconstruction and Dense Face Alignment Method Based on Improved 3D Morphable Model”, Journal of Computer Applications, Jun. 2, 2020, ISSN: 1001-9081, pp. 3-6. [cited by applicant]
Notice of Reasons for Refusal of Japanese application No. 2023-523272 issued on Jan. 29, 2024. [cited by applicant]
China National Intellectual Property Administration, First Office Action in Chinese Patent Application No. 202011211255.4 issued on Jun. 13, 2024, which is a foreign counterpart application corresponding to this U.S. Pa… [cited by applicant]
Tran, Luan et al., “Towards High-fidelity Nonlinear 3D Face Morphable Model”, CVPR 2019, Apr. 30, 2019, entire document. [cited by applicant]