Method and apparatus for training parameter estimation models, device, and storage medium
Provided is a method for training parameter estimation models. The method includes: estimating reconstruction parameters specified for three-dimensional face reconstruction by inputting training samples in a face image training set to a pre-constructed neural network model, and reconstructing three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to a pre-constructed three-dimensional morphable model; calculating a plurality of loss functions of a plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, and adjusting weights corresponding to the plurality of loss functions; and generating fitting loss functions based on the plurality of loss functions and the weights corresponding to the plurality of loss functions, and acquiring a trained parameter estimation model by performing an inverse correction on the neural network model using the fitting loss functions.
1 . A method for training parameter estimation models, comprising:
estimating reconstruction parameters specified for three-dimensional face reconstruction by inputting training samples in a face image training set to a pre-constructed neural network model, and reconstructing three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to a pre-constructed three-dimensional morphable model;
calculating a plurality of loss functions of a plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, and adjusting weights corresponding to the plurality of loss functions; and
generating fitting loss functions based on the plurality of loss functions and the weights corresponding to the plurality of loss functions, and acquiring a trained parameter estimation model by performing an inverse correction on the pre-constructed neural network model using the fitting loss functions;
wherein the three-dimensional morphable model comprises a dual principal component analysis model and a single principal component analysis model, wherein the dual principal component analysis model comprises a three-dimensional average face, a first principal component base indicating a change of face appearance, and a second principal component base indicating a change of face expression, and the single principal component analysis model comprises an average face albedo and a third principal component base indicating a change of a face albedo; and
reconstructing the three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model comprises:
inputting reconstruction parameters matched with the first principal component base and reconstruction parameters matched with the second principal component base to the dual principal component analysis model, and acquiring a three-dimensional morphable face by deforming the three-dimensional average face; and
inputting the three-dimensional morphable face and reconstruction parameters matched with the third principal component base to the single principal component analysis model, and acquiring the reconstructed three-dimensional faces by performing an albedo correction on the three-dimensional morphable face based on the average face albedo.
2 . The method according to claim 1 , wherein the plurality of loss functions of the plurality of pieces of two-dimensional supervision information comprise an image pixel loss function, a key point loss function, an identity feature loss function, an albedo penalty function, and a regular term corresponding to a target reconstruction parameter in the reconstruction parameters specified for the three-dimensional face reconstruction.
3 . The method according to claim 2 , wherein
in a case of calculating the image pixel loss function between the three-dimensional faces and the training samples, the method further comprises:
segmenting skin masks from the training samples; and
calculating the image pixel loss function between the three-dimensional faces and the training samples comprises:
acquiring the image pixel loss function by calculating, based on the skin masks, pixel errors of same pixel points in facial skin regions in the three-dimensional faces and the training samples.
4 . The method according to claim 2 , wherein
in a case of calculating the key point loss function between the three-dimensional faces and the training samples, the method further comprises:
extracting a plurality of key feature points at a predetermined position from the training samples, and determining visibility of the plurality of key feature points; and
determining a plurality of visible key feature points based on the visibility of the plurality of key feature points; and
calculating the key point loss function between the three-dimensional faces and the training samples comprises:
acquiring the key point loss function by calculating position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples.
5 . The method according to claim 4 , wherein prior to calculating the position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples, the method further comprises:
dynamically selecting, based on head gestures in the three-dimensional faces, three-dimensional mesh vertexes matched with the plurality of visible key feature points from the three-dimensional faces, and calculating the position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples by determining position information of the three-dimensional mesh vertexes in the three-dimensional faces as reconstruction positions of the plurality of visible key feature points.
6 . The method according to claim 2 , wherein
in a case of calculating the identity feature loss function between the three-dimensional faces and the training samples, the method further comprises:
acquiring first identity features corresponding to the training samples and second identity features corresponding to the three-dimensional faces by inputting the training samples and the reconstructed three-dimensional faces to a pre-constructed face recognition model; and
calculating the identity feature loss function between the three-dimensional faces and the training samples comprises:
calculating the identity feature loss function based on similarities between the first identity features and the second identity features.
7 . The method according to claim 6 , wherein the first identity features corresponding to the training samples are further calculated by:
collecting a plurality of face images, containing same faces as the training samples, shot from a plurality of angles, and extracting a plurality of identity sub-features corresponding to the plurality of face images by inputting the plurality of face images to the pre-constructed three-dimensional morphable model; and
acquiring the first identity features corresponding to the training samples by integrating the extracted plurality of identity sub-features.
8 . The method according to claim 2 , wherein
in a case of calculating the albedo penalty function of the three-dimensional faces, the method further comprises:
calculating albedos of a plurality of vertexes in the three-dimensional faces; and
calculating the albedo penalty function of the three-dimensional faces comprises:
calculating the albedo penalty function based on the albedos of the plurality of vertexes in the three-dimensional faces and a predetermined albedo interval.
9 . The method according to claim 1 , wherein prior to reconstructing the three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model, the method further comprises:
collecting three-dimensional face scan data with uniform illumination under a multi-dimensional data source, and acquiring the three-dimensional average face, the average face albedo, the first principal component base, the second principal component base, and the third principal component base by performing a deformation analysis, an expression change analysis, and an albedo analysis on the three-dimensional face scan data.
10 . The method according to claim 1 , wherein the three-dimensional morphable model further comprises an illumination parameter indicating a change of face illumination, a position parameter indicating face translation, and a rotation parameter indicating a head gesture.
11 . The method according to claim 1 , wherein prior to calculating the plurality of loss functions of the plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, the method further comprises:
rendering the three-dimensional faces by a differentiable renderer, and training the parameter estimation model based on the rendered three-dimensional faces.
12 . The method according to claim 1 , wherein upon acquiring the trained parameter estimation models by performing the inverse correction on the pre-constructed neural network model using the fitting loss functions, the method further comprises:
inputting to-be-reconstructed two-dimensional face images to the parameter estimation model, estimating the reconstruction parameters specified for the three-dimensional face reconstruction, and reconstructing three-dimensional faces corresponding to the two-dimensional face images by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model.
13 . A computer device for training parameter estimation models, comprising:
at least one processor; and
a memory, configured to store at least one program;
wherein the at least one processor, when running the at least one program, is caused to perform:
estimating reconstruction parameters specified for three-dimensional face reconstruction by inputting training samples in a face image training set to a pre-constructed neural network model, and reconstructing three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to a pre-constructed three-dimensional morphable model;
calculating a plurality of loss functions of a plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, and adjusting weights corresponding to the plurality of loss functions; and
generating fitting loss functions based on the plurality of loss functions and the weights corresponding to the plurality of loss functions, and acquiring a trained parameter estimation model by performing an inverse correction on the pre-constructed neural network model using the fitting loss functions;
wherein the three-dimensional morphable model comprises a dual principal component analysis model and a single principal component analysis model, wherein the dual principal component analysis model comprises a three-dimensional average face, a first principal component base indicating a change of face appearance, and a second principal component base indicating a change of face expression, and the single principal component analysis model comprises an average face albedo and a third principal component base indicating a change of a face albedo; and
reconstructing the three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model comprises:
inputting reconstruction parameters matched with the first principal component base and reconstruction parameters matched with the second principal component base to the dual principal component analysis model, and acquiring a three-dimensional morphable face by deforming the three-dimensional average face; and
inputting the three-dimensional morphable face and reconstruction parameters matched with the third principal component base to the single principal component analysis model, and acquiring the reconstructed three-dimensional faces by performing an albedo correction on the three-dimensional morphable face based on the average face albedo.
14 . The computer device for training parameter estimation models according to claim 13 , wherein the plurality of loss functions of the plurality of pieces of two-dimensional supervision information comprise an image pixel loss function, a key point loss function, an identity feature loss function, an albedo penalty function, and a regular term corresponding to a target reconstruction parameter in the reconstruction parameters specified for the three-dimensional face reconstruction.
15 . The computer device for training parameter estimation models according to claim 14 , wherein the at least one processor, when running the at least one program, is caused to perform:
segmenting skin masks from the training samples; and
acquiring the image pixel loss function by calculating, based on the skin masks, pixel errors of same pixel points in facial skin regions in the three-dimensional faces and the training samples.
16 . The computer device for training parameter estimation models according to claim 14 , wherein the at least one processor, when running the at least one program, is caused to perform:
extracting a plurality of key feature points at a predetermined position from the training samples, and determining visibility of the plurality of key feature points;
determining a plurality of visible key feature points based on the visibility of the plurality of key feature points; and
acquiring the key point loss function by calculating position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples.
17 . The computer device for training parameter estimation models according to claim 16 , wherein the at least one processor, when running the at least one program, is caused to perform:
dynamically selecting, based on head gestures in the three-dimensional faces, three-dimensional mesh vertexes matched with the plurality of visible key feature points from the three-dimensional faces, and calculating the position reconstruction errors of the plurality of visible key feature points between the three-dimensional faces and the training samples by determining position information of the three-dimensional mesh vertexes in the three-dimensional faces as reconstruction positions of the plurality of visible key feature points.
18 . The computer device for training parameter estimation models according to claim 14 , wherein the at least one processor, when running the at least one program, is caused to perform:
acquiring first identity features corresponding to the training samples and second identity features corresponding to the three-dimensional faces by inputting the training samples and the reconstructed three-dimensional faces to a pre-constructed face recognition model; and
calculating the identity feature loss function based on similarities between the first identity features and the second identity features.
19 . A non-transitory computer-readable storage medium, storing a computer program thereon, wherein the program, when run by a processor, causes the processor to perform:
estimating reconstruction parameters specified for three-dimensional face reconstruction by inputting training samples in a face image training set to a pre-constructed neural network model, and reconstructing three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to a pre-constructed three-dimensional morphable model;
calculating a plurality of loss functions of a plurality of pieces of two-dimensional supervision information between the three-dimensional faces and the training samples, and adjusting weights corresponding to the plurality of loss functions; and
generating fitting loss functions based on the plurality of loss functions and the weights corresponding to the plurality of loss functions, and acquiring a trained parameter estimation model by performing an inverse correction on the pre-constructed neural network model using the fitting loss functions;
wherein the three-dimensional morphable model comprises a dual principal component analysis model and a single principal component analysis model, wherein the dual principal component analysis model comprises a three-dimensional average face, a first principal component base indicating a change of face appearance, and a second principal component base indicating a change of face expression, and the single principal component analysis model comprises an average face albedo and a third principal component base indicating a change of a face albedo; and
reconstructing the three-dimensional faces corresponding to the training samples by inputting the reconstruction parameters to the pre-constructed three-dimensional morphable model comprises:
inputting reconstruction parameters matched with the first principal component base and reconstruction parameters matched with the second principal component base to the dual principal component analysis model, and acquiring a three-dimensional morphable face by deforming the three-dimensional average face; and
inputting the three-dimensional morphable face and reconstruction parameters matched with the third principal component base to the single principal component analysis model, and acquiring the reconstructed three-dimensional faces by performing an albedo correction on the three-dimensional morphable face based on the average face albedo.