Three-dimensional human head reconstruction method, electronic device and non-transient computer-readable storage medium
View Patent ↗A three-dimensional human head reconstruction method, an electronic device and a non-transient computer-readable storage medium are provided. The method includes: acquiring a target portrait image; inputting the target portrait image into a target model to obtain an output result of the target model; wherein the target model is obtained by pre-training with a plurality of training samples which are generated according to a sample portrait image and a sample three-dimensional human head model, the sample three-dimensional human head model is obtained by iteratively fitting a standard three-dimensional human face statistical model according to two-dimensional feature information related to a portrait in the sample portrait image, and the two-dimensional feature information includes human face feature points and a human head projection contour line; and generating, according to the output result, a target three-dimensional human head model corresponding to the target portrait image.
1 . A three-dimensional human head reconstruction method, comprising:
acquiring a portrait image;
inputting the portrait image into a target model to obtain an output result from the target model, wherein the target model pre-trained with a plurality of training samples which are generated according to a sample portrait image and a sample three-dimensional human head model;
generating, according to the output result, a three-dimensional human head model corresponding to the portrait image;
wherein before acquiring the portrait image, the method further comprises:
acquiring the sample portrait image;
extracting two-dimensional feature information from the sample portrait image;
plane projecting a standard three-dimensional human face statistical model to obtain projection feature information corresponding to the two-dimensional feature information; and
iteratively fitting, based on the two-dimensional feature information and the projection feature information, the standard three-dimensional human face statistical model through a regression loss function to obtain the sample three-dimensional human head model.
2 . The method according to claim 1 , wherein the training samples include the sample portrait image and sample statistical model parameters corresponding to the sample three-dimensional human head model, and the output result includes target statistical model parameters.
3 . The method according to claim 2 , wherein before acquiring the portrait image, the method further comprises:
acquiring the plurality of training samples;
learning a mapping relationship between the sample portrait image and the sample statistical model parameters in each of the training samples through a first regression loss function, to obtain the target model; and
continuing to learn the mapping relationship between the sample portrait image and the sample statistical model parameters in each of the training samples through a target loss function, to obtain an optimized target model; wherein the target loss function includes a second regression loss function and a projection loss function, and weight values of identity coefficients in the second regression loss function are greater than weight values of identity coefficients in the first regression loss function.
4 . The method according to claim 2 , wherein the generating, according to the output result, the three-dimensional human head model corresponding to the portrait image comprises:
generating the three-dimensional human head model according to the target statistical model parameters and the standard three-dimensional human face statistical model.
5 . The method according to claim 1 , wherein the projection feature information includes human face projection feature points;
wherein the plane projecting the standard three-dimensional human face statistical model to obtain the projection feature information corresponding to the two-dimensional feature information comprises:
projecting respective vertexes of the standard three-dimensional human face statistical model into a two-dimensional space to obtain respective vertex projection points corresponding to the standard three-dimensional human face statistical model; and
determining, according to the respective vertex projection points, the human face projection feature points.
6 . The method according to claim 5 , wherein the projection feature information further includes a human head projection contour line;
wherein the plane projecting the standard three-dimensional human face statistical model to obtain the projection feature information corresponding to the two-dimensional feature information further comprises:
dilating the respective vertex projection points to obtain a first head region image;
eroding the first head region image to obtain a second head region image; and
edge extracting the second head region image to obtain the human head projection contour line.
7 . The method according to claim 5 , wherein the two-dimensional feature information further includes shoulder feature points, and the projection feature information further includes shoulder projection feature points;
wherein the plane projecting the standard three-dimensional human face statistical model to obtain the projection feature information corresponding to the two-dimensional feature information further comprises:
determining, according to the respective vertex projection points, the shoulder projection feature points.
8 . The method according to claim 1 , wherein the projection feature information includes human face projection feature points and a human head projection contour line;
wherein before the iteratively fitting, based on the two-dimensional feature information and the projection feature information, the standard three-dimensional human face statistical model through the regression loss function to obtain the sample three-dimensional human head model, the method further comprises:
randomly sampling the human head contour line and the human head projection contour line respectively to obtain human head contour feature points and human head projection contour feature points;
wherein the iteratively fitting, based on the two-dimensional feature information and the projection feature information, the standard three-dimensional human face statistical model through the third regression loss function to obtain the sample three-dimensional human head model comprises:
iteratively fitting the standard three-dimensional human face statistical model through the regression loss function to obtain the sample three-dimensional human head model, based on the human face feature points, the human head contour feature points, the human face projection feature points and the human head projection contour feature points.
9 . The method according to claim 1 , wherein before the plane projecting the standard three-dimensional human face statistical model to obtain the projection feature information corresponding to the two-dimensional feature information, the method further comprises:
detecting a human head posture in the sample portrait image;
wherein the plane projecting the standard three-dimensional human face statistical model to obtain the projection feature information corresponding to the two-dimensional feature information comprises:
projecting the standard three-dimensional human face statistical model onto an imaging plane of a capturing device to obtain the projection feature information, according to projection parameters of the capturing device for the sample portrait image in the case that the standard three-dimensional human face statistical model is in the human head posture.
10 . An electronic device, comprising:
a processor;
a memory for storing executable instructions;
wherein the processor is configured to read the executable instructions from the memory and executing the executable instructions to implement operations comprising:
acquiring a portrait image;
inputting the portrait image into a target model to obtain an output result from the target model, wherein the target model pre-trained with a plurality of training samples which are generated according to a sample portrait image and a sample three-dimensional human head model;
generating, according to the output result, a three-dimensional human head model corresponding to the portrait image;
wherein before acquiring the portrait image, the method further comprises:
acquiring the sample portrait image;
extracting two-dimensional feature information from the sample portrait image, wherein the two-dimensional feature information includes human face feature points and a human head projection contour line;
plane projecting a standard three-dimensional human face statistical model to obtain projection feature information corresponding to the two-dimensional feature information; and
iteratively fitting, based on the two-dimensional feature information and the projection feature information, the standard three-dimensional human face statistical model through a regression loss function to obtain the sample three-dimensional human head model.
11 . The electronic device according to claim 10 , wherein the training samples include the sample portrait image and sample statistical model parameters corresponding to the sample three-dimensional human head model, and the output result includes target statistical model parameters.
12 . The electronic device according to claim 11 , wherein before acquiring the portrait image, the operations further comprise:
acquiring the plurality of training samples;
learning a mapping relationship between the sample portrait image and the sample statistical model parameters in each of the training samples through a first regression loss function, to obtain the target model; and
continuing to learn the mapping relationship between the sample portrait image and the sample statistical model parameters in each of the training samples through a target loss function, to obtain an optimized target model; wherein the target loss function includes a second regression loss function and a projection loss function, and weight values of identity coefficients in the second regression loss function are greater than weight values of identity coefficients in the first regression loss function.
13 . The electronic device according to claim 11 , wherein the generating, according to the output result, the three-dimensional human head model corresponding to the portrait image comprises:
generating the three-dimensional human head model according to the target statistical model parameters and the standard three-dimensional human face statistical model.
14 . The electronic device according to claim 10 , wherein the projection feature information includes human face projection feature points, and wherein the plane projecting the standard three-dimensional human face statistical model to obtain the projection feature information corresponding to the two-dimensional feature information comprises:
projecting respective vertexes of the standard three-dimensional human face statistical model into a two-dimensional space to obtain respective vertex projection points corresponding to the standard three-dimensional human face statistical model; and
determining, according to the respective vertex projection points, the human face projection feature points.
15 . The electronic device according to claim 10 ,
wherein the projection feature information includes human face projection feature points and a human head projection contour line;
wherein before iteratively fitting, based on the two-dimensional feature information and the projection feature information, the standard three-dimensional human face statistical model through the regression loss function to obtain the sample three-dimensional human head model, the operations further comprise randomly sampling the human head contour line and the human head projection contour line respectively to obtain human head contour feature points and human head projection contour feature points; and
wherein the iteratively fitting, based on the two-dimensional feature information and the projection feature information, the standard three-dimensional human face statistical model through the regression loss function to obtain the sample three-dimensional human head model comprises iteratively fitting the standard three-dimensional human face statistical model through the regression loss function to obtain the sample three-dimensional human head model, based on the human face feature points, the human head contour feature points, the human face projection feature points and the human head projection contour feature points.
16 . The electronic device according to claim 10 ,
wherein before the plane projecting the standard three-dimensional human face statistical model to obtain the projection feature information corresponding to the two-dimensional feature information, the operations further comprise detecting a human head posture in the sample portrait image; and
wherein the plane projecting the standard three-dimensional human face statistical model to obtain the projection feature information corresponding to the two-dimensional feature information comprises projecting the standard three-dimensional human face statistical model onto an imaging plane of a capturing device to obtain the projection feature information according to projection parameters of the capturing device for the sample portrait image in response to determining that the standard three-dimensional human face statistical model is in the human head posture.
17 . A non-transitory computer-readable storage medium having stored computer programs which, when executed by a processor, cause the processor to implement operations comprising:
acquiring a portrait image;
inputting the portrait image into a target model to obtain an output result from the target model, wherein the target model pre-trained with a plurality of training samples which are generated according to a sample portrait image and a sample three-dimensional human head model;
generating, according to the output result, a three-dimensional human head model corresponding to the portrait image;
wherein before acquiring the portrait image, the method further comprises:
acquiring the sample portrait image;
extracting two-dimensional feature information from the sample portrait image, wherein the two-dimensional feature information includes human face feature points and a human head projection contour line;
plane projecting a standard three-dimensional human face statistical model to obtain projection feature information corresponding to the two-dimensional feature information; and
iteratively fitting, based on the two-dimensional feature information and the projection feature information, the standard three-dimensional human face statistical model through a regression loss function to obtain the sample three-dimensional human head model.
18 . The non-transitory computer-readable storage medium according to claim 17 , wherein the training samples include the sample portrait image and sample statistical model parameters corresponding to the sample three-dimensional human head model, and the output result includes target statistical model parameters.
19 . The non-transitory computer-readable storage medium according to claim 18 , wherein before acquiring the portrait image, the operations further comprise:
acquiring the plurality of training samples;
learning a mapping relationship between the sample portrait image and the sample statistical model parameters in each of the training samples through a first regression loss function, to obtain the target model; and
continuing to learn the mapping relationship between the sample portrait image and the sample statistical model parameters in each of the training samples through a target loss function, to obtain an optimized target model; wherein the target loss function includes a second regression loss function and a projection loss function, and weight values of identity coefficients in the second regression loss function are greater than weight values of identity coefficients in the first regression loss function.
20 . The non-transitory computer-readable storage medium according to claim 18 , wherein the generating, according to the output result, the three-dimensional human head model corresponding to the portrait image comprises:
generating the three-dimensional human head model according to the target statistical model parameters and the standard three-dimensional human face statistical model.