Systems, methods and media for deep shape prediction
Exemplary embodiments include a computer-implemented method of training a neural network for facial reconstruction including collecting a set of 3D head scans, combining each feature of each 3D head scan with a weight to create a modified set of 3D head scans, training the neural network using the modified set of head scans, and inputting a real digital facial image into the neural network for facial reconstruction. Further exemplary embodiments include the set of 3D head scans comprising approximately a tenth or less in quantity in comparison to a quantity of the modified set of 3D head scans. The modified set of 3D head scans may comprise features found in the set of 3D head scans or the modified set of 3D head scans may consist of features found in the set of 3D head scans.
1 . A computer-implemented method of training a neural network for facial reconstruction comprising:
collecting a synthetic dataset of non-existing 3D head scans produced by randomly sampling a 3D morphable model and rendered under diverse lighting conditions, camera settings, and pose conditions, wherein the diverse lighting conditions comprise rendering using high dynamic range images (HDRIs);
combining each feature of each 3D head scan with a projected weight to create a modified set of 3D head scans;
measuring an error between the projected weight and an actual weight;
adjusting neural network weights for the error and repeating the measuring and adjusting until the error converges or is near or at zero;
training the neural network using the modified set of 3D head scans; and
inputting a real digital facial image into the neural network for the facial reconstruction.
2 . The computer-implemented method of claim 1 , the synthetic dataset of non-existing 3D head scans comprising approximately a tenth or less in quantity in comparison to a quantity of the modified set of 3D head scans.
3 . The computer-implemented method of claim 1 , wherein the modified set of 3D head scans comprises features found in the synthetic dataset of non-existing 3D head scans.
4 . The computer-implemented method of claim 1 , the facial reconstruction resulting in an estimate of a subject's head geometry based on a weighted sum of a plurality of individual modified 3D head scans.
5 . The computer-implemented method of claim 1 , the facial reconstruction performed without including a face of an actual human in the modified set of 3D head scans.
6 . The computer-implemented method of claim 1 , the facial reconstruction including recognition of a feature on the modified set of 3D head scans.
7 . The computer-implemented method of claim 6 , the feature being a dimension of a nose.
8 . The computer-implemented method of claim 6 , the feature being a dimension of an ear.
9 . The computer-implemented method of claim 1 , the facial reconstruction resulting in an estimate of a subject's jawline shape.
10 . The computer-implemented method of claim 1 , the facial reconstruction resulting in an estimate of a thickness of a subject's lip.
11 . The computer-implemented method of claim 1 , further comprising combining each feature of each 3D head scan with the projected weight, wherein the projected weight is randomly sampled from a 3D morphable-model parameter distribution derived from principal-component analysis of a training dataset, to create a modified set of 3D head scans.
12 . The computer-implemented method of claim 11 , wherein the error is computed using a loss function that combines prediction error of the 3D morphable-model parameter distribution and prediction error of a concurrently predicted normal map.
13 . The computer-implemented method of claim 12 , further comprising adjusting the neural network's weights for the error, wherein the adjustment is performed by an Adam optimizer executing with a fixed learning rate.
14 . The computer-implemented method of claim 13 , further comprising stopping the method when the error converges.
15 . The computer-implemented method of claim 13 , further comprising stopping the method when the error is near or at zero.
16 . The computer-implemented method of claim 1 , the facial reconstruction resulting in an estimate of a subject's shape of a face.