IP Library Granted Patent US 12,243,349
Granted Patent B2
US 12,243,349 · App. 17/697,774 · Granted Mar 4, 2025

Face reconstruction using a mesh convolution network

Inventors: Derek Edward Bradley (Zurich, CH); Prashanth Chandran (Zurich, CH); Simone Foti (London, GB); Paulo Fabiano Urnau Gotardo (Zurich, CH); Gaspard Zoss (Zurich, CH)
Assignees: Disney Enterprises, INC.; ETH Zürich (Eidgenössische Technische Hochschule Zürich)
G06V40/176G06T17/20G06V40/166G06V40/172G06V10/454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,349
App. No.
17/697,774
Granted
Mar 4, 2025
Kind
B2
Abstract

Embodiment of the present invention sets forth techniques for performing face reconstruction. The techniques include generating an identity mesh based on an identity encoding that represents an identity associated with a face in one or more images. The techniques also include generating an expression mesh based on an expression encoding that represents an expression associated with the face in the one or more images. The techniques also include generating, by a machine learning model, an output mesh of the face based on the identity mesh and the expression mesh.

Claims (44)

1. A computer-implemented method for performing reconstruction of a face, the computer-implemented method comprising:

generating an identity mesh based on an identity encoding that represents an identity associated with a face in one or more images;

generating an expression mesh based on an expression encoding that represents an expression associated with the face in the one or more images; and

generating, by a machine learning model, an output mesh of the face via an upsampling operation associated with at least one of the identity mesh or the expression mesh.

2. The computer-implemented method of claim 1 , wherein the expression mesh associates one or more expression features with one or more locations of a mesh topology, and the identity mesh associates one or more identity features with one or more locations of the mesh topology.

3. The computer-implemented method of claim 1 , further comprising generating, based on the one or more images of the face, one or more camera parameters associated with the one or more images.

4. The computer-implemented method of claim 3 , further comprising adjusting one or both of the identity mesh or the expression mesh based on a feature selection, the feature selection being based on the one or more camera parameters.

5. The computer-implemented method of claim 3 , further comprising training the machine learning model based on one or more losses, the one or more losses including one or more of,

an identity loss based on the generated identity mesh and a ground truth identity mesh of the face,

an expression loss based on the generated expression mesh and a ground truth expression mesh of the face,

an output mesh loss based on the generated output mesh and a ground truth mesh, or

a camera parameter loss based on the one or more camera parameters and one or more ground truth camera parameters.

6. The computer-implemented method of claim 1 , wherein each of the expression mesh and the identity mesh includes one or both of a set of vertex coordinates or a set of vertex displacement vectors.

7. The computer-implemented method of claim 1 , wherein a resolution of the output mesh is higher than a resolution of one or both of the identity mesh or the expression mesh.

8. The computer-implemented method of claim 1 , further comprising training the machine learning model based on an identity consistency loss, wherein the identity consistency loss is based on identity encodings associated with each of the one or more images.

9. The computer-implemented method of claim 1 , further comprising normalizing a set of vertices in one or both of the identity mesh or the expression mesh, wherein the normalizing is based on a difference between the one or both of the identity mesh or the expression mesh and an average mesh.

10. One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:

generating an identity mesh based on an identity encoding that represents an identity associated with a face in one or more images;

generating an expression mesh based on an expression encoding that represents an expression associated with the face in the one or more images; and

generating, by a machine learning model, an output mesh of the face via an upsampling operation associated with at least one of the identity mesh or the expression mesh.

11. The one or more non-transitory computer readable media of claim 10 , wherein the expression mesh associates one or more expression features with one or more locations of a mesh topology, and the identity mesh associates one or more identity features with one or more locations of the mesh topology.

12. The one or more non-transitory computer readable media of claim 10 , further comprising generating, based on the one or more images of the face, one or more camera parameters associated with the one or more images.

13. The one or more non-transitory computer readable media of claim 12 , further comprising adjusting one or both of the identity mesh or the expression mesh based on a feature selection, the feature selection being based on the one or more camera parameters.

14. The one or more non-transitory computer readable media of claim 12 , wherein the instructions further cause the one or more processors to perform the step of training the machine learning model based on one or more losses, the one or more losses including one or more of,

an identity loss based on the generated identity mesh and a ground truth identity mesh of the face,

an expression loss based on the generated expression mesh and a ground truth expression mesh of the face,

an output mesh loss based on the generated output mesh and a ground truth mesh, or

a camera parameter loss based on the one or more camera parameters and one or more ground truth camera parameters.

15. The one or more non-transitory computer readable media of claim 10 , wherein each of the expression mesh and the identity mesh includes one or both of a set of vertex coordinates or a set of vertex displacement vectors.

16. The one or more non-transitory computer readable media of claim 10 , wherein a resolution of the output mesh is higher than a resolution of one or both of the identity mesh or the expression mesh.

17. The one or more non-transitory computer readable media of claim 10 , wherein the instructions further cause the one or more processors to perform the step of training the machine learning model based on an identity consistency loss, wherein the identity consistency loss is based on identity encodings associated with each of the one or more images.

18. The one or more non-transitory computer readable media of claim 10 , wherein the instructions further cause the one or more processors to perform the step of normalizing a set of vertices in one or both of the identity mesh or the expression mesh, wherein the normalizing is based on a difference between the one or both of the identity mesh or the expression mesh and an average mesh.

19. A system, comprising:

one or more memories that store instructions, and

one or more processors that are coupled to the one or more memories and,

when executing the instructions, are configured to:

generate an identity mesh based on an identity encoding that represents an identity of a face in one or more images;

generate an expression mesh based on an expression encoding that represents an expression of the face in the one or more images of the face; and

generate, by a machine learning model, an output mesh of the face via an upsampling operation associated with at least one of the identity mesh or the expression mesh.

20. The system of claim 19 , wherein the one or more processors, when executing the instructions, are configured to train the machine learning model based on one or more losses, the one or more losses including one or more of,

an identity loss based on the generated identity mesh and a ground truth identity mesh of the face,

an expression loss based on the generated expression mesh and a ground truth expression mesh of the face,

an output mesh loss based on the generated output mesh and a ground truth mesh, or

a camera parameter loss based on one or more camera parameters and one or more ground truth camera parameters.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2022
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 059338/0847 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2022
From: BRADLEY, DEREK EDWARD; CHANDRAN, PRASHANTH; FOTI, SIMONE; URNAU GOTARDO, PAULO FABIANO; ZOSS, GASPARD
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH; ETH ZÜRICH (EIDGENÖSSISCHE TECHNISCHE HOCHSCHULE ZÜRICH)
Reel/Frame 059323/0657 →
Continuity (2)
Provisional Application 63162204 · Mar 17, 2021
Related Publication 20220301348A1 · Sep 22, 2022
References Cited (22)
US 10331941B2 · Rhee et al. · 2019 [cited by applicant]
US 10692265B2 · Hadap et al. · 2020 [cited by applicant]
US 11074733B2 · Petriv et al. · 2021 [cited by applicant]
US 11869150B1 · Mason · 2024 [cited by examiner]
US 20210279956A1 · Chandran · 2021 [cited by examiner]
CN 109255831B · 2020 [cited by applicant]
CN 111968191A · 2020 [cited by applicant]
CN 112085836A · 2020 [cited by applicant]
Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications”, arXiv:1704.04861, Apr. 17, 2017, 9 pages. [cited by applicant]
Zhou et al., “On the Continuity of Rotation Representations in Neural Networks”, IEEE, CVPR, 2019, pp. 5745-5753. [cited by applicant]
Gong et al., “SpiralNet++: A Fast and Highly Efficient Mesh Convolution Operator”, https://github.com/sw-gong/spiralnet_plus, alarXiv:1911.05856, Nov. 13, 2019, 8 pages. [cited by applicant]
Ranjan et al., “Generating 3D faces using Convolutional Mesh Autoencoders”, ECCV, arxiv:1807.10267, Jul. 31, 2018, 21 pages. [cited by applicant]
Zhou et al., “Fully Convolutional Mesh Autoencoder using Efficient Spatially Varying Kernels”, 34th Conference on Neural Information Processing Systems, arXiv:2006.04325, Oct. 21, 2020, 14 pages. [cited by applicant]
Wang et al., “Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images”, ECCV, arXiv:1804.01654, Aug. 3, 2018, 16 pages. [cited by applicant]
Blanz et al., “A Morphable Model for the Synthesis of the 3D Faces”, SIGGRAPH 99, https://doi.org/10.1145/311535.311556, 1999, pp. 187-194. [cited by applicant]
Gecer et al., “GANFIT: Generative Adversarial Network Fitting for High Fidelity 3D Face Reconstruction”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, arXiv:1902.05978, Apr. 6, 2019, … [cited by applicant]
Tewari et al., “MoFA: Model-based Deep Convolutional Face Autoencoder for Unsupervised Monocular Reconstruction”, In Proceedings of the IEEE International Conference on Computer Vision Workshops, arXiv:1703.10580, Dec. … [cited by applicant]
Cheng et al., “Faster, Better and More Detailed: 3D Face Reconstruction with Graph Convolutional Networks”, In Proceedings of the Asian Conference on Computer Vision, 2020, 18 pages. [cited by applicant]
Lee et al., “Uncertainty-Aware Mesh Decoder for High Fidelity 3D Face Reconstruction”, DOI 10.1109/CVPR42600.2020.00614, CVPR, 2020, pp. 6100-6109. [cited by applicant]
Lin et al., “Towards High-Fidelity 3D Face Reconstruction from In-the-Wild Images Using Graph Convolutional Networks”, CVPR, 2020, pp. 5891-5900. [cited by applicant]
Zhou et al., “Dense 3D Face Decoding over 2500FPS: Joint Texture & Shape Convolutional Mesh Decoders”, CVPR, arXiv:1904.03525, Apr. 6, 2019, pp. 1097-1106. [cited by applicant]
Shu et al., “Neural Face Editing with Intrinsic Image Disentangling”, arXiv:1704.04131, Apr. 13, 2017, 22 pages. [cited by applicant]
Cited By (1)
US 12,620,260