IP Library Granted Patent US 12,249,178
Granted Patent B2
US 12,249,178 · App. 17/745,158 · Granted Mar 11, 2025

Face reconstruction from a learned embedding

Inventors: Forrester H. Cole (Cambridge, MA); Dilip Krishnan (Arlington, MA); William T. Freeman (Acton, MA); David Benjamin Belanger (Cambridge, MA)
Assignee: GOOGLE LLC
G06V40/16G06F18/2148G06N20/00G06T11/60G06T15/02G06T17/00G06V10/82G06V40/168G06T2210/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,249,178
App. No.
17/745,158
Granted
Mar 11, 2025
Kind
B2
Abstract

The present disclosure provides systems and methods that perform face reconstruction based on an image of a face. In particular, one example system of the present disclosure combines a machine-learned image recognition model with a face modeler that uses a morphable model of a human's facial appearance. The image recognition model can be a deep learning model that generates an embedding in response to receipt of an image (e.g., an uncontrolled image of a face). The example system can further include a small, lightweight, translation model structurally positioned between the image recognition model and the face modeler. The translation model can be a machine-learned model that is trained to receive the embedding generated by the image recognition model and, in response, output a plurality of facial modeling parameter values usable by the face modeler to generate a model of the face.

Claims (42)

1. A computing system comprising:

one or more processors;

one or more non-transitory computer readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining an embedding, wherein the embedding was generated by processing an image of a face with a machine-learned image recognition model;

processing the embedding with a machine-learned translation model to generate a plurality of facial modeling parameter values, wherein the plurality of facial modeling parameter values are descriptive of a plurality of facial attributes of the face in the image, and wherein the machine-learned translation model was trained on example embeddings generated from training data comprising a plurality of face morphs, wherein the plurality of face morphs were generated by: processing a plurality of training images to determine a plurality of facial landmarks, determining an average set of facial landmarks, and generating the plurality of face morphs based on warping the plurality of training images to the average set of facial landmarks;

processing the plurality of facial modeling parameter values with a face modeler to generate a model of the face; and

processing the model of the face with a face renderer to generate a rendering of the face.

2. The computing system of claim 1 , wherein the model of the face comprises a three-dimensional model of the face, and wherein the rendering of the face comprises a two-dimensional rendering of the face.

3. The computing system of claim 1 , wherein the machine-learned translation model was trained on a set of training data that includes a plurality of example embeddings labeled with a plurality of example facial modeling parameter values usable by a face modeler to generate a model of a face.

4. The computing system of claim 3 , wherein the example embeddings are obtained by inputting example faces into the machine-learned image recognition model, said example faces being generated by the face modeler using the respective example facial modeling parameter values.

5. The computing system of claim 1 , wherein the operations further comprise:

processing the rendering of the face with an attribute measurer to determine one or more measured facial attributes; and

processing the one or more measured facial attributes with an artistic renderer to generate an artistic rendering of the face.

6. The computing system of claim 1 , wherein the operations further comprise:

providing the rendering of the face as an avatar in a virtual reality environment.

7. The computing system of claim 1 , wherein the machine-learned translation model has been trained on a training dataset comprising a plurality of example images of faces respectively associated with a plurality of example sets of facial modeling parameter values, wherein at least a portion of the training dataset was generated with a morphable model training set to generate faces that have different facial modeling parameter values.

8. The computing system of claim 1 , wherein obtaining an embedding comprises:

obtaining the image of the face;

inputting the image of the face into a face reconstruction system, wherein the face reconstruction system comprises:

a machine-learned convolutional neural network configured to receive and process the image of the face to generate the embedding for the face.

9. The computing system of claim 1 , wherein the plurality of facial modeling parameter values comprises one or more facial modeling parameter values descriptive of landmark positions for the face.

10. The computing system of claim 1 , further comprising:

a machine-learned face reconstruction model that is operable to receive the embedding for the face and, in response to receipt of the embedding, output a reconstructed representation of the face.

11. The computing system of claim 10 , wherein the machine-learned face reconstruction model comprises the machine-learned translation model and the face renderer.

12. A computer-implemented method, the method comprising:

obtaining, by a computing system comprising one or more processors, an embedding, wherein the embedding was generated by processing an image of a face with a machine-learned image recognition model;

processing, by the computing system, the embedding with a face reconstruction model to generate a model of the face, wherein processing the embedding with the face reconstruction model to generate the face comprises:

processing, by the computing system, the embedding with a machine-learned translation model to generate a plurality of facial modeling parameter values, wherein the plurality of facial modeling parameter values are descriptive of a plurality of facial attributes of the face in the image, and wherein the machine-learned translation model was trained on example embeddings generated from training data comprising a plurality of face morphs, wherein the plurality of face morphs were generated by: processing a plurality of training images to determine a plurality of facial landmarks, determining an average set of facial landmarks, and generating the plurality of face morphs based on warping the plurality of training images to the average set of facial landmarks; and

processing, by the computing system, the plurality of facial modeling parameter values with a face modeler to generate a model of the face.

13. The method of claim 12 , further comprising:

processing, by the computing system, the model of the face with a face renderer to generate a rendering of the face.

14. The method of claim 12 , wherein the machine-learned translation model comprises a regression neural network.

15. The method of claim 12 , wherein the machine-learned translation model comprises one or more fully connected layers.

16. The method of claim 12 , wherein the face reconstruction model comprises a machine-learned variational autoencoder.

17. One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations comprising:

obtaining an embedding, wherein the embedding was generated by processing an image of a face with a machine-learned image recognition model;

processing the embedding with a machine-learned translation model to generate a plurality of facial modeling parameter values, wherein the plurality of facial modeling parameter values are descriptive of a plurality of facial attributes of the face in the image, and wherein the machine-learned translation model was trained on example embeddings generated from training data comprising a plurality of face morphs, wherein the plurality of face morphs were generated by: processing a plurality of training images to determine a plurality of facial landmarks, determining an average set of facial landmarks, and generating the plurality of face morphs based on warping the plurality of training images to the average set of facial landmarks;

processing the plurality of facial modeling parameter values with a face modeler to generate a model of the face; and

processing the model of the face with a face renderer to generate a rendering of the face.

18. The one or more non-transitory computer-readable media of claim 17 , wherein the rendering comprises one or more personalized emojis.

19. The one or more non-transitory computer-readable media of claim 17 , wherein the image is descriptive of a side of the face with uneven lighting, and wherein the rendering comprises a front-facing, evenly-lit rendering of the face.

20. The one or more non-transitory computer-readable media of claim 17 , wherein the image of the face comprises a controlled image of the face with front-facing pose.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2022
From: COLE, FORRESTER H.; KRISHMAN, DILIP; FREEMAN, WILLIAM T.; BELANGER, DAVID BENJAMIN
To: GOOGLE INC.
Reel/Frame 059928/0814 →
CHANGE OF NAME Recorded May 17, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 060081/0318 →
Continuity (4)
Continuation 16857219 · Apr 24, 2020
Continuation 16061344
Provisional Application 62414944 · Oct 31, 2016
Related Publication 20220270402A1 · Aug 25, 2022
References Cited (40)
US 7848548B1 · Moon · 2010 [cited by examiner]
US 8824808B2 · Brandt · 2014 [cited by examiner]
US 9384384B1 · Tyagi · 2016 [cited by examiner]
US 10055879B2 · Wang et al. · 2018 [cited by applicant]
US 10474883B2 · Yu · 2019 [cited by examiner]
US 10650227B2 · Cole · 2020 [cited by examiner]
US 20080192991A1 · Gremse · 2008 [cited by examiner]
US 20100134487A1 · Lai et al. · 2010 [cited by applicant]
US 20130169621A1 · Mei · 2013 [cited by examiner]
US 20150052084A1 · Kolluru · 2015 [cited by examiner]
US 20150347822A1 · Zhou · 2015 [cited by examiner]
US 20170140214A1 · Matas · 2017 [cited by examiner]
US 20170161635A1 · Oono · 2017 [cited by examiner]
US 20170256086A1 · Park · 2017 [cited by examiner]
US 20180137678A1 · Kaehler · 2018 [cited by examiner]
US 20190019012A1 · Huang · 2019 [cited by examiner]
US 20200202111A1 · Yuan · 2020 [cited by examiner]
US 20200334867A1 · Chen · 2020 [cited by examiner]
CN 102054291 · 2011 [cited by applicant]
CN 104063890 · 2014 [cited by applicant]
CN 105359166 · 2016 [cited by applicant]
CN 105404877 · 2016 [cited by applicant]
CN 105447473 · 2016 [cited by applicant]
CN 105760821 · 2016 [cited by applicant]
JP 2006107145 · 2006 [cited by applicant]
JP 2007026088 · 2007 [cited by applicant]
JP 2009075880 · 2009 [cited by applicant]
KR 20090092473 · 2009 [cited by applicant]
Chinese Search Report Corresponding to Application No. 2022103114418 on Dec. 26, 2023. [cited by applicant]
Blanz et al., “A Morphable Model for the Synthesis of 3D Faces”, 26th Annual Conference on Computer Graphics and Interactive Techniques, Los Angeles, California, Aug. 8-13, 1999, pp. 187-194. [cited by applicant]
Cootes et al., “Active Appearance Models”, 5th European Conference on Computer Vision, Freiburg, Germany, Jun. 2-6, 1998, 16 pages. [cited by applicant]
GitHub, “Morphing Faces: Repository for the Morphing Faces Demo”, https://vdumoulin.github.io/morphing faces/, retrieved on Jun. 11, 2018, 6 pages. [cited by applicant]
International Preliminary Report on Patentability for PCT/US2017/053681, mailed on May 2, 2018, 16 pages. [cited by applicant]
International Search Report and Written Opinion from PCT/US2017/053681, mailed on Dec. 6, 2017, 14 pages. [cited by applicant]
Kulkarni et al., “Deep Convolutional Inverse Graphics Network”, Advances in Neural Information Processing Systems, Mar. 11, 2015, 10 pages. [cited by applicant]
Liu et al., “Deep Learning Face Attributes in the Wild” International Conference on Computer Vision, Santiago, Chile, Dec. 11-18, 2015, pp. 3730-3738. [cited by applicant]
Richardson et al., “3D Face Reconstruction by Learning from Synthetic Data”, Fourth International Conference on 3D Vision, Stanford, California, Oct. 25-28, 2016, pp. 460-469. [cited by applicant]
Xue et al., “Face Reconstruction Based on Shape Matching Deformed Model”, Oct. 30, 2016, Electronics Journal, Period 10, DD. 1896-1899 (Abstract Only). [cited by applicant]
Zhmoginov et al., “Inverting Face Embeddings with Convolutional Neural Networks”, arXiv:1606.04189v2, Jun. 7, 2016, 12 pages. [cited by applicant]
Zhong et al., “Leveraging Mid-Level Deep Representations for Predicting Face Attributes in the Wild”, International Conference on Image Processing, Phoenix, Arizona, Sep. 25-28, 2016, pp. 3239-3243. [cited by applicant]