IP Library › Granted Patent US 10,559,111
Granted Patent B2
US 10,559,111 · App. 16/221,321 · Granted Feb 11, 2020

Systems and methods for generating computer ready animation models of a human head from captured data images

Inventors: Ian Sachs (Larkspur, CA); Kiran Bhat (San Francisco, CA); Dominic Monn (Mels, CH); Senthil Radhakrishnan (Oakland, CA); Will Welch (San Francisco, CA)
Assignee: LoomAi, Inc.
G06T13/40G06F3/005G06K9/00208G06K9/00234G06K9/00255G06K9/00281G06K9/00308G06K9/4628G06K9/6218G06K9/6273G06T7/11G06T7/13G06T7/246G06T7/248G06T7/73G06T7/90G06T13/205G06T15/04G06T17/20G06T19/20H04N7/157G06T2207/10016G06T2207/10024G06T2207/20021G06T2207/20036G06T2207/20081G06T2207/20084G06T2207/30201G06T2219/2012G06T2219/2021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,559,111
App. No.
16/221,321
Granted
Feb 11, 2020
Kind
B2
Abstract

System and methods for computer animations of 3D models of heads generated from images of faces is disclosed. A 2D captured image that includes an image of a face can be received and used to generate a static 3D model of a head. A rig can be fit to the static 3D model to generate an animation-ready 3D generative model. Sets of rigs can be parameters that each map to particular sounds or particular facial movement observed in a video. These mappings can be used to generate a playlists of sets of rig parameters based upon received audio or video content. The playlist may be played in synchronization with an audio rendition of the audio content. Methods can receive a captured image, identify taxonomy attributes from the captured image, select a template model for the captured image, and perform a shape solve for the selected template model based on the identified taxonomy attributes.

Claims (43)

1. A method for generating a three dimensional (3D) head model from a captured image, the method comprising:

receiving a captured image;

identifying a set of taxonomy attributes from the captured image by:

using a single encoder model to generate an embedding of the captured image;

using a set of one or more terminal models to analyze the generated embedding; and

identifying taxonomy attributes based on the classifications of the terminal models;

selecting a template model for the captured image; and

performing a shape solve for the selected template model based on the identified taxonomy attributes.

2. The method of claim 1 , wherein the set of taxonomy attributes comprises a set of one or more local taxonomy attributes and a set of one or more global taxonomy attributes.

3. The method of claim 2 further comprising performing a texture synthesis based on one or more of the taxonomy attributes.

4. The method of claim 2 , wherein the template model is selected based on the set of global taxonomy attributes.

5. The method of claim 2 , wherein identifying the set of global taxonomy attributes comprises using a classifier model to classify the captured image.

6. The method of claim 2 , wherein the set of global taxonomy attributes comprises at least one of an ethnicity and a gender.

7. The method of claim 2 , wherein the set of local taxonomy attributes comprise one or more of head shape, eyelid turn, eyelid height, nose width, nose turn, mustache, beard, sideburns, eye rotation, jaw angle, lip thickness, and chin shape.

8. The method of claim 1 , wherein each terminal model of the set of terminal models produces a score for an associated local taxonomy attribute, wherein performing the shape solve for a given local taxonomy attribute is based on the score calculated for the local taxonomy attribute.

9. The method of claim 1 further comprising:

training the single encoder model and the set of terminal models, wherein each terminal model is trained to classify a different taxonomy attribute.

10. The method of claim 9 , wherein training the encoder comprises:

generating an embedding from the single encoder model;

calculating a first loss for a first terminal model based on the generated embedding;

back propagating the first calculated loss through the encoder;

calculating a second loss for a second terminal model based on the generated embedding; and

back propagating the second calculated loss through the single encoder model.

11. The method of claim 1 , wherein receiving the captured image comprises capturing an image with a camera of a mobile device, wherein performing the shape solve is performed on the same mobile device.

12. A non-transitory machine readable medium containing processor instructions for generating a three dimensional (3D) head model from a captured image, where execution of the instructions by a processor causes the processor to perform a process that comprises:

receiving a captured image;

identifying a set of taxonomy attributes from the captured image by:

using a single encoder model to generate an embedding of the captured image;

using a set of one or more terminal models to analyze the generated embedding; and

identifying taxonomy attributes based on the classifications of the terminal models;

selecting a template model for the captured image; and

performing a shape solve for the selected template model based on the identified taxonomy attributes.

13. The non-transitory machine readable medium of claim 12 , wherein the set of taxonomy attributes comprises a set of one or more local taxonomy attributes and a set of one or more global taxonomy attributes, wherein the template model is selected based on the set of global attributes.

14. The non-transitory machine readable medium of claim 13 , wherein the set of global taxonomy attributes comprises at least one of an ethnicity and a gender.

15. The non-transitory machine readable medium of claim 13 , wherein execution of the instruction further causes the processor to perform a texture synthesis based on one or more of the taxonomy attributes.

16. The non-transitory machine readable medium of claim 12 , wherein the process further comprises training the single encoder model and the set of terminal models, wherein each terminal model is trained to classify a different taxonomy attribute.

17. The non-transitory machine readable medium of claim 16 , wherein training the encoder comprises:

generating an embedding from the single encoder model;

calculating a first loss for a first terminal model based on the generated embedding;

back propagating the first calculated loss through the encoder;

calculating a second loss for a second terminal model based on the generated embedding; and

back propagating the second calculated loss through the single encoder model.

18. The non-transitory machine readable medium of claim 12 , wherein receiving the captured image comprises capturing an image with a camera of a mobile device, wherein performing the shape solve is performed on the same mobile device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2019
From: SACHS, IAN; BHAT, KIRAN; MONN, DOMINIC; RADHAKRISHNAN, SENTHIL; WELCH, WILL
To: LOOMAI, INC.
Reel/Frame 048242/0912 →
Continuity (6)
Continuation In Part 15885667 · Jan 31, 2018
Continuation 15632251 · Jun 23, 2017
Provisional Application 62767393 · Nov 14, 2018
Provisional Application 62353944 · Jun 23, 2016
Provisional Application 62367233 · Jul 27, 2016
Related Publication 20190122411A1 · Apr 25, 2019
Cited By (11)
US 12,198,389 US 12,277,665 US 12,380,622 US 12,401,822 US 12,423,861 US 12,439,083 US 12,548,229 US 12,594,030 US 12,614,364 US 12,626,432 US 12,749,190