IP Library › Granted Patent US 9,378,576
Granted Patent B2
US 9,378,576 · App. 13/912,378 · Granted Jun 28, 2016

Online modeling for real-time facial animation

Inventors: Sofien Bouaziz (Lausanne, CH); Mark Pauly (Lausanne, CH)
Assignee: faceshift AG
G06T13/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,378,576
App. No.
13/912,378
Filed
Jun 7, 2013
Granted
Jun 28, 2016
Kind
B2
Art Unit
2665
USPC
382/215
Abstract

Embodiments relate to a method for real-time facial animation, and a processing device for real-time facial animation. The method includes providing a dynamic expression model, receiving tracking data corresponding to a facial expression of a user, estimating tracking parameters based on the dynamic expression model and the tracking data, and refining the dynamic expression model based on the tracking data and estimated tracking parameters. The method may further include generating a graphical representation corresponding to the facial expression of the user based on the tracking parameters. Embodiments pertain to a real-time facial animation system.

Claims (62)

1. A method for real-time facial animation, comprising:

providing a dynamic expression model that includes a plurality of blendshapes and one or more corrective deformation fields;

receiving tracking data corresponding to a facial expression of a user;

estimating tracking parameters based on the dynamic expression model and the tracking data, wherein the tracking parameters include a plurality of weights for the plurality of blendshapes; and

refining the dynamic expression model based on the tracking data and the estimated tracking parameters to produce a refined dynamic expression model, wherein the refining includes applying at least one of the one or more corrective deformation fields to each of the plurality of blendshapes.

2. The method of claim 1 , wherein estimating tracking parameters and refining the dynamic expression model are performed in real-time.

3. The method of claim 1 , wherein estimating tracking parameters is performed in a first stage and refining the dynamic expression model is performed in a second stage, wherein the first stage and the second stage are iteratively repeated.

4. The method of claim 1 , further comprising generating a graphical representation corresponding to the facial expression of the user based on the tracking parameters.

5. The method of claim 1 , further comprising:

receiving further tracking data corresponding to a further facial expression of the user;

estimating updated tracking parameters based on the refined dynamic expression model and the further tracking data; and

generating a graphical representation corresponding to the further facial expression of the user based on the updated tracking parameters.

6. The method of claim 1 , wherein the plurality of blendshapes includes a blendshape b 0 representing a neutral facial expression and the dynamic expression model further includes an identity model, wherein the method further comprises matching the blendshape b 0 representing the neutral facial expression to a neutral expression of the user based on the identity model.

7. The method of claim 6 , wherein the plurality of blendshapes further includes one or more blendshapes b i , each representing a corresponding facial expression of one or more additional facial expressions, wherein the dynamic expression model further includes a template blendshape model, wherein the method further comprises approximating the one or more blendshapes b i based on the template blendshape model and the blendshape b 0 representing the neutral facial expression.

8. The method of claim 7 , further comprising parameterizing the one or more blendshapes b i as b i =T* i b 0 +Ez i , wherein T* i is an expression transfer operator derived from the template blendshape model, and Ez i is a corrective deformation field for the blendshape b i .

9. The method of claim 7 , wherein the plurality of blendshapes comprises n blendshapes, wherein refining the dynamic expression model includes:

determining a coverage coefficient σ i , for i=0, 1, . . . , n−1, for each blendshape b i , for i=0, 1, . . . , n−1, of the dynamic expression model, wherein the coverage coefficient σ i is indicative of the applicability of past tracking data corresponding to past facial expressions of the user for the blendshape b i ; and refining each blendshape b i , for i=0, 1, . . . , n−1, having a coverage coefficient σ i , for i=0, 1, . . . , n−1, below a pre-determined threshold.

10. The method of claim 1 , further comprising:

receiving neutral facial expression tracking data corresponding to a neutral facial expression of the user; and

initializing the dynamic expression model using the tracking data corresponding to the neutral facial expression of the user.

11. The method of claim 1 , wherein refining the dynamic expression model based on the tracking data and the estimated tracking parameters comprises refining the dynamic expression model based on the tracking data, the estimated tracking parameters, and aggregated tracking data, wherein the aggregated tracking data is based on additional received tracking data corresponding to one or more past facial expressions of the user.

12. The method of claim 11 , further comprising aggregating the additional received tracking data corresponding to one or more past facial expressions of the user subject to a decay over time to produce the aggregated tracking data.

13. The method of claim 1 , further comprising:

automatically creating a virtual avatar of the user based on the dynamic expression model; and

using the virtual avatar for generating a graphical representation corresponding to the facial expression of the user.

14. The method of claim 13 , wherein the graphical representation corresponding to the facial expression of the user is generated based on one or more blendshapes of the plurality of blendshapes representing a virtual avatar.

15. A processing device, comprising:

an input interface configured to receive tracking data corresponding to a facial expression of a user;

a memory configured to store a dynamic expression model that includes a plurality of blendshapes and one or more corrective deformation fields; and

a processing component coupled to the input interface and the memory, wherein the processing component is configured to:

estimate tracking parameters based on the dynamic expression model and the tracking data, wherein the tracking parameters include a plurality of weights for the plurality of blendshapes and

refine the dynamic expression model based on the tracking data and the estimated tracking parameters to produce a refined dynamic expression model, wherein the refining the dynamic model includes applying at least one of the one or more corrective deformation fields to each of the plurality of blendshapes.

16. The device of claim 15 , wherein the processing component is further configured to estimate the tracking parameters and refine the dynamic expression model in real-time.

17. The device of claim 15 , wherein the processing component is further configured to generate a graphical representation corresponding to the facial expression of the user based on the tracking parameters.

18. The device of claim 15 , wherein the plurality of blendshapes includes a blendshape b 0 representing a neutral facial expression and the dynamic expression model further includes an identity model, wherein the processing component is further configured to match the blendshape b 0 representing the neutral facial expression to a neutral expression of the user based on the identity model.

19. The device of claim 18 , wherein the plurality of blendshapes further includes one or more blendshapes b i , each representing a corresponding facial expression of one or more additional facial expressions, wherein the dynamic expression model further includes a template blendshape model, wherein the processing component is further configured to approximate the one or more blendshapes b i based on the template blendshape model and the blendshape b 0 representing the neutral facial expression.

20. The device of claim 19 , wherein the plurality of blendshapes has n blendshapes, wherein, in order to refine the dynamic expression model, the processing component is further configured to:

determine a coverage coefficient σ i , for 1=0, 1, . . . , n−1, for each blendshape b i , for i=0, 1, . . . , n−1, of the dynamic expression model, wherein the coverage coefficient σ i for i=0, 1, . . . , n−1, is indicative of the applicability of past tracking data corresponding to past facial expressions for the blendshape b i , for i=0, 1, . . . , n−1, and only refine each blendshape b i , for i=0, 1, . . . , n−1, having a coverage coefficient σ i , for i=0, 1, . . . , n−1, below a pre-determined threshold.

21. The device of claim 15 , wherein the processing component is further configured to refine the dynamic expression model based on the tracking data, the estimated tracking parameters, and aggregated tracking data, wherein the aggregated tracking data is based on additional received tracking data corresponding to one or more past facial expressions of the user, wherein the aggregated tracking data is produced by aggregating the additional received tracking data corresponding to one or more past facial expressions of the user subject to a decay over time.

22. A real-time facial animation system, comprising:

a camera device, wherein the camera device is configured to generate tracking data corresponding to a facial expression of a user; and

a processing device, wherein the processing device comprises:

an input interface coupled to the camera device and configured to receive the tracking data,

a memory configured to store a dynamic expression model, wherein the dynamic expression model includes a plurality of blendshapes and one or more corrective deformation fields for each of the blendshapes, and

a processing component coupled to the input interface and the memory, wherein the processing component is configured to—

estimate a corresponding plurality of weights for the plurality of blendshapes of the dynamic expression model based on the tracking data,

generate a graphical representation corresponding to the facial expression of the user based on the plurality of weights, and

refine the dynamic expression model based on the tracking data, the plurality of weights for the blendshapes, and the one or more corrective deformation fields.

23. The system of claim 22 , wherein the camera device is configured to generate tracking data including video data and depth information.

24. A non-transitory program storage device comprising instructions stored thereon to cause one or more processors to:

provide a dynamic expression model that includes a plurality of blendshapes and one or more corrective deformation fields;

receive tracking data corresponding to a facial expression of a user;

estimate tracking parameters based on the dynamic expression model and the tracking data, wherein the tracking parameters include a plurality of weights for the plurality of blendshapes; and

refine the dynamic expression model based on the tracking data and the estimated tracking parameters to produce a refined dynamic expression model, wherein the refining includes applying at least one of the one or more corrective deformation fields to each of the plurality of blendshapes.

25. The non-transitory program storage device of claim 24 , wherein the instructions to cause the one or more processors to estimate comprise instructions to cause the one or more processors to estimate tracking parameters in a first stage and refine the dynamic expression model in a second stage, wherein the first stage and the second stage are iteratively repeated.

26. The non-transitory program storage device of claim 24 , further comprising instructions to cause the one or more processors to:

receive further tracking data corresponding to a further facial expression of the user;

estimate updated tracking parameters based on the refined dynamic expression model and the further tracking data; and

generate a graphical representation corresponding to the further facial expression of the user based on the updated tracking parameters.

27. The non-transitory program storage device of claim 24 , further comprising instructions to cause the one or more processors to:

receive neutral facial expression tracking data corresponding to a neutral facial expression of the user; and

initialize the dynamic expression model using the tracking data corresponding to the neutral facial expression of the user.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2019
From: FACESHIFT GMBH
To: APPLE INC.
Reel/Frame 050170/0136 →
CHANGE OF NAME Recorded Aug 26, 2019
From: FACESHIFT AG
To: FACESHIFT GMBH
Reel/Frame 050170/0061 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2015
From: ECOLE POLYTECHNIQUE FEDERALE DE LAUSANNE
To: FACESHIFT AG
Reel/Frame 036445/0429 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2013
From: BOUAZIZ, SOFIEN; PAULY, MARK
To: ECOLE POLYTECHNIQUE FEDERALE DE LAUSANNE
Reel/Frame 030607/0515 →
Continuity (1)
Related Publication 20140362091A1 · Dec 11, 2014