IP Library Patent Application 17082895
Patent Application
App. No. 17/082,895

Systems and Methods for Building a Skin-to-Muscle Transformation in Computer Animation

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/082,895
Abstract

An animation system wherein a machine learning model is adopted to learn a transformation relationship between facial muscle movements and skin surface movements. For example, for the skin surface representing “smile,” the transformation model derives movement vectors relating to what facial muscles are activated, what are the muscle strains, what is the joint movement, and/or the like. Such derived movement vectors may be used to simulate the skin surface “smile.”

Claims (50)

1 . A computer-implemented method for learning a skin deformation of a facial action in an animation system, the method comprising:

receiving a plurality of facial scans of a face of an actor representing a plurality of facial movements over a data bundle time period;

generating, from the plurality of facial scans, a data bundle comprising a first cache of a time-varying facial muscle strain vector over the data bundle time period, a second cache of a time-varying skin surface vector over the data bundle time period;

inputting the data bundle and anatomical data corresponding to the actor to a machine learning model;

generating, by the machine learning model, a predicted muscle strain vector based on the second cache of the time-varying skin surface vector over the data bundle time period and the anatomical data corresponding to the actor;

computing a loss objective based on the predicted muscle strain vector and a ground truth strain from the first cache of the time-varying facial muscle strain vector over the data bundle time period; and

updating parameters of the machine learning model by minimizing the computed loss objective.

2 . The method of claim 1 , wherein the machine learning model is a deep neural network, and the loss objective is any of a cross-entropy loss, a L2 loss and a KL-distance loss.

3 . The method of claim 1 , further comprising:

generating a training dataset of a plurality of data bundles, wherein the plurality of data bundles corresponding to multiple types of facial movements including any combination of a facial action, a dialog, and a depiction of an emotion.

4 . The method of claim 1 , further comprising:

deriving a transformational relationship between the skin surface vector and the facial muscle strain vectors from the updated learning model.

5 . The method of claim 4 , further comprising:

determining, based on the transformational relationship, at least one derived facial muscle strain vector given a target skin surface vector; and

sending the at least one derived facial muscle strain vector to an animation creation system for creating an animated skin surface based on the derived facial muscle strain vector.

6 . The method of claim 1 , wherein the plurality of facial scans includes a first facial scan of a neutral pose of the actor and a second facial scan of a non-neural pose of the actor, and wherein the second facial scan is a period of time apart from the first facial scan.

7 . The method of claim 1 , wherein the plurality of facial scans are collected from the actor across a period of time, and the ground truth strain is obtained from facial scans that are averaged out among similar facial scans collected from the period of time.

8 . The method of claim 1 , wherein the plurality of facial movements include one or more of a facial action, a dialog, and/or a depiction of an emotion.

9 . A system for learning a skin deformation of a facial action in an animation system, the system comprising:

a communication interface that receives a plurality of facial scans of a face of an actor representing a plurality of facial movements over a data bundle time period;

a memory storing a plurality of processor-executable instructions; and

one or more hardware processors that read the plurality of processor-executable instructions to perform:

generating, from the plurality of facial scans, a data bundle comprising a first cache of a time-varying facial muscle strain vector over the data bundle time period, a second cache of a time-varying skin surface vector over the data bundle time period;

inputting the data bundle and anatomical data corresponding to the actor to a machine learning model;

generating, by the machine learning model, a predicted muscle strain vector based on the second cache of the time-varying skin surface vector over the data bundle time period and the anatomical data corresponding to the actor;

computing a loss objective based on the predicted muscle strain vector and a ground truth strain from the first cache of the time-varying facial muscle strain vector over the data bundle time period; and

updating parameters of the machine learning model by minimizing the computed loss objective.

10 . The system of claim 9 , wherein the machine learning model is a deep neural network, and the loss objective is any of a cross-entropy loss, a L2 loss and a KL-distance loss.

11 . The system of claim 10 , wherein the one or more hardware processors read the plurality of processor-executable instructions to further perform:

generating a training dataset of a plurality of data bundles, wherein the plurality of data bundles corresponding to multiple types of facial movements including any combination of a facial action, a dialog, and a depiction of an emotion.

12 . The system of claim 9 , wherein the one or more hardware processors read the plurality of processor-executable instructions to further perform:

deriving a transformational relationship between the skin surface vector and the facial muscle strain vectors from the updated learning model.

13 . The system of claim 12 , wherein the one or more hardware processors read the plurality of processor-executable instructions to further perform:

determining, based on the transformational relationship, at least one derived facial muscle strain vector given a target skin surface vector; and

sending the at least one derived facial muscle strain vector to an animation creation system for creating an animated skin surface based on the derived facial muscle strain vector.

14 . The system of claim 9 , wherein the plurality of facial scans includes a first facial scan of a neutral pose of the actor and a second facial scan of a non-neural pose of the actor, and wherein the second facial scan is a period of time apart from the first facial scan.

15 . The system of claim 9 , wherein the plurality of facial scans are collected from the actor across a period of time, and the ground truth strain is obtained from facial scans that are averaged out among similar facial scans collected from the period of time.

16 . The system of claim 9 , wherein the plurality of facial movements include one or more of a facial action, a dialog, and/or a depiction of an emotion.

17 . A processor-readable non-transitory medium storing processor-executable instructions for learning a skin deformation of a facial action in an animation system, the processor-executable instructions being executed by one or more hardware processors to perform:

receiving a plurality of facial scans of a face of an actor representing a plurality of facial movements over a data bundle time period;

generating, from the plurality of facial scans, a data bundle comprising a first cache of a time-varying facial muscle strain vector over the data bundle time period, a second cache of a time-varying skin surface vector over the data bundle time period;

inputting the data bundle and anatomical data corresponding to the actor to a machine learning model;

generating, by the machine learning model, a predicted muscle strain vector based on the second cache of the time-varying skin surface vector over the data bundle time period and the anatomical data corresponding to the actor;

computing a loss objective based on the predicted muscle strain vector and a ground truth strain from the first cache of the time-varying facial muscle strain vector over the data bundle time period; and

updating parameters of the machine learning model by minimizing the computed loss objective.

18 . The medium of claim 17 , wherein the machine learning model is a deep neural network, and the loss objective is any of a cross-entropy loss, a L2 loss and a KL-distance loss.

19 . The medium of claim 17 , wherein the one or more hardware processors read the plurality of processor-executable instructions to further perform:

generating a training dataset of a plurality of data bundles, wherein the plurality of data bundles corresponding to multiple types of facial movements including any combination of a facial action, a dialog, and a depiction of an emotion.

20 . The medium of claim 17 , wherein the one or more hardware processors read the plurality of processor-executable instructions to further perform:

deriving a transformational relationship between the skin surface vector and the facial muscle strain vectors from the updated learning model.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2022
From: UNITY SOFTWARE INC.
To: UNITY TECHNOLOGIES SF
Reel/Frame 058980/0369 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2022
From: WETA DIGITAL LIMITED
To: UNITY SOFTWARE INC.
Reel/Frame 058978/0905 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2021
From: CHOI, BYUNG KUK
To: WETA DIGITAL LIMITED
Reel/Frame 055408/0693 →