IP Library Granted Patent US 11,683,448
Granted Patent B2
US 11,683,448 · App. 17/543,519 · Granted Jun 20, 2023

System, method, and computer program for transmitting face models based on face data points

Inventors: William Guie Rivard (Menlo Park, CA); Brian J. Kindle (Sunnyvale, CA); Adam Barry Feder (Mountain View, CA)
Assignee: DUELIGHT LLC
H04N7/157G06T13/40G06T19/00G10L21/10H04L12/1827H04L51/04H04L51/046H04L65/403H04L65/80H04N7/147G06T2219/024G10L25/63G10L2021/105H04L51/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,683,448
App. No.
17/543,519
Granted
Jun 20, 2023
Kind
B2
Abstract

A system, method, and computer program are provided for receiving face models based on face nodal points. In use, a real-time face model is received, wherein the real-time face model includes one or more face nodal points. Real-time face nodal points are received, including additional one or more face nodal points. The real-time face model is manipulated based on the real-time face nodal points.

Claims (56)

1. A method comprising:

receiving, by one or more processors through an audio input, an audio output of a person over a plurality of time instances, as an audio stream;

acquiring, by the one or more processors through an imaging device, images of facial expressions of the person over the plurality of time instances, as corresponding image data;

inferring, by the one or more processors, phonetic elements of the audio output according to the audio stream;

translating, by the one or more processors, the inferred phonetic elements of the audio output into a correlation map;

determining, by the one or more processors, a plurality of geometric shapes and face data points, according to the corresponding image data, to form a three-dimensional (3D) model of an avatar of the person incorporating the facial expressions of the person, the plurality of geometric shapes comprising 3D structures, the face data points indicating an amount of transformation applied to the 3D structures of the 3D model of the avatar;

combining, by the one or more processors, the correlation map with the 3D model of the avatar to form a 3D representation of the avatar, by synchronizing the correlation map with the 3D model of the avatar over the plurality of time instances; and

rendering, by the one or more processors, the 3D representation of the avatar with the incorporated facial expressions of the person in synchronization with rendering of the audio stream.

2. The method of claim 1 , wherein inferring the phonetic elements of the audio output comprises determining, by the one or more processors, probabilities of the phonetic elements of the audio output according to the audio stream.

3. The method of claim 1 , comprising determining, by the one or more processors, the plurality of geometric shapes and face data points according to animation parameters.

4. The method of claim 1 , comprising determining the face data points by:

determining, by the one or more processors, a base set of main data points of a face of the person from the corresponding image data;

determining, by the one or more processors, a first set of facial depths;

determining, by the one or more processors, a first candidate set of main data points using a 3D model of the face formed according to the first set of facial depths;

comparing, by the one or more processors, the base set of main data points and the first candidate set of main data points; and

determining, by the one or more processors, a second set of facial depths that results in a second candidate set of main data points, and a difference between the second candidate set of main data points and the base set of main data points that is less than a difference between the first candidate set of main data points and the base set of main data points.

5. The method of claim 1 , comprising generating the 3D model of the avatar by combining, by the one or more processors, the plurality of geometric shapes according to the face data points.

6. The method of claim 5 , wherein combining the a correlation map with the 3D model of the avatar includes animating or rendering, by the one or more processors, at least a portion of a mouth of the 3D model corresponding to a first time instance, according to one of the correlation map corresponding to the first time instance.

7. A system comprising:

an audio input;

an imaging device;

one or more processors coupled to the audio input and the imaging device; and

a non-transitory computer readable medium storing instructions when executed by the one or more processors cause the one or more processors to:

receive, through the audio input, an audio output of a person over a plurality of time instances, as an audio stream;

acquire, through the imaging device, images of facial expressions of the person over the plurality of time instances, as corresponding image data;

infer phonetic elements of the audio output according to the audio stream;

translate the inferred phonetic elements of the audio output into a correlation map;

determine a plurality of geometric shapes and face data points, according to the corresponding image data of the face, to form a three-dimensional (3D) model of an avatar of the person incorporating the facial expressions of the person, the plurality of geometric shapes comprising 3D structures, the face data points indicating an amount of transformation applied to the 3D structures of the 3D model of the avatar;

combine the correlation map with the 3D model of the avatar to form a 3D representation of the avatar, by synchronizing the correlation map with the 3D model of the avatar over the plurality of time instances; and

render the 3D representation of the avatar with the incorporated facial expressions of the person in synchronization with rendering of the audio stream.

8. The system of claim 7 , wherein the one or more processors infer the phonetic elements of the audio output by determining probabilities of the phonetic elements of the audio output according to the audio stream.

9. The system of claim 7 , wherein the one or more processors determine the plurality of geometric shapes and face data points according to animation parameters.

10. The system of claim 7 , wherein the one or more processors determine the face data points by:

determining a base set of main data points of a face of the person from the corresponding image data;

determining a first set of facial depths;

determining a first candidate set of main data points using a 3D model of the face formed according to the first set of facial depths;

comparing the base set of main data points and the first candidate set of main data points; and

determining a second set of facial depths that results in a second candidate set of main data points, and a difference between the second candidate set of main data points and the base set of main data points that is less than a difference between the first candidate set of main data points and the base set of main data points.

11. The system of claim 7 , wherein the one or more processors generate the 3D model of the avatar by combining the plurality of geometric shapes according to the face data points.

12. The system of claim 11 , wherein the one or more processors combine the correlation map with the 3D model of the avatar by animating or rendering at least a portion of a mouth of the 3D model corresponding to a first time instance, according to one of the correlation map corresponding to the first time instance.

13. A non-transitory computer readable medium storing instructions when executed by one or more processors cause the one or more processors to:

receive, through an audio input, an audio output of a person over a plurality of time instances, as an audio stream;

acquire, through an imaging device, images of facial expressions of the person over the plurality of time instances, as corresponding image data;

infer phonetic elements of the audio output according to the audio stream;

translate the inferred phonetic elements of the audio output into a correlation map;

determine a plurality of geometric shapes and face data points, according to the corresponding image data of the face, to form a three-dimensional (3D) model of an avatar of the person incorporating the facial expressions of the person, the plurality of geometric shapes comprising 3D structures, the face data points indicating an amount of transformation applied to the 3D structures of the 3D model of the avatar;

combine the correlation map with the 3D model of the avatar to form a 3D representation of the avatar, by synchronizing the correlation map with the 3D model of the avatar over the plurality of time instances; and

render the 3D representation of the avatar with the incorporated facial expressions of the person in synchronization with rendering of the audio stream.

14. The non-transitory computer readable medium of claim 13 , wherein the instructions that cause the one or more processors to infer the phonetic elements of the audio output further comprise instructions when executed by the one or more processors cause the one or more processors to determine probabilities of the phonetic elements of the audio output according to the audio stream.

15. The non-transitory computer readable medium of claim 13 , further comprising instructions when executed by the one or more processors cause the one or more processors to determine the plurality of geometric shapes and face data points according to animation parameters.

16. The non-transitory computer readable medium of claim 13 , wherein the instructions that cause the one or more processors to determine the face data points further comprise instructions when executed by the one or more processors cause the one or more processors to:

determine a base set of main data points of a face of the person from the corresponding image data;

determine a first set of facial depths;

determine a first candidate set of main data points using a 3D model of the face formed according to the first set of facial depths;

compare the base set of main data points and the first candidate set of main data points; and

determine a second set of facial depths that results in a second candidate set of main data points, and a difference between the second candidate set of main data points and the base set of main data points that is less than a difference between the first candidate set of main data points and the base set of main data points.

Assignments (2)
SECURITY INTEREST Recorded Jul 5, 2023
From: DUELIGHT, LLC
To: ZILKA-KOTAB PC
Reel/Frame 064207/0585 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2021
From: RIVARD, WILLIAM GUIE; KINDLE, BRIAN J.; FEDER, ADAM BARRY
To: DUELIGHT LLC
Reel/Frame 058328/0424 →
Continuity (5)
Continuation 17108867 · Dec 1, 2020
Continuation 16547358 · Aug 21, 2019
Continuation 16206241 · Nov 30, 2018
Provisional Application 62618520 · Jan 17, 2018
Related Publication 20220239866A1 · Jul 28, 2022
Cited By (5)
US 12,401,911 US 12,401,912 US 12,418,727 US 12,445,736 US 12,666,159