AVATAR GENERATION IN A VIDEO COMMUNICATIONS PLATFORM
Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for generating an avatar within a video communication platform. The system may receive a selection of an avatar model from a group of one or more avatar models. The system receives a first video stream and audio data of a first video conference participant. The system analyzes image frames of the first video stream to determine a group of pixels representing the first video conference participant. The system determines a plurality of facial expression parameter associated with the determined group of pixels. Based on the determined plurality of facial expression parameter values, the system generates a first modified video stream depicting a digital representation of the first video conference participant in an avatar form.
1 . A computer-implemented method comprising:
receiving a selection of an avatar model from a group of one or more avatar models;
receiving a first video stream comprising multiple image frames of a first video conference participant;
inputting at least a group of pixels of each of the multiple image frames into a trained machine learning network;
determining by the trained machine learning network, a plurality of facial expression parameter values associated with the multiple images;
generating a first modified video stream by:
based on the plurality of facial expression parameter values, morphing a three-dimensional head mesh of the selected avatar model; and
rendering a digital representation of the first video conference participant in an avatar form; and
providing for display, in a user interface of a video conferencing environment, the first modified video stream.
2 . The method of claim 1 , wherein the morphing a three-dimensional head mesh comprises:
selecting one or more blendshapes based on the determined plurality of facial expression parameter values; and
applying the selected one or more blendshapes to modify a mesh geometry of the selected avatar model.
3 . The method of claim 1 , further comprising:
receiving a second modified video stream of a video conference participant, the second modified video stream comprising a digital representation of a second video conference participant in an avatar form; and
providing for display, in a user interface of the video conferencing environment, the second modified video stream.
4 . The method of claim 1 , further comprising:
receiving a selection of a first virtual background for use with the selected avatar model; and
wherein the first modified video stream depicts the digital representation of the first video conference participant in avatar form overlayed on the selected first virtual background.
5 . The method of claim 4 , further comprising:
determining that the first video conference participant is not being captured in the first video stream; and
changing the first modified video stream to depict the selected first virtual background without the digital representation of the first video conference participant in an avatar form; and
providing for display, in the user interface of a video conferencing environment, the changed first modified video stream.
6 . The method of claim 1 wherein the plurality of facial expression parameter associated include one or more action unit values and associated intensity values.
7 . The method of claim 1 , wherein the plurality of facial expression parameter values comprises at least 51 different action unit values.
8 . A non-transitory computer readable medium that stores executable program instructions that when executed by one or more computing devices configure the one or more computing devices to perform operations comprising:
receiving a selection of an avatar model from a group of one or more avatar models;
receiving a first video stream comprising multiple image frames of a first video conference participant;
inputting at least a group of pixels of each of the multiple image frames into a trained machine learning network;
determining by the trained machine learning network, a plurality of facial expression parameter values associated with the multiple images;
generating a first modified video stream by:
based on the determined plurality of facial expression parameter values, morphing a three-dimensional head mesh of the selected avatar model; and
rendering a digital representation of the first video conference participant in an avatar form; and
providing for display, in a user interface of a video conferencing environment, the first modified video stream.
9 . The non-transitory computer readable medium of claim 8 , wherein the operation of morphing a three-dimensional head mesh comprises the operations of:
selecting one or more blendshapes based on the determined plurality of facial expression parameter values; and
applying the one or more blendshapes to modify a mesh geometry of the selected avatar model.
10 . The non-transitory computer readable medium of claim 8 , further comprising the operations of:
receiving a second modified video stream of a video conference participant, the second modified video stream comprising a digital representation of a second video conference participant in an avatar form; and
providing for display, in a user interface of the video conferencing environment, the second modified video stream.
11 . The non-transitory computer readable medium of claim 8 , further comprising the operations of:
receiving a selection of a first virtual background for use with the selected avatar model; and
wherein the first modified video stream depicts the digital representation of the first video conference participant in avatar form overlayed on the selected first virtual background.
12 . The non-transitory computer readable medium of claim 8 , further comprising the operations of:
determining that the first video conference participant is not being captured in the first video stream; and
changing the first modified video stream to depict the selected first virtual background without the digital representation of the first video conference participant in an avatar form; and
providing for display, in the user interface of a video conferencing environment, the changed first modified video stream.
13 . The non-transitory computer readable medium of claim 8 , wherein the plurality of facial expression parameter associated include one or more action unit values and associated intensity values.
14 . The non-transitory computer readable medium of claim 8 , wherein the plurality of facial expression parameter values comprises at least 51 different action unit values.
15 . A system comprising one or more processors configured to perform the operations of:
receiving a selection of an avatar model from a group of one or more avatar models;
receiving a first video stream comprising multiple image frames of a first video conference participant;
inputting at least a group of pixels of each of the multiple image frames into a trained machine learning network;
determining by the trained machine learning network, a plurality of facial expression parameter values associated with the multiple image frames;
generating a first modified video stream by:
based on the determined plurality of facial expression parameter values, morphing a three-dimensional head mesh of the selected avatar model; and
rendering a digital representation of the first video conference participant in an avatar form; and
providing for display, in a user interface of a video conferencing environment, the first modified video stream.
16 . The system of claim 15 , wherein morphing a three-dimensional head mesh comprises:
selecting one or more blendshapes based on the generated plurality of facial expression parameter values; and
applying the one or more blendshapes to modify a mesh geometry of the selected avatar model.
17 . The system of claim 15 , further comprising the operations of:
receiving a second modified video stream of a video conference participant, the second modified video stream comprising a digital representation of a second video conference participant in an avatar form; and
providing for display, in a user interface of the video conferencing environment, the second modified video stream.
18 . The system of claim 15 , further comprising the operations of:
receiving a selection of a first virtual background for use with the selected avatar model; and
wherein the first modified video stream depicts the digital representation of the first video conference participant in avatar form overlayed on the selected first virtual background.
19 . The system of claim 18 , further comprising the operations of:
determining that the first video conference participant is not being captured in the first video stream; and
changing the first modified video stream to depict the selected first virtual background without the digital representation of the first video conference participant in an avatar form; and
providing for display, in the user interface of a video conferencing environment, the changed first modified video stream.
20 . The system of claim 18 , wherein the plurality of facial expression parameter associated include one or more action unit values and associated intensity values.