IP Library Patent Application 17589771
Patent Application
App. No. 17/589,771

AVATAR GENERATION IN A VIDEO COMMUNICATIONS PLATFORM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/589,771
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for generating an avatar within a video communication platform. The system may receive a selection of an avatar model from a group of one or more avatar models. The system receives a first video stream and audio data of a first video conference participant. The system analyzes image frames of the first video stream to determine a group of pixels representing the first video conference participant. The system determines a plurality of facial expression parameter associated with the determined group of pixels. Based on the determined plurality of facial expression parameter values, the system generates a first modified video stream depicting a digital representation of the first video conference participant in an avatar form.

Claims (71)

1 . A computer-implemented method comprising:

receiving a selection of an avatar model from a group of one or more avatar models;

receiving a first video stream comprising multiple image frames of a first video conference participant;

inputting at least a group of pixels of each of the multiple image frames into a trained machine learning network;

determining by the trained machine learning network, a plurality of facial expression parameter values associated with the multiple images;

generating a first modified video stream by:

based on the plurality of facial expression parameter values, morphing a three-dimensional head mesh of the selected avatar model; and

rendering a digital representation of the first video conference participant in an avatar form; and

providing for display, in a user interface of a video conferencing environment, the first modified video stream.

2 . The method of claim 1 , wherein the morphing a three-dimensional head mesh comprises:

selecting one or more blendshapes based on the determined plurality of facial expression parameter values; and

applying the selected one or more blendshapes to modify a mesh geometry of the selected avatar model.

3 . The method of claim 1 , further comprising:

receiving a second modified video stream of a video conference participant, the second modified video stream comprising a digital representation of a second video conference participant in an avatar form; and

providing for display, in a user interface of the video conferencing environment, the second modified video stream.

4 . The method of claim 1 , further comprising:

receiving a selection of a first virtual background for use with the selected avatar model; and

wherein the first modified video stream depicts the digital representation of the first video conference participant in avatar form overlayed on the selected first virtual background.

5 . The method of claim 4 , further comprising:

determining that the first video conference participant is not being captured in the first video stream; and

changing the first modified video stream to depict the selected first virtual background without the digital representation of the first video conference participant in an avatar form; and

providing for display, in the user interface of a video conferencing environment, the changed first modified video stream.

6 . The method of claim 1 wherein the plurality of facial expression parameter associated include one or more action unit values and associated intensity values.

7 . The method of claim 1 , wherein the plurality of facial expression parameter values comprises at least 51 different action unit values.

8 . A non-transitory computer readable medium that stores executable program instructions that when executed by one or more computing devices configure the one or more computing devices to perform operations comprising:

receiving a selection of an avatar model from a group of one or more avatar models;

receiving a first video stream comprising multiple image frames of a first video conference participant;

inputting at least a group of pixels of each of the multiple image frames into a trained machine learning network;

determining by the trained machine learning network, a plurality of facial expression parameter values associated with the multiple images;

generating a first modified video stream by:

based on the determined plurality of facial expression parameter values, morphing a three-dimensional head mesh of the selected avatar model; and

rendering a digital representation of the first video conference participant in an avatar form; and

providing for display, in a user interface of a video conferencing environment, the first modified video stream.

9 . The non-transitory computer readable medium of claim 8 , wherein the operation of morphing a three-dimensional head mesh comprises the operations of:

selecting one or more blendshapes based on the determined plurality of facial expression parameter values; and

applying the one or more blendshapes to modify a mesh geometry of the selected avatar model.

10 . The non-transitory computer readable medium of claim 8 , further comprising the operations of:

receiving a second modified video stream of a video conference participant, the second modified video stream comprising a digital representation of a second video conference participant in an avatar form; and

providing for display, in a user interface of the video conferencing environment, the second modified video stream.

11 . The non-transitory computer readable medium of claim 8 , further comprising the operations of:

receiving a selection of a first virtual background for use with the selected avatar model; and

wherein the first modified video stream depicts the digital representation of the first video conference participant in avatar form overlayed on the selected first virtual background.

12 . The non-transitory computer readable medium of claim 8 , further comprising the operations of:

determining that the first video conference participant is not being captured in the first video stream; and

changing the first modified video stream to depict the selected first virtual background without the digital representation of the first video conference participant in an avatar form; and

providing for display, in the user interface of a video conferencing environment, the changed first modified video stream.

13 . The non-transitory computer readable medium of claim 8 , wherein the plurality of facial expression parameter associated include one or more action unit values and associated intensity values.

14 . The non-transitory computer readable medium of claim 8 , wherein the plurality of facial expression parameter values comprises at least 51 different action unit values.

15 . A system comprising one or more processors configured to perform the operations of:

receiving a selection of an avatar model from a group of one or more avatar models;

receiving a first video stream comprising multiple image frames of a first video conference participant;

inputting at least a group of pixels of each of the multiple image frames into a trained machine learning network;

determining by the trained machine learning network, a plurality of facial expression parameter values associated with the multiple image frames;

generating a first modified video stream by:

based on the determined plurality of facial expression parameter values, morphing a three-dimensional head mesh of the selected avatar model; and

rendering a digital representation of the first video conference participant in an avatar form; and

providing for display, in a user interface of a video conferencing environment, the first modified video stream.

16 . The system of claim 15 , wherein morphing a three-dimensional head mesh comprises:

selecting one or more blendshapes based on the generated plurality of facial expression parameter values; and

applying the one or more blendshapes to modify a mesh geometry of the selected avatar model.

17 . The system of claim 15 , further comprising the operations of:

receiving a second modified video stream of a video conference participant, the second modified video stream comprising a digital representation of a second video conference participant in an avatar form; and

providing for display, in a user interface of the video conferencing environment, the second modified video stream.

18 . The system of claim 15 , further comprising the operations of:

receiving a selection of a first virtual background for use with the selected avatar model; and

wherein the first modified video stream depicts the digital representation of the first video conference participant in avatar form overlayed on the selected first virtual background.

19 . The system of claim 18 , further comprising the operations of:

determining that the first video conference participant is not being captured in the first video stream; and

changing the first modified video stream to depict the selected first virtual background without the digital representation of the first video conference participant in an avatar form; and

providing for display, in the user interface of a video conferencing environment, the changed first modified video stream.

20 . The system of claim 18 , wherein the plurality of facial expression parameter associated include one or more action unit values and associated intensity values.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2023
From: CHEN, WENYU; FU, CHICHEN; HU, GUOZHU; LIN, WENCHONG; LING, BO; LIU, GENGDAI; LI, QIANG; WANG, GENG; WEI, KAI; ZHU, YIAN; LI, WENHAO
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 063148/0655 →