System and method for generating avatar of an active speaker in a meeting
A method includes receiving a facial data associated with a participant user; generating an avatar data of the participant user based on the facial data using a first machine learning (ML) model; receiving data associated with facial movement; and training the generated avatar based on the data associated with facial movement using a second ML model to generate a trained avatar data, wherein the trained avatar data mimics appropriate facial movements associated with audio data.
1 . A computer-implemented method comprising:
receiving at least one image of a face of a participant user;
generating a static avatar data of the participant user based on the at least one image using a first machine learning (ML) model;
receiving video data associated with the participant user, wherein the video data includes lips movement, eyebrows movement, jaw movement, and eyes movement of the participant user;
using a second ML model to generate animated avatar data for generating an avatar video stream, wherein input for the second ML model includes the static avatar data, the video data, and a corresponding audio stream of the participant user, wherein the second ML model is trained using the participant user's own speech to mimic the participant user's specific lips movements and accent;
wherein the animated avatar data mimics appropriate facial movements associated with the corresponding audio stream;
determining, from a meeting audio stream associated with a meeting, that the participant user is an active speaker at the meeting;
in response to determining that the participant user is the active speaker at the meeting, determining whether a meeting video stream associated with the meeting contains facial content of the active speaker;
in response to determining that the participant user is the active speaker and that the meeting video stream does not include facial content of the active speaker, retrieving the animated avatar data associated with the active speaker; and
generating, based on the animated avatar data, an avatar video stream of the active speaker mimicking the meeting audio stream, wherein the active speaker is the participant user.
2 . The computer-implemented method of claim 1 , wherein determining that the meeting video stream does not include facial content of the active speaker comprises determining that the meeting video stream contains a whiteboard or shared screen.
3 . The computer-implemented method of claim 1 further comprising:
injecting the generated avatar video stream to data being transmitted to participants of the meeting when the participant user is speaking as the active speaker.
4 . The computer-implemented method of claim 3 , wherein the injecting complements data being transmitted during the meeting session without video stream of the participant user.
5 . The computer-implemented method of claim 1 , wherein the determining that the participant user has become the active speaker comprises performing voice recognition processing of the meeting audio stream.
6 . The computer-implemented method of claim 1 , wherein the video data comprises video streams of one or more users speaking and wherein the one or more users are different from the participant user.
7 . The computer-implemented method of claim 1 , wherein the at least one image comprises at least one static image of the participant user from an application that is different from an application that facilitates an online meeting for the participant user.
8 . The computer-implemented method of claim 1 , wherein the at least one image comprises at least one static image of the participant user from an that facilitates an online meeting for the participant user.
9 . The computer-implemented method of claim 1 , wherein the at least one image comprises at least a portion of a video stream.
10 . One or more non-transitory computer readable storage medium storing instructions that when executed by one or more processors, cause:
receiving at least one image of a face of a participant user;
generating a static avatar data of the participant user based on the at least one image using a first machine learning (ML) model;
receiving video data associated with the participant user, wherein the video data includes lips movement, eyebrows movement, jaw movement, and eyes movement of the participant user;
using a second ML model to generate animated avatar data for generating an avatar video stream, wherein input for the second ML model includes the static avatar data, the video data, and a corresponding audio stream of the participant user, wherein the second ML model is trained using the participant user's own speech to mimic the participant user's specific lips movements and accent;
wherein the animated avatar data mimics appropriate facial movements associated with the corresponding audio stream;
determining, from a meeting audio stream associated with a meeting, that the participant user is an active speaker at the meeting;
in response to determining that the participant user is the active speaker at the meeting, determining whether a meeting video stream associated with the meeting contains facial content of the active speaker;
in response to determining that the participant user is the active speaker and that the meeting video stream does not include facial content of the active speaker, retrieving the animated avatar data associated with the active speaker; and
generating, based on the animated avatar data, an avatar video stream of the active speaker mimicking the meeting audio stream, wherein the active speaker is the participant user.
11 . The non-transitory computer readable storage medium of claim 10 , wherein the determining that the participant user has become the active speaker comprises performing voice recognition processing of the meeting audio stream.
12 . The non-transitory computer readable storage medium of claim 10 , further comprising storing instructions that when executed by one or more processors, cause:
injecting the generated avatar video stream to data being transmitted to participants of the meeting when the participant user is speaking as the active speaker.
13 . The non-transitory computer readable storage medium of claim 12 , wherein the injecting complements data being transmitted during the meeting session without video stream of the participant user.
14 . A system, comprising:
a processor;
a memory operatively connected to the processor and storing instructions that, when executed by the processor, cause:
receiving at least one image of a face of a participant user;
generating a static avatar data of the participant user based on the at least one image using a first machine learning (ML) model;
receiving video data associated with the participant user, wherein the video data includes lips movement, eyebrows movement, jaw movement, and eyes movement of the participant user; and
using a second ML model to generate animated avatar data for generating an avatar video stream, wherein input for the second ML model includes the static avatar data, the video data, and a corresponding audio stream of the participant user, wherein the second ML model is trained using the participant user's own speech to mimic the participant user's specific lips movements and accent;
wherein the animated avatar data mimics appropriate facial movements associated with the corresponding audio stream
determining, from a meeting audio stream associated with a meeting, that the participant user is an active speaker at the meeting;
in response to determining that the participant user is the active speaker at the meeting, determining whether a meeting video stream associated with the meeting contains facial content of the active speaker;
in response to determining that the participant user is the active speaker and that the meeting video stream does not include facial content of the active speaker, retrieving the animated avatar data associated with the active speaker; and
generating, based on the animated avatar data, an avatar video stream of the active speaker mimicking the meeting audio stream, wherein the active speaker is the participant user.
15 . The system of claim 14 , wherein determining that the meeting video stream does not include facial content of the active speaker comprises determining that the video stream contains a whiteboard or shared screen.
16 . The system of claim 14 , wherein the instructions when executed by the process further cause:
injecting the generated avatar video stream into the data being transmitted to participants of the meeting when the participant user is speaking as the active speaker.