IP Library Granted Patent US 12,731,316
Granted Patent B2
US 12,731,316 · App. 18/976,898 · Granted Sep 8, 2026

Generating a realistic animated avatar of a user in real-time during a teleconference

Inventors: Alexander Tormasov (Hochrhein, DE); Serg Bell (Singapore, SG); Stanislav Protasov (Singapore, SG); Nikolay Dobrovolskiy (Alanya, TR); Laurent Dedenis (Geneva, CH)
Assignee: Constructor Technology AG
G06T13/205G06F3/012G06F3/017G06T13/40G10L13/047G10L13/10G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,316
App. No.
18/976,898
Granted
Sep 8, 2026
Kind
B2
Abstract

Disclosed herein are systems and method for generating animated avatars of users in real-time during a teleconference. The method includes training AI avatar generation models to create an avatar of a first user, deploying an AI avatar generation agent on a communication device, collecting sensor data from different sensors associated with the first user and sending the collected sensor data to the communication device of the second user, activating a data processing model and to identify the types of sensor data received from the communication device of the first user and activating the AI avatar generation agent of the second user to execute, on the communication device of the second user, the plurality of AI avatar generation models, and displaying, on the communication device of second user, the animated avatar of the first user synced with the real-time audio of the voice of the first user during the teleconference.

Claims (63)

1 . A method for generating animated avatars of users in real-time during a teleconference, the method comprising:

training a plurality of AI avatar generation models to create an avatar of a first user, wherein each AI avatar generation model is trained using different types of sensor data;

deploying an AI avatar generation agent on a communication device of a second user or on a cloud server;

in response to a teleconference call being initiated between a communication device of the first user and a communication device of the second user, collecting sensor data from a plurality of different sensors associated with the first user, wherein the sensor data comprises at least real-time audio of voice of the first user;

activating a data processing model to identify the types of sensor data received from the communication device of the first user and activating the AI avatar generation agent to execute one or more of the plurality of AI avatar generation models, corresponding to the identified types of sensor data, for generating, based at least on the received sensor data, an animated avatar of the first user, wherein the animated avatar simulates at least one of: physical likeness, facial expressions, speech mannerisms, co-speech gestures, and voice of the first user during the teleconference; and

displaying, on the communication device of second user, the animated avatar of the first user synced with the real-time audio of the voice of the first user during the teleconference.

2 . The method of claim 1 , wherein the plurality of different sensors comprises one or more of:

a wearable sensor configured to measure a head position or a head movement of the first user,

a wearable sensor with interior-facing cameras configured to capture face movement or lip-sync movement of the first user,

a wearable Wi-Fi signal strength measurement device configured to measure Wi-Fi strength in accordance with gestures of the first user,

a microphone configured to capture the real-time audio of the voice of the first user, and

an input device configured to capture text from the first user.

3 . The method of claim 2 , wherein the plurality of AI avatar generation models to generate the animated avatar of the first user comprises at least one or more of:

a head position AI recognition model to predict a head position or head movement of the first user based on using the wearable sensor to measure the head position of the first user in relation with a body of the first user when the first user is speaking, wherein the head position AI recognition model is trained to predict the head position using a head position training set comprising of a sequence of images of users speaking and a head position label identifying each head position in the sequence of images;

a mimic AI recognition model to predict facial expressions or lip-sync of the first user based on using the wearable sensor with interior-facing cameras to capture face movement when the first user is speaking, wherein the mimic AI recognition model is trained to predict the facial expressions of the first user using a mimic head position training set comprising of a sequence of images of users speaking and a facial expression label identifying a facial expression in the sequence of images;

a gesture AI recognition model to predict gestures of the first user based on using the wearable Wi-Fi signal strength measurement device to detect changes in a Wi-Fi field around the first user when the first user is speaking, wherein the gesture AI recognition model is trained to predict the gestures of the first user using a gesture training set comprising of a sequence of images of users and a gesture label identifying a gesture in the sequence of images;

a lip-sync AI recognition model to predict a lip-sync of the first user based on using the wearable sensor with interior-facing cameras or the microphone to detect speech patterns when the first user is speaking, wherein the lip-sync AI recognition model is trained to predict lip-sync of the first user using audio files matched to sequence of images of users and a lip-sync label identifying a lip-sync movement audio files matched to the sequence of images; or

an emotion AI recognition model to predict emotions of the first user based on using the microphone to capture the voice of the first user, wherein the emotion AI recognition model is trained to predict emotions of the first user using audio files of users and an emotion label identifying an emotion in the audio files; or

a voice generation model to generate computer-generated speech for the first user based on using text obtained from an input device of the first user in real-time, wherein the voice generation model is trained to predict speech of the first user using audio files of the first user and a text-to-speech (TTS) model.

4 . The method of claim 1 , wherein the plurality of AI avatar generation models are trained, stored, and executed on the cloud server.

5 . The method of claim 1 , wherein the plurality of AI avatar generation models are trained, stored, and executed on a wearable device, invasive implant, non-invasive implant, teleconference device, or edge device.

6 . The method of claim 1 , further comprising:

based on a determination that a Wi-Fi strength of a wearable Wi-Fi signal strength measurement device of the first user does not pass a threshold, displaying, on the communication device of the second user, a basic avatar of the first user without animations along with the real-time audio of the voice of the first user during the teleconference.

7 . The method of claim 6 , further comprising:

based on a determination that the Wi-Fi strength of the wearable Wi-Fi signal strength measurement device of the first user passes the threshold, updating, on the communication device of the second user, the display of the basic avatar to a display of the animated avatar of the first user along with the real-time audio of the voice of the first user during the teleconference.

8 . The method of claim 1 , wherein the speech mannerisms comprises one or more of: frequency of pauses, length of the pauses, talking speed, tone, or diction.

9 . The method of claim 1 , wherein the co-speech gestures comprises at least head movement, facial feature movement, gestures, lip-sync movement, and body part movement of the first user.

10 . The method of claim 1 , wherein the voice comprise at least one of gender, tone, emphasis, emotions, speech defects, and prosody of the first user.

11 . The method of claim 1 , wherein based on one or more types of sensor data not being available from the communication device of the first user, the AI avatar generation agent uses one or more of the AI avatar generation models to predict the one or more of the physical likeness, facial expressions, speech mannerisms, co-speech gestures, and audio of the first user based on available sensor data or previously collected sensor data.

12 . The method of claim 1 , further comprising:

collecting the sensor data from a plurality of different sensors associated with the first user and sending the collected sensor data to the AI avatar generation agent based on a determination that a video of the first user is not available.

13 . The method of claim 1 , wherein the data processing model is deployed on a cloud fog and the AI avatar generation agent are deployed on a cloud.

14 . A method for generating animated avatars of users in real-time during a teleconference, the method comprising:

training a plurality of AI avatar generation models to create an avatar of a first user, wherein each AI avatar generation model is trained using different types of sensor data;

deploying an AI avatar generation agent on a cloud server;

in response to a teleconference call being initiated between a communication device of the first user and a communication device of the second user, collecting sensor data from a plurality of different sensors associated with the first user, wherein the sensor data comprises at least text from the first user;

activating a data processing model and of the second user to identify the types of sensor data received from the communication device of the first user and activating the AI avatar generation agent to execute one or more of the plurality of AI avatar generation models, corresponding to the identified types of sensor data, for generating, based at least on the received sensor data, an animated avatar of the first user, wherein the animated avatar simulates at least one of: physical likeness, facial expressions, speech mannerisms, co-speech gestures, and voice of the first user during the teleconference,

wherein the one or more of the plurality of AI avatar generation models comprises at least a voice generation model configured to generate computer-generated speech for the first user from text of the first user using the AI avatar generation agent on the cloud server; and

displaying, on the communication device of second user, the animated avatar of the first user with the computer-generated speech for the first user synced with the text obtained from the first user in real-time during the teleconference.

15 . A system for generating animated avatars of users in real-time during a teleconference, comprising:

at least one memory; and

at least one hardware processor coupled with the at least one memory and configured, individually or in combination, to:

train a plurality of AI avatar generation models to create an avatar of a first user, wherein each AI avatar generation model is trained using different types of sensor data;

deploy an AI avatar generation agent on a communication device of a second user or on a cloud server;

in response to a teleconference call being initiated between a communication device of the first user and a communication device of the second user, collect sensor data from a plurality of different sensors associated with the first user, wherein the sensor data comprises at least real-time audio of voice of the first user;

activate a data processing model to identify the types of sensor data received from the communication device of the first user and activating the AI avatar generation agent to execute one or more of the plurality of AI avatar generation models, corresponding to the identified types of sensor data, for generating, based at least on the received sensor data, an animated avatar of the first user, wherein the animated avatar simulates at least one of: physical likeness, facial expressions, speech mannerisms, co-speech gestures, and voice of the first user during the teleconference; and

display, on the communication device of second user, the animated avatar of the first user synced with the real-time audio of the voice of the first user during the teleconference.

16 . The system of claim 15 , wherein the plurality of different sensors comprises one or more of:

a wearable sensor configured to measure a head position or a head movement of the first user,

a wearable sensor with interior-facing cameras configured to capture face movement or lip-sync movement of the first user,

a wearable Wi-Fi signal strength measurement device configured to measure Wi-Fi strength in accordance with gestures of the first user,

a microphone configured to capture the real-time audio of the voice of the first user, and

an input device configured to capture text from the first user.

17 . The system of claim 16 , wherein the plurality of AI avatar generation models to generate the animated avatar of the first user comprises at least one or more of:

a head position AI recognition model to predict a head position or head movement of the first user based on using the wearable sensor to measure the head position of the first user in relation with a body of the first user when the first user is speaking, wherein the head position AI recognition model is trained to predict the head position using a head position training set comprising of a sequence of images of users speaking and a head position label identifying each head position in the sequence of images;

a mimic AI recognition model to predict facial expressions or lip-sync of the first user based on using the wearable sensor with interior-facing cameras to capture face movement when the first user is speaking, wherein the mimic AI recognition model is trained to predict the facial expressions of the first user using a mimic head position training set comprising of a sequence of images of users speaking and a facial expression label identifying a facial expression in the sequence of images;

a gesture AI recognition model to predict gestures of the first user based on using the wearable Wi-Fi signal strength measurement device to detect changes in a Wi-Fi field around the first user when the first user is speaking, wherein the gesture AI recognition model is trained to predict the gestures of the first user using a gesture training set comprising of a sequence of images of users and a gesture label identifying a gesture in the sequence of images;

a lip-sync AI recognition model to predict a lip-sync of the first user based on using the wearable sensor with interior-facing cameras or the microphone to detect speech patterns when the first user is speaking, wherein the lip-sync AI recognition model is trained to predict lip-sync of the first user using audio files matched to sequence of images of users and a lip-sync label identifying a lip-sync movement audio files matched to the sequence of images; or

an emotion AI recognition model to predict emotions of the first user based on using the microphone to capture the voice of the first user, wherein the emotion AI recognition model is trained to predict emotions of the first user using audio files of users and an emotion label identifying an emotion in the audio files; or

a voice generation model to generate computer-generated speech for the first user based on using text obtained from an input device of the first user in real-time, wherein the voice generation model is trained to predict speech of the first user using audio files of the first user and a text-to-speech (TTS) model.

18 . The system of claim 15 , wherein the plurality of AI avatar generation models are trained, stored, and executed on the cloud server.

19 . The system of claim 15 , wherein the plurality of AI avatar generation models are trained, stored, and executed on a teleconference device or edge device.

20 . The system of claim 15 , based on one or more types of sensor data not being available from the communication device of the first user, the AI avatar generation agent uses one or more of the AI avatar generation models to predict the one or more of the physical likeness, facial expressions, speech mannerisms, co-speech gestures, and audio of the first user based on available sensor data or previously collected sensor data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2024
From: TORMASOV, ALEXANDER; BELL, SERG; PROTASOV, STANISLAV; DOBROVOLSKIY, NIKOLAY; DEDENIS, LAURENT
To: CONSTRUCTOR TECHNOLOGY AG
Reel/Frame 069552/0761 →
Continuity (1)
Related Publication 20260162342A1 · Jun 11, 2026
References Cited (12)
US 11582424B1 · Kasaba · 2023 [cited by examiner]
US 11620780B2 · Lee · 2023 [cited by examiner]
US 11741651B2 · Stewart · 2023 [cited by examiner]
US 20130155169A1 · Hoover · 2013 [cited by examiner]
US 20170206797A1 · Solomon · 2017 [cited by examiner]
US 20170365084A1 · Hayashida · 2017 [cited by examiner]
US 20200306640A1 · Kolen · 2020 [cited by examiner]
US 20230130287A1 · Zhao · 2023 [cited by examiner]
US 20230223022A1 · Singh · 2023 [cited by examiner]
US 20240046536A1 · Prasad · 2024 [cited by examiner]
US 20240078731A1 · Beith · 2024 [cited by examiner]
US 20250061634A1 · Huang · 2025 [cited by examiner]