Three-dimensional avatar generation system
A computing system receives video data and non-video data from a user device for generation of an avatar corresponding to a user of the user device. The computing system generates an avatar reflective of the appearance and the behavior of the user by inputting the video data and non-video data to a generative machine learning model trained to generate highly realistic avatars for users. The avatar may look, behave, sound, and interact like the user. The video data is pre-processed to remove background information from the video data and isolate the user within the video data. The computing system clones a voice of the user for use with the avatar. The computing system stores the avatar and the cloned voice of the user in a network accessible location. The computing system deploys the avatar in a third-party application for real-time or near real-time interaction with a second user.
1 . A method, comprising:
receiving, by a computing system, video data and non-video data from a user device for generation of an avatar corresponding to a user of the user device, the video data comprising a video of the user of the user device, the non-video data comprising information associated with a desired appearance and behavior of the avatar;
generating, by the computing system, the avatar reflective of the appearance and the behavior of the user by inputting the video data and non-video data to a generative machine learning model trained to generate highly realistic avatars for users, wherein the avatar looks, behaves, sounds, and interacts like the user, wherein the video data is pre-processed to remove background information from the video data and isolate the user within the video data;
cloning, by the computing system, a voice of the user for use with the avatar, wherein the cloning comprises extracting audio data from the video data and generating a synthetic voice for the avatar using a model trained on voice characteristics of the user, wherein speech output of the avatar matches one or more of a tone, pace, inflection, or accent of the user;
storing, by the computing system, the avatar and the cloned voice of the user in a network accessible location;
deploying, by the computing system, the avatar in a third-party application for real-time or near real-time interaction with a second user, wherein responses of the avatar are generated using a large language model and are delivered with synchronized lip movements generated by a lip-sync module based on learned speech patterns of the user; and
defining, by the computing system, one or more session-specific guardrails for the avatar while interacting in the third-party application, the one or more session-specific guardrails comprising domain-based, administrator-defined restrictions on avatar output, wherein the one or more session-specific guardrails limit the avatar to interfacing with one or more knowledge sources, the one or more knowledge sources comprising pre-approved, session-specific database or retrieval-augmented generation (RAG) sources.
2 . The method of claim 1 , wherein storing, by the computing system, the avatar and the cloned voice of the user in the network accessible location comprises:
storing the avatar and the cloned voice of the avatar on a blockchain, wherein the avatar is minted as a non-fungible token linked to a verified identity of the user.
3 . The method of claim 1 , wherein generating, by the computing system, the avatar reflective of the appearance and the behavior of the user comprises:
extracting one or more gestures of the user from the video data;
extracting one or more facial expressions of the user from the video data; and
causing the avatar to replicate the one or more gestures and the one or more facial expressions during interaction with the avatar.
4 . The method of claim 1 , further comprising:
dynamically transitioning, by the computing system, the avatar between the user device, one or more of edge components, and cloud servers based on device capability.
5 . The method of claim 1 , wherein storing, by the computing system, the avatar and the cloned voice of the user in the network accessible location comprises:
applying one or more watermarks to the avatar, wherein the one or more watermarks are bound to an interaction session in which the avatar exists, the one or more watermarks providing authentication to the avatar.
6 . A system comprising:
a non-transitory storage medium storing computer program instructions; and
a processor configured to execute the computer program instructions to cause operations comprising:
receiving video data and non-video data from a user device for generation of an avatar corresponding to a user of the user device, the video data comprising a video of the user of the user device, the non-video data comprising information associated with a desired appearance and behavior of the avatar;
generating the avatar reflective of the appearance and the behavior of the user by inputting the video data and non-video data to a generative machine learning model trained to generate highly realistic avatars for users, wherein the avatar looks, behaves, sounds, and interacts like the user, wherein the video data is pre-processed to remove background information from the video data and isolate the user within the video data;
cloning a voice of the user for use with the avatar, wherein the cloning comprises extracting audio data from the video data and generating a synthetic voice for the avatar using a model trained on voice characteristics of the user, wherein speech output of the avatar matches one or more of a tone, pace, inflection, or accent of the user;
storing the avatar and the cloned voice of the user in a network accessible location;
deploying the avatar in a third-party application for real-time or near real-time interaction with a second user, wherein responses of the avatar are generated using a large language model and are delivered with synchronized lip movements generated by a lip-sync module based on learned speech patterns of the user; and
defining one or more session-specific guardrails for the avatar while interacting in the third-party application, the one or more session-specific guardrails comprising domain-based, administrator-defined restrictions on avatar output, wherein the one or more session-specific guardrails limit the avatar to interfacing with one or more knowledge sources, the one or more knowledge sources comprising pre-approved, session-specific database or retrieval-augmented generation (RAG) sources.
7 . The system of claim 6 , wherein storing the avatar and the cloned voice of the user in the network accessible location comprises:
storing the avatar and the cloned voice of the avatar on a blockchain, wherein the avatar is minted as a non-fungible token linked to a verified identity of the user.
8 . The system of claim 6 , wherein generating the avatar reflective of the appearance and the behavior of the user comprises:
extracting one or more gestures of the user from the video data;
extracting one or more facial expressions of the user from the video data; and
causing the avatar to replicate the one or more gestures and the one or more facial expressions during interaction with the avatar.
9 . The system of claim 6 , the operations further comprising:
dynamically transitioning, the avatar between the user device, one or more of edge components, and cloud servers based on device capability.
10 . The system of claim 6 , wherein storing the avatar and the cloned voice of the user in the network accessible location comprises:
applying one or more watermarks to the avatar, wherein the one or more watermarks are bound to an interaction session in which the avatar exists, the one or more watermarks providing authentication to the avatar.
11 . A method, comprising:
receiving, by a computing system, video data and non-video data from a user device for generation of an avatar corresponding to a user of the user device, the video data comprising a video of the user of the user device, the non-video data comprising information associated with a desired appearance and behavior of the avatar;
generating, by the computing system, the avatar reflective of the appearance and the behavior of the user by inputting the video data and non-video data to a generative machine learning model trained to generate highly realistic avatars for users, wherein the avatar looks, behaves, sounds, and interacts like the user, wherein the video data is pre-processed to remove background information from the video data and isolate the user within the video data;
cloning, by the computing system, a voice of the user for use with the avatar, wherein the cloning comprises extracting audio data from the video data and generating a synthetic voice for the avatar using a model trained on voice characteristics of the user, wherein speech output of the avatar matches one or more of a tone, pace, inflection, or accent of the user;
storing, by the computing system, the avatar and the cloned voice of the user in a network accessible location;
deploying, by the computing system, the avatar in a third-party application for real-time or near real-time interaction with a second user, wherein responses of the avatar are generated using a large language model and are delivered with synchronized lip movements generated by a lip-sync module based on learned speech patterns of the user; and
dynamically transitioning, by the computing system, the avatar between the user device, one or more of edge components, and cloud servers based on device capability.
12 . The method of claim 11 , further comprising:
defining, by the computing system, one or more session-specific guardrails for the avatar while interacting in the third-party application, the one or more session-specific guardrails comprising domain-based, administrator-defined restrictions on avatar output.
13 . The method of claim 12 , wherein the one or more session-specific guardrails limit the avatar to interfacing with one or more knowledge sources, the one or more knowledge sources comprising pre-approved, session-specific database or retrieval-augmented generation (RAG) sources.
14 . The method of claim 11 , wherein storing, by the computing system, the avatar and the cloned voice of the user in the network accessible location comprises:
storing the avatar and the cloned voice of the avatar on a blockchain, wherein the avatar is minted as a non-fungible token linked to a verified identity of the user.
15 . The method of claim 11 , wherein generating, by the computing system, the avatar reflective of the appearance and the behavior of the user comprises:
extracting one or more gestures of the user from the video data;
extracting one or more facial expressions of the user from the video data; and
causing the avatar to replicate the one or more gestures and the one or more facial expressions during interaction with the avatar.
16 . The method of claim 11 , wherein storing, by the computing system, the avatar and the cloned voice of the user in the network accessible location comprises:
applying one or more watermarks to the avatar, wherein the one or more watermarks are bound to an interaction session in which the avatar exists, the one or more watermarks providing authentication to the avatar.
17 . A system comprising:
a non-transitory storage medium storing computer program instructions; and
a processor configured to execute the computer program instructions to cause operations comprising:
receiving video data and non-video data from a user device for generation of an avatar corresponding to a user of the user device, the video data comprising a video of the user of the user device, the non-video data comprising information associated with a desired appearance and behavior of the avatar;
generating the avatar reflective of the appearance and the behavior of the user by inputting the video data and non-video data to a generative machine learning model trained to generate highly realistic avatars for users, wherein the avatar looks, behaves, sounds, and interacts like the user, wherein the video data is pre-processed to remove background information from the video data and isolate the user within the video data;
cloning a voice of the user for use with the avatar, wherein the cloning comprises extracting audio data from the video data and generating a synthetic voice for the avatar using a model trained on voice characteristics of the user, wherein speech output of the avatar matches one or more of a tone, pace, inflection, or accent of the user;
storing the avatar and the cloned voice of the user in a network accessible location;
deploying the avatar in a third-party application for real-time or near real-time interaction with a second user, wherein responses of the avatar are generated using a large language model and are delivered with synchronized lip movements generated by a lip-sync module based on learned speech patterns of the user; and
dynamically transitioning the avatar between the user device, one or more of edge components, and cloud servers based on device capability.
18 . The system of claim 17 , the operations further comprising:
defining one or more session-specific guardrails for the avatar while interacting in the third-party application, the one or more session-specific guardrails comprising domain-based, administrator-defined restrictions on avatar output.
19 . The system of claim 18 , wherein the one or more session-specific guardrails limit the avatar to interfacing with one or more knowledge sources, the one or more knowledge sources comprising pre-approved, session-specific database or retrieval-augmented generation (RAG) sources.
20 . The system of claim 17 , wherein storing the avatar and the cloned voice of the user in the network accessible location comprises:
storing the avatar and the cloned voice of the avatar on a blockchain, wherein the avatar is minted as a non-fungible token linked to a verified identity of the user.
21 . The system of claim 17 , wherein generating the avatar reflective of the appearance and the behavior of the user comprises:
extracting one or more gestures of the user from the video data;
extracting one or more facial expressions of the user from the video data; and
causing the avatar to replicate the one or more gestures and the one or more facial expressions during interaction with the avatar.
22 . The system of claim 17 , wherein storing the avatar and the cloned voice of the user in the network accessible location comprises:
applying one or more watermarks to the avatar, wherein the one or more watermarks are bound to an interaction session in which the avatar exists, the one or more watermarks providing authentication to the avatar.