Sentiment-based interactive avatar system for sign language
Systems and methods for doing presenting an avatar that speaks sign language based on sentiment of a speaker is disclosed herein. A translation application running on a device receives a content item comprising a video and an audio, wherein the audio comprises a first plurality of spoken words in a first language. The video comprises a character speaking the first plurality of spoken words in the first language. The translation application translates the first plurality of spoken words of the first language into a first sign of a first sign language. The translation application determines an emotional state expressed by the character based on sentiment analysis. The translation application generates an avatar that speaks the first sign of the first sign language where the avatar exhibits the determined emotional state. The content item and the avatar are presented for display on the device.
1. A method comprising:
capturing video data using a camera of a device and capturing audio data using a microphone of the device;
extracting a spoken word from the audio data and an image of a speaker who uttered the spoken word from the video data;
querying a sign language database to determine a translation of the spoken word to a sign language gesture;
identifying visual characteristics of the speaker based on the extracted image of the speaker who uttered the spoken word;
generating an avatar based on the identified visual characteristics of the speaker who uttered the spoken word;
identifying in a model database, a skeleton model representing the sign language gesture; and
generating for display an animation of the avatar that was generated based on the identified visual characteristics of the speaker who uttered the spoken word performing the sign language gesture by applying the skeleton model to the avatar.
2. The method of claim 1 , further comprising:
determining an emotional state expressed of the speaker who uttered the spoken word based on sentiment analysis; and
wherein the generated avatar exhibits the determined emotional state.
3. The method of claim 2 , wherein the sentiment analysis is performed by:
determining an emotion identifier contained in the spoken word;
determining a facial expression or a body expression of the speaker who uttered the spoken word using one or more expression recognition algorithms; and
determining a vocal tone of the speaker who uttered the spoken word using one or more voice recognition algorithms.
4. The method of claim 1 , wherein the animation of the avatar comprises a movement of a hand, a finger, an arm, or a face of the avatar.
5. The method of claim 1 , further comprising:
converting the spoken word into a text using one or more speech recognition algorithms, the text comprising a word corresponding to the spoken word.
6. The method of claim 1 , wherein the speaker who uttered the spoken word is a live individual.
7. The method of claim 1 , further comprising:
receiving user input specifying a visual characteristic of the avatar, wherein the avatar is modified based on the specified visual characteristic.
8. The method of claim 1 , further comprising:
receiving a user request to transmit the avatar to a second device;
transmitting a configuration file that includes a visual characteristic of the avatar to the second device; and
causing display of the avatar based on the transmitted configuration file on the second device.
9. The method of claim 1 , wherein the avatar is automatically generated in real time.
10. A system comprising:
control circuitry configured to:
capture video data using a camera of a device and capture audio data using a microphone of the device;
extract a spoken word from the audio data and an image of a speaker who uttered the spoken word from the video data;
query a sign language database to determine a translation of the spoken word to a sign language gesture;
identify visual characteristics of the speaker based on the extracted image of the speaker who uttered the spoken word;
generate an avatar based on the identified visual characteristics of the speaker who uttered the spoken word;
identify in a model database, a skeleton model representing the sign language gesture; and
generate for display an animation of the avatar that was generated based on the identified visual characteristics of the speaker who uttered the spoken word performing the sign language gesture by applying the skeleton model to the avatar.
11. The system of claim 10 , the control circuitry further configured to determine an emotional state expressed of the speaker who uttered the spoken word based on sentiment analysis and wherein the generated avatar exhibits the determined emotional state.
12. The system of claim 11 , the control circuitry further configured to perform the sentiment analysis by:
determining an emotion identifier contained in the spoken word;
determining a facial expression or a body expression of the speaker who uttered the spoken word using one or more expression recognition algorithms; and
determining a vocal tone of the speaker who uttered the spoken word using one or more voice recognition algorithms.
13. The system of claim 10 , wherein the animation of the avatar comprises a movement of a hand, a finger, an arm, or a face of the avatar.
14. The system of claim 10 , the control circuitry further configured to:
convert the spoken word into text using one or more speech recognition algorithms, the text comprising a text word corresponding to the spoken word.
15. The system of claim 10 , wherein the speaker who uttered the spoken word is a live individual.
16. The system of claim 10 , the control circuitry further configured to receive user input specifying a visual characteristic of the avatar, wherein the avatar is modified based on the specified visual characteristic.
17. The system of claim 10 , the control circuitry further configured to:
receive a user request to transmit the avatar to a second device;
transmit a configuration file that includes a visual characteristic of the avatar to the second device; and
cause display of the avatar based on the transmitted configuration file on the second device.
18. The system of claim 10 , wherein the avatar is automatically generated in real time.