IP Library Granted Patent US 12,347,012
Granted Patent B2
US 12,347,012 · App. 18/411,611 · Granted Jul 1, 2025

Sentiment-based interactive avatar system for sign language

Inventors: Yusuf AbdElhakam AbdElkader Marey (Tulsa, OK); Reda Harb (Tampa, FL)
Assignee: Adeia Guides Inc.
G06T13/40G06F40/47G06T13/205G06V40/174G06V40/20G09B21/009G10L15/1815G10L15/22G10L21/10G10L25/63G10L2021/065
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,012
App. No.
18/411,611
Granted
Jul 1, 2025
Kind
B2
Abstract

Systems and methods for doing presenting an avatar that speaks sign language based on sentiment of a speaker is disclosed herein. A translation application running on a device receives a content item comprising a video and an audio, wherein the audio comprises a first plurality of spoken words in a first language. The video comprises a character speaking the first plurality of spoken words in the first language. The translation application translates the first plurality of spoken words of the first language into a first sign of a first sign language. The translation application determines an emotional state expressed by the character based on sentiment analysis. The translation application generates an avatar that speaks the first sign of the first sign language where the avatar exhibits the determined emotional state. The content item and the avatar are presented for display on the device.

Claims (50)

1. A method comprising:

capturing video data using a camera of a device and capturing audio data using a microphone of the device;

extracting a spoken word from the audio data and an image of a speaker who uttered the spoken word from the video data;

querying a sign language database to determine a translation of the spoken word to a sign language gesture;

identifying visual characteristics of the speaker based on the extracted image of the speaker who uttered the spoken word;

generating an avatar based on the identified visual characteristics of the speaker who uttered the spoken word;

identifying in a model database, a skeleton model representing the sign language gesture; and

generating for display an animation of the avatar that was generated based on the identified visual characteristics of the speaker who uttered the spoken word performing the sign language gesture by applying the skeleton model to the avatar.

2. The method of claim 1 , further comprising:

determining an emotional state expressed of the speaker who uttered the spoken word based on sentiment analysis; and

wherein the generated avatar exhibits the determined emotional state.

3. The method of claim 2 , wherein the sentiment analysis is performed by:

determining an emotion identifier contained in the spoken word;

determining a facial expression or a body expression of the speaker who uttered the spoken word using one or more expression recognition algorithms; and

determining a vocal tone of the speaker who uttered the spoken word using one or more voice recognition algorithms.

4. The method of claim 1 , wherein the animation of the avatar comprises a movement of a hand, a finger, an arm, or a face of the avatar.

5. The method of claim 1 , further comprising:

converting the spoken word into a text using one or more speech recognition algorithms, the text comprising a word corresponding to the spoken word.

6. The method of claim 1 , wherein the speaker who uttered the spoken word is a live individual.

7. The method of claim 1 , further comprising:

receiving user input specifying a visual characteristic of the avatar, wherein the avatar is modified based on the specified visual characteristic.

8. The method of claim 1 , further comprising:

receiving a user request to transmit the avatar to a second device;

transmitting a configuration file that includes a visual characteristic of the avatar to the second device; and

causing display of the avatar based on the transmitted configuration file on the second device.

9. The method of claim 1 , wherein the avatar is automatically generated in real time.

10. A system comprising:

control circuitry configured to:

capture video data using a camera of a device and capture audio data using a microphone of the device;

extract a spoken word from the audio data and an image of a speaker who uttered the spoken word from the video data;

query a sign language database to determine a translation of the spoken word to a sign language gesture;

identify visual characteristics of the speaker based on the extracted image of the speaker who uttered the spoken word;

generate an avatar based on the identified visual characteristics of the speaker who uttered the spoken word;

identify in a model database, a skeleton model representing the sign language gesture; and

generate for display an animation of the avatar that was generated based on the identified visual characteristics of the speaker who uttered the spoken word performing the sign language gesture by applying the skeleton model to the avatar.

11. The system of claim 10 , the control circuitry further configured to determine an emotional state expressed of the speaker who uttered the spoken word based on sentiment analysis and wherein the generated avatar exhibits the determined emotional state.

12. The system of claim 11 , the control circuitry further configured to perform the sentiment analysis by:

determining an emotion identifier contained in the spoken word;

determining a facial expression or a body expression of the speaker who uttered the spoken word using one or more expression recognition algorithms; and

determining a vocal tone of the speaker who uttered the spoken word using one or more voice recognition algorithms.

13. The system of claim 10 , wherein the animation of the avatar comprises a movement of a hand, a finger, an arm, or a face of the avatar.

14. The system of claim 10 , the control circuitry further configured to:

convert the spoken word into text using one or more speech recognition algorithms, the text comprising a text word corresponding to the spoken word.

15. The system of claim 10 , wherein the speaker who uttered the spoken word is a live individual.

16. The system of claim 10 , the control circuitry further configured to receive user input specifying a visual characteristic of the avatar, wherein the avatar is modified based on the specified visual characteristic.

17. The system of claim 10 , the control circuitry further configured to:

receive a user request to transmit the avatar to a second device;

transmit a configuration file that includes a visual characteristic of the avatar to the second device; and

cause display of the avatar based on the transmitted configuration file on the second device.

18. The system of claim 10 , wherein the avatar is automatically generated in real time.

Assignments (2)
CHANGE OF NAME Recorded Oct 4, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069113/0348 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2024
From: MAREY, YUSUF ABDELHAKAM ABDELKADER; HARB, REDA
To: ROVI GUIDES, INC.
Reel/Frame 066834/0882 →
Continuity (2)
Continuation 17240128 · Apr 26, 2021
Related Publication 20240153186A1 · May 9, 2024
References Cited (10)
US 7761892B2 · Ellis et al. · 2010 [cited by applicant]
US 10885692B2 · Wedig · 2021 [cited by examiner]
US 11438669B2 · Janugani et al. · 2022 [cited by applicant]
US 20140171036A1 · Simmons · 2014 [cited by applicant]
US 20190138607A1 · Zhang et al. · 2019 [cited by applicant]
US 20200294525A1 · Santos et al. · 2020 [cited by applicant]
US 20220335971A1 · Gruszka et al. · 2022 [cited by applicant]
US 20220343576A1 · Marey et al. · 2022 [cited by applicant]
Darcey Pittman, “Students create app to translate sign language”, Washington Square News Apr. 2018, 4. [cited by applicant]
Lyuba Azbel, “How do the deaf read?”, The paradox of performing a phonemic task without sound, 2004, 21. [cited by applicant]