IP Library › Granted Patent US 11,521,516
Granted Patent B2
US 11,521,516 · App. 16/875,974 · Granted Dec 6, 2022

Nuance-based augmentation of sign language communication

Inventors: Kyle Johnson (McLean, VA); Eric Loucks (McLean, VA); Gaurang Bhatt (McLean, VA); Joshua Edwards (McLean, VA)
Assignee: Capital One Services, LLC
G09B21/009G06N3/084G06V40/28G10L13/00G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,521,516
App. No.
16/875,974
Filed
May 15, 2020
Granted
Dec 6, 2022
Kind
B2
Examiner
SADIO, INSA
Art Unit
2628
USPC
434/116
Abstract

In certain embodiments, nuance-based augmentation of gesture may be facilitated. In some embodiments, a video stream depicting sign language gestures of an individual may be obtained via a wearable device associated with a user. A textual translation of the sign language gestures in the video stream may be determined. Emphasis information related to the sign language gestures may be identified based on an intensity of the sign language gestures. One or more display characteristics may be determined based on the emphasis information. The textual translation may be caused to be displayed to the user via the wearable device according to the one or more display characteristics. In some embodiments, a unique voice profile for the individual may be determined. A spoken translation of the sign language gestures may be generated according to the textual translation, the unique voice profile, and the emphasis information.

Claims (55)

1. A system for facilitating nuance-based augmentation of gestures, the system comprising:

a computer system that comprises one or more processors programmed with computer program instructions that, when executed, cause the computer system to:

obtain, via a wearable device associated with a user, a video stream depicting sign language gestures of a person;

determine a textual translation of the sign language gestures in the video stream;

identify emphasis information related to the sign language gestures based on an intensity of the sign language gestures;

determine one or more display characteristics based on the emphasis information;

cause, via the wearable device, the textual translation of the sign language gestures to be displayed to the user according to the one or more display characteristics;

determine a unique voice profile for the person; and

generate a spoken translation of the sign language gestures according to the textual translation, the unique voice profile, and the emphasis information.

2. The system of claim 1 , wherein the computer system is further caused to:

provide a training video depicting a sign language gesture as input to a machine learning model to cause the machine learning model to generate a predicted emphasis of the sign language gesture;

obtain feedback indicating an emphasis of the sign language gesture; and

provide the feedback as reference feedback to the machine learning model to cause the machine learning model to assess the feedback against the predicted emphasis, the machine learning model being updated based on the assessment of the feedback,

wherein to identify the emphasis information, the computer system is caused to use the updated machine learning model to identify the emphasis information.

3. The system of claim 1 , wherein the unique voice profile is determined based on demographic factors or voice characteristics of the person.

4. A method being implemented by one or more processors executing computer program instructions that, when executed, perform the method, the method comprising:

obtaining, via a user device of a user, an image stream of sign language gestures of a person;

determining a textual translation of the sign language gestures in the image stream;

identifying emphasis information related to the sign language gestures based on an intensity of the sign language gestures;

determining one or more display characteristics based on the emphasis information; and

causing, based on the one or more display characteristics, the textual translation to be displayed on a graphical user interface of the user device.

5. The method of claim 4 , wherein the intensity of the sign language gestures comprises speed, vigor, or size of the sign language gestures.

6. The method of claim 4 , wherein the one or more display characteristics comprise font size, font color, font style, or font type.

7. The method of claim 4 , further comprising identifying emotion information based on one or more facial expressions of the person, and wherein the one or more display characteristics are further based on the emotion information.

8. The method of claim 4 , further comprising:

identifying, in the image stream, second sign language gestures of a second person; and

causing an identifier associated with the second person to be displayed with a second textual translation of the second sign language gestures on the graphical user interface.

9. The method of claim 4 , further comprising:

determining a unique voice profile for the person based on attributes of the person;

determining one or more audio characteristics based on the unique voice profile and the emphasis information; and

generate, based on the one or more audio characteristics, a spoken translation of the sign language gestures via the user device.

10. The method of claim 9 , wherein the attributes of the person comprise demographic factors or voice characteristics of the person.

11. The method of claim 9 , wherein the one or more audio characteristics comprise volume, speed, or pitch.

12. The method of claim 9 , further comprising:

determining, based on an interaction history of the user, an interaction frequency between the user and the person;

determining whether the interaction frequency satisfies a threshold; and

in response to determining that the interaction frequency satisfies the threshold, storing the unique voice profile with an identifier associated with the person.

13. The method of claim 9 , further comprising identifying emotion information based on one or more facial expressions of the person, and wherein the one or more audio characteristics are further based on the emotion information.

14. One or more non-transitory, computer-readable media having instructions that, when executed by one or more processors, cause operations comprising:

obtaining, via a user device of a user, an image stream of gestures of a person;

determining a translation of the gestures in the image stream;

identifying emphasis information related to the gestures based on an intensity of the gestures;

determining one or more presentation characteristics based on the emphasis information; and

causing, based on the one or more presentation characteristics, the translation to be presented via the user device.

15. The one or more non-transitory, computer-readable media of claim 14 , wherein the intensity of the gestures comprises speed, vigor, or size of the gestures.

16. The one or more non-transitory, computer-readable media of claim 14 , wherein the one or more presentation characteristics comprise font size, font color, font style, font type, volume, speed, or pitch.

17. The one or more non-transitory, computer-readable media of claim 14 , further comprising identifying emotion information based on one or more facial expressions of the person, and wherein the one or more presentation characteristics are further based on the emotion information.

18. The one or more non-transitory, computer-readable media of claim 14 , further comprising:

identifying, in the image stream, second gestures of a second person; and

causing an identifier associated with the second person to be presented by the user device with a second translation of the second gestures.

19. The one or more non-transitory, computer-readable media of claim 14 , wherein the one or more presentation characteristics are further based on demographic factors or voice characteristics of the person.

20. The one or more non-transitory, computer-readable media of claim 19 , further comprising:

determining, based on an interaction history of the user, an interaction frequency between the user and the person;

determining whether the interaction frequency satisfies a threshold; and

in response to determining that the interaction frequency satisfies the threshold, storing a profile with an identifier associated with the person and the demographic factors or the voice characteristics of the person.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2020
From: JOHNSON, KYLE; LOUCKS, ERIC; BHATT, GAURANG; EDWARDS, JOSHUA
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 052679/0781 →
Continuity (1)
Related Publication 20210358330A1 · Nov 18, 2021
Cited By (1)
US 12,738,260