IP Library Granted Patent US 11,847,426
Granted Patent B2
US 11,847,426 · App. 16/762,302 · Granted Dec 19, 2023

Computer vision based sign language interpreter

Inventors: David Retek (Dunaujvaros, HU); David Palhazi (Tata, HU); Marton Kajtar (Felsopakony, HU); Attila Alvarez (Budapest, HU); Peter Poscsi (Budapest, HU); Andras Nemeth (Budapest, HU); Matyas Trosztel (Budapest, HU); Zsolt Robotka (Budapest, HU); Janos Rovnyai (Leanyfalu, HU)
Assignee: Snap Inc.
G06F40/58G06F3/014G06F3/017G06F40/205G06V40/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,847,426
App. No.
16/762,302
Granted
Dec 19, 2023
Kind
B2
Abstract

A system and method for translating sign language utterances into a target language, including: receiving motion capture data; producing phonemes/sign fragments from the received motion capture data; producing a plurality of sign sequences from the phonemes/sign fragments; parsing these sign sequences to produce grammatically parsed sign utterances; translating the grammatically parsed sign utterances into grammatical representations in the target language; and generating output utterances in the target language based upon the grammatical representations.

Claims (69)

1. A method of translating sign language utterances into a target language, comprising:

receiving motion capture data of a signing user making signs in the sign language;

determining one or more manual sign features in parallel with one or more eye markers based on an independent temporal segmentation of the motion capture data for the manual sign features and the one or more eye markers;

producing phonemes based on the one or more manual sign features and the one or more eye markers;

producing a plurality of sign sequences from the phonemes;

parsing the sign sequences to produce grammatically parsed sign utterances;

translating the grammatically parsed sign utterances into grammatical representations in the target language; and

generating output utterances in the target language based upon the grammatical representations.

2. The method of claim 1 , wherein confidence values are produced for each generated output utterance.

3. The method of claim 1 , wherein producing the phonemes from the motion capture data includes producing a confidence value for each produced phoneme .

4. The method of claim 1 , wherein producing the phonemes includes producing a plurality of segments as time intervals matching the phonemes, where these intervals of the plurality of segments may overlap.

5. The method of claim 4 , wherein producing the phonemes and their intervals includes determining of a set of possible succeeding phonemes for each phoneme.

6. The method of claim 1 , wherein producing the sign sequences from the phonemes includes matching potential paths in a graph of phonemes to each sign in each sign sequence.

7. The method of claim 1 , wherein producing the grammatically parsed sign utterances includes producing a grammatical context and using the grammatical context of previous utterances.

8. The method of claim 1 , wherein producing the grammatically parsed sign utterances includes producing a confidence value based on the confidences of the signs, the confidence of the parsing and confidence of the parse matching a grammatical context for each parse of each sign sequence.

9. The method of claim 1 , further comprising

generating a plurality of output utterances in the target language based upon the plurality of sign sequences;

displaying the plurality output utterances to the signing user; and

receiving an indication from the signing user selecting one of the plurality of displayed output utterances as a correct translation.

10. The method of claim 1 , further comprising detecting an end of a sign language utterance before parsing the sign sequence to produce a grammatically parsed sign utterance.

11. The method of claim 1 , wherein the motion capture data includes data captured using marked gloves used by the signing user.

12. The method of claim 1 , wherein user specific parameter data is used for one of:

producing the phonemes;

producing the plurality of sign sequences;

parsing the sign sequences; and

translating the grammatically parsed sign utterances.

13. The method of claim 1 , further comprising:

in response to detecting that the signing user is using fingerspelling based on the phonemes, performing operations comprising:

generating fingerspelling phonemes based on the phonemes;

translating the fingerspelling phonemes to translated letters in the target language;

generating an output to the signing user showing the translated letters to the signing user; and

receiving an input from the signing user indicating a correctness of the translated letters.

14. A system configured to translate sign language utterances into a target language, comprising:

an input interface configured to receive motion capture data from a signing user making signs in the sign language;

a memory; and

a processor in communication with the input interface and the memory, the processor being configured to:

determine one or more manual sign features in parallel with one or more eye markers based on an independent temporal segmentation of the motion capture data for the manual sign features and the one or more eye markers;

produce phonemes based on the one or more manual sign features and the one or more eye markers;

produce a plurality of sign sequences from the phonemes;

parse the sign sequences to produce grammatically parsed sign utterances;

translate the grammatically parsed sign utterances into a grammatical representation in the target language; and

generate output utterances in the target language based upon the grammatical representation.

15. The system of claim 14 , wherein confidence values are produced for each generated output utterance.

16. The system of claim 14 , wherein producing the phonemes from motion capture data includes producing a confidence value for each produced phoneme.

17. The system of claim 14 , wherein producing the phonemes includes producing a plurality of segments as time intervals matching the phonemes, where the time intervals of the plurality of segments may overlap.

18. The system of claim 17 , wherein producing the phonemes and their intervals includes determining of a set of possible succeeding phonemes for each phoneme.

19. The system of claim 14 , wherein producing the sign sequences from the phonemes includes matching potential paths in a graph of phonemes to each sign in each sign sequence.

20. The system of claim 14 , wherein producing the grammatically parsed sign utterances includes producing a grammatical context and using the grammatical context of previous utterances.

21. The system of claim 14 , wherein producing the grammatically parsed sign utterances includes producing a confidence value based on the confidences of the signs, the confidence of the parsing and confidence of the parse matching a grammatical context for each parse of each sign sequence.

22. The system of claim 14 , wherein the processor is further configured to:

generate a plurality of output utterances in the target language based upon the plurality of sign sequences;

display the plurality output utterances to the user; and

receive an indication from the signing user selecting one of the plurality of displayed output utterances as a correct translation.

23. The system of claim 14 , wherein the processor is further configured to detect an end of a sign language utterance before parsing the sign sequence to produce a grammatically parsed sign utterance.

24. The system of claim 14 , wherein the motion capture data includes data captured using marked gloves used by the signing user.

25. The system of claim 14 , wherein user specific parameter data is used for one of:

producing the phonemes;

producing the plurality of sign sequences;

parsing the sign sequences; and

translating the grammatically parsed sign utterances.

26. The system of claim 14 , wherein the processor is further configured to:

in response to detecting that the signing user is using fingerspelling based on the phonemes, performing operations comprising:

generating fingerspelling phonemes based on the phonemes detecting that the signing user is using fingerspelling, after producing phonemes;

translating the fingerspelling phonemes to letters in the target language;

generating an output to the signing user showing translated letters to the signing user; and

receiving an input from the signing user indicating a correctness of the translated letters.

27. The system of claim 14 , further comprising sensors producing the motion capture data.

28. The system of claim 14 , further comprising marked gloves used by the signing.

29. The system of claim 14 , wherein the input interface further receives communication input from a second user to facilitate a conversation between the signing user and the second user.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 20, 2022
From: SIGNALL TECHNOLOGIES ZRT
To: SNAP INC.
Reel/Frame 059974/0086 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2021
From: RETEK, DAVID; PALHAZI, DAVID; KAJTAR, MARTON; ALVAREZ, ATILLA; POCSI, PETER; NEMETH, ANDRAS; ROBOTKA, ZSOLT; ROVNYAI, JANOS
To: SIGNALL TRECNOLOGIES ZRT
Reel/Frame 054925/0827 →
EMPLOYMENT AGREEMENT Recorded Jan 14, 2021
From: MATYAS, TROSZTEL
To: SIGNALL TRECNOLOGIES ZRT
Reel/Frame 055007/0320 →
EMPLOYMENT AGREEMENT Recorded Jan 5, 2021
From: TROSZTEL, MATYAS
To: SIGNALL TRECNOLOGIES ZRT
Reel/Frame 054900/0634 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2021
From: RETEK, DAVID; PALHAZI, DAVID; KAJTAR, MARTON; ALVAREZ, ATILLA; POCSI, PETER; NEMETH, ANDRAS; ROBOTKA, ZSOLT; ROVNYAI, JANOS
To: SIGNALL TRECNOLOGIES ZRT
Reel/Frame 054801/0474 →
Continuity (2)
Provisional Application 62583026 · Nov 8, 2017
Related Publication 20210174034A1 · Jun 10, 2021
Cited By (3)
US 12,318,148 US 12,368,691 US 12,597,365