IP Library Granted Patent US 12,518,653
Granted Patent B2
US 12,518,653 · App. 17/302,699 · Granted Jan 6, 2026

Realtime AI sign language recognition

Inventor: Nikolas Anthony Kelly (Victor, NY)
Assignee: SIGN-SPEAK INC.
G09B21/04G06F40/58G06V40/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,653
App. No.
17/302,699
Granted
Jan 6, 2026
Kind
B2
Abstract

A real time sign language recognition method that allows Deaf and Hard of Hearing individuals to sign into any apparatus with a camera to extract target information (such as a translation in a target language) is proposed.

Claims (56)

1 . A method, comprising:

capturing a sequence of images;

for an image within the sequence of images, detecting pose information by:

applying a first pose network to detect a body pose configuration in the image;

applying a second pose network to detect a hand pose configuration in the image;

generating a feature vector including the body pose configuration and the hand pose configuration; and

converting the feature vector into a flattened feature vector having a predefined size;

generating a feature queue by collecting the flattened feature vector for the image of the sequence of images; and

converting the feature queue into a target language output.

2 . The method of claim 1 , wherein converting the feature queue comprises:

splitting the feature queue into individual regions; and

processing the individual regions into a sign language string.

3 . The method of claim 2 , wherein processing the individual regions comprises determining whether the individual regions are one of a pre-recorded sentence or an individual sign in one or more databases.

4 . The method of claim 2 , wherein processing the individual regions comprises applying a binary classifier to determine whether one or more of the individual regions is fingerspelled.

5 . The method of claim 2 , wherein processing the individual regions comprises:

comparing the individual regions to signs in one or more databases to generate comparison results; and

choosing a sign based on a K Nearest Neighbor function or a Dynamic Time Warping function applied to the comparison results.

6 . The method of claim 1 , wherein converting the feature queue comprises applying a Convolutional Neural Network (CNN) configured to output one or more flag values associated with an intrasign region, an intersign region, or a non-signing region, and wherein the one or more flag values correspond to an individual sign.

7 . A method, comprising:

detecting hand pose information from a sequence of image frames;

converting the hand pose information into a flattened feature vector;

normalizing the flattened feature vector into a resultant feature vector;

applying a convolutional neural network (CNN) to split the resultant feature vector into a plurality of individual regions by highlighting sign transition periods;

applying the CNN to output a respective flag corresponding to each individual region of the plurality of individual regions, wherein the respective flag indicates whether the corresponding individual region is an intrasign region, intersign region, or a nonsigning region, and wherein the intrasign and intersign regions correspond to signing regions of the plurality of individual regions;

processing the signing regions of the plurality of individual regions into a sign language string based on language information in one or more databases; and

translating the sign language string into a target language output.

8 . The method of claim 7 , wherein normalizing the flattened feature vector comprises:

setting head coordinates to be (0,0) in a pose and shoulders to be an average of 1 unit via a first affine transform; and

setting mean coordinates of one or more hands to be (0,0,0), wherein a standard deviation in the mean coordinates is an average of 1 unit via a second affine transform.

9 . The method of claim 7 , wherein processing the signing regions comprises determining whether the signing regions are one of a pre-recorded sentence or an individual sign in the one or more databases by applying a K Nearest Neighbor function or a Dynamic Time Warping function on the individual regions.

10 . The method of claim 7 , wherein processing the signing regions comprises applying a binary classifier to determine whether one or more of the signing regions is fingerspelled.

11 . The method of claim 7 , wherein processing the signing regions comprises:

comparing one or more of the signing regions to signs in the one or more databases to generate comparison results; and

choosing a sign based on a K Nearest Neighbor function or a Dynamic Time Warping function applied to the comparison results.

12 . The method of claim 7 , further comprising outputting a plurality of selectable translations associated with the target language output.

13 . The method of claim 7 , wherein normalizing the flattened feature vector comprises setting coordinates corresponding to a user head to be at an origin and coordinates corresponding to user shoulders to be at a unit distance from the origin.

14 . A device, comprising:

a camera configured to capture a sequence of image frames;

a computing device coupled to the camera, the computing device comprising:

a processor; and

a memory, wherein the memory contains instructions stored thereon that when executed by the processor cause the processor to:

detect hand pose information from the sequence of image frames;

convert the hand pose information into a flattened feature vector;

normalize the flattened feature vector into a resultant feature vector;

apply a convolutional neural network (CNN) to split the resultant feature vector into a plurality of individual regions by highlighting sign transition periods;

apply the CNN to output a respective flag corresponding to each individual region of the plurality of individual regions, wherein the respective flag indicates whether the corresponding individual region is an intrasign region, intersign region, or a nonsigning region, and wherein the intrasign and intersign regions correspond to signing regions of the plurality of individual regions;

process the signing regions of the plurality of individual regions into a sign language string based on language information in one or more databases; and

translate the sign language string into a target language output.

15 . The device of claim 14 , wherein to process the signing regions, the processor is configured to determine whether the signing regions are one of a pre-recorded sentence or an individual sign in the one or more databases by applying a K Nearest Neighbor function or a Dynamic Time Warping function on the individual regions.

16 . The device of claim 14 , wherein to process the signing regions, the processor is configured to apply a binary classifier to determine whether one or more of the signing regions is fingerspelled.

17 . The device of claim 14 , wherein to process the signing regions, the processor is configured to:

compare the signing regions to signs in the one or more databases to generate comparison results; and

choose a sign based on a K Nearest Neighbor function or a Dynamic Time Warping function applied to the comparison results.

18 . The device of claim 14 , wherein the processor is further configured to output a plurality of selectable translations associated with the target language output.

19 . The device of claim 14 , wherein the processor is further configured to use an affine transformation to normalize the flattened feature vector.

20 . The device of claim 14 , wherein to normalize the flattened feature vector, the processor is configured to set coordinates corresponding to a user head to be at an origin and coordinates corresponding to user shoulders to be at a unit distance from the origin.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2022
From: KELLY, NIKOLAS ANTHONY
To: SIGN-SPEAK INC.
Reel/Frame 059117/0240 →
Continuity (2)
Provisional Application 63101716 · May 11, 2020
Related Publication 20220327961A1 · Oct 13, 2022
References Cited (31)
US 5473705A · Abe · 1995 [cited by applicant]
US 5659764A · Sakiyama · 1997 [cited by applicant]
US 7746986B2 · Bucchieri · 2010 [cited by applicant]
US 8428643B2 · Lin · 2013 [cited by applicant]
US 8890813B2 · Minnen · 2014 [cited by applicant]
US 9098493B2 · Tardif · 2015 [cited by applicant]
US 9734730B2 · Divakaran · 2017 [cited by examiner]
US 10192105B1 · Mahmoud · 2019 [cited by applicant]
US 10261596B2 · Burr · 2019 [cited by applicant]
US 10262198B2 · Mahmoud · 2019 [cited by applicant]
US 10268879B2 · Mahmoud · 2019 [cited by applicant]
US 10289903B1 · Chandler · 2019 [cited by applicant]
US 10304208B1 · Chandler · 2019 [cited by applicant]
US 10521264B2 · Chandler · 2019 [cited by applicant]
US 10762340B2 · Kaur · 2020 [cited by applicant]
US 10885318B2 · Maxwell · 2021 [cited by applicant]
US 10977452B2 · Wang · 2021 [cited by examiner]
US 10991380B2 · Santos · 2021 [cited by applicant]
US 11163373B2 · Liu · 2021 [cited by examiner]
US 11323663B1 · Kasaba · 2022 [cited by applicant]
US 11741755B2 · Ko · 2023 [cited by examiner]
US 20060174315A1 · Kim et al. · 2006 [cited by applicant]
US 20160307469A1 · Zhou · 2016 [cited by applicant]
US 20170351910A1 · Elwazer · 2017 [cited by applicant]
US 20190050637A1 · Mahmoud · 2019 [cited by examiner]
US 20190171716A1 · Weber · 2019 [cited by applicant]
US 20220327961A1 · Kelly · 2022 [cited by applicant]
Habibie, I., et al., “Learning Speech-driven 3D Conversational Gestures from Video,” arXiv, Computer Vision and Pattern Recognition, (Feb. 2021), 15 Pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2024/036228, mailed on Dec. 2, 2024, 5 pages. [cited by applicant]
Kotha, H., et al., Audio to Sign language Using NLTK, International Journal of Emerging Technologies and Innovative Research 10(6):e404-e408, (Jun. 2023). (retrieved on Sep. 4, 2024). URL: [https://www.jetir.org/view?pa… [cited by applicant]
Sharma, P., et al., “Translating Speech to Indian Sign Language Using Natural Language Processing,” Future Internet 14(9):253, (Aug. 2022), 17 Pages. [cited by applicant]