IP Library › Granted Patent US 12,738,260
Granted Patent B2
US 12,738,260 · App. 18/793,807 · Granted Sep 15, 2026

Gesture Vox

Inventor: Harivatsan Selvam (Alpharetta, GA)
G10L13/027G06V40/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,260
App. No.
18/793,807
Filed
Aug 3, 2024
Granted
Sep 15, 2026
Kind
B2
Art Unit
2659
USPC
704/200
Abstract

GestureVox is an innovative AI-powered software system designed to convert sign language into spoken words in real-time. Utilizing advanced machine learning techniques, including frameworks such as TensorFlow, PyTorch, Keras, and Scikit-learn, GestureVox offers a seamless and accurate gesture recognition and speech synthesis process. The system's architecture includes modules for data collection, pre-processing, model training, testing, hyperparameter tuning, and deployment. Key features include the ability to process live video feeds, a user-friendly interface, and scalability to handle a large number of concurrent users, potentially utilizing cloud services such as AWS, Azure, and Google Cloud. GestureVox significantly enhances communication for individuals with speech impairments, providing an inclusive and accessible solution.

Claims (41)

1 . A computer system for converting sign language gestures into spoken audio output in real time, the computer system comprising:

one or more processors; and

one or more non-transitory computer-readable memories storing instructions that, when executed by the one or more processors, cause the computer system to perform a model-deployment preparation stage followed by a calibration-gated live inference stage, wherein the model-deployment preparation stage comprises:

generating pre-processed training gesture image data from labeled sign-language gesture images by resizing the labeled sign-language gesture images to a uniform model-input size, normalizing pixel values of the labeled sign-language gesture images to a model-input range, applying noise-reduction processing to reduce background noise or non-gesture image information, and augmenting the labeled sign-language gesture images using at least one of rotation, flipping, or scaling;

dividing the pre-processed training gesture image data into a training set, a validation set, and an independent test set;

training a plurality of candidate convolutional neural-network models using the training set;

evaluating the plurality of candidate convolutional neural-network models using the validation set;

measuring a final accuracy of each of the plurality of candidate convolutional neural-network models using the independent test set;

calculating precision, recall, and F1-score for each of the plurality of candidate convolutional neural-network models;

selecting, after the precision, recall, and F1-score have been calculated, a candidate convolutional neural-network model having a highest final accuracy on the independent test set among the plurality of candidate convolutional neural-network models; and

embedding the selected candidate convolutional neural-network model into a software application as a deployed convolutional neural network configured for real-time gesture-to-speech conversion;

wherein the calibration-gated live inference stage comprises:

guiding a user, through a user interface of the software application, through a setup sequence comprising camera calibration and an initial gesture-recognition test;

only after completion of the setup sequence, receiving, from a camera input calibrated during the setup sequence, a live video feed comprising sequential image frames of the user performing sign language gestures;

generating standardized live gesture image data from the sequential image frames by resizing the sequential image frames to the uniform model-input size, normalizing pixel values of the sequential image frames to the model-input range, and applying noise-reduction processing to reduce background noise or non-gesture image information before the standardized live gesture image data is provided to the deployed convolutional neural network;

processing the standardized live gesture image data using the deployed convolutional neural network, wherein the deployed convolutional neural network comprises:

an input layer configured to receive the standardized live gesture image data while preserving height, width, and color-channel information;

one or more convolutional layers configured to apply convolutional filters to the standardized live gesture image data to generate feature maps representing gesture-related edges, textures, or image patterns;

one or more rectified linear unit activation functions configured to introduce non-linearity into the feature maps;

one or more pooling layers configured to reduce dimensionality of the feature maps;

a flattening operation configured to convert the reduced-dimensionality feature maps into a one-dimensional vector;

one or more fully connected layers configured to classify the one-dimensional vector into a plurality of predefined sign-language gesture classes;

a softmax output layer configured to generate a probability value for each of the plurality of predefined sign-language gesture classes;

selecting, as a predicted sign-language gesture, one of the plurality of predefined sign-language gesture classes having a highest probability value generated by the softmax output layer; and

generating spoken audio output corresponding to the predicted sign-language gesture.

2 . The computer system of claim 1 , wherein the setup sequence of the calibration-gated live inference stage comprises:

displaying, through the user interface of the software application, a setup prompt instructing the user to perform one or more sign language gestures within a field of view of the camera input;

capturing, from the camera input during the setup sequence and before receiving the live video feed for real-time gesture-to-speech conversion, one or more test image frames corresponding to the one or more sign language gestures;

generating standardized setup-test gesture image data from the one or more test image frames by resizing the one or more test image frames to the uniform model-input size, normalizing pixel values of the one or more test image frames to the model-input range, and applying noise-reduction processing to reduce background noise or non-gesture image information;

providing the standardized setup-test gesture image data to the input layer of the deployed convolutional neural network during the setup sequence;

generating, using the softmax output layer of the deployed convolutional neural network during the setup sequence, one or more test probability values for the initial gesture-recognition test, the one or more test probability values corresponding to the plurality of predefined sign-language gesture classes; and

receiving the live video feed for real-time gesture-to-speech conversion after the deployed convolutional neural network generates the one or more test probability values during the setup sequence.

3 . The computer system of claim 2 , wherein the pre-processed training gesture image data, the standardized setup-test gesture image data, and the standardized live gesture image data are generated according to a common gesture-image input-conformity protocol that imposes the uniform model-input size, the model-input range, and the noise-reduction processing across the model-deployment preparation stage, the setup sequence, and the calibration-gated live inference stage, and wherein the deployed convolutional neural network receives both the standardized setup-test gesture image data and the standardized live gesture image data through the input layer while preserving height, width, and color-channel information.

4 . The computer system of claim 3 , wherein, during the calibration-gated live inference stage, the deployed convolutional neural network operates as an embedded inference model within the software application after completion of the model-deployment preparation stage, and wherein real-time gesture-to-speech conversion of the live video feed is performed through a runtime inference path comprising:

generating the standardized live gesture image data according to the common gesture-image input-conformity protocol;

providing the standardized live gesture image data to the input layer of the deployed convolutional neural network;

processing the standardized live gesture image data through the one or more convolutional layers, the one or more rectified linear unit activation functions, the one or more pooling layers, the flattening operation, and the one or more fully connected layers of the deployed convolutional neural network;

generating, by the softmax output layer, probability values corresponding to the plurality of predefined sign-language gesture classes;

selecting the predicted sign-language gesture as the predefined sign-language gesture class having the highest probability value; and

generating the spoken audio output corresponding to the predicted sign-language gesture,

wherein the runtime inference path is performed using the deployed convolutional neural network embedded in the software application without performing training, validation-set evaluation, independent-test-set measurement, or candidate-model selection among the plurality of candidate convolutional neural-network models during receipt of the live video feed.

Continuity (1)
Related Publication 20260038478A1 · Feb 5, 2026
References Cited (12)
US 11521516B2 · Johnson · 2022 [cited by examiner]
US 11741755B2 · Ko · 2023 [cited by examiner]
US 11854308B1 · Marsden · 2023 [cited by examiner]
US 20090115721A1 · Aull · 2009 [cited by examiner]
US 20190251702A1 · Chandler · 2019 [cited by examiner]
US 20200005028A1 · Gu · 2020 [cited by examiner]
US 20210279453A1 · Luqman · 2021 [cited by examiner]
US 20220188538A1 · Vieira Rocha · 2022 [cited by examiner]
US 20230085161A1 · Rahmani · 2023 [cited by examiner]
US 20250292191A1 · Shankar · 2025 [cited by examiner]
US 20250350584A1 · Hadad · 2025 [cited by examiner]
US 20260080802A1 · Jadhav · 2026 [cited by examiner]