IP Library Granted Patent US 12,562,006
Granted Patent B2
US 12,562,006 · App. 18/787,620 · Granted Feb 24, 2026

System(s) and method(s) for training a sign language captioning model and subsequent use thereof

Inventors: Garrett Tanzer (Boston, MA); Sepehr Sam Sepah (Pleasanton, CA)
Assignee: GOOGLE LLC
G06V40/28G06N3/08G09B21/009H04N21/4884
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,006
App. No.
18/787,620
Granted
Feb 24, 2026
Kind
B2
Abstract

Implementations are directed to training and subsequently utilizing a sign language captioning model. Initially, processor(s) of a system can obtain a plurality of training instances that are generated based on processing sign language video content, sign language conversations, etc. Each of the plurality of training instances can include at least corresponding sign language feature tokens for a sign language video content segment, ground truth caption tokens associated with ground truth sign language captions for the sign language video content segment and ground truth timestamp tokens that align the ground truth sign language captions with respect to the sign language video content segment. Further, the processor(s), can train the sign language captioning model based on the plurality of training instances, and can cause the sign language captioning model to be deployed in an offline manner and/or in an online manner for processing sign language content.

Claims (78)

1 . A method implemented by one or more processors, the method comprising:

obtaining a plurality of training instances for training a sign language captioning model,

each of the plurality of training instances including a corresponding training instance input and a corresponding training instance output,

the corresponding training instance inputs including at least corresponding sign language feature tokens for a sign language video content segment, and

corresponding alignment indicators that indicate whether ground truth sign language captions for the sign language video content segment are well-aligned or misaligned; and

the corresponding training instance outputs including ground truth caption tokens associated with the ground truth sign language captions for the sign language video content segment and ground truth timestamp tokens that align the ground truth sign language captions with respect to the sign language video content segment;

training, based on the plurality of training instances, the sign language captioning model, wherein training the sign language captioning model based on a given training instance, of the plurality of training instances, comprises:

processing, using the sign language captioning model, at least the corresponding sign language feature tokens for the sign language video content segment and the corresponding alignment indicator, included in the corresponding training instance input for the given training instance, to generate sign language captioning model output;

determining, based on the sign language captioning model output, predicted caption tokens associated with predicted sign language captions for the sign language video content segment and predicted timestamp tokens that are predicted to align the predicted sign language captions with respect to the sign language video content segment;

generating, based on a comparison of (i) the predicted caption tokens associated with the predicted sign language captions for the sign language video content segment and the ground truth caption tokens associated with the ground truth sign language captions for the sign language video content segment, and/or (ii) the predicted timestamp tokens that are predicted to align the predicted sign language captions with respect to the sign language video content segment and the ground truth timestamp tokens that align the ground truth sign language captions with respect to the sign language video content segment, one or more losses; and

updating, based on one or more of the losses, the sign language captioning model; and

subsequent to training the sign language captioning model:

causing the sign language captioning model to be deployed.

2 . The method of claim 1 , wherein the corresponding training instance inputs further include corresponding current caption tokens for current captions from the sign language video content segment.

3 . The method of claim 1 , wherein the corresponding training instance inputs further include corresponding previous caption tokens for previous captions from a previous sign language video content segment that precedes the sign language video content segment.

4 . The method of claim 1 , wherein the corresponding training instance inputs further include corresponding next caption tokens for next captions from a next sign language video content segment that follows the sign language video content segment.

5 . The method of claim 1 , wherein the corresponding sign language feature tokens for the sign language video content segment comprise one or more of: a corresponding video embedding for the sign language video content segment, corresponding image embeddings for the sign language video content segment, or corresponding vectors for skeletonized representations of the sign language video content segment.

6 . The method of claim 1 , wherein the corresponding training instance outputs further include a ground truth language token associated with a ground truth language for the sign language video content segment.

7 . The method of claim 6 , further comprising:

determining, based on the sign language captioning model output, a predicted language token associated with a predicted language for the sign language video content segment,

wherein generating one or more of the losses is further based on a comparison of: (iii) the predicted language token associated with the predicted language for the sign language video content segment and the ground truth language token associated with the ground truth language for the sign language video content segment.

8 . The method of claim 1 , wherein the corresponding training instance inputs further include corresponding instructions for the sign language captioning model to determine the predicted sign language captions for the sign language video content segment and to determine the predicted timestamp tokens that are predicted to align the predicted sign language captions with respect to the sign language video content segment.

9 . The method of claim 1 , further comprising:

prior to obtaining the plurality of training instances for training the sign language captioning model:

generating the plurality of training instances for training the sign language captioning model.

10 . The method of claim 9 , wherein generating the given training instance, of the plurality of training instances, comprises:

obtaining sign language video content that includes a plurality of signs being performed by a user and a caption track for the plurality of signs being performed by the user;

segmenting the sign language video content into a plurality of sign language video content segments,

the plurality of sign language video content segments including the sign language video content segment for the given training instance, and

the sign language video content segment including a corresponding subset of the plurality of signs being performed by the user, in the sign language video content, and a corresponding subset of the caption track for the plurality of signs being performed by the user;

determining, based on processing the corresponding subset of the plurality of signs being performed by the user in the sign language video content segment for the given training instance, the corresponding sign language feature tokens for the sign language video content segment and for the given training instance; and

determining, based on processing the corresponding subset of the caption track for the plurality of signs being performed by the user in the sign language video content segment for the given training instance, the ground truth caption tokens associated with the ground truth sign language captions for the sign language video content segment and the ground truth timestamp tokens that align the ground truth sign language captions with respect to the sign language video content segment.

11 . The method of claim 10 , wherein the caption track for the plurality of signs being performed by the user comprises corresponding captions associated with the plurality of signs and corresponding caption timestamps for the corresponding captions associated with the plurality of signs.

12 . The method of claim 1 , wherein causing the sign language captioning model to be deployed is in response to determining that one or more conditions are satisfied.

13 . The method of claim 12 , wherein the one or more conditions comprise one or more of: whether the sign language captioning model has been updated based on a threshold quantity of training instances, or whether performance of the sign language captioning model satisfies a threshold quality of performance.

14 . The method of claim 1 , wherein causing the sign language captioning model to be deployed comprises:

identifying newly added sign language video content that is newly added to a repository of sign language video content;

processing, using the sign language captioning model, the newly added sign language video content to determine a timestamped caption track associated with the newly added sign language video content; and

storing, in association with the newly added sign language video content, the timestamped caption track.

15 . The method of claim 14 , further comprising:

in response to receiving a request for playback of the newly added sign language video content:

causing the timestamped caption track to be played back along with the playback of the newly added sign language video content.

16 . The method of claim 1 , wherein causing the sign language captioning model to be deployed comprises:

identifying an ongoing conversation between a given user of a client device and an automated assistant;

processing, using the sign language captioning model, vision data that captures a plurality of sign language signs of the given user that are directed to the automated assistant and a dialog history of the ongoing conversation between the given user and the automated assistant to determine a timestamped caption track associated with the ongoing conversation; and

causing the timestamped caption track to be visually rendered for presentation to the given user, via a display of the client device, throughout the ongoing dialog.

17 . The method of claim 1 , wherein causing the sign language captioning model to be deployed comprises:

identifying an ongoing conversation between a given user and an additional user;

processing, using the sign language captioning model, vision data that captures a plurality of sign language signs of a given user that are directed to the additional user and a dialog history of the ongoing conversation between the given user and the additional user to determine a timestamped caption track associated with the ongoing conversation; and

causing the timestamped caption track to be visually rendered for presentation to the given user and/or the additional user throughout the ongoing dialog.

18 . A system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the at least one processor to be operable to:

obtain a plurality of training instances for training a sign language captioning model,

each of the plurality of training instances including a corresponding training instance input and a corresponding training instance output,

the corresponding training instance inputs including at least corresponding sign language feature tokens for a sign language video content segment, and

corresponding alignment indicators that indicate whether ground truth sign language captions for the sign language video content segment are well-aligned or misaligned; and

the corresponding training instance outputs including ground truth caption tokens associated with the ground truth sign language captions for the sign language video content segment and ground truth timestamp tokens that align the ground truth sign language captions with respect to the sign language video content segment;

train, based on the plurality of training instances, the sign language captioning model, wherein the instructions to train the sign language captioning model based on a given training instance, of the plurality of training instances, comprise instructions to:

process, using the sign language captioning model, at least the corresponding sign language feature tokens for the sign language video content segment and the corresponding alignment indicator, included in the corresponding training instance input for the given training instance, to generate sign language captioning model output;

determine, based on the sign language captioning model output, predicted caption tokens associated with predicted sign language captions for the sign language video content segment and predicted timestamp tokens that are predicted to align the predicted sign language captions with respect to the sign language video content segment;

generate, based on a comparison of (i) the predicted caption tokens associated with the predicted sign language captions for the sign language video content segment and the ground truth caption tokens associated with the ground truth sign language captions for the sign language video content segment, and/or (ii) the predicted timestamp tokens that are predicted to align the predicted sign language captions with respect to the sign language video content segment and the ground truth timestamp tokens that align the ground truth sign language captions with respect to the sign language video content segment, one or more losses; and

update, based on one or more of the losses, the sign language captioning model; and

subsequent to training the sign language captioning model:

cause the sign language captioning model to be deployed.

19 . A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to be operable to perform operations, the operations comprising:

obtaining a plurality of training instances for training a sign language captioning model,

each of the plurality of training instances including a corresponding training instance input and a corresponding training instance output,

the corresponding training instance inputs including at least corresponding sign language feature tokens for a sign language video content segment, and

corresponding alignment indicators that indicate whether ground truth sign language captions for the sign language video content segment are well-aligned or misaligned; and

the corresponding training instance outputs including ground truth caption tokens associated with the ground truth sign language captions for the sign language video content segment and ground truth timestamp tokens that align the ground truth sign language captions with respect to the sign language video content segment;

training, based on the plurality of training instances, the sign language captioning model, wherein training the sign language captioning model based on a given training instance, of the plurality of training instances, comprises:

processing, using the sign language captioning model, at least the corresponding sign language feature tokens for the sign language video content segment and the corresponding alignment indicator, included in the corresponding training instance input for the given training instance, to generate sign language captioning model output;

determining, based on the sign language captioning model output, predicted caption tokens associated with predicted sign language captions for the sign language video content segment and predicted timestamp tokens that are predicted to align the predicted sign language captions with respect to the sign language video content segment;

generating, based on a comparison of (i) the predicted caption tokens associated with the predicted sign language captions for the sign language video content segment and the ground truth caption tokens associated with the ground truth sign language captions for the sign language video content segment, and/or (ii) the predicted timestamp tokens that are predicted to align the predicted sign language captions with respect to the sign language video content segment and the ground truth timestamp tokens that align the ground truth sign language captions with respect to the sign language video content segment, one or more losses; and

updating, based on one or more of the losses, the sign language captioning model; and

subsequent to training the sign language captioning model:

causing the sign language captioning model to be deployed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2024
From: TANZER, GARRETT; SEPAH, SEPEHR SAM
To: GOOGLE LLC
Reel/Frame 068125/0132 →
Continuity (2)
Provisional Application 63660284 · Jun 14, 2024
Related Publication 20250384717A1 · Dec 18, 2025
References Cited (38)
US 10268879B2 · Mahmoud · 2019 [cited by examiner]
US 10289903B1 · Chandler · 2019 [cited by examiner]
US 11263409B2 · Zhang · 2022 [cited by examiner]
US 11741755B2 · Ko · 2023 [cited by examiner]
US 20190130176A1 · Maxwell · 2019 [cited by examiner]
US 20200005028A1 · Gu · 2020 [cited by examiner]
US 20220139417A1 · Maxwell · 2022 [cited by examiner]
US 20220327309A1 · Carlock · 2022 [cited by examiner]
US 20220327961A1 · Kelly · 2022 [cited by examiner]
US 20220391612A1 · Chakrabarty · 2022 [cited by examiner]
US 20230290371A1 · Jawahar · 2023 [cited by examiner]
US 20240233745A1 · Maxwell · 2024 [cited by examiner]
US 20240320449A1 · Kumar · 2024 [cited by examiner]
US 20250078574A1 · Thomson · 2025 [cited by examiner]
US 20250078818A1 · Sridhar · 2025 [cited by examiner]
US 20250165722A1 · Thomson · 2025 [cited by examiner]
US 20250165730A1 · Thomson · 2025 [cited by examiner]
US 20250201263A1 · Maxwell · 2025 [cited by examiner]
Anthropic; “The Claude 3 Model Family: Opus, Sonnet, Haiku”; retrieved from internet at URL https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf; 42 pages; dated 2024. [cited by applicant]
Bandarkar, L. et al., “The Belebele Benchmark: A Parallel Reading Comprehension Dataset in 122 Language Variants”; arXiv.org, Cornell University; arXiv:2308.16884v1; 25 pages; dated Aug. 31, 2023. [cited by applicant]
Conneau, A. et al., “FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech”; arXiv.org, Cornell University; arXiv:2205.12446v1; 10 pages; dated May 25, 2022. [cited by applicant]
Desai, A. et al., “Systemic Biases in Sign Language AI Research: A Deaf-Led Call to Reevaluate Research Agendas”; arXiv.org, Cornell University; arXiv:2403.02563v1; 12 pages; dated Mar. 5, 2024. [cited by applicant]
Duarte, A. et al., “How2Sign: A Large-scale Multimodal Dataset for Continuous American Sign Language”; in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); pp. 2735-2744; dated 20… [cited by applicant]
Goyal, N. et al., “The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation”; arXiv.org, Cornell University; arXiv:2106.03193v1; 26 pages; dated Jun. 6, 2021. [cited by applicant]
Grishchenko, I. et al., “MediaPipe Holistic—Simultaneous Face, Hand and Pose Prediction, on Device”; retrieved from the internet: URL https://ai.googleblog.com/2020/12/mediapipe-holistic-simultaneous-face.html; 6 pages;… [cited by applicant]
Guzmán, F. et al., “Two New Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-English”; arXiv.org, Cornell University; arXiv:1902.01382v1; 13 pages; dated Feb. 4, 2019. [cited by applicant]
Lugaresi, C. et al., “MediaPipe: A Framework for Building Perception Pipelines”; arXiv.org, Cornell University; arXiv:1906.08172v1; 9 pages; dated Jun. 14, 2019. [cited by applicant]
NLLB Team; “No Language Left Behind: Scaling Human-Centered Machine Translation”; arXiv.org, Cornell University; arXiv:2207.04672; 192 pages; dated Aug. 25, 2022. [cited by applicant]
OpenAI; “Hello Gpt-4o”; retrieved from the internet: URL https://openai.com/index/hello-gpt-40/; 19 pages; dated May 13, 2024. [cited by applicant]
Radford, A. et al., “Robust Speech Recognition via Large-Scale Weak Supervision”; arXiv.org, Cornell University; arXiv:2212.04356v1; 28 pages; dated Dec. 6, 2022. [cited by applicant]
Raffel, C. et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”; Journal of Machine Learning Research; 67 pages; dated Jun. 2020. [cited by applicant]
Sellam, T. et al., “BLEURT: Learning Robust Metrics for Text Generation”; arXiv.org, Cornell University; arXiv:2004-04696; 12 pages; dated May 21, 2020. [cited by applicant]
Sincan, O.M. et al., “Is context all you need? Scaling Neural Sign Language Translation to Large Domains of Discourse”; Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); pp. 1955-1965; date… [cited by applicant]
Tanzer, G. et al., “Reconsidering Sentence-Level Sign Language Translation”; arXiv.org, Cornell University; arXiv:2406.11049; 26 pages; dated Jun. 16, 2024. [cited by applicant]
Gemini Team; “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”; arXiv.org, Cornell University; arXiv:2403.05530; 154 pages; dated Jun. 14, 2024. [cited by applicant]
Uthus, D. et al., “Youtube-ASL: A Large-Scale, Open-Domain American Sign Language-English Parallel Corpus”; Conference on Neural Information Processing Systems (NeurIPS); 19 pages; dated 2023. [cited by applicant]
Yin, K. et al., “Including Signed Languages in Natural Language Processing”; In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natu… [cited by applicant]
Zhang, B. et al., “SLTUnet: A Simple Unified Model for Sign Language Translation”; in the 11th International Conference on Learning Representations; 18 pages; dated 2023. [cited by applicant]