IP Library › Granted Patent US 12,633,027
Granted Patent B2
US 12,633,027 · App. 18/626,166 · Granted May 19, 2026

Systems and methods for gesture generation from text

Inventors: Gwantae Kim (Seoul, KR); Hanseok Ko (Seoul, KR)
Assignee: Datum Point Labs Inc.
G06T13/40G06F40/284G06N3/044G06N3/045G06N3/08G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,027
App. No.
18/626,166
Granted
May 19, 2026
Kind
B2
Abstract

Embodiments described herein provide systems and methods for gesture generation from text. A method for gesture generation includes receiving an input text. The method may further include generating, via an encoder, an action representation in an action representation space based on the input text. The method may further include generating, via a first motion decoder, a first body configuration based on the action representation. The method may further include generating, via a second motion decoder, a second body configuration based on the first body configuration. The method may further include generating, via a token decoder, a first stop token based on the first body configuration.

Claims (30)

1 . A method for gesture generation from text, the method comprising:

receiving, via a data interface, an input text;

generating, via an encoder, an action representation in an action representation space based on the input text;

generating, via a first motion decoder, a first body configuration based on the action representation;

generating, via a second motion decoder, a second body configuration based on the first body configuration; and

generating, via a token decoder, a first stop token based on the first body configuration.

2 . The method of claim 1 , wherein generating the action representation includes first generating, via a language model, an intermediate representation in an intermediate representation space based on the input text, wherein the intermediate representation includes a first action representation.

3 . The method of claim 1 , wherein the second motion decoder includes a sequence of a first neural network, a second neural network, and a third neural network.

4 . The method of claim 3 , wherein a hidden state of the first motion decoder is shared with the first neural network.

5 . The method of claim 3 , wherein the second motion decoder includes a residual connection between the first neural network and third neural network.

6 . The method of claim 1 , wherein the first stop token meets a threshold and further comprising:

generating, via the second motion decoder, a third body configuration based on the second body configuration; and

generating via the token decoder, a second stop token based on the second body configuration.

7 . The method of claim 3 , wherein generating the second body configuration further comprises combining the first body configuration with an output of the third neural network.

8 . A system for gesture generation from text, the system comprising:

a memory that stores a plurality of processor-executable instructions;

a data interface that receives an input text; and

one or more processors that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising:

receiving, via a data interface, an input text;

generating, via an encoder, an action representation in an action representation space based on the input text;

generating, via a first motion decoder, a first body configuration based on the action representation;

generating, via a second motion decoder, a second body configuration based on the first body configuration; and

generating, via a token decoder, a first stop token based on the first body configuration.

9 . The system of claim 8 , wherein operations for generating the action representation include first generating, via a language model, an intermediate representation in an intermediate representation space based on the input text, wherein the intermediate representation includes a first action representation.

10 . The system of claim 8 , wherein the second motion decoder includes a sequence of a first neural network, a second neural network, and a third neural network.

11 . The system of claim 10 , wherein a hidden state of the first motion decoder is shared with the first neural network.

12 . The system of claim 10 , wherein the second motion decoder includes a residual connection between the first neural network and third neural network.

13 . The system of claim 8 , wherein the first stop token meets a threshold and the operations further comprising:

generating, via the second motion decoder, a third body configuration based on the second body configuration; and

generating via the token decoder, a second stop token based on the second body configuration.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2024
From: KIM, GWANTAE; KO, HANSEOK
To: DATUM POINT LABS INC.
Reel/Frame 067507/0665 →
Continuity (2)
Provisional Application 63457551 · Apr 6, 2023
Related Publication 20240338874A1 · Oct 10, 2024
References Cited (22)
US 11648480B2 · Akhoundi · 2023 [cited by examiner]
US 11908180B1 · Ho · 2024 [cited by examiner]
US 20190304157A1 · Amer · 2019 [cited by examiner]
US 20210220739A1 · Zinno · 2021 [cited by examiner]
US 20210248804A1 · Hussen Abdelaziz · 2021 [cited by examiner]
US 20230082830A1 · Fan · 2023 [cited by examiner]
US 20230135769A1 · Bhattacharya · 2023 [cited by examiner]
US 20230260182A1 · Saito · 2023 [cited by examiner]
US 20230290371A1 · Jawahar · 2023 [cited by examiner]
US 20230316616A1 · Akhoundi · 2023 [cited by examiner]
Ahn H, Ha T, Choi Y, Yoo H, Oh S. Text2action: Generative adversarial synthesis from language to action. In2018 IEEE International Conference on Robotics and Automation (ICRA) May 21, 2018 (pp. 5915-5920). IEEE. (Year: … [cited by examiner]
Ahuja C, Morency LP. Language2pose: Natural language grounded pose forecasting. In2019 International conference on 3D vision (3DV) Sep. 16, 2019 (pp. 719-728). IEEE. (Year: 2019). [cited by examiner]
Kucherenko T, Jonell P, Van Waveren S, Henter GE, Alexandersson S, Leite I, Kjellström H. Gesticulator: A framework for semantically-aware speech-driven gesture generation. In Proceedings of the 2020 international confe… [cited by examiner]
Hong F, Zhang M, Pan L, Cai Z, Yang L, Liu Z. Avatarclip: Zero-shot text-driven generation and animation of 3d avatars. arXiv preprint arXiv:2205.08535. May 17, 2022. (Year: 2022). [cited by examiner]
Guo C, Zuo X, Wang S, Cheng L. Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts. InEuropean Conference on Computer Vision Oct. 23, 2022 (pp. 580-597). Cham: Springer Na… [cited by examiner]
Tevet G, Gordon B, Hertz A, Bermano AH, Cohen-Or D. Motionclip: Exposing human motion generation to clip space. In European Conference on Computer Vision Oct. 23, 2022 (pp. 358-374). Cham: Springer Nature Switzerland. (… [cited by examiner]
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee, “Speech gesture generation from the trimodal context of text, audio, and speaker identity,” ACM Transactions on Graphics (TOG… [cited by applicant]
Julieta Martinez, Michael J Black, and Javier Romero, “On human motion prediction using recurrent neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2891-2900. [cited by applicant]
Sijie Yan, Zhizhong Li, Yuanjun Xiong, Huahan Yan, and Dahua Lin, “Convolutional sequence generation for skeleton-based action synthesis,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019… [cited by applicant]
Haoye Cai, Chunyan Bai, Yu-Wing Tai, and Chi-Keung Tang, “Deep video generation, prediction and completion of human action sequences,” in Proceedings of the European conference on computer vision (ECCV), 2018. [cited by applicant]
Ping Yu, Yang Zhao, Chunyuan Li, Junsong Yuan, and Changyou Chen, “Structure-aware human-action generation,” in European Conference on Computer Vision. Springer, 2020. [cited by applicant]
Chuan Guo, Xinxin Zuo, SenWang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng, “Action2motion: Conditioned generation of 3d human motions,” in Proceedings of the 28th ACM International Conference on Mu… [cited by applicant]