Systems and methods for gesture generation from text
Embodiments described herein provide systems and methods for gesture generation from text. A method for gesture generation includes receiving an input text. The method may further include generating, via an encoder, an action representation in an action representation space based on the input text. The method may further include generating, via a first motion decoder, a first body configuration based on the action representation. The method may further include generating, via a second motion decoder, a second body configuration based on the first body configuration. The method may further include generating, via a token decoder, a first stop token based on the first body configuration.
1 . A method for gesture generation from text, the method comprising:
receiving, via a data interface, an input text;
generating, via an encoder, an action representation in an action representation space based on the input text;
generating, via a first motion decoder, a first body configuration based on the action representation;
generating, via a second motion decoder, a second body configuration based on the first body configuration; and
generating, via a token decoder, a first stop token based on the first body configuration.
2 . The method of claim 1 , wherein generating the action representation includes first generating, via a language model, an intermediate representation in an intermediate representation space based on the input text, wherein the intermediate representation includes a first action representation.
3 . The method of claim 1 , wherein the second motion decoder includes a sequence of a first neural network, a second neural network, and a third neural network.
4 . The method of claim 3 , wherein a hidden state of the first motion decoder is shared with the first neural network.
5 . The method of claim 3 , wherein the second motion decoder includes a residual connection between the first neural network and third neural network.
6 . The method of claim 1 , wherein the first stop token meets a threshold and further comprising:
generating, via the second motion decoder, a third body configuration based on the second body configuration; and
generating via the token decoder, a second stop token based on the second body configuration.
7 . The method of claim 3 , wherein generating the second body configuration further comprises combining the first body configuration with an output of the third neural network.
8 . A system for gesture generation from text, the system comprising:
a memory that stores a plurality of processor-executable instructions;
a data interface that receives an input text; and
one or more processors that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising:
receiving, via a data interface, an input text;
generating, via an encoder, an action representation in an action representation space based on the input text;
generating, via a first motion decoder, a first body configuration based on the action representation;
generating, via a second motion decoder, a second body configuration based on the first body configuration; and
generating, via a token decoder, a first stop token based on the first body configuration.
9 . The system of claim 8 , wherein operations for generating the action representation include first generating, via a language model, an intermediate representation in an intermediate representation space based on the input text, wherein the intermediate representation includes a first action representation.
10 . The system of claim 8 , wherein the second motion decoder includes a sequence of a first neural network, a second neural network, and a third neural network.
11 . The system of claim 10 , wherein a hidden state of the first motion decoder is shared with the first neural network.
12 . The system of claim 10 , wherein the second motion decoder includes a residual connection between the first neural network and third neural network.
13 . The system of claim 8 , wherein the first stop token meets a threshold and the operations further comprising:
generating, via the second motion decoder, a third body configuration based on the second body configuration; and
generating via the token decoder, a second stop token based on the second body configuration.