IP Library Granted Patent US 12682538
Granted Patent B2
US 12682538 · App. 19/191,143 · Granted Jul 14, 2026

Device including motion generation artificial intelligence algorithm based on spatiotemporal feature of motion and operation method thereof

Inventors: Jungmin Chung (Yongin-si, KR); Kyoungchin Seo (Bucheon-si, KR); Dohee Lee (Seongnam-si, KR); Jihun Kim (Gunpo-si, KR)
Assignee: AILIVE INC.
G06T13/80G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682538
App. No.
19/191,143
Granted
Jul 14, 2026
Kind
B2
Abstract

Disclosed is a device including a motion generation artificial intelligence algorithm based on spatiotemporal feature of motion and an operating method thereof, and the device may include a memory configured to store at least one process for executing the motion generation artificial intelligence algorithm based on spatiotemporal features of motion; and a processor configured to execute the motion generation artificial intelligence algorithm based on spatiotemporal feature of motion according to the process.

Claims (38)

1 . A device including a motion generation artificial intelligence algorithm based on spatiotemporal feature of motion, comprising:

a memory configured to store at least one process for executing the motion generation artificial intelligence algorithm based on spatiotemporal features of motion; and

a processor configured to execute the motion generation artificial intelligence algorithm based on spatiotemporal feature of motion according to the process, wherein the processor comprises:

a sentence separation module configured to receive sentence information, separate an entire sentence included in the sentence information into single sentences including one subject and one verb, output time ratio information indicating whether each of movements describing the single sentences has a certain time ratio in an entire movement describing the entire sentence, and output single sentence information including the single sentences, the time ratio information indicating an influence of the movements corresponding to the single sentences in the entire movement;

a text feature extraction module configured to extract a text feature of each of the single sentences included in the single sentence information, and output text feature information including text feature values extracted from the single sentences;

a motion feature search module configured to search for most suitable motion feature values for each of the single sentences based on the text feature information in a database storing motion feature values for each motion item, and output motion feature information including the most suitable motion feature values for each single sentence;

a motion information integration module configured to extract final motion feature values for motion describing the entire sentence based on the sentence information, the time ratio information, and the motion feature information, and output final motion feature information including the final motion feature values; and

a motion reconstruction module configured to output motion information representing the motion describing the entire sentence based on a motion reconstruction model and the final motion feature information,

wherein the motion feature search module is configured to return, among all motion feature values of motion items searched from the database, first motion feature values representing some frames for an entire frame of the searched motion item or some periods for an entire period of the searched motion item as the most suitable motion feature values,

wherein the motion feature search module is configured to additionally return second motion feature values representing most important and most relevant body structures in the movement of the searched motion item within the some frames or the some periods, as the most suitable motion feature values, and

wherein the motion feature search module is configured to divide a human movement into nodes corresponding to a torso, a left arm, a right arm, a left leg, and a right leg, and extract motion feature values of the respective nodes, and is further configured to extract the second motion feature values representing the most important and the most relevant body structures, based on joint information in the movement of the motion item within the some frames or the some periods.

2 . The device according to claim 1 ,

wherein the processor further comprises:

a motion feature extraction module configured to receive first motion information representing a first motion describing a first single sentence and first text feature information including a first text feature value of the first single sentence, extract motion features of the first motion based on the first text feature value, output first motion feature information including the motion features of the first motion, and store the first motion feature information in the database,

wherein the text feature extraction module is configured to receive first single sentence information including the first single sentence, map the first single sentence information to the first text feature value using a pre-trained text feature extraction model, and output the first text feature information, and

wherein the motion reconstruction module is configured to output the first motion information reconstructed based on the first motion feature information.

3 . The device according to claim 2 ,

wherein the motion feature extraction module comprises: a spatial information extraction module configured to extract third motion feature values representing nodes corresponding to body parts and important nodes corresponding to most important body part in the first motion that changes over time based on the first text feature values, and output spatial information including the third motion feature values; and

a temporal information extraction module configured to extract fourth motion feature values representing temporal importance of the first motion over time based on the first text feature values and the third motion feature values, and output temporal information including the fourth motion feature values.

4 . The device according to claim 1 ,

wherein the motion feature search module comprises:

a temporal relevance-based searching module configured to calculate temporal relevances for each motion item and each single sentence based on the correlation between the motion feature vector including the motion feature values of each motion item and the text feature value of each single sentence, and select a selection area including the temporal relevance greater than a reference relevance among the temporal relevances for each of the single sentences; and

a spatial relevance-based information extraction module configured to extract nodes corresponding to body parts based on the temporal relevances within the selection area and the text feature values of each single sentence.

5 . The device according to claim 4 ,

wherein the spatial relevance-based information extraction module is configured to: return the nodes corresponding to the body parts and the subsequent candidate nodes corresponding to the body parts according to the temporal relevance within the selection area.

6 . The device according to claim 1 ,

wherein the motion information integration module is configured to utilize a detailed motion part feature value in the process of generating the motion feature value based on the entire sentence of the sentence information.

7 . The device according to claim 6 ,

wherein the motion information integration module is configured to generate the motion feature value based on a motion usage ratio corresponding to the time ratio.

8 . An operation method performed by a device including a motion generation artificial intelligence algorithm based on spatiotemporal feature of motion, comprising:

a sentence separation step of separating an entire sentence included in sentence information input from an outside into single sentences including one subject and one verb, outputting time ratio information indicating whether each of movements describing the single sentences has a certain time ratio in an entire movement describing the entire sentence, and outputting single sentence information including the single sentences, the time ratio information indicating an influence of the movements corresponding to the single sentences in the entire movement;

a text feature extraction step of extracting a text feature of each of the single sentences included in the single sentence information, and outputting text feature information including text feature values extracted from the single sentences;

a motion feature search step of searching for most suitable motion feature values for each of the single sentences based on the text feature information in a database storing motion feature values for each motion item, and outputting motion feature information including the most suitable motion feature values;

a motion information integration step of extracting final motion feature values for motion describing the entire sentence based on the sentence information, the time ratio information, and the motion feature information, and outputting final motion feature information including the final motion feature values; and

a motion reconstruction step of outputting motion information representing the motion describing the entire sentence based on a motion reconstruction model and the final motion feature information,

wherein the motion feature search step includes returning, among all motion feature values of motion items searched from the database, first motion feature values representing some frames for an entire frame of the searched motion item or some periods for an entire period of the searched motion item as the most suitable motion feature values,

wherein the motion feature search step includes additionally returning second motion feature values representing most important and most relevant body structures in the movement of the searched motion item within the some frames or the some periods, as the most suitable motion feature values, and

wherein the motion feature search step includes dividing a human movement into nodes corresponding to a torso, a left arm, a right arm, a left leg, and a right leg, and extract motion feature values of the respective nodes, and extracting the second motion feature values representing most important and most relevant body structures, based on joint information, in the movement of the motion item within the some frames or the some periods.