Motion generation device for generating motion based on input information including text and operation method thereof
Disclosed is a motion generation device for generating motion based on input information including a text and an operation method thereof, and the motion generation device may include a memory configured to store at least one instruction; and at least one processor configured to execute the at least one instruction stored in the memory, wherein the at least one processor is configured to: obtain the input information including the text representing a movement of a character, extract feature information from the obtained input information, generate motion data for the movement of the character using the extracted feature information, obtain correction motion data by post-processing the generated motion data, and generate motion animation of the character based on the obtained correction motion data.
1 . A motion generation device for generating motion based on input information including text, comprising:
a memory configured to store at least one instruction; and
at least one processor configured to execute the at least one instruction stored in the memory,
wherein the processor is configured to:
obtain input information for a control element having audio, emotion, style, image, pose, and video related to a movement of a character including the text;
when extracting feature information after removing noise included in the input information through an input information pre-processing module included in the memory, extract a text feature value obtained based on a characteristic of the character, a behavior characteristic of the character, and a generation intent of the text through a text feature extraction model of the input information pre-processing module;
extract a control feature value corresponding to types of elements included in the control element through a feature extraction model for each control element of the input information pre-processing module;
extract a synthetic feature value generated by fusing the text feature value and the control feature value through a feature fusion pre-processing model of the input information pre-processing module;
when generating motion data for the movement of the character from the feature information using an artificial intelligence model trained to infer the motion data for the movement of the character from the feature information through a motion generation module included in the memory, convert the synthetic feature value into a motion feature value for generating motion data using the control feature value through a motion feature mapping model of the motion generation module;
generate the motion data based on the motion feature value through a motion feature-based motion generation model of the motion generation module;
obtain correction motion data by post-processing the motion data through a motion post-processing module included in the memory;
generate motion animation of the character based on the correction motion data; and
when obtaining the correction motion data, obtain the correction motion data by removing errors of a jittering phenomenon, a phenomenon of feet not being grounded on a floor or slipping, and a phenomenon of meshes overlapping due to different ratios of a retargeted character, which are included in the motion data, through a filtering-based unnatural motion improvement module of the motion post-processing module.
2 . The device according to claim 1 , wherein the motion post-processing module includes an artificial intelligence-based unnatural motion improvement module, or a physics engine-based unnatural motion improvement module.
3 . The device according to claim 2 , wherein the processor is configured to:
obtain post-processing information including a degree of post-processing of the motion data by analyzing the generated motion data, and
obtain the correction motion data by post-processing the generated motion data using the obtained post-processing information.
4 . A method of operating a motion generation device for generating motion based on input information including text, performed by a processor, comprising:
obtaining, by the processor, input information for a control element having audio, emotion, style, image, pose, and video related to a movement of a character including the text;
when extracting feature information after removing noise included in the input information through an input information pre-processing module included in the memory, extracting, by the processor, a text feature value obtained based on a characteristic of the character, a behavior characteristic of the character, and a generation intent of the text through a text feature extraction model of the input information pre-processing module, extracting a control feature value corresponding to types of elements included in the control element through a feature extraction model for each control element of the input information pre-processing module, extracting a synthetic feature value generated by fusing the text feature value and the control feature value through a feature fusion pre-processing model of the input information pre-processing module;
when generating motion data for the movement of the character from the feature information using an artificial intelligence model trained to infer the motion data for the movement of the character from the feature information through a motion generation module included in the memory, converting, by the processor, the synthetic feature value into a motion feature value for generating motion data using the control feature value through a motion feature mapping model of the motion generation module, generating the motion data based on the motion feature value through a motion feature-based motion generation model of the motion generation module;
obtaining, by the processor, correction motion data by post-processing the motion data through a motion post-processing module included in the memory; and
generating, by the processor, motion animation of the character based on the correction motion data; and
when obtaining the correction motion data, obtaining, by the processor, the correction motion data by removing errors of a jittering phenomenon, a phenomenon of feet not being grounded on a floor or slipping, and a phenomenon of meshes overlapping due to different ratios of a retargeted character, which are included in the motion data, through a filtering-based unnatural motion improvement module of the motion post-processing module.
5 . The method according to claim 4 , wherein the motion post-processing module includes an artificial intelligence-based unnatural motion improvement module, or a physics engine-based unnatural motion improvement module.