IP Library › Granted Patent US 12,437,457
Granted Patent B2
US 12,437,457 · App. 18/312,149 · Granted Oct 7, 2025

Apparatus and method for generating dancing avatar

Inventor: Sang Hoon Lee (Seoul, KR)
Assignee: UIF (University Industry Foundation, Yonsei University
G06T13/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,457
App. No.
18/312,149
Granted
Oct 7, 2025
Kind
B2
Abstract

The present disclosure provides an apparatus and method for generating a dancing avatar, that receives a latent code and map it using a neural network operation to obtain a plurality of genre-specific style codes for each of a plurality of dance genres, and decodes seed motion data and music data, which are motion data that must be referred to when generating an avatar's dance motion, using a genre-specific style code for a dance genre selected among the plurality of genre-specific style codes as a guide, thereby obtaining a dance vector representing a dance motion feature of the avatar in the selected dance genre. According to the present disclosure, it is possible to continuously generate various dance motions of the avatar in relation to previous dance motions, and freely change the dance genre according to user commands or music.

Claims (63)

1. An apparatus for generating a dancing avatar comprising: one or more processors; and a memory that stores one or more programs executed by the one or more processors,

wherein the processors

receive a latent code and map it using a neural network operation to obtain a plurality of genre-specific style codes for each of a plurality of dance genres, and

decode seed motion data and music data, which are motion data that must be referred to when generating an avatar's dance motion, using a genre-specific style code for a dance genre selected among the plurality of genre-specific style codes as a guide, thereby obtaining a dance vector representing a dance motion feature of the avatar in the selected dance genre,

wherein the processors

receive the dance vector, convert it into a format of the motion data to obtain dance data, and

apply an avatar skin to the obtained dance data, thereby generating a dancing avatar,

wherein during training, the processors

receive the dance data and the music data and project them into a virtual common feature vector space to obtain a motion vector and a music vector,

obtain a feature map by transformer encoding the motion vector and the music vector using a neural network operation,

determine the dance genre of the dance data from a genre score obtained by pooling the feature map, and

calculate a loss according to a difference between the determined dance genre and the selected dance genre, and back-propagate it.

2. The apparatus for generating a dancing avatar according to claim 1 ,

wherein the processors obtain the plurality of genre-specific style codes by mapping the latent code to each area divided according to each of a plurality of dance genres in a virtual style space.

3. The apparatus for generating a dancing avatar according to claim 1 ,

wherein the processors randomly and repeatedly generate the latent code.

4. The apparatus for generating a dancing avatar according to claim 1 ,

wherein the processors select one dance genre from the plurality of dance genres in response to a user command.

5. The apparatus for generating a dancing avatar according to claim 1 ,

wherein the processors

project each of the seed motion data and the music data into a virtual common feature vector space to obtain a motion vector and a music vector, and

decode the obtained motion vector and the music vector using a transformer decoder so that features designated by the selected genre-specific style code stand out, thereby obtaining the dance vector.

6. The apparatus for generating a dancing avatar according to claim 1 ,

wherein the processors obtain previously obtained dance data as the seed motion data.

7. The apparatus for generating a dancing avatar according to claim 1 ,

wherein the processors obtain a random value, capture data obtained by capturing user's motions, and 3D dance motion data extracted from a 2D or 3D dance video as an initial value of the seed motion data.

8. The apparatus for generating a dancing avatar according to claim 1 ,

wherein the processors calculate

a style focus loss according to the difference between the dance genre of the seed motion data and the determined dance genre, and

when the dance genre of the seed motion data and the selected dance genre are the same, a dance genre-specific loss calculated as the difference between the seed motion data and the dance data, and

a style diversity loss that maximizes the difference between dance data by obtaining dance data previously obtained from the same selected dance genre as the seed motion data, so that repeatedly generated dance data represents various motions,

and further apply them to the loss.

9. A method for generating a dancing avatar, performed by a computing device having one or more processors and a memory storing one or more programs executed by the one or more processors, comprising the steps of:

receiving a latent code and mapping it using a neural network operation to obtain a plurality of genre-specific style codes for each of a plurality of dance genres; and

decoding seed motion data and music data that must be referred to when generating an avatar's dance motion, using a genre-specific style code for a dance genre selected among the plurality of genre-specific style codes as a guide, thereby obtaining a dance vector representing a dance motion feature of the avatar in the selected dance genre,

wherein the method further includes the steps of

receiving the dance vector, converting it into a format of the motion data to obtain dance data, and applying an avatar skin to the obtained dance data, thereby generating a dancing avatar, and

during training, receiving the dance data and the music data and projecting them into a virtual common feature vector space to obtain a motion vector and a music vector, obtaining a feature map by transformer encoding the motion vector and the music vector using a neural network operation, determining the dance genre of the dance data from a genre score obtained by pooling the feature map, and calculating a loss according to a difference between the determined dance genre and the selected dance genre, and back-propagating it.

10. The method for generating a dancing avatar according to claim 9 ,

wherein the step of obtaining the style codes includes

obtaining the plurality of genre-specific style codes by mapping the latent code to each area divided according to each of a plurality of dance genres in a virtual style space.

11. The method for generating a dancing avatar according to claim 9 ,

wherein the step of obtaining the style codes includes

randomly and repeatedly generating the latent code.

12. The method for generating a dancing avatar according to claim 9 ,

wherein the step of obtaining the style codes includes

selecting one dance genre from the plurality of dance genres in response to a user command.

13. The method for generating a dancing avatar according to claim 9 ,

wherein the step of obtaining a dance vector includes

projecting each of the seed motion data and the music data into a virtual common feature vector space to obtain a motion vector and a music vector, and

decoding the obtained motion vector and the music vector using a transformer decoder so that features designated by the selected genre-specific style code stand out, thereby obtaining the dance vector.

14. The method for generating a dancing avatar according to claim 9 ,

wherein the step of obtaining a dance vector includes

obtaining previously obtained dance data as the seed motion data.

15. The method for generating a dancing avatar according to claim 9 ,

wherein the step of obtaining the feature vector includes

obtaining a random value, capture data obtained by capturing user's motions, and 3D dance motion data extracted from a 2D or 3D dance video as an initial value of the seed motion data.

16. The method for generating a dancing avatar according to claim 9 ,

wherein the step of propagating includes calculating

a style focus loss according to the difference between the dance genre of the seed motion data and the determined dance genre, and

when the dance genre of the seed motion data and the selected dance genre are the same, a dance genre-specific loss calculated as the difference between the seed motion data and the dance data, and

a style diversity loss that maximizes the difference between dance data by obtaining dance data previously obtained from the same selected dance genre as the seed motion data, so that repeatedly generated dance data represents various motions,

and back-propagating by adding them to the loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 4, 2023
From: LEE, SANG HOON
To: UIF (UNIVERSITY INDUSTRY FOUNDATION), YONSEI UNIVERSITY
Reel/Frame 063537/0066 →
Priority Claims (1)
KR 10-2022-0056684 · May 9, 2022 · national
Continuity (1)
Related Publication 20240161378A1 · May 16, 2024
References Cited (15)
US 7297860B2 · Decuir · 2007 [cited by examiner]
US 11816773B2 · Krishnan Gorumkonda · 2023 [cited by examiner]
US 20120139830A1 · Hwang · 2012 [cited by examiner]
US 20180214777A1 · Hingorani · 2018 [cited by examiner]
CN 111986295A · 2020 [cited by examiner]
CN 112330779A · 2021 [cited by examiner]
KR 101270151B1 · 2013 [cited by applicant]
Ferreira, J. P., Coutinho, T. M., Gomes, T. L., Neto, J. F., Azevedo, R., Martins, R., & Nascimento, E. R. (2021). Learning to dance: A graph convolutional adversarial network to generate realistic dance motions from au… [cited by examiner]
Zhang, X., Xu, Y., Yang, S., Gao, L., & Sun, H. (2021). Dance generation with style embedding: Learning and transferring latent representations of dance styles. arXiv preprint arXiv:2104.14802. (Year: 2021). [cited by examiner]
Li, R., Yang, S., Ross, D. A., & Kanazawa, A. (2021). Ai choreographer: Music conditioned 3d dance generation with aist++. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 13401-13412). (Y… [cited by examiner]
Aberman, K., Weng, Y., Lischinski, D., Cohen-Or, D., & Chen, B. (2020). Unpaired motion style transfer from video to animation. ACM Transactions on Graphics (TOG), 39(4), 64-1. (Year: 2020). [cited by examiner]
Guo, X., Zhao, Y., & Li, J. (2021). Dancelt: music-inspired dancing video synthesis. IEEE Transactions on Image Processing, 30, 5559-5572. (Year: 2021). [cited by examiner]
Yuhang Huang et al., “Genre-Conditioned Long-Term 3D Dance Generation Driven by Music,” 2022 IEEE International Conference on Acoustics, Speech and Signal Processing, [Date Added to IEEE Xplore Apr. 27, 2022]. [cited by applicant]
Soomin Park et al., “Diverse Motion Stylization for Multiple Style Domains via Spatial-Temporal Graph-Based Generative Model”, Proceedings of the ACM on Computer Graphics and Interactive Techniques col. 4, Issue 3, [Sep… [cited by applicant]
Ruilong Li et al., “AI Choreographer: Music Conditioned 3D Dance Generation with AIST++”, Proceedings of the IEEE/CVF International Conference on Computer Vision, [Oct. 17, 2021]. [cited by applicant]