IP Library › Granted Patent US 11,816,773
Granted Patent B2
US 11,816,773 · App. 17/487,558 · Granted Nov 14, 2023

Music reactive animation of human characters

Inventors: Gurunandan Krishnan Gorumkonda (Seattle, WA); Hsin-Ying Lee (Sunnyvale, CA); Jie Xu (Cambridge, MA)
Assignee: Snap Inc.
G06T13/205G06N3/044G06N3/045G06N3/08G06T13/40G06T13/80G10H2210/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,816,773
App. No.
17/487,558
Granted
Nov 14, 2023
Kind
B2
Abstract

Example methods for generating an animated character in dance poses to music may include generating, by at least one processor, a music input signal based on an acoustic signal associated with the music, and receiving, by the at least one processor, a model output signal from an encoding neural network. A current generated pose data is generated using a decoding neural network, the current generated pose data being based on previous generated pose data of a previous generated pose, the music input signal, and the model output signal. An animated character is generated based on a current generated pose data; and the animated character caused to be displayed by a display device.

Claims (54)

1. A method for generating an animated character in dance poses to music, the method comprising:

generating, by at least one processor, a music input signal based on an acoustic signal associated with the music;

training, by the at least one processor, an encoding neural network to generate a model output signal including test previous pose data, the music input signal, and test current pose data;

receiving, by the at least one processor, the model output signal from the encoding neural network;

generating, by the at least one processor, current generated pose data using a decoding neural network, the current generated pose data being based on previous generated pose data of a previous generated pose, the music input signal, and the model output signal;

generating, by the at least one processor, an animated character based on the current generated pose data; and

causing, by the at least one processor, the animated character to be displayed by a display device.

2. The method of claim 1 , wherein the encoding neural network and the decoding neural network are based on a conditional recursive neural network model.

3. The method of claim 1 , wherein the model output signal is a unit normal distribution or Gaussian distribution.

4. The method of claim 3 , wherein the generating of the current generated pose data further comprises:

generating a mean and a standard deviation of the model output signal.

5. The method of claim 4 , wherein the generating of the current generated pose data further comprises:

sampling a plurality of pose data from poses generated during the training step from of the model output signal.

6. The method of claim 1 , wherein the generating of the animated character further comprises:

associating one or more portions, wherein at least one portion of the one or more portions comprise a set of joints connected to a set of frames used to represent, of the animated character using the current generated pose data.

7. The method of claim 1 , wherein generating the music input signal further comprises:

generating real-time music features based on at least one of Mel Frequency Cepstral Coefficients (MFCC), chromogram, onset, delta values, and beat indicators,

wherein the music input signal is a music input vector.

8. The method of claim 1 , wherein the generating of the animated character further comprises:

generating a plurality of characters based on the current generated pose data; and

causing the plurality of characters to be displayed on a display device in an animated sequence.

9. The method of claim 1 , further comprising:

wherein the previous pose data is based on the current pose data.

10. A computing apparatus comprising:

at least one processor; and

a memory storing instructions that, when executed by the at least one processor, configure the apparatus to:

generate a music input signal based on an associated music acoustic signal;

train, by at least one processor, an encoding neural network to generate a model output signal using test previous pose data, the music input signal, and test current pose data;

receive the model output signal from the encoding neural network;

generate a current generated pose data using a decoding neural network, the current generated pose data being based on previous generated pose data of a previous generated pose, the music input signal, and the model output signal;

generate an animated character based on the current generated pose data; and

cause the animated character to be displayed by a display device.

11. The computing apparatus of claim 10 , wherein the encoding neural network and the decoding neural network are based on a conditional recursive neural network model.

12. The computing apparatus of claim 10 , wherein the generating of the animated character further comprises:

associate one or more portions, wherein at least one portion of the one or more portions comprise a set of joints connected to a set of frames used to represent, of the animated character using the current generated pose data.

13. The computing apparatus of claim 10 ,

wherein the test previous pose data is based on the current pose data.

14. The computing apparatus of claim 10 , wherein the training of the computing apparatus further jointly trains both a 2D dataset and a 3D dataset; wherein the 2D dataset has at least 25 joints representation for a humanoid figure and the 3D dataset has at least 26 joints representation for the humanoid figure.

15. The computing apparatus of claim 10 , wherein the generating of the animated character further comprises:

generating a plurality of characters based on the current generated pose data; and

causing the plurality of characters to be displayed on a display device in an animated sequence.

16. A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium including instructions that when executed by a computer, cause the computer to:

generate, by at least one processor, a music input signal based on an associated music acoustic;

train, by at least one processor, an encoding neural network to generate a model output signal using test previous pose data, the music input signal, and test current pose data;

receive, by at least one processor, the model output signal from the encoding neural network;

generate, by at least one processor, current generated pose data using a decoding neural network, the current generated pose data being based on a previous generated pose of a previous generated pose, the music input signal, and the model output signal;

generate, by at least one processor, current generated pose based on the current generated pose data; and

cause, by at least one processor, the animated character to be displayed by a display device.

17. The non-transitory computer-readable storage medium of claim 16 ,

wherein the test previous pose data is based on the current pose data, a mean of the model output signal and a standard deviation of the model output signal.

18. The non-transitory computer-readable storage medium of claim 16 , wherein the generating of the animated character further comprises:

associating one or more portions, wherein at least one portion of the one or more portions comprise a set of joints connected to a set of frames used to represent, of the animated character using the current generated pose data.

19. The non-transitory computer-readable storage medium of claim 16 , wherein generate the animated character further comprises, generating a plurality of characters based on the current generated pose data and causing the plurality of characters to be displayed on a display device in an animated sequence.

20. The non-transitory computer-readable storage medium of claim 16 , wherein the model output signal is generated by transitioning from initially a forced training to an auto-regressive training over a predetermined number of iterations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2023
From: GORUMKONDA, GURUNANDAN KRISHNAN; LEE, HSIN-YING; XU, JIE
To: SNAP INC.
Reel/Frame 065090/0270 →
Continuity (2)
Provisional Application 63085755 · Sep 30, 2020
Related Publication 20220101586A1 · Mar 31, 2022
Cited By (8)
US 12,293,444 US 12,299,793 US 12,437,457 US 12,488,525 US 12,573,122 US 12,620,157 US 12,646,240 US 12,675,930