IP Library › Granted Patent US 12,394,250
Granted Patent B2
US 12,394,250 · App. 18/018,879 · Granted Aug 19, 2025

Movement extraction method and apparatus for dance video, computer device, and storage medium

Inventors: Shichen Zhao (Shanghai, CN); Weijia Li (Shanghai, CN); Chaoran Li (Shanghai, CN); Peng Wang (Shanghai, CN); Zhihui Chen (Shanghai, CN)
Assignee: SHANGHAI BILIBILI TECHNOLOGY CO., LTD.
G06V40/23G06F18/23G06T7/248G06V10/762G06V20/41A63F13/65A63F13/816A63F2300/6607G06T2207/10016G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,250
App. No.
18/018,879
Granted
Aug 19, 2025
Kind
B2
Abstract

This disclosure provides techniques of extracting movements from a dance video. The techniques comprise receiving a dance video that includes one or more dancers and that is uploaded by a user, and obtaining a video frame in the dance video; recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames corresponding to a target dancer, where the target dancer is selected from the one or more dancers; performing cluster analysis on the skeleton node images recognized from the dance video and corresponding to each target dancer to obtain a plurality of cluster sets; determining a cluster center in each of the plurality of cluster sets as a key skeleton node image; and sequentially outputting the key skeleton node images to obtain a standard movement sequence corresponding to each target dancer in the dance video.

Claims (78)

1. A method of extracting movements from a dance video, comprising:

receiving a dance video that comprises one or more dancers and that is uploaded by a user, and obtaining video frames in the dance video;

recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames corresponding to each target dancer selected from the one or more dancers in the dance video, each of the skeleton node images is associated with an identification number, and identification numbers of the skeleton node images indicate a chronological order of video frames in the dance video from which the skeleton node images are generated;

performing cluster analysis on the skeleton node images recognized from the dance video and corresponding to each target dancer to obtain a plurality of cluster sets, wherein the performing cluster analysis on the skeleton node images corresponding to each target dancer to obtain a plurality of cluster sets comprises:

selecting K target skeleton node images from all the skeleton node images,

calculating similarities between each of remaining skeleton node images and the target skeleton node images,

grouping each of the remaining skeleton node images and a corresponding target skeleton node image based on the calculated similarities to obtain K candidate sets,

determining a new target skeleton node image based on identification numbers of skeleton node images comprised in each of the K candidate sets,

repeating calculation of similarities based on K new target skeleton node images and obtaining K new candidate sets, and

in response to determining that identification numbers of K new target skeleton node images no longer change or a preset number of repetitions is reached, identifying K new candidate sets corresponding to the K new target skeleton node images as the plurality of cluster sets;

determining a cluster center in each of the plurality of cluster sets as a key skeleton node image; and

sequentially outputting key skeleton node images to obtain a standard movement sequence corresponding to each target dancer in the dance video.

2. The method according to claim 1 , wherein the dance video comprises a plurality of dancers, and the target dancer is one selected from the plurality of dancers.

3. The method according to claim 1 , wherein the recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames further comprises:

determining a skeleton node model; and

extracting the skeleton node images corresponding to the target dancer from the video frames based on the determined skeleton node model, wherein each of the skeleton node images comprises data associated with a plurality of skeleton nodes.

4. The method according to claim 3 , wherein the recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames further comprises:

determining whether a number of skeleton nodes comprised in each of the skeleton node images is within a preset range, and deleting one or more skeleton node images when a number of skeleton nodes in each of the one or more skeleton node images is not within the preset range.

5. The method according to claim 1 , wherein the K target skeleton node images are determined based on a rhythm of the dance video.

6. The method according to claim 1 , wherein the similarities are determined based on Euclidean distances, and the calculating similarities between each of remaining skeleton node images and the target skeleton node images comprises:

calculating a Euclidean distance between each skeleton node in each of the remaining skeleton node images and a corresponding skeleton node in each of the target skeleton node images; and

determining a sum of Euclidean distances for all skeleton nodes as a Euclidean distance between each of the remaining skeleton node images and each of the target skeleton node images.

7. The method according to claim 1 , wherein the determining a new target skeleton node image based on identification numbers of skeleton node images comprised in each of the K candidate sets comprises:

obtaining an identification number of each skeleton node image comprised in each of the K candidate sets;

averaging and rounding the identification numbers of all skeleton node images comprised in each of the K candidate sets to obtain a target identification number; and

identifying a skeleton node image corresponding to the target identification number as the new target skeleton node image.

8. The method according to claim 1 , further comprising:

calculating a difference between identification numbers of any two adjacent key skeleton node images; and

in response to determining that the difference is less than a first threshold, deleting one of the two adjacent key skeleton, the one of the two adjacent key skeleton node images associated with a later identification number.

9. A computing system, comprising a memory, a processor, and computer-readable instructions stored on the memory and executable on the processor, wherein when executing the computer-readable instructions, the processor implements operations comprising:

receiving a dance video that comprises one or more dancers and that is uploaded by a user, and obtaining video frames in the dance video;

recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames corresponding to each target dancer selected from the one or more dancers in the dance video, each of the skeleton node images is associated with an identification number, and identification numbers of the skeleton node images indicate a chronological order of video frames in the dance video from which the skeleton node images are generated;

performing cluster analysis on the skeleton node images recognized from the dance video and corresponding to each target dancer to obtain a plurality of cluster sets, wherein the performing cluster analysis on the skeleton node images corresponding to each target dancer to obtain a plurality of cluster sets comprises:

selecting K target skeleton node images from all the skeleton node images,

calculating similarities between each of remaining skeleton node images and the target skeleton node images,

grouping each of the remaining skeleton node images and a corresponding target skeleton node image based on the calculated similarities to obtain K candidate sets,

determining a new target skeleton node image based on identification numbers of skeleton node images comprised in each of the K candidate sets,

repeating calculation of similarities based on K new target skeleton node images and obtaining K new candidate sets, and

in response to determining that identification numbers of K new target skeleton node images no longer change or a preset number of repetitions is reached, identifying K new candidate sets corresponding to the K new target skeleton node images as the plurality of cluster sets;

determining a cluster center in each of the plurality of cluster sets as a key skeleton node image; and

sequentially outputting key skeleton node images to obtain a standard movement sequence corresponding to each target dancer in the dance video.

10. The computing system according to claim 9 , wherein the dance video comprises a plurality of dancers, and the target dancer is one selected from the plurality of dancers.

11. The computing system according to claim 9 , wherein the recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames further comprises:

determining a skeleton node model; and

extracting the skeleton node images corresponding to the target dancer from the video frames based on the determined skeleton node model, wherein each of the skeleton node images comprises data associated with a plurality of skeleton nodes.

12. The computing system according to claim 11 , wherein the recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames further comprises:

determining whether a number of skeleton nodes comprised in each of the skeleton node images is within a preset range, and deleting one or more skeleton node images when a number of skeleton nodes in each of the one or more skeleton node images is not within the preset range.

13. The computing system according to claim 9 , wherein the K target skeleton node images are determined based on a rhythm of the dance video.

14. The computing system according to claim 9 , wherein the similarities are determined based on Euclidean distances, and the calculating similarities between each of remaining skeleton node images and the target skeleton node images comprises:

calculating a Euclidean distance between each skeleton node in each of the remaining skeleton node images and a corresponding skeleton node in each of the target skeleton node images; and

determining a sum of Euclidean distances for all skeleton nodes as a Euclidean distance between each of the remaining skeleton node images and each of the target skeleton node images.

15. The computing system according to claim 9 , wherein the determining a new target skeleton node image based on identification numbers of skeleton node images comprised in each of the K candidate sets comprises:

obtaining an identification number of each skeleton node image comprised in each of the K candidate sets;

averaging and rounding the identification numbers of all skeleton node images comprised in each of the K candidate sets to obtain a target identification number; and

identifying a skeleton node image corresponding to the target identification number as the new target skeleton node image.

16. The computing system according to claim 9 , the operations further comprising:

calculating a difference between identification numbers of any two adjacent key skeleton node images; and

in response to determining that the difference is less than a first threshold, deleting one of the two adjacent key skeleton, the one of the two adjacent key skeleton node images associated with a later identification number.

17. A non-transitory computer-readable storage medium having stored thereon computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the processor implements operations comprising:

receiving a dance video that comprises one or more dancers and that is uploaded by a user, and obtaining video frames in the dance video;

recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames corresponding to a target dancer, wherein the target dancer is selected from the one or more dancers in the dance video, each of the skeleton node images is associated with an identification number, and identification numbers of the skeleton node images indicate a chronological order of video frames in the dance video from which the skeleton node images are generated;

performing cluster analysis on the skeleton node images recognized from the dance video and corresponding to each target dancer to obtain a plurality of cluster sets, wherein the performing cluster analysis on the skeleton node images corresponding to each target dancer to obtain a plurality of cluster sets comprises:

selecting K target skeleton node images from all the skeleton node images,

calculating similarities between each of remaining skeleton node images and the target skeleton node images,

clustering each of the remaining skeleton node images and a corresponding target skeleton node image based on the calculated similarities to obtain K candidate sets,

determining a new target skeleton node image based on identification numbers of skeleton node images comprised in each of the K candidate sets,

repeating calculation of similarities based on K new target skeleton node images and obtaining K new candidate sets, and

in response to determining that identification numbers of K new target skeleton node images no longer change or a preset number of repetitions is reached, identifying K new candidate sets corresponding to the K new target skeleton node images as the plurality of cluster sets;

determining a cluster center in each of the plurality of cluster sets as a key skeleton node image; and

sequentially outputting key skeleton node images to obtain a standard movement sequence corresponding to each target dancer in the dance video.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames further comprises:

determining a skeleton node model; and

extracting the skeleton node images corresponding to each target dancer from the video frames based on the determined skeleton node model, wherein each of the skeleton node images comprises data associated with a plurality of skeleton nodes.

19. The non-transitory computer-readable storage medium of claim 17 , wherein the recognizing, based on a skeleton node recognition algorithm, skeleton node images from the video frames further comprises:

determining whether a number of skeleton nodes comprised in each of the skeleton node images is within a preset range, and deleting one or more skeleton node images when a number of skeleton nodes in each of the one or more skeleton node images is not within the preset range.

20. The non-transitory computer-readable storage medium of claim 17 , the operations further comprising:

calculating a difference between identification numbers of any two adjacent key skeleton node images; and

in response to determining that the difference is less than a first threshold, deleting one of the two adjacent key skeleton, the one of the two adjacent key skeleton node images associated with a later identification number.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2023
From: ZHAO, SHICHEN; LI, WEIJIA; LI, CHAORAN; WANG, PENG; CHEN, ZHIHUI
To: SHANGHAI BILIBILI TECHNOLOGY CO., LTD.
Reel/Frame 062538/0691 →
Priority Claims (1)
CN 202010784431.7 · Aug 6, 2020 · national
Continuity (1)
Related Publication 20230306787A1 · Sep 28, 2023
References Cited (19)
US 8448056B2 · Pulsipher · 2013 [cited by examiner]
US 10643492B2 · Lee · 2020 [cited by examiner]
US 20070040836A1 · Schickler · 2007 [cited by examiner]
US 20190392729A1 · Lee et al. · 2019 [cited by applicant]
CN 108665492A · 2018 [cited by applicant]
CN 109151501A · 2019 [cited by applicant]
CN 109308438A · 2019 [cited by applicant]
CN 109508656A · 2019 [cited by applicant]
CN 110096950A · 2019 [cited by applicant]
CN 110245638A · 2019 [cited by applicant]
CN 110448870A · 2019 [cited by applicant]
CN 110728220A · 2020 [cited by applicant]
CN 111144217A · 2020 [cited by applicant]
WO WO2016019973A1 · 2016 [cited by applicant]
Classification of K-Pop Dance Movements Based on Skeleton Information Obtained by a Kinect Sensor, by Kim et al., Sensors 2017, 17, 1261; doi:10.3390/s17061261 (Year: 2017). [cited by examiner]
International Patent Application No. PCT/CN2021/101384; Int'l Search Report; dated Sep. 26, 2021; 2 pages. [cited by applicant]
Zhao et al.; “Optimization and Behavior Identification of Keyframes in Human Action Video”; Journal of Graphics; vol. 39 No. 3; Jun. 2018; p. 463-469 (contains English Abstract). [cited by applicant]
China Patent Application No. 202010784431.7; First Office Action; dated Apr. 26, 2024; 24 pages. [cited by applicant]
China Patent Application No. 202010784431.7; Second Office Action; dated Aug. 13, 2024; 20 pages. [cited by applicant]