IP Library Granted Patent US 12675526
Granted Patent B2
US 12675526 · App. 18/570,310 · Granted Jul 7, 2026

Method, device, storage medium and program product for music screening

Inventors: Ding Liu (Los Angeles, CA); Xiaojie Jin (Los Angeles, CA); Yan Wang (Beijing, CN); Weibo Gong (Beijing, CN)
Assignee: LEMON INC.
G06F16/68G06F16/55G06F16/65
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12675526
App. No.
18/570,310
Granted
Jul 7, 2026
Kind
B2
Abstract

Embodiments of the disclosure provide a method, device, storage medium and program product for music screening. The method includes: obtaining at least one image and at least one piece of to-be-selected music; determining an analysis result of the at least one image corresponding to an image classification tag based on N predetermined image classification tags, N being an integer greater than or equal to 1; determining attribute information for each piece of to-be-selected music based on the at least one image and the at least one piece of to-be-selected music; determining target music that matches the at least one image among the at least one piece of to-be-selected music based on the analysis result and the attribute information of each piece of to-be-selected music.

Claims (96)

1 . A method of music screening for automatic soundtrack generation in a video editing system, comprising:

obtaining, by a processor, at least one image and at least one piece of to-be-selected music;

determining, by an image analysis model, an analysis result of the at least one image corresponding to an image classification tag based on N predetermined image classification tags, N being an integer greater than or equal to 1;

determining, by a music matching model, attribute information for each of the at least one piece of to-be-selected music based on the at least one image and the at least one piece of to-be-selected music; and

determining, by the processor, a target music that matches the at least one image among the at least one piece of to-be-selected music based on the analysis result obtained from the image analysis model and the attribute information of each of the at least one piece of to-be-selected music obtained from the music matching model;

wherein the determining the target music that matches the at least one image among the at least one piece of to-be-selected music comprises:

for a piece of music of the at least one piece of to-be-selected music, determining at least one first score of the piece of music corresponding to at least one image classification tag based on the analysis result and the attribute information of the piece of music;

determining, by the processor, a target score of the piece of music based on the at least one first score corresponding to the at least one image classification tag and an initial score of the piece of music, the initial score of the piece of music being included in the attribute information of the piece of music, comprising:

obtaining, by the processor, respective weights of the N image classification tags, the weights being preset and periodically updated based on accumulated matching data; and

determining, by the processor, the target score of the piece of music based on N first scores corresponding to the N image classification tags, the respective weights of the N image classification tags and the initial score of the piece of music; and

determining, by the processor, the target music based on at least one respective target score of the at least one piece of to-be-selected music, comprising:

obtaining, by the processor, a music sequence by ranking the at least one piece of to-be-selected music in an order of the at least one respective target score of the at least one piece of to-be-selected music; and

determining, by the processor, a predetermined number of pieces of to-be-selected music ranking at a top of the music sequence as the target music that matches the at least one image, and automatically inserting the target music as a synchronized soundtrack into a user-generated video comprising the at least one image.

2 . The method according to claim 1 , wherein the N image classification tags comprise at least one of: an image emotion, an image style, or an image theme; the attribute information further comprises M music classification tags of the to-be-selected music, M being an integer greater than or equal to 1; and the M music classification tags comprise at least one of: a music style, a music emotion, or a music scene.

3 . The method according to claim 2 , wherein the analysis result is an emotion analysis result of the at least one image corresponding to the image emotion, the emotion analysis result comprising at least one first image emotion and a confidence level of the at least one first image emotion; the attribute information comprises a music emotion of the to-be-selected music, the music emotion comprising the at least one first music emotion;

the determining a first score of the to-be-selected music corresponding to the image emotion based on the emotion analysis result and the music emotion comprises:

determining a score of the at least one first music emotion corresponding to the image emotion according to the at least one first image emotion, the confidence level of the at least one first image emotion and the at least one first music emotion; and

determining a ratio of a sum of the scores of the at least one first music emotion corresponding to the image emotion to a total number of emotions of the at least one first music emotion as a first score of the to-be-selected music corresponding to the image emotion.

4 . The method according to claim 3 , wherein the determining a score of the at least one first music emotion corresponding to the image emotion based on the at least one first image emotion, the confidence level of the at least one first image emotion and the at least one first music emotion comprises:

step 1: obtaining an ith first music emotion of the at least one first music emotion;

step 2: obtaining a jth first image emotion of the at least one first image emotion;

step 3: looking up a jth correlation value corresponding to the ith first music emotion and the jth first image emotion in a prestored correlation list; the correlation list comprising a plurality of correlation values corresponding to the first music emotion and the first image emotion;

step 4: determining a sum of a product of the jth correlation value and the confidence level of the jth first image emotion and a (j−1)th score of the jth first music emotion corresponding to a (j−1)th first image emotion as a jth score of the ith first music emotion corresponding to the jth first image emotion;

adding j by 1 and repeating steps 2, 3, and 4 until j is equal to Y, to obtain a Yth score of the ith first music emotion corresponding to the Yth first image emotion; and

determining a ratio of the Yth score to a sum of the confidence levels of the at least one first image emotion as a score of the ith first music emotion corresponding to the image emotion,

wherein i is an integer between 1 and X, j is an integer between 1 and Y, X is a total number of emotions of the at least one first music emotion, and Y is a total number of emotions of the at least one first image emotion.

5 . The method according to claim 2 , wherein the analysis result is a style analysis result of the at least one image corresponding to the image style, the style analysis result comprising at least one first image style,

the attribute information comprises a music emotion and a music genre of the to-be-selected music, the music genre comprising at least one first music genre, and the music emotion comprising at least one first music emotion,

the determining a first score of the to-be-selected music corresponding to the image style based on the style analysis result, the music emotion and the music genre comprises:

determining a third score of the music genre corresponding to the image style based on the at least one first image style, the at least one first music genre and a prestored first predetermined list; the first predetermined list comprising a plurality of first image styles and a first music genre corresponding to each first image style;

determining a fourth score of the music emotion corresponding to the image style based on at least one first image style, the at least one first music emotion and a prestored second predetermined list; the second predetermined list comprising first image styles and a first music emotion corresponding to each first image style; and

determining a sum of the third score and the fourth score as a first score of the to-be-selected music corresponding to the image style.

6 . The method according to claim 5 , wherein the determining a third score of the music genre corresponding to the image style based on the at least one first image style, the at least one first music genre and a prestored first predetermined list comprises:

for each first image style, looking up a first music genre corresponding to the first image style in the first predetermined list;

if a found first music genre corresponding to the first image style exists among the at least one first music genre, obtaining a score of the found first music genre corresponding to the first image style;

determining a sum of the scores of the found first music genre corresponding to the first image style as a score of the music genre corresponding to the first image style; and

determining a largest one among the scores of the music genre corresponding to each of the first image styles as a third score of the music genre corresponding to the image style.

7 . The method according to claim 2 , wherein the analysis result is a theme analysis result of the at least one image corresponding to the image theme, the theme analysis result comprising at least one first image theme,

the attribute information comprises a music scene, a music emotion and a music genre of to-be-selected music, the music scene comprising at least one first music scene, the music emotion comprising at least one first music emotion, and the music genre comprising at least one first music genre,

the determining a first score of the to-be-selected music corresponding to the image theme based on the theme analysis result, the music scene, the music emotion and the music genre comprises:

determining a fifth score of the music scene corresponding to the image theme based on the at least one first image theme, the at least one first music scene, and a prestored third predetermined list; the third predetermined list comprising a plurality of first image themes and a first music scene corresponding to each first image style;

determining a sixth score of the music emotion corresponding to the image theme based on the at least one first image theme, the at least one first music emotion and a prestored fourth predetermined list; the fourth predetermined list comprising a plurality of first image themes and a first music emotion corresponding to each first image style;

determining a seventh score of the music genre corresponding to the image theme based on the at least one first image theme, the at least one first music genre and a prestored fifth predetermined list; the fifth predetermined list comprising a plurality of first image themes and a first music genre corresponding to each first image style; and

determining a sum of the fifth score, the sixth score and the seventh score as a first score of the to-be-selected music corresponding to the image theme.

8 . The method according to claim 1 , wherein the determining the target score of the piece of music based on N first scores corresponding to the N image classification tags, respective weights corresponding to the N image classification tags and an initial score of the piece of music comprises:

for each image classification tag, determining a product of a first score of the to-be-selected music corresponding to each image classification tag and a weight corresponding to the image classification tag to obtain a first product corresponding to the image classification tag; and

determining a sum of the first product corresponding to the N image classification tags and the initial score of the to-be-selected music as a target score of the to-be-selected music.

9 . The method according to claim 1 , wherein the determining an analysis result of the at least one image corresponding to the image classification tag based on the predetermined N image classification tags comprises:

analyzing and processing the at least one image based on the predetermined N image classification tags and with respective image analysis models corresponding to the N image classification tags, to obtain an analysis result of the at least one image corresponding to the image classification tag,

wherein the respective image analysis models corresponding to the N image classification tags are obtained by training a respective plurality of sample images corresponding to the N image classification tags.

10 . The method according to claim 1 , wherein the determining attribute information of each of the at least one piece of to-be-selected music based on the at least one image and the at least one piece of to-be-selected music comprises:

processing the at least one image and each of the at least one piece of to-be-selected music with a pre-trained music matching model, to obtain the attribute information of each of the at least one piece of to-be-selected music, the music matching model being obtained by training a plurality of sample images and a plurality of pieces of sample music.

11 . The method according to claim 1 , wherein the obtaining at least one image comprises:

obtaining at least one frame of image from at least one to-be-processed video and determining the at least one frame of image as the at least one image; or

obtaining at least one frame of image from at least one to-be-processed video and determining the at least one frame of image and a prestored image as the at least one image.

12 . A terminal device, comprising:

a processor and a memory; the memory storing computer-executed instructions;

the processor executing the computer-executed instructions stored in the memory, causing the processor to perform a method of music screening for automatic soundtrack generation in a video editing system, the method comprising:

obtaining at least one image and at least one piece of to-be-selected music;

determining, by an image analysis model, an analysis result of the at least one image corresponding to an image classification tag based on N predetermined image classification tags, N being an integer greater than or equal to 1;

determining, by a music matching model, attribute information for each of the at least one piece of to-be-selected music based on the at least one image and the at least one piece of to-be-selected music; and

determining a target music that matches the at least one image among the at least one piece of to-be-selected music based on the analysis result obtained from the image analysis model and the attribute information of each of the at least one piece of to-be-selected music obtained from the music matching model;

wherein the determining the target music that matches the at least one image among the at least one piece of to-be-selected music comprises:

for a piece of music of the at least one piece of to-be-selected music, determining at least one first score of the piece of music corresponding to at least one image classification tag based on the analysis result and the attribute information of the piece of music;

determining a target score of the piece of music based on the at least one first score corresponding to the at least one image classification tag and an initial score of the piece of music, the initial score of the piece of music being included in the attribute information of the piece of music, comprising:

obtaining, by the processor, respective weights of the N image classification tags, the weights being preset and periodically updated based on accumulated matching data; and

determining, by the processor, the target score of the piece of music based on N first scores corresponding to the N image classification tags, the respective weights of the N image classification tags and the initial score of the piece of music; and

determining the target music based on at least one respective target score of the at least one piece of to-be-selected music, comprising:

obtaining, by the processor, a music sequence by ranking the at least one piece of to-be-selected music in an order of the at least one respective target score of the at least one piece of to-be-selected music; and

determining, by the processor, a predetermined number of pieces of to-be-selected music ranking at a top of the music sequence as the target music that matches the at least one image, and automatically inserting the target music as a synchronized soundtrack into a user-generated video comprising the at least one image.

13 . The device according to claim 12 , wherein the N image classification tags comprise at least one of: an image emotion, an image style, or an image theme; the attribute information further comprises M music classification tags of the to-be-selected music, M being an integer greater than or equal to 1; and the M music classification tags comprise at least one of: a music style, a music emotion, or a music scene.

14 . The device according to claim 13 , wherein the analysis result is an emotion analysis result of the at least one image corresponding to the image emotion, the emotion analysis result comprising at least one first image emotion and a confidence level of the at least one first image emotion; the attribute information comprises a music emotion of the to-be-selected music, the music emotion comprising the at least one first music emotion,

the determining a first score of the to-be-selected music corresponding to the image emotion based on the emotion analysis result and the music emotion comprises:

determining a score of the at least one first music emotion corresponding to the image emotion according to the at least one first image emotion, the confidence level of the at least one first image emotion and the at least one first music emotion; and

determining a ratio of a sum of the scores of the at least one first music emotion corresponding to the image emotion to a total number of emotions of the at least one first music emotion as a first score of the to-be-selected music corresponding to the image emotion.

15 . The device according to claim 14 , wherein the determining a score of the at least one first music emotion corresponding to the image emotion based on the at least one first image emotion, the confidence level of the at least one first image emotion and the at least one first music emotion comprises:

step 1: obtaining an ith first music emotion of the at least one first music emotion;

step 2: obtaining a jth first image emotion of the at least one first image emotion;

step 3: looking up a jth correlation value corresponding to the ith first music emotion and the jth first image emotion in a prestored correlation list; the correlation list comprising a plurality of correlation values corresponding to the first music emotion and the first image emotion;

step 4: determining a sum of a product of the jth correlation value and the confidence level of the jth first image emotion and a (j−1)th score of the jth first music emotion corresponding to a (j−1)th first image emotion as a jth score of the ith first music emotion corresponding to the jth first image emotion;

adding j by 1 and repeating steps 2, 3, and 4 until j is equal to Y, to obtain a Yth score of the ith first music emotion corresponding to the Yth first image emotion; and

determining a ratio of the Yth score to a sum of the confidence levels of the at least one first image emotion as a score of the ith first music emotion corresponding to the image emotion,

wherein i is an integer between 1 and X, j is an integer between 1 and Y, X is a total number of emotions of the at least one first music emotion, and Y is a total number of emotions of the at least one first image emotion.

16 . A non-transitory computer-readable storage medium, wherein computer-executed instructions are stored in the computer-readable storage medium, the computer-executed instructions, when executed by a processor, performing a method of music screening for automatic soundtrack generation in a video editing system, the method comprising:

obtaining, by a processor, at least one image and at least one piece of to-be-selected music;

determining, by an image analysis model, an analysis result of the at least one image corresponding to an image classification tag based on N predetermined image classification tags, N being an integer greater than or equal to 1;

determining, by a music matching model, attribute information for each of the at least one piece of to-be-selected music based on the at least one image and the at least one piece of to-be-selected music; and

determining, by the processor, a target music that matches the at least one image among the at least one piece of to-be-selected music based on the analysis result obtained from the image analysis model and the attribute information of each of the at least one piece of to-be-selected music obtained from the music matching model;

wherein the determining the target music that matches the at least one image among the at least one piece of to-be-selected music comprises:

for a piece of music of the at least one piece of to-be-selected music, determining at least one first score of the piece of music corresponding to at least one image classification tag based on the analysis result and the attribute information of the piece of music;

determining, by the processor, a target score of the piece of music based on the at least one first score corresponding to the at least one image classification tag and an initial score of the piece of music, the initial score of the piece of music being included in the attribute information of the piece of music, comprising:

obtaining, by the processor, respective weights of the N image classification tags, the weights being preset and periodically updated based on accumulated matching data; and

determining, by the processor, the target score of the piece of music based on N first scores corresponding to the N image classification tags, the respective weights of the N image classification tags and the initial score of the piece of music; and

determining, by the processor, the target music based on at least one respective target score of the at least one piece of to-be-selected music, comprising:

obtaining, by the processor, a music sequence by ranking the at least one piece of to-be-selected music in an order of the at least one respective target score of the at least one piece of to-be-selected music; and

determining, by the processor, a predetermined number of pieces of to-be-selected music ranking at a top of the music sequence as the target music that matches the at least one image, and automatically inserting the target music as a synchronized soundtrack into a user-generated video comprising the at least one image.