IP Library Granted Patent US 11,763,849
Granted Patent B1
US 11,763,849 · App. 17/815,402 · Granted Sep 19, 2023

Automatic and fast generation of music audio content for videos

Inventors: Zhihao Ouyang (Los Angeles, CA); Daiyu Zhang (Los Angeles, CA); Bochen Li (Los Angeles, CA); Baoman Liu (Los Angeles, CA); Liuqing Yang (Los Angeles, CA)
Assignee: LEMON INC.
G11B27/031G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,763,849
App. No.
17/815,402
Granted
Sep 19, 2023
Kind
B1
Abstract

The present disclosure describes techniques for automatically and fast generating music for videos. The techniques comprise receiving a video from a user. The video may comprise a plurality of segments of frames. Information may be extracted from the video, wherein the extracted information comprises information indicating motion speed in the video, information indicating motion saliency in the video, information indicating scene transition in the video, and timing information associated with the video. A plurality of sets of music notes matching the plurality of segments of frames may be generated based at least in part on the extracted information. A plurality of vectors corresponding to the plurality of sets of music notes may be generated. The plurality of pieces of music audio corresponding to the plurality of segments of frames may be generated based at least in part on the plurality of vectors.

Claims (56)

1. A method of automatically and efficiently generating music audio content for videos, comprising:

receiving a video from a user, the video comprising a plurality of segments of frames;

extracting information from the video, wherein the extracted information comprises information indicating motion speed in the video, information indicating motion saliency in the video, information indicating scene transition in the video, and timing information associated with the video; and

generating a plurality of sets of music notes matching the plurality of segments of frames based at least in part on the extracted information using a model, wherein the model is pre-trained and learns to correlate video motion speed with music note density, correlate video motion saliency with music note strength, correlate video scene transition with music structure, and correlate video timing with music beat;

generating, by a vector generation model, a plurality of vectors utilizing the plurality of sets of music notes, wherein each of the plurality of vectors indicates at least one music feature of one of a plurality of pieces of music audio, and wherein the vector generation model is configured to determine music characteristics based on music notes; and

generating the plurality of pieces of music audio corresponding to the plurality of segments of frames by inputting the plurality of vectors into an audio generation model, wherein the generating the plurality of pieces of music audio corresponding to the plurality of segments of frames further comprises:

determining at least one template based on each of the plurality of vectors by the audio generation model,

acquiring the at least one template from a pre-stored database by the audio generation model, wherein the at least one template comprises at least one audio file, and

generating a piece of music audio corresponding to each of the plurality of segments of frames based at least in part on the at least one template by the audio generation model.

2. The method of claim 1 ,

wherein the pre-stored database comprises a plurality of templates, and each of the plurality of templates comprises an audio file with a particular music feature.

3. The method of claim 1 , wherein the generating a piece of music audio corresponding to each of the plurality of segments of frames based at least in part on the at least one template further comprises:

modifying the at least one template by adding a melody based on one of the plurality of sets of music notes.

4. The method of claim 1 , wherein the at least one music feature of one of the plurality of pieces of music audio comprises at least one of music style, bar structure, or music instrument.

5. The method of claim 1 , further comprising:

generating music audio content for the video based at least in part on synthesizing the plurality of pieces of music audio, wherein the music audio content matches motion, intensity, and transition in the video.

6. The method of claim 5 , further comprising:

determining whether the user likes the music audio content based on user input.

7. The method of claim 6 , further comprising:

presenting a plurality of options in response to determining that the user does not like the music audio content; and

updating the music audio content based on one or more options selected by the user, wherein the one or more options are among the plurality of options.

8. The method of claim 5 , further comprising:

finetuning video-audio matching between the video and the music audio content by applying video warping during the generating the music audio content.

9. A system of automatically and efficiently generating music audio content for videos, comprising:

at least one processor; and

at least one memory communicatively coupled to the at least one processor and comprising computer-readable instructions that upon execution by the at least one processor cause the at least one processor to perform operations comprising:

receiving a video from a user, the video comprising a plurality of segments of frames;

extracting information from the video, wherein the extracted information comprises information indicating motion speed in the video, information indicating motion saliency in the video, information indicating scene transition in the video, and timing information associated with the video; and

generating a plurality of sets of music notes matching the plurality of segments of frames based at least in part on the extracted information using a model, wherein the model is pre-trained and learns to correlate video motion speed with music note density, correlate video motion saliency with music note strength, correlate video scene transition with music structure, and correlate video timing with music beat;

generating, by a vector generation model, a plurality of vectors utilizing the plurality of sets of music notes, wherein each of the plurality of vectors indicates at least one music feature of one of a plurality of pieces of music audio, and wherein the vector generation model is configured to determine music characteristics based on music notes; and

generating the plurality of pieces of music audio corresponding to the plurality of segments of frames by inputting the plurality of vectors into an audio generation model, wherein the generating the plurality of pieces of music audio corresponding to the plurality of segments of frames further comprises:

determining at least one template based on each of the plurality of vectors by the audio generation model,

acquiring the at least one template from a pre-stored database by the audio generation model, wherein the at least one template comprises at least one audio file, and

generating a piece of music audio corresponding to each of the plurality of segments of frames based at least in part on the at least one template by the audio generation model.

10. The system of claim 9 ,

wherein the pre-stored database comprises a plurality of templates, and each of the plurality of templates comprises an audio file with a particular music feature.

11. The system of claim 9 , wherein the generating a piece of music audio corresponding to each of the plurality of segments of frames based at least in part on the at least one template further comprises:

modifying the at least one template by adding a melody based on one of the plurality of sets of music notes.

12. The system of claim 9 , wherein the at least one music feature of one of the plurality of pieces of music audio comprises at least one of music style, bar structure, or music instrument.

13. The system of claim 9 , the operations further comprising:

generating music audio content for the video based at least in part on synthesizing the plurality of pieces of music audio, wherein the music audio content matches motion, intensity, and transition in the video.

14. A non-transitory computer-readable storage medium, storing computer-readable instructions that upon execution by a processor cause the processor to implement operations comprising:

receiving a video from a user, the video comprising a plurality of segments of frames;

extracting information from the video, wherein the extracted information comprises information indicating motion speed in the video, information indicating motion saliency in the video, information indicating scene transition in the video, and timing information associated with the video; and

generating a plurality of sets of music notes matching the plurality of segments of frames based at least in part on the extracted information using a model, wherein the model is pre-trained and learns to correlate video motion speed with music note density, correlate video motion saliency with music note strength, correlate video scene transition with music structure, and correlate video timing with music beat;

generating, by a vector generation model, a plurality of vectors utilizing the plurality of sets of music notes, wherein each of the plurality of vectors indicates at least one music feature of one of a plurality of pieces of music audio, and wherein the vector generation model is configured to determine music characteristics based on music notes; and

generating the plurality of pieces of music audio corresponding to the plurality of segments of frames by inputting the plurality of vectors into an audio generation model, wherein the generating the plurality of pieces of music audio corresponding to the plurality of segments of frames further comprises:

determining at least one template based on each of the plurality of vectors by the audio generation model,

acquiring the at least one template from a pre-stored database by the audio generation model, wherein the at least one template comprises at least one audio file, and

generating a piece of music audio corresponding to each of the plurality of segments of frames based at least in part on the at least one template by the audio generation model.

15. The non-transitory computer-readable storage medium of claim 14 ,

wherein the pre-stored database comprises a plurality of templates, and each of the plurality of templates comprises an audio file with a particular music feature.

16. The non-transitory computer-readable storage medium of claim 14 , wherein the generating a piece of music audio corresponding to each of the plurality of segments of frames based at least in part on the at least one template further comprises:

modifying the at least one template by adding a melody based on one of the plurality of sets of music notes.

17. The non-transitory computer-readable storage medium of claim 14 , the operations further comprising:

generating music audio content for the video based at least in part on synthesizing the plurality of pieces of music audio, wherein the music audio content matches motion, intensity, and transition in the video.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2023
From: OUYANG, ZHIHAO; ZHANG, DAIYU; LI, BOCHEN; LIU, BAOMAN; YANG, LIUQING
To: BYTEDANCE INC.
Reel/Frame 064459/0966 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2023
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 064459/0996 →
Cited By (1)
US 12,713,107