IP Library Granted Patent US 10,891,928
Granted Patent B2
US 10,891,928 · App. 16/500,995 · Granted Jan 12, 2021

Automatic song generation

Inventors: Jian Luan (Redmond, WA); Qinying Liao (Redmond, WA); Zhen Liu (Redmond, WA); Nan Yang (Redmond, WA); Furu Wei (Redmond, WA)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G10H1/0025G06N20/00G10H1/368G10H2210/111G10H2210/151G10H2220/011G10H2250/455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,891,928
App. No.
16/500,995
Granted
Jan 12, 2021
Kind
B2
Abstract

In accordance with implementations of the subject matter described herein, there is provided a solution for supporting a machine to automatically generate a song. In this solution, an input from a user is used to determine a creation intention of the user with respect to a song to be generated. Lyrics of the song are generated based on the creation intention. Then, a template for the song is generated based at least in part on the lyrics. The template indicates a melody matching with the lyrics. In this way, it is feasible to automatically create the melody and lyrics which not only conform to the creation intention of the user but also match with each other.

Claims (53)

1. A computer-implemented method, comprising:

in response to reception of an input from a user, determining, based on the input, a creation intention of the user with respect to a song to be generated;

generating lyrics of the song that are unique and different than the input and based on the creation intention; and

generating a template for the song that is unique and different than the input and based at least in part on the lyrics, the template indicating a melody matching with the lyrics.

2. The method of claim 1 , further comprising:

combining the lyrics and the melody indicated by the template to generate the song.

3. The method of claim 1 , wherein generating the template comprises:

dividing the lyrics into a plurality of lyrics segments;

for each of the plurality of lyrics segments, selecting, from a plurality of candidate melody segments, at least one candidate melody segment matching with the lyrics segment;

determining respective candidate melody segments corresponding to the plurality of lyrics segments based on smoothness among the candidate melody segments selected for adjacent lyrics segments in the plurality of lyrics segments; and

concatenating the determined candidate melody segments into the melody indicated by the template.

4. The method of claim 1 , wherein generating the template further comprises:

representing at least a theme or a key element in the generation of the template as indicated by the creation intention determined and based on the input of the user.

5. The method of claim 1 , wherein generating the lyrics based on the creation intention comprises:

generating candidate lyrics based on the creation intention; and

modifying the candidate lyrics based on a further input received from the user to obtain the lyrics.

6. The method of claim 1 , wherein generating the lyrics based on the creation intention comprises:

obtaining a predefined lyrics generation model, the predefined lyrics generation model being obtained with at least one of the following: existing lyrics and documents including words; and

generating the lyrics based on the creation intention using the lyrics generation model.

7. The method of claim 1 , further comprising:

obtaining a voice model representing a voice characteristic of a singer;

generating a voice spectrum trajectory for the lyrics using the voice model;

synthesizing the voice spectrum trajectory and the melody indicated by the template into a singing waveform of the song; and

playing the song based on the singing waveform.

8. The method of claim 7 , wherein obtaining the voice model comprises:

receiving a voice segment of the singer; and

training the voice model by using the received voice segment of the singer to adjust a predefined average voice model with the received voice segment, the average voice model being obtained with voice segments of a plurality of different singers.

9. The method of claim 1 , wherein the input includes a video.

10. The method of claim 1 , wherein determining the creation intention includes performing image recognition, human face recognition, or emotion detection to the input wherein the input is an image.

11. The method of claim 10 , wherein the creation intention is further based on applying at least one of a posture recognition, a gender recognition, or an age recognition to the image.

12. The method of claim 10 , wherein the creation intention is further based on determining a size, a format, or a type of the image.

13. The method of claim 1 , wherein the method further includes, upon generating the template, causing the template and the melody to match the lyrics by at least determining a distribution of the lyrics based on at least one of a duration of phonemes for words in the lyrics, a pitch trajectory for the lyrics, or a sound intensity trajectory for the lyrics.

14. The method of claim 1 , wherein the input comprises words and wherein determining the creation intention includes performing natural language processing on the input to determine an emotion associated with the words.

15. The method of claim 14 , wherein the method further includes selecting the melody for the lyrics based on the emotion.

16. A device, comprising:

a processing unit; and

a memory coupled to the processing unit and including instructions stored thereon which, when executed by the processing unit, cause the device to perform acts including:

in response to reception of an input from a user, determining, based on the input, a creation intention of the user with respect to a song to be generated;

generating lyrics of the song that are unique and different than the input and based on the creation intention; and

generating a template for the song that is unique and different than the input and based at least in part on the lyrics, the template indicating a melody matching with the lyrics.

17. The device of claim 16 , wherein the acts further include:

combining the lyrics and the melody indicated by the template to generate the song.

18. The device of claim 16 , wherein generating the template comprises:

dividing the lyrics into a plurality of lyrics segments;

for each of the plurality of lyrics segments, selecting, from a plurality of candidate melody segments, at least one candidate melody segment matching with the lyrics segment;

determining respective candidate melody segments corresponding to the plurality of lyrics segments based on smoothness among the candidate melody segments selected for adjacent lyrics segments in the plurality of lyrics segments; and

concatenating the determined candidate melody segments into the melody indicated by the template.

19. The device of claim 16 , wherein generating the lyrics based on the creation intention comprises:

generating candidate lyrics based on the creation intention; and

modifying the candidate lyrics based on a further input received from the user to obtain the lyrics.

20. The device of claim 16 , wherein generating the lyrics based on the creation intention comprises:

obtaining a predefined lyrics generation model, the predefined lyrics generation model being obtained with at least one of the following: existing lyrics and documents including words; and

generating the lyrics based on the creation intention using the lyrics generation model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2019
From: LUAN, JIAN; LIAO, QUINYING; LIU, ZHEN; YANG, NAN; WEI, FURU
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 050627/0959 →
Priority Claims (1)
CN 2017 1 0284177 · Apr 26, 2017 · national
Continuity (1)
Related Publication 20200035209A1 · Jan 30, 2020
Cited By (1)
US 12,731,600