IP Library › Granted Patent US 12,567,394
Granted Patent B2
US 12,567,394 · App. 17/737,216 · Granted Mar 3, 2026

Converting audio samples to full song arrangements

Inventors: Bochen Li (Los Angeles, CA); Andrew Shaw (Los Angeles, CA); Jitong Chen (Los Angeles, CA)
Assignee: LEMON INC.
G10H1/0025G10H1/38G10H2250/311
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,394
App. No.
17/737,216
Granted
Mar 3, 2026
Kind
B2
Abstract

In examples, a method for converting audio samples to full song arrangements is provided. The method includes receiving audio sample data, determining a melodic transcription, based on the audio sample data, and determining a sequence of music chords, based on the melodic transcription. The method further includes generating a full song arrangement, based on the sequence of music chords, and the audio sample data.

Claims (66)

1 . A method for converting audio samples to full song arrangements, the method comprising:

receiving audio sample data;

determining a melodic transcription with a plurality of bars, based on the audio sample data;

determining a sequence of music chords, based on the melodic transcription, wherein the determining of the sequence of music chords comprises:

inputting the melodic transcription to a machine learning model, the machine learning model being trained based on a dataset of paired melody bars and chords;

receiving from the machine learning model, a plurality of chord candidates for each bar of the plurality of bars of the melodic transcription;

determining, from pre-defined chord progressions, one or more chord progressions corresponding to the plurality of chord candidates for each bar of the plurality of bars of the melodic transcription; and

selecting the sequence of music chords from the determined one or more chord progressions;

performing vocal processing on the audio sample data by dynamically time warping the audio sample data to fit the determined sequence of music chords; and

generating a full song arrangement, based on the sequence of music chords and the vocally processed audio sample data.

2 . The method of claim 1 , wherein the pre-defined chord progressions are 4-bar chord progressions.

3 . The method of claim 1 , wherein the trained machine learning model is a neural network.

4 . The method of claim 3 , wherein the chords in the data set include maj, min, 7, min7, min7b5, aug, and sus4.

5 . The method of claim 1 , further comprising:

displaying a user-interface;

receiving, via the user-interface, a user-input corresponding to a selection of an accompaniment style of the full song arrangement; and

re-generating the full song arrangement, based on the user-input.

6 . The method of claim 1 , wherein the audio sample data includes a subset of data corresponding to auditory words.

7 . The method of claim 1 , wherein the vocal processing further comprises:

removing a subset of the audio sample data corresponding to ambient noise.

8 . The method of claim 7 , wherein the generating of the full song arrangement is based on the sequence of music chords, and the vocally processed audio sample data.

9 . The method of claim 7 , wherein the vocal processing further comprises:

performing autotuning on the audio sample data; and

normalizing a volume of the audio sample data.

10 . The method of claim 9 , wherein the vocal processing further comprises:

beautifying the audio sample data, by applying one or more vocal effects from the group of: compressor adjustment, reverb adjustment, and chorus adjustment.

11 . The method of claim 1 , further comprising:

receiving the audio sample data from an application on a mobile computing device; and

transmitting the full song arrangement to the mobile computing device.

12 . A system for converting audio samples to full song arrangements, the system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, causes the system to perform a set of operations, the set of operations including:

receiving audio sample data;

determining a melodic transcription with a plurality of bars, based on the audio sample data;

determining a sequence of music chords, based on the melodic transcription, wherein the determining of the sequence of music chords comprises:

inputting the melodic transcription to a machine learning model, the machine learning model being trained based on a dataset of paired melody bars and chords;

receiving from the machine learning model, a plurality of chord candidates for each bar of the plurality of bars of the melodic transcription;

determining, from pre-defined chord progressions, one or more chord progressions corresponding to the plurality of chord candidates for each bar of the plurality of bars of the melodic transcription; and

selecting the sequence of music chords from the determined one or more chord progressions;

performing vocal processing on the audio sample data by dynamically time warping the audio sample data to fit the determined sequence of music chords; and

generating a full song arrangement, based on the sequence of music chords and the vocally processed audio sample data.

13 . The system of claim 12 , wherein the pre-defined chord progressions are 4-bar chord progressions.

14 . The method of claim 12 , wherein the trained machine learning model is a neural network.

15 . The method of claim 12 , wherein the vocal processing further comprises:

removing a subset of the audio sample data corresponding to ambient noise; and

performing autotuning on the audio sample data.

16 . The method of claim 15 , wherein the generating of the full song arrangement is based on the sequence of music chords, and the vocally processed audio sample data.

17 . The method of claim 15 , wherein the vocal processing further comprises:

normalizing a volume of the audio sample data; and

beautifying the audio sample data, by applying one or more vocal effects from the group of: compressor adjustment, reverb adjustment, and chorus adjustment.

18 . One or more computer readable non-transitory storage media embodying software that is operable when executed, by at least one processor of a device, to:

receive audio sample data;

determine a melodic transcription with a plurality of bars, based on the audio sample data;

determine a sequence of music chords, based on the melodic transcription, wherein to determine the sequence of music chords comprises:

inputting the melodic transcription to a machine learning model, the machine learning model being trained based on a dataset of paired melody bars and chords;

receiving from the machine learning model, a plurality of chord candidates for each bar of the plurality of bars of the melodic transcription;

determining, from pre-defined chord progressions, one or more chord progressions corresponding to the plurality of chord candidates for each bar of the plurality of bars of the melodic transcription; and

selecting the sequence of music chords from the determined one or more chord progressions;

perform vocal processing on the audio sample data by dynamically time warping the audio sample data to fit the determined sequence of music chords; and

generate a full song arrangement, based on the sequence of music chords and the vocally processed audio sample data.

19 . The method of claim 1 , further comprising:

estimating a beats per minute of the audio sample data,

wherein the dynamic time warping is performed based on the estimates beats per minute.

20 . The method of claim 1 , wherein the selecting the sequence of music chords from the determined one or more chord progressions comprises:

ranking the determined one or more chord progressions based on a probability of how well each chord progression of the one or more chord progressions matches the plurality of bars; and

selecting the sequence of music chords based on their ranking.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2023
From: LI, BOCHEN; SHAW, ANDREW; CHEN, JITONG
To: BYTEDANCE INC.
Reel/Frame 064063/0741 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2023
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 064063/0773 →
Continuity (1)
Related Publication 20230360620A1 · Nov 9, 2023
References Cited (13)
US 5719346A · Yoshida · 1998 [cited by examiner]
US 6441289B1 · Shinsky · 2002 [cited by examiner]
US 10818308B1 · Chu · 2020 [cited by examiner]
US 20100192755A1 · Morris · 2010 [cited by examiner]
US 20140053711A1 · Serletic, II · 2014 [cited by examiner]
US 20140074459A1 · Chordia · 2014 [cited by examiner]
US 20170213534A1 · Braasch · 2017 [cited by examiner]
US 20210208842A1 · Cassidy · 2021 [cited by examiner]
WO WO2019182390A2 · 2019 [cited by examiner]
Carsault T. et al., Combining Real-Time Extraction and Prediction of Musical Chord Progressions for Creative Applications, Oct. 28, 2021, vol. 10, No. 2634, pp. 1-33 [Retrieved on Sep. 27, 2023] <DOI: /10.3390/ELECTRONI… [cited by applicant]
Hum2Song Multi-track Polyphonic Music Generation from Voice Melody Transcription with Neural Networks. Dec. 17, 2018 [Retrieved on Sep. 27, 2023 from https://medium.com/@carlostoxtli/hum2songmulti-track-polyphonic-music… [cited by applicant]
Automatic Chord Recognition with Fully Convolutional Neural Networks. Sep. 25, 2020 [Retrieved on Sep. 27, 2023 from https://www.static.tu.berlin/fileadmin/www/10002020/Dokumente/Abschlussarbeiten/Masterarbeit_Hamed_Far… [cited by applicant]
International Search Report mailed Sep. 28, 2023 in International Application No. PCT/SG2023/050307. [cited by applicant]