IP Library › Granted Patent US 12,254,855
Granted Patent B2
US 12,254,855 · App. 17/930,933 · Granted Mar 18, 2025

Method, system, and computer-readable medium for creating song mashups

Inventors: Juan José Bosch Vicente (Paris, FR); Youn Jin Kim (Brooklyn, NY); Peter Milan Thomson Sobot (Brooklyn, NY); Angus William Sackfield (Brooklyn, NY)
Assignee: Spotify AB
G10H1/0008G10H2210/056G10H2210/076G10H2210/081G10H2210/561G10H2240/325
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,855
App. No.
17/930,933
Granted
Mar 18, 2025
Kind
B2
Abstract

A system, method and computer product for combining audio tracks. In one example embodiment herein, the method comprises determining at least one music track that is musically compatible with a base music track, aligning those tracks in time, and combining the tracks. In one example embodiment herein, the tracks may be music tracks of different songs, the base music track can be an instrumental accompaniment track, and the at least one music track can be a vocal track. Also in one example embodiment herein, the determining is based on musical characteristics associated with at least one of the tracks, such as an acoustic feature vector distance between tracks, a likelihood of at least one track including a vocal component, a tempo, or musical key. Also, determining of musical compatibility can include determining at least one of a vertical musical compatibility or a horizontal musical compatibility among tracks.

Claims (41)

1. A system for combining audio tracks, the system comprising:

a memory storing a computer program; and

a computer processor, controllable by the computer program to:

determine at least one music track from a plurality of music tracks that is musically compatible with a base music track by determining respective cosine distances between an acoustic feature vector of the base music track and an acoustic feature vector of each of the plurality of music tracks;

align the at least one music track and the base music track;

separate the at least one music track into an accompaniment component and a vocal component; and

add the vocal component of the at least one music track to the base music track.

2. The system of claim 1 , wherein to determine the at least one music track from the plurality of music tracks includes to determine at least one segment of the at least one music track that is musically compatible with at least one segment of the base music track.

3. The system of claim 1 , wherein the base music track and the at least one music track are music tracks of different songs.

4. The system of claim 1 , wherein the computer processor is further controllable by the computer program to:

transform a musical key of the base music track and a corresponding musical key of the at least one music track, so that keys of the base music track and the corresponding musical key of the at least one music track are compatible.

5. The system of claim 1 , wherein the computer processor is further controllable by the computer program to:

identify music tracks for which a plurality of users have an affinity; and

identify those ones of the identified music tracks for which one of the plurality of users has an affinity, wherein at least one of the identified music tracks for which one of the plurality of users has an affinity is used as the base music track.

6. The system of claim 5 , wherein at least another one of the identified music tracks for which one of the plurality of users has an affinity is used as the at least one music track.

7. The system of claim 1 , wherein to determine the at least one music track from the plurality of music tracks includes to determine at least one of a vertical musical compatibility between segments of the base music track and the at least one music track, or a horizontal musical compatibility among tracks.

8. The system of claim 7 , wherein the vertical musical compatibility is based on at least one of a tempo compatibility, a harmonic compatibility, a loudness compatibility, vocal activity, beat stability, or a segment length.

9. The system of claim 7 , wherein to determine the horizontal musical compatibility includes to determine at least one of a distance between acoustic feature vectors among the plurality of music tracks, and a repetition of a segment of one of the plurality of music tracks being selected as a candidate for being mixed with the base music track.

10. The system of claim 7 , wherein to determine the at least one music track from the plurality of music tracks further includes to determine a compatibility score based on a key distance score associated with at least one of the tracks, an acoustic feature vector distance associated with at least one of the tracks, the vertical musical compatibility, and the horizontal musical compatibility.

11. A method for combining audio tracks, comprising:

determining at least one music track from a plurality of music tracks that is musically compatible with a base music track by determining respective cosine distances between an acoustic feature vector of the base music track and an acoustic feature vector of each of the plurality of music tracks;

aligning the at least one music track and the base music track;

separating the at least one music track into an accompaniment component and a vocal component; and

adding the vocal component of the at least one music track to the base music track.

12. The method of claim 11 , wherein determining the at least one music track from the plurality of music tracks includes determining at least one segment of the at least one music track that is musically compatible with at least one segment of the base music track.

13. The method of claim 11 , wherein the base music track and the at least one music track are music tracks of different songs.

14. The method of claim 11 , further comprising determining whether to keep a vocal component of the base music track, or replace the vocal component of the base music track with the vocal component of the at least one music track before adding the vocal component of the at least one music track to the base music track.

15. The method of claim 11 , further comprising:

identifying music tracks for which a plurality of users have an affinity; and

identifying those ones of the identified music tracks for which one of the plurality of users has an affinity, wherein at least one of the identified music tracks for which one of the plurality of users has an affinity is used as the base music track.

16. The method of claim 15 , wherein at least another one of the identified music tracks for which one of the plurality of users has an affinity is used as the at least one music track.

17. The method of claim 11 , wherein determining the at least one music track from the plurality of music tracks includes calculating a respective musical compatibility score between the base music track and each of the plurality of music tracks.

18. The method of claim 11 , wherein determining the at least one music track from the plurality of music tracks includes determining at least one of: a vertical musical compatibility between segments of the base music track and the at least one music track, and a horizontal musical compatibility among tracks.

19. The method of claim 18 , wherein the vertical musical compatibility is based on at least one of a tempo compatibility, a harmonic compatibility, a loudness compatibility, vocal activity, beat stability, or a segment length.

20. A system for combining audio tracks, the system comprising:

a memory storing a computer program; and

a computer processor, controllable by the computer program to:

determine a first segment of at least one music track from a plurality of music tracks that is musically compatible with a second segment of a base music track by determining respective cosine distances between an acoustic feature vector of the second segment of the base music track and acoustic feature vectors of segments of the plurality of music tracks;

align the first segment of the at least one music track and the second segment of the base music track;

separate the first segment of the at least one music track into an accompaniment component and a vocal component; and

add the vocal component of the first segment of the at least one music track to the second segment of the base music track.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2022
From: BOSCH VICENTE, JUAN JOSÉ; KIM, YOUN JIN; SOBOT, PETER MILAN THOMSON; SACKFIELD, ANGUS WILLIAM
To: SPOTIFY AB
Reel/Frame 061940/0437 →
Continuity (2)
Continuation 16728953 · Dec 27, 2019
Related Publication 20230075074A1 · Mar 9, 2023
References Cited (80)
US 8258390B1 · Gossweiler · 2012 [cited by examiner]
US 8660845B1 · Sodeifi · 2014 [cited by examiner]
US 8855334B1 · Lavine et al. · 2014 [cited by applicant]
US 9257954B2 · Ball · 2016 [cited by applicant]
US 9280313B2 · Ball · 2016 [cited by applicant]
US 9286877B1 · Dabby · 2016 [cited by applicant]
US 9372925B2 · Ball · 2016 [cited by applicant]
US 9412390B1 · Chaudhary · 2016 [cited by applicant]
US 9798974B2 · Ball · 2017 [cited by applicant]
US 9852745B1 · Tootill · 2017 [cited by applicant]
US 10284985B1 · Chaudhary · 2019 [cited by applicant]
US 10446126B1 · Kaye · 2019 [cited by applicant]
US 10614785B1 · Dabby · 2020 [cited by examiner]
US 10803118B2 · Jehan · 2020 [cited by examiner]
US 11475867B2 · Bosch Vicente · 2022 [cited by examiner]
US 20040027369A1 · Kellock · 2004 [cited by examiner]
US 20070083558A1 · Martinez · 2007 [cited by applicant]
US 20070292106A1 · Finkelstein · 2007 [cited by applicant]
US 20080271592A1 · Beckford · 2008 [cited by applicant]
US 20090038467A1 · Brennan · 2009 [cited by examiner]
US 20100305732A1 · Serletic · 2010 [cited by examiner]
US 20110112672A1 · Brown · 2011 [cited by applicant]
US 20130139057A1 · Vlassopulos · 2013 [cited by applicant]
US 20130170670A1 · Casey · 2013 [cited by examiner]
US 20140018947A1 · Ales · 2014 [cited by applicant]
US 20140039891A1 · Sodeifi · 2014 [cited by examiner]
US 20140121797A1 · Ales · 2014 [cited by applicant]
US 20150067512A1 · Roswell · 2015 [cited by applicant]
US 20150302009A1 · Henderson · 2015 [cited by applicant]
US 20160012853A1 · Cabanilla · 2016 [cited by applicant]
US 20160042761A1 · Motta · 2016 [cited by applicant]
US 20160239876A1 · Ales · 2016 [cited by applicant]
US 20160372095A1 · Lyske · 2016 [cited by applicant]
US 20160372096A1 · Lyske · 2016 [cited by applicant]
US 20170214963A1 · Di Franco · 2017 [cited by applicant]
US 20180181730A1 · Lyske · 2018 [cited by examiner]
US 20180374462A1 · Steinwedel · 2018 [cited by applicant]
US 20190043528A1 · Humphrey · 2019 [cited by applicant]
US 20190066641A1 · Nazer · 2019 [cited by examiner]
US 20190066643A1 · Packouz · 2019 [cited by applicant]
US 20190378482A1 · Vorobyev · 2019 [cited by applicant]
US 20200042879A1 · Jansson · 2020 [cited by applicant]
US 20200043517A1 · Jansson · 2020 [cited by applicant]
US 20200043518A1 · Jansson · 2020 [cited by applicant]
US 20200082019A1 · Allen · 2020 [cited by examiner]
US 20200089705A1 · Roswell · 2020 [cited by applicant]
US 20200133620A1 · Boumi · 2020 [cited by applicant]
US 20200135176A1 · Stoller · 2020 [cited by applicant]
US 20200135237A1 · Gauvin · 2020 [cited by applicant]
US 20200410968A1 · Mahdavi · 2020 [cited by applicant]
US 20210090536A1 · Pachet · 2021 [cited by examiner]
US 20210201863A1 · Bosch Vicente · 2021 [cited by examiner]
US 20210279030A1 · Morsy · 2021 [cited by examiner]
US 20230057082A1 · Fabbro · 2023 [cited by examiner]
US 20230075074A1 · Bosch Vicente · 2023 [cited by examiner]
AU 2020432954A1 · 2022 [cited by examiner]
CN 108022604 · 2018 [cited by applicant]
EP 3796306A1 · 2021 [cited by examiner]
EP 3843083A1 · 2021 [cited by examiner]
WO WO2006079813A1 · 2006 [cited by examiner]
WO WO2011103498A2 · 2011 [cited by examiner]
“Nearest Neighbour Algorithm”, found at Wikipedia.org, last edited Mar. 10, 2020. Available at: https://en.wikipedia.org/wiki/Nearest_neighbour_algorithm. [cited by applicant]
C. Macas et al., “MixMash: A Visualisation System for Musical Mashup Creation”, 22nd International Conference Information Visualisation (IV), Fisciano, pp. 471-477 (2018). [cited by applicant]
Davies et al. “AutoMashUpper: Automatic Creation of Multi-Song Music Mashups.” IEEE/ACM Transactions on Audio, Speech, and Language Processing vol. 22, No. 12, (2014), pp. 1726-1737. [cited by applicant]
De Roure et al. “Music SOFA: An architecture for semantically informed recomposition of Digital Music Objects.” SAAM, 2018, pp. 1-9. [cited by applicant]
Durand et al., “Downbeat Tracking with Multiple Features and Deep Neural Networks”, 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, QLD, pp. 409-413 (2015). [cited by applicant]
Durand et al., “Robust Downbeat Tracking Using an Ensemble of Convolutional Networks”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, Issue 1 (Jan. 2017). [cited by applicant]
Erik Bernhardsson, “Annoy”, available at: github.com/spotify/annoy (2017). <last accessed Sep. 25, 2020>. [cited by applicant]
Erik Bernhardsson, “Nearest Neighbors and vector models—part 2—algorithms and data structures” (2015), available at: https://erikbern.com/2015/10/01/nearest-neighbors-and-vector-models-part-2-how-to-search-in-high-dimen… [cited by applicant]
European Search Report for EP Application No. 20213406.0 mailed May 31, 2021 (19 pages). [cited by applicant]
Jehan et al., “Analyzer Documentation”, The Echo Nest analyzer developed by the Echo Nest of Somerville, MA (2011). [cited by applicant]
Jehan, Tristan. Creating music by listening. Diss. Massachusetts Institute of Technology, School of Architecture and Planning, Program in Media Arts and Sciences (2005). [cited by applicant]
M. Davies et al., “Improvasher: a real-time mashup system for live musical input”, NIME (2014). [cited by applicant]
O. Nieto et al., “Systematic Exploration of Computational Music Structure Research, Music and Performance Arts Professions, Urban Initiative”, Proceedings of the 17th International Society for Music Information Retrieva… [cited by applicant]
U.S. Appl. No. 16/521,756, filed Jul. 25, 2019, entitled “Automatic Isolation of Multiple Instruments From Musical Mixtures”, by A. Jansson et al. [cited by applicant]
U.S. Appl. No. 15/974,767, filed May 8, 2018, entitled “Extracting Signals From Paired Recordings”, by E. Humphrey et al. (hereinafter the “Humphrey application”). [cited by applicant]
U.S. Appl. No. 16/055,870, filed Aug. 6, 2018, entitled “Singing Voice Separation With Deep U-Net Convolutional Networks”, by A. Jansson et al. [cited by applicant]
U.S. Appl. No. 16/165,498, filed Oct. 19, 2018, 2018, entitled “Singing Voice Separation With Deep U-Net Convolutional Networks”, by A. Jansson. [cited by applicant]
U.S. Appl. No. 16/242,525, filed Jan. 8, 2019, entitled “Singing Voice Separation With Deep U-Net Convolutional Networks”, by A. Jansson et al. [cited by applicant]
Van den Oord, Aaron, Sander Dieleman, and Benjamin Schrauwen. “Deep content-based music recommendation.” Advances in neural information processing systems (2013). [cited by applicant]