IP Library › Granted Patent US 12,394,400
Granted Patent B2
US 12,394,400 · App. 17/671,099 · Granted Aug 19, 2025

Relations between music items

Inventor: Juan José Bosch Vicente (Paris, FR)
Assignee: Spotify AB
G10H7/00G06F16/634G06F16/65G06N3/045G10G1/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,400
App. No.
17/671,099
Granted
Aug 19, 2025
Kind
B2
Abstract

A method of determining relations between music items, the method comprising determining a first input representation for a symbolic representation of a first music item, mapping the first input representation onto to one or more subspaces derived from a vector space using a first model, wherein each subspace models a characteristic of the music items, determining a second input representation for music data representing a second music item, mapping the second input representation onto the one or more subspaces using a second model, determining a distance between the mappings of the first and second input representation in each subspace, wherein the distance represents the degree of relation between the first and second input representation with respect to the characteristic modelled by the subspace.

Claims (50)

1. A method of identifying related tracks for song assembly, the method comprising:

determining a first input representation for a symbolic representation of a first musical track stored in a database of musical tracks;

mapping the first input representation to a particular subspace derived from a vector space using a first model, wherein the particular subspace models a particular characteristic;

determining a second input representation for music data representing a second musical track stored in the database of musical tracks;

mapping the second input representation to the particular subspace using a second model;

determining a distance between the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace, wherein the distance represents a degree of relation between the first input representation and the second input representation with respect to the particular characteristic; and

identifying, based on the degree of relation, the first musical track and the second musical track as candidate musical tracks for song assembly.

2. The method of claim 1 , wherein the first model comprises a first encoder and a first set of one or more mapping functions, wherein the second model comprises a second encoder and a second set of one or more mapping functions, wherein:

the first encoder is configured to map the first input representation to the vector space;

the second encoder is configured to map the second input representation to the vector space;

the first set of mapping functions is configured to map the first input representation from the vector space to the particular subspace; and

the second set of mapping functions is configured to map the second input representation from the vector space to the particular subspace.

3. The method of claim 2 , wherein the first encoder comprises a first neural network, and wherein the second encoder comprises a second neural network.

4. The method of claim 1 , further comprising storing a musical track use record in the database of musical tracks, wherein the musical track use record indicates:

whether the first musical track has been previously retrieved from the database of musical tracks to create a song; and

whether the second musical track has been previously retrieved from the database of musical tracks to create a song.

5. The method of claim 2 , wherein the first set of mapping functions comprises at least one neural network, and wherein the second set of mapping functions comprises at least one neural network.

6. The method of claim 1 , wherein the music data is an audio representation of the second musical track.

7. The method of claim 1 , wherein the music data is a symbolic representation of the second musical track.

8. The method of claim 1 , wherein the first model and the second model are different.

9. The method of claim 1 , wherein the first model and the second model are the same.

10. The method of claim 1 , wherein the particular characteristic comprise a genre, a rhythm, a mood, or a sound.

11. The method of claim 1 , wherein the symbolic representation of the first musical track is a MIDI file, a MusicXML file, a list of events, or a piano-roll representation.

12. The method of claim 1 , wherein a musical track is at least part of a musical composition.

13. The method of claim 1 , wherein a musical track is part of a sub-mix comprising a number of musical tracks.

14. The method of claim 1 , wherein a musical track represents an instrumental or vocal part of a musical composition.

15. The method of claim 12 , wherein a larger distance between t the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace represents a lower degree of relation between the first input representation and the second input representation with respect to the particular characteristic.

16. The method of claim 1 , wherein a smaller distance between the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace represents a higher degree of relation between the first input representation and the second input representation with respect to the particular characteristic.

17. A non-transitory computer-readable medium having instructions stored thereon that, when executed by a computing device, cause the computing device to perform operations comprising:

determining a first input representation for a symbolic representation of a first musical track stored in a database of musical tracks;

mapping the first input representation to a particular subspace derived from a vector space using a first model, wherein the particular subspace models a particular characteristic;

determining a second input representation for music data representing a second musical track stored in the database of musical tracks;

mapping the second input representation to the particular subspace using a second model;

determining a distance between the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace, wherein the distance represents a degree of relation between the first input representation and the second input representation with respect to the particular characteristic; and

identifying, based on the degree of relation, the first musical track and the second musical track as candidate musical tracks for song assembly.

18. The non-transitory computer-readable medium of claim 17 , wherein the operations further comprise storing a musical track use record in the database of musical tracks, wherein the musical track use record indicates:

whether the first musical track has been previously retrieved from the database of musical tracks to create a song; and

whether the second musical track has been previously retrieved from the database of musical tracks to create a song.

19. A system comprising:

a memory; and

one or more processors coupled to the memory, the one or more processor configured to:

determine a first input representation for a symbolic representation of a first musical track stored in a database of musical tracks;

map the first input representation to a particular subspace derived from a vector space using a first model, wherein the particular subspace models a particular characteristic;

determine a second input representation for music data representing a second musical track stored in the database of musical tracks;

map the second input representation to the particular subspace using a second model;

determine a distance between the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace, wherein the distance represents a degree of relation between the first input representation and the second input representation with respect to the particular characteristic; and

identify, based on the degree of relation, the first musical track and the second musical track as candidate musical tracks for song assembly.

20. The system of claim 19 , wherein the one or more processors are further configured to store a musical track use record in the database of musical tracks, wherein the musical track use record indicates:

whether the first musical track has been previously retrieved from the database of musical tracks to create a song; and

whether the second musical track has been previously retrieved from the database of musical tracks to create a song.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2022
From: BOSCH VICENTE, JUAN JOSÉ
To: SPOTIFY AB
Reel/Frame 059269/0229 →
Continuity (1)
Related Publication 20230260492A1 · Aug 17, 2023
References Cited (67)
US 7024424B1 · Platt · 2006 [cited by applicant]
US 7812241B2 · Ellis · 2010 [cited by applicant]
US 8073854B2 · Whitman · 2011 [cited by applicant]
US 10303771B1 · Jezewski · 2019 [cited by applicant]
US 10671666B2 · Jin · 2020 [cited by applicant]
US 10997986B2 · Vincente · 2021 [cited by applicant]
US 11238839B2 · Pachet · 2022 [cited by applicant]
US 20040093202A1 · Fischer · 2004 [cited by applicant]
US 20040215447A1 · Sundareson · 2004 [cited by applicant]
US 20050247185A1 · Uhle · 2005 [cited by examiner]
US 20080021851A1 · Alcalde · 2008 [cited by applicant]
US 20080288255A1 · Carin · 2008 [cited by applicant]
US 20090044689A1 · Komori · 2009 [cited by applicant]
US 20090049082A1 · Slaney · 2009 [cited by applicant]
US 20100199833A1 · McNaboe · 2010 [cited by applicant]
US 20110004642A1 · Schnitzer · 2011 [cited by applicant]
US 20120237041A1 · Pohle · 2012 [cited by applicant]
US 20120300950A1 · Usui · 2012 [cited by applicant]
US 20130226957A1 · Ellis · 2013 [cited by applicant]
US 20150242750A1 · Anderson · 2015 [cited by applicant]
US 20170116533A1 · Jehan · 2017 [cited by applicant]
US 20170154216A1 · Kennedy · 2017 [cited by applicant]
US 20170236504A1 · Brooker · 2017 [cited by applicant]
US 20180137845A1 · Prokop · 2018 [cited by applicant]
US 20180341704A1 · Barkan · 2018 [cited by applicant]
US 20190318060A1 · Brenner · 2019 [cited by applicant]
US 20200074982A1 · McCallum · 2020 [cited by applicant]
US 20200320388A1 · Lyske · 2020 [cited by applicant]
US 20210049989A1 · Bretan · 2021 [cited by examiner]
US 20210090536A1 · Pachet · 2021 [cited by applicant]
US 20210090590A1 · Vincente · 2021 [cited by applicant]
US 20210294840A1 · Lee · 2021 [cited by examiner]
US 20210312941A1 · Vincente · 2021 [cited by applicant]
US 20230223037A1 · Vicente · 2023 [cited by applicant]
US 20230260488A1 · Vicente · 2023 [cited by applicant]
EP 751471 · 1996 [cited by applicant]
EP 3796306 · 2021 [cited by applicant]
WO 2015035492 · 2015 [cited by applicant]
WO 2015154159 · 2015 [cited by applicant]
WO 2016189307 · 2016 [cited by applicant]
WO 2017030661 · 2017 [cited by applicant]
WO 2019084419 · 2019 [cited by applicant]
Aucouturier; J.-J., Pachet, F and Sandler, M. “The Way It Sounds : Timbre Models for Analysis and Retrieval of Polyphonic Music Signals.” IEEE Transactions of Multimedia, 7(6):1028-1035 (Dec. 2005). [cited by applicant]
Ellis, Daniel et al., “Identifying ‘Cover Songs’ with Chroma Features and Dynamic Programming Beat Tracking”, 2007, 4 pages. [cited by applicant]
European Communication in Application 20174092.5, mailed Dec. 18, 2020, 8 pages. [cited by applicant]
European Communication in Application 20174092.5, mailed Aug. 23, 2021, 9 pages. [cited by applicant]
European Communication in Application 20174093.3, mailed Jun. 2, 2021, 13 pages. [cited by applicant]
European Communication in Application 20174093.3, mailed Dec. 3, 2020, 11 pages. [cited by applicant]
European Communication in Application 20205650.3, mailed Feb. 14, 2022, 7 pages. [cited by applicant]
European Extended Search Report in Application 20174092.5, mailed Sep. 1, 2020, 8 pages. [cited by applicant]
European Extended Search Report in Application 20174093.3, mailed Sep. 1, 2020, 12 pages. [cited by applicant]
European Extended Search Report in Application 20205650.3, mailed May 7, 2021, 16 pages. [cited by applicant]
European Extended Search Report in Application 20205651.1, mailed May 3, 2021, 13 pages. [cited by applicant]
European Minutes of the Oral Proceedings in Application 20174093.3, mailed Nov. 3, 2021, 13 pages. [cited by applicant]
European Result of Consulation in Application 20174093.3, mailed Oct. 20, 2021, 5 pages. [cited by applicant]
European Result of Consultation in Application 20174092.5, mailed Jan. 24, 2022, 3 pgs. [cited by applicant]
European Result of Consultation in Application 20174092.5, mailed Dec. 16, 2021, 7 pgs. [cited by applicant]
European Written Submission in Preparation to Oral Proceedings in Application 20174093.3, mailed Oct. 8, 2021, 5 pages. [cited by applicant]
Jehan, T. “Creating music by listening,” PhD, MIT Media Lab (2005), 7 pages. [cited by applicant]
Lee, Jongpil, et al., “Disentangled Multidimensional Metric Learning for Music Similarity”, ARXIV.org, Aug. 9, 2020, 5 pages. [cited by applicant]
Marco A. Martinez Ramirez, et al., “Deep Learning and Intelligent Audio Mixing.” Proceedings of the 3rd Workshop on Intelligent Music Production, Salford, UK (Sep. 15, 2017), 4 pages. [cited by applicant]
Marolt, M., “A Mid-Level Representation for Melody-Based Retrieval in Audio Collections”, IEEE Transactions on Multimedia, vol. 10, No. 8, Dec. 1, 2008, 9 pages. [cited by applicant]
Meinard Muller et al., “Multimodal Music Processing,” DFU, vol. 3 (2012). Available at: https://drops.dagstuhl.de/opus/volltexte/dfu-complete/dfu-vo13-complete.pdf, 258 pages. [cited by applicant]
Oderkerken, Daphne, et al., “Decibel: Improving Audio Chord Estimation for Popular Music by Alignment and Integration of Crowd-Sourced Symbolic Representations”, ARXIV.org, Feb. 22, 2020, 81 pages. [cited by applicant]
Van den Oord, Aaron, et al., “Deep content-based music recommendation.” Advances in neural information processing systems (2013), 9 pages. [cited by applicant]
Veit, Andreas, et al., “Conditional Similarity Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21, 2017,9 pages. [cited by applicant]
European Communication in Application 20205651.1, mailed Jan. 10, 2024, 5 pages. [cited by applicant]