IP Library Granted Patent US 12,394,399
Granted Patent B2
US 12,394,399 · App. 17/671,096 · Granted Aug 19, 2025

Relations between music items

Inventor: Juan José Bosch Vicente (Paris, FR)
Assignee: Spotify AB
G10H1/0025G10H1/0066G10H2250/311
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,399
App. No.
17/671,096
Granted
Aug 19, 2025
Kind
B2
Abstract

A method of determining relations between music items, wherein a music item is a submix of a musical composition comprising one or more music tracks, the method comprising determining a first input representation for at least part of a first music item, mapping the first input representation onto to one or more subspaces derived from a vector space using a first model, wherein each subspace models a characteristic of the music items, determining a second input representation for at least part of a second music item, mapping the second input representation onto the one or more subspaces using a second model, and determining a distance between the mappings of the first and second input representations in each subspace, wherein the distance represents the degree of relation between the first and second input representations with respect to the characteristic modelled by the subspace.

Claims (50)

1. A method of identifying related sub-mixes of musical tracks for song assembly, the method comprising:

determining a first input representation for a first sub-mix of musical tracks stored in a database of sub-mixes;

mapping the first input representation to a particular subspace derived from a vector space using a first model, wherein the particular subspace models a particular characteristic;

determining a second input representation for a second sub-mix of musical tracks stored in the database of sub-mixes;

mapping the second input representation to the particular subspace using a second model;

determining a distance between the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace, wherein the distance represents a degree of relation between the first input representation and the second input representation with respect to the particular characteristic; and

identifying, based on the degree of relation, the first sub-mix of musical tracks and the second sub-mix of musical tracks as candidate sub-mixes for the song assembly.

2. The method of claim 1 , wherein the first model comprises a first encoder and a first set of one or more mapping functions, wherein the second model comprises a second encoder and a second set of one or more mapping functions, wherein:

the first encoder is configured to map the first input representation to the vector space;

the second encoder is configured to map the second input representation to the vector space;

the first set of mapping functions is configured to map the first input representation from the vector space to the particular subspace; and

the second set of mapping functions is configured to map the second input representation from the vector space to the particular subspace.

3. The method of claim 2 , wherein the first encoder comprises a first neural network, and wherein the second encoder comprises a second neural network.

4. The method of claim 1 , further comprising storing a sub-mix use record in the database of sub-mixes, wherein the sub-mix use record indicates:

whether the first sub-mix of musical tracks has been previously retrieved from the database of sub-mixes to create a song and;

whether the second sub-mix of musical tracks has been previously retrieved from the database of sub-mixes to create a song.

5. The method of claim 2 , wherein the first set of mapping functions comprises at least one neural network, and wherein the second set of mapping functions comprises at least one neural network.

6. The method of claim 1 , wherein the first sub-mix of musical tracks and second sub-mix of musical tracks are audio representations of music.

7. The method of claim 1 , wherein the first sub-mix of musical tracks and second sub-mix of musical tracks are symbolic representations of music.

8. The method of claim 1 , wherein the first model and the second model are the same.

9. The method of claim 1 , wherein the first model and the second model are different.

10. The method of claim 1 , wherein the first sub-mix of musical tracks is an audio representation of music, the second sub-mix of musical tracks is a symbolic representation of music, and the first model and the second model are different.

11. The method of claim 1 , wherein the particular characteristic comprises a genre, a rhythm, a mood, or a sound.

12. The method of claim 1 , wherein a musical track represents an instrumental or vocal part of a musical composition.

13. The method of claim 1 , wherein a larger distance between the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace represents a lower degree of relation between the first input representation and the second input representation with respect to the particular characteristic.

14. The method of claim 1 , wherein a smaller distance between the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace represents a higher degree of relation between the first input representation and the second input representation with respect to the particular characteristic.

15. A non-transitory computer-readable medium having instructions stored thereon that, when executed by a computing device, cause the computing device to perform operations comprising:

determining a first input representation for a first sub-mix of musical tracks stored in a database of sub-mixes;

mapping the first input representation to a particular subspace derived from a vector space using a first model, wherein the particular subspace models a particular characteristic;

determining a second input representation for a second sub-mix of musical tracks stored in the database of sub-mixes;

mapping the second input representation to the particular subspace using a second model;

determining a distance between the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace, wherein the distance represents a degree of relation between the first input representation and the second input representation with respect to the particular characteristic; and

identifying, based on the degree of relation, the first sub-mix of musical tracks and the second sub-mix of musical tracks as candidate sub-mixes for a song assembly.

16. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise storing a sub-mix use record in the database of sub-mixes, wherein the sub-mix use record indicates:

whether the first sub-mix of musical tracks has been previously retrieved from the database of sub-mixes to create a song and;

whether the second sub-mix of musical tracks has been previously retrieved from the database of sub-mixes to create a song.

17. The non-transitory computer-readable medium of claim 15 , wherein the first sub-mix of musical tracks and second sub-mix of musical tracks are audio representations of music.

18. A system comprising:

a memory; and

one or more processors coupled to the memory, the one or more processors configured to:

determine a first input representation for a first sub-mix of musical tracks stored in a database of sub-mixes;

map the first input representation to a particular subspace derived from a vector space using a first model, wherein the particular subspace models a particular characteristic;

determine a second input representation for a second sub-mix of musical tracks stored in the database of sub-mixes;

map the second input representation to the particular subspace using a second model;

determine a distance between the mapping of the first input representation in the particular subspace and the mapping of the second input representation in the particular subspace, wherein the distance represents a degree of relation between the first input representation and the second input representation with respect to the particular characteristic; and

identify, based on the degree of relation, the first sub-mix of musical tracks and the second sub-mix of musical tracks as candidate sub-mixes for a song assembly.

19. The system of claim 18 , wherein the one or more processors are further configured to store a sub-mix use record in the database of sub-mixes, wherein the sub-mix use record indicates:

whether the first sub-mix of musical tracks has been previously retrieved from the database of sub-mixes to create a song and;

whether the second sub-mix of musical tracks has been previously retrieved from the database of sub-mixes to create a song.

20. The system of claim 18 , wherein the first sub-mix of musical tracks and second sub-mix of musical tracks are symbolic representations of music.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2022
From: BOSCH VICENTE, JUAN JOSÉ
To: SPOTIFY AB
Reel/Frame 059269/0244 →
Continuity (1)
Related Publication 20230260488A1 · Aug 17, 2023
References Cited (67)
US 7024424B1 · Platt · 2006 [cited by applicant]
US 7812241B2 · Ellis · 2010 [cited by applicant]
US 8073854B2 · Whitman · 2011 [cited by applicant]
US 10303771B1 · Jezewski · 2019 [cited by applicant]
US 10671666B2 · Jin · 2020 [cited by applicant]
US 10997986B2 · Vincente · 2021 [cited by applicant]
US 11238839B2 · Pachet · 2022 [cited by applicant]
US 20040093202A1 · Fischer · 2004 [cited by applicant]
US 20040215447A1 · Sundareson · 2004 [cited by applicant]
US 20050247185A1 · Uhle · 2005 [cited by examiner]
US 20080021851A1 · Alcalde · 2008 [cited by applicant]
US 20080288255A1 · Carin · 2008 [cited by applicant]
US 20090044689A1 · Komori · 2009 [cited by applicant]
US 20090049082A1 · Slaney · 2009 [cited by applicant]
US 20100199833A1 · McNaboe · 2010 [cited by applicant]
US 20110004642A1 · Schnitzer · 2011 [cited by applicant]
US 20120237041A1 · Pohle · 2012 [cited by applicant]
US 20120300950A1 · Usui · 2012 [cited by applicant]
US 20130226957A1 · Ellis · 2013 [cited by applicant]
US 20150242750A1 · Anderson · 2015 [cited by applicant]
US 20170116533A1 · Jehan · 2017 [cited by applicant]
US 20170154216A1 · Kennedy · 2017 [cited by applicant]
US 20170236504A1 · Brooker · 2017 [cited by applicant]
US 20180137845A1 · Prokop · 2018 [cited by applicant]
US 20180341704A1 · Barkan · 2018 [cited by applicant]
US 20190318060A1 · Brenner · 2019 [cited by applicant]
US 20200074982A1 · McCallum · 2020 [cited by applicant]
US 20200320388A1 · Lyske · 2020 [cited by applicant]
US 20210049989A1 · Bretan · 2021 [cited by examiner]
US 20210090536A1 · Pachet · 2021 [cited by applicant]
US 20210090590A1 · Vincente · 2021 [cited by applicant]
US 20210294840A1 · Lee · 2021 [cited by examiner]
US 20210312941A1 · Vincente · 2021 [cited by applicant]
US 20230223037A1 · Vicente · 2023 [cited by applicant]
US 20230260492A1 · Vicente · 2023 [cited by applicant]
EP 751471 · 1996 [cited by applicant]
EP 3796306 · 2021 [cited by applicant]
WO 2015035492 · 2015 [cited by applicant]
WO 2015154159 · 2015 [cited by applicant]
WO 2016189307 · 2016 [cited by applicant]
WO 2017030661 · 2017 [cited by applicant]
WO 2019084419 · 2019 [cited by applicant]
European Communication in Application 20205651.1, mailed Jan. 10, 2024, 5 pages. [cited by applicant]
Aucouturier; J.-J., Pachet, F and Sandler, M. “The Way It Sounds : Timbre Models for Analysis and Retrieval of Polyphonic Music Signals.” IEEE Transactions of Multimedia, 7(6):1028-1035 (Dec. 2005). [cited by applicant]
Ellis, Daniel et al., “Identifying ‘Cover Songs’ with Chroma Features and Dynamic Programming Beat Tracking”, 2007, 4 pages. [cited by applicant]
European Communication in Application 20174092.5, mailed Dec. 18, 2020, 8 pages. [cited by applicant]
European Communication in Application 20174092.5, mailed Aug. 23, 2021, 9 pages. [cited by applicant]
European Communication in Application 20174093.3, mailed Jun. 2, 2021, 13 pages. [cited by applicant]
European Communication in Application 20174093.3, mailed Dec. 3, 2020, 11 pages. [cited by applicant]
European Communication in Application 20205650.3, mailed Feb. 14, 2022, 7 pages. [cited by applicant]
European Extended Search Report in Application 20174092.5, mailed Sep. 1, 2020, 8 pages. [cited by applicant]
European Extended Search Report in Application 20174093.3, mailed Sep. 1, 2020, 12 pages. [cited by applicant]
European Extended Search Report in Application 20205650.3, mailed May 7, 2021, 16 pages. [cited by applicant]
European Extended Search Report in Application 20205651.1, mailed May 3, 2021, 13 pages. [cited by applicant]
European Minutes of the Oral Proceedings in Application 20174093.3, mailed Nov. 3, 2021, 13 pages. [cited by applicant]
European Result of Consulation in Application 20174093.3, mailed Oct. 20, 2021, 5 pages. [cited by applicant]
European Result of Consultation in Application 20174092.5, mailed Jan. 24, 2022, 3 pgs. [cited by applicant]
European Result of Consultation in Application 20174092.5, mailed Dec. 16, 2021, 7 pgs. [cited by applicant]
European Written Submission in Preparation to Oral Proceedings in Application 20174093.3, mailed Oct. 8, 2021, 5 pages. [cited by applicant]
Jehan, T. “Creating music by listening,” PhD, MIT Media Lab (2005). [cited by applicant]
Lee, Jongpil, et al., “Disentangled Multidimensional Metric Learning for Music Similarity”, ARXIV.org, Aug. 9, 2020, 5 pages. [cited by applicant]
Marco A. Martinez Ramirez et al., “Deep Learning and Intelligent Audio Mixing.” Proceedings of the 3rd Workshop on Intelligent Music Production, Salford, UK (Sep. 15, 2017), 4 pages. [cited by applicant]
Marolt, M., “A Mid-Level Representation for Melody-Based Retrieval in Audio Collections”, IEEE Transactions on Multimedia, vol. 10, No. 8, Dec. 1, 2008, 9 pages. [cited by applicant]
Meinard Muller et al., “Multimodal Music Processing,” DFU, vol. 3 (2012). Available at: https://drops.dagstuhl.de/opus/volltexte/dfu-complete/dfu-vo13-complete.pdf, 258 pages. [cited by applicant]
Oderkerken, Daphne, et al., “Decibel: Improving Audio Chord Estimation for Popular Music by Alignment and Integration of Crowd-Sourced Symbolic Representations”, ARXIV.org, Feb. 22, 2020, 81 pages. [cited by applicant]
Van den Oord, Aaron, et al., “Deep content-based music recommendation.” Advances in neural information processing systems (2013), 9 pages. [cited by applicant]
Veit, Andreas, et al., “Conditional Similarity Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21, 2017,9 pages. [cited by applicant]