IP Library Granted Patent US 12,283,287
Granted Patent B2
US 12,283,287 · App. 18/090,228 · Granted Apr 22, 2025

Audio stem identification systems and methods

Inventors: Juan José Bosch Vicente (Paris, FR); François Pachet (Paris, FR); Pierre Roy (Paris, FR); Mathieu Ramona (Paris, FR); Tristan Jehan (Brooklyn, NY)
Assignee: Spotify AB
G10L25/51G06F16/634G06N3/04G06N3/08G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,283,287
App. No.
18/090,228
Granted
Apr 22, 2025
Kind
B2
Abstract

Methods, systems and computer program products are provided for determining acoustic feature vectors of query and target items in a first vector space, and mapping the acoustic feature vectors to a second vector space having a lower dimension. The distribution of vectors in the second vector space can then be used to identify items from the same songs, and/or items that are complementary. A mapping function is trained using a machine learning algorithm, such that complementary audio items are closer in the second vector space than the first, according to a given distance metric.

Claims (48)

1. An audio content item identifier, comprising:

at least one processor; and

at least one memory storing instructions which when executed by the at least one processor cause the at least one processor to:

receive, by a client device, a query corresponding to a query audio content item;

determine a query vector corresponding to the query audio content item;

compare, in a vector space, the query vector and a plurality of target vectors corresponding to a plurality of target audio content items, to determine likelihood values indicating, for each respective target audio content item of the plurality of target audio content items, a probability that the respective target audio content item is a match for the query audio content item, wherein the query audio content item and the target audio content items are audio stems that are configured to be inserted into an audio content item during an audio content item creation process; and

output, to an editor tool, information identifying one of the plurality of target audio content items having a highest of the likelihood values.

2. The audio content item identifier of claim 1 ,

wherein the instructions, when executed by the at least one processor further cause the at least one processor to:

generate, by the editor tool, another audio content item based on the query audio content item and the information; and

play back, by the client device, the another audio content item.

3. The audio content item identifier of claim 2 , wherein the editor tool assembles together the audio stems of the query audio content item and the one of the plurality of target audio content items having the highest of the likelihood values to generate the another audio content item.

4. The audio content item identifier of claim 1 ,

wherein the query includes a first audio type for the query audio content item;

wherein the plurality of target audio content items have at least one second audio type; and

wherein the first audio type is different from the at least one second audio type.

5. The audio content item identifier of claim 4 , wherein the query includes the at least one second audio type for the plurality of target audio content items.

6. The audio content item identifier of claim 4 ,

wherein the first audio type is one of vocals or instrumentals; and

wherein the at least one second audio type is the other of vocals or instrumentals.

7. The audio content item identifier of claim 1 , wherein the query vector and the plurality of target vectors are acoustic feature vectors.

8. The audio content item identifier of claim 7 , wherein acoustic features of the acoustic feature vectors include one or more of: vibration, distortion, presence of a vocoder, energy, valance, signal amplitude, and time-frequency progression.

9. The audio content item identifier of claim 1 , wherein the information identifies two or more of the plurality of target audio content items having the highest of the likelihood values.

10. A method of identifying an audio content item, comprising:

receiving, by a client device, a query corresponding to a query audio content item;

determining a query vector corresponding to the query audio content item;

comparing, in a vector space, the query vector and a plurality of target vectors corresponding to a plurality of target audio content items, to determine likelihood values indicating, for each respective target audio content item of the plurality of target audio content items, a probability that the respective target audio content item is a match for the query audio content item, wherein the query audio content item and the target audio content items are audio stems that are configured to be inserted into an audio content item during an audio content item creation process; and

outputting, to an editor tool, information identifying one of the plurality of target audio content items having a highest of the likelihood values.

11. The method of claim 10 , further comprising:

generating, by the editor tool, another audio content item based on the query audio content item and the information; and

playing back, by the client device, the another audio content item.

12. The method of claim 11 , wherein the editor tool assembles together the audio stems of the query audio content item and the one of the plurality of target audio content items having the highest of the likelihood values to generate the another audio content item.

13. The method of claim 10 ,

wherein the query includes a first audio type for the query audio content item;

wherein the plurality of target audio content items have at least one second audio type;

wherein the first audio type is different from the at least one second audio type; and

wherein the query includes the at least one second audio type for the plurality of target audio content items.

14. The method of claim 13 ,

wherein the first audio type is one of vocals or instrumentals; and

wherein the at least one second audio type is the other of vocals or instrumentals.

15. The method of claim 10 , wherein the query vector and the plurality of target vectors are acoustic feature vectors.

16. The method of claim 15 , wherein acoustic features of the acoustic feature vectors include one or more of: vibration, distortion, presence of a vocoder, energy, valance, signal amplitude, and time-frequency progression.

17. The method of claim 10 , wherein the information identifies two or more of the plurality of target audio content items having the highest of the likelihood values.

18. A non-transitory computer-readable medium having stored thereon one or more sequences of instructions for causing one or more processors to perform:

receiving, by a client device, a query corresponding to a query audio content item;

determining a query vector corresponding to the query audio content item;

comparing, in a vector space, the query vector and a plurality of target vectors corresponding to a plurality of target audio content items, to determine likelihood values indicating, for each respective target audio content item of the plurality of target audio content items, a probability that the respective target audio content item is a match for the query audio content item, wherein the query audio content item and the target audio content items are audio stems that are configured to be inserted into an audio content item during an audio content item creation process; and

outputting, to an editor tool, information identifying one of the plurality of target audio content items having a highest of the likelihood values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2024
From: BOSCH VICENTE, JUAN JOSÉ; PACHET, FRANÇOIS; ROY, PIERRE; RAMONA, MATHIEU; JEHAN, TRISTAN
To: SPOTIFY AB
Reel/Frame 066429/0342 →
Continuity (3)
Continuation 17202841 · Mar 16, 2021
Continuation 16575926 · Sep 19, 2019
Related Publication 20230223037A1 · Jul 13, 2023
References Cited (69)
US 7024424B1 · Platt · 2006 [cited by examiner]
US 7812241B2 · Ellis · 2010 [cited by examiner]
US 8073854B2 · Whitman · 2011 [cited by examiner]
US 10236006B1 · Gurijala · 2019 [cited by examiner]
US 10303771B1 · Jezewski · 2019 [cited by examiner]
US 10671666B2 · Jin · 2020 [cited by examiner]
US 10997986B2 · Vincente · 2021 [cited by applicant]
US 11568886B2 · Vincente · 2023 [cited by applicant]
US 20040093202A1 · Fischer · 2004 [cited by examiner]
US 20040215447A1 · Sundareson · 2004 [cited by examiner]
US 20050247185A1 · Uhle · 2005 [cited by applicant]
US 20080021851A1 · Alcalde · 2008 [cited by examiner]
US 20080288255A1 · Carin · 2008 [cited by examiner]
US 20090044689A1 · Komori · 2009 [cited by examiner]
US 20090049082A1 · Slaney · 2009 [cited by examiner]
US 20100199833A1 · McNaboe · 2010 [cited by examiner]
US 20110004642A1 · Schnitzer · 2011 [cited by examiner]
US 20120237041A1 · Pohle · 2012 [cited by examiner]
US 20120300950A1 · Usui · 2012 [cited by examiner]
US 20130226957A1 · Ellis · 2013 [cited by examiner]
US 20140270263A1 · Fejzo · 2014 [cited by examiner]
US 20150242750A1 · Anderson · 2015 [cited by examiner]
US 20170116533A1 · Jehan · 2017 [cited by examiner]
US 20170154216A1 · Kennedy · 2017 [cited by examiner]
US 20170236504A1 · Brooker · 2017 [cited by examiner]
US 20180137845A1 · Prokop · 2018 [cited by applicant]
US 20180341704A1 · Barkan · 2018 [cited by examiner]
US 20190318060A1 · Brenner · 2019 [cited by examiner]
US 20200074982A1 · McCallum · 2020 [cited by examiner]
US 20200320388A1 · Lyske · 2020 [cited by applicant]
US 20210049989A1 · Bretan · 2021 [cited by applicant]
US 20210090536A1 · Pachet · 2021 [cited by applicant]
US 20210090590A1 · Vincente · 2021 [cited by applicant]
US 20210294840A1 · Lee · 2021 [cited by applicant]
US 20210350778A1 · Kokkinis · 2021 [cited by examiner]
US 20230260488A1 · Vicente · 2023 [cited by applicant]
US 20230260492A1 · Vicente · 2023 [cited by applicant]
EP 751471 · 1996 [cited by applicant]
EP 3796306 · 2021 [cited by applicant]
WO 2015035492 · 2015 [cited by applicant]
WO 2015154159 · 2015 [cited by applicant]
WO 2016189307 · 2016 [cited by applicant]
WO 2017030661 · 2017 [cited by applicant]
WO 2019084419 · 2019 [cited by applicant]
Aucouturier; J.-J., Pachet, F. and Sandler, M. “The Way It Sounds: Timbre Models for Analysis and Retrieval of Polyphonic Music Signals.” IEEE Transactions of Multimedia, 7(6):1028-1035 (Dec. 2005). [cited by applicant]
Ellis, Daniel et al., “Identifying ‘Cover Songs’ with Chroma Features and Dynamic Programming Beat Tracking”, 2007, 4 pages. [cited by applicant]
European Communication in Application 20174092.5, mailed Dec. 18, 2020, 8 pages. [cited by applicant]
European Communication in Application 20174092.5, mailed Aug. 23, 2021, 9 pages. [cited by applicant]
European Communication in Application 20174093.3, mailed Jun. 2, 2021, 13 pages. [cited by applicant]
European Communication in Application 20174093.3, mailed Dec. 3, 2020, 11 pages. [cited by applicant]
European Communication in Application 20205650.3, mailed Feb. 14, 2022, 7 pages. [cited by applicant]
European Extended Search Report in Application 20174092.5, mailed Sep. 1, 2020, 8 pages. [cited by applicant]
European Extended Search Report in Application 20174093.3, mailed Sep. 1, 2020, 12 pages. [cited by applicant]
European Extended Search Report in Application 20205650.3, mailed May 7, 2021, 16 pages. [cited by applicant]
European Extended Search Report in Application 20205651.1, mailed May 3, 2021, 13 pages. [cited by applicant]
European Minutes of the Oral Proceedings in Application 20174093.3, mailed Nov. 3, 2021, 13 pages. [cited by applicant]
European Result of Consultation in Application 20174093.3, mailed Oct. 20, 2021, 5 pages. [cited by applicant]
European Result of Consultation in Application 20174092.5, mailed Jan. 24, 2022, 3 pgs. [cited by applicant]
European Result of Consultation in Application 20174092.5, mailed Dec. 16, 2021, 7 pgs. [cited by applicant]
European Written Submission in Preparation to Oral Proceedings in Application 20174093.3, mailed Oct. 8, 2021, 5 pages. [cited by applicant]
Jehan, T. “Creating music by listening,” PhD, MIT Media Lab (2005). [cited by applicant]
Lee, Jongpil, et al., “Disentangled Multidimensional Metric Learning for Music Similarity”, ARXIV.org, Aug. 9, 2020, 5 pages. [cited by applicant]
Marco A. Martinez Ramirez, Joshua D. Reiss. “Deep Learning and Intelligent Audio Mixing.” Proceedings of the 3rd Workshop on Intelligent Music Production, Salford, UK (Sep. 15, 2017). [cited by applicant]
Marolt, M., “A Mid-Level Representation for Melody-Based Retrieval in Audio Collections”, IEEE Transactions on Multimedia, vol. 10, No. 8, Dec. 1, 2008, 9 pages. [cited by applicant]
Meinard Muller et al., “Multimodal Music Processing,” DFU, vol. 3 (2012). Available at: https://drops.dagstuhl.de/opus/volltexte/dfu-complete/dfu-vo13-complete.p- df. [cited by applicant]
Oderkerken, Daphne, et al., “Decibel: Improving Audio Chord Estimation for Popular Music by Alignment and Integration of Crowd-Sourced Symbolic Representations”, ARXIV.org, Feb. 22, 2020, 81 pages. [cited by applicant]
Van den Oord, Aaron, Sander Dieleman, and Benjamin Schrauwen. “Deep content-based music recommendation.” Advances in neural information processing systems (2013). [cited by applicant]
Veit, Andreas, et al., “Conditional Similarity Networks”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21, 2017,9 pages. [cited by applicant]
European Communication in Application 20205651.1, mailed Jan. 10, 2024, 5 pages. [cited by applicant]