IP Library › Granted Patent US 11,636,835
Granted Patent B2
US 11,636,835 · App. 17/003,614 · Granted Apr 25, 2023

Spoken words analyzer

Inventors: Tahora H. Nazer (Tempe, AZ); Tristan Jehan (Brooklyn, NY)
Assignee: Spotify AB
G10H1/0008G06F16/4387G06F16/685G06F16/686G06F17/18G06F40/242G06F40/279G06F40/30G06N7/005G06N20/00G06N20/20G10L15/14G06N5/003G10H2210/031G10H2210/056G10H2220/005G10H2220/011G10H2240/081G10H2240/085G10H2240/141
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,636,835
App. No.
17/003,614
Granted
Apr 25, 2023
Kind
B2
Abstract

A lyrics analyzer generates tags and explicitness indicators for a set of tracks. These tags may indicate the genre, mood, occasion, or other features of each track. The lyrics analyzer does so by generating an n-dimensional vector relating to a set of topics extracted from the lyrics and then using those vectors to train a classifier to determine whether each tag applies to each track. The lyrics analyzer may also generate playlists for a user based on a single seed song by comparing the lyrics vector or the lyrics and acoustics vectors of the seed song to other songs to select songs that closely match the seed song. Such a playlist generator may also take into account the tags generated for each track.

Claims (114)

1. A method, comprising:

receiving a plurality of tracks at an information storage and retrieval platform via an electronic communication from a datastore of tracks, the plurality of tracks including spoken words;

extracting n topics summarizing the spoken words of the plurality of tracks, each topic consisting of a plurality of words found together within the spoken words of the plurality of tracks, where n is an integer;

receiving a seed track from among the plurality of tracks;

receiving, for at least one of the plurality of tracks, a seed track lyrics vector;

calculating a similarity score for each of the plurality of tracks to the seed track, based on their respective n-dimensional lyrics vectors, thereby generating a plurality of similarity scores, wherein calculating the similarity score for at least one of the plurality of tracks further includes calculating the similarity between the lyrics vector of the seed track and the lyrics vector of at least one of the plurality of tracks;

predicting, based on the n-dimensional lyrics vectors, a set of track tags for each of the plurality of tracks;

identifying one or more similar track tags from among the set of track tags for each of the plurality of tracks; and

generating a playlist of tracks based on the plurality of similarity scores and the one or more similar track tags.

2. The method according to claim 1 , wherein extracting n topics summarizing the spoken words of the plurality of tracks comprises using a generative statistical model; and

the generative statistical model is a Latent Dirichlet Allocation (LDA) model.

3. The method according to claim 1 , wherein calculating a similarity score includes calculating a cosine distance.

4. The method according to claim 1 , further comprising:

processing the spoken words of the plurality of tracks, wherein the processing includes: (i) white-space standardizing, (ii) lowercasing, (iii) removing stopwords, (iv) removing punctuation, (v) lemmatizing, (vi) removing character repetition based on a dictionary, or (vii) any combination of (i), (ii), (iii), (iv), (v), and (vi).

5. The method according to claim 1 , wherein calculating a similarity score further includes comparing a hybrid vector of at least one of the plurality of tracks to a hybrid vector of the seed track.

6. A method, comprising:

receiving a plurality of tracks at an information storage and retrieval platform via an electronic communication from a datastore of tracks, the plurality of tracks including spoken words;

extracting n topics summarizing the spoken words of the plurality of tracks, each topic consisting of a plurality of words found together within the spoken words of the plurality of tracks, where n is an integer;

generating, for each of the plurality of tracks, an n-dimensional lyrics vector using a generative statistical model based on the association of the spoken words of the track with the n topics, thereby generating a plurality of n-dimensional lyrics vectors;

receiving a seed track from among the plurality of tracks;

calculating a similarity score for each of the plurality of tracks to the seed track, based on their respective n-dimensional lyrics vectors, thereby generating a plurality of similarity scores;

predicting, based on the n-dimensional lyrics vectors, a set of track tags for each of the plurality of tracks;

identifying one or more similar track tags from among the set of track tags for each of the plurality of tracks;

obtaining, for each track of the plurality of tracks, track metadata including: (i) at least one tag based on a playlist title, (ii) a title, (iii) an album identifier, (iv) an artist name, (v) a value representing popularity, (vi) a plurality of audio features, (vii) one or more genres, (viii) a sentiment score, or (ix) any combination of (i), (ii), (iii), (iv), (v), (vi), (vii) and (viii); and

generating a playlist of tracks based on the plurality of similarity scores and the one or more similar track tags.

7. The method according to claim 6 , further comprising:

displaying the track metadata for at least one of the plurality of tracks.

8. The method according to claim 6 , wherein extracting n topics summarizing the spoken words of the plurality of tracks comprises using a generative statistical model; and

the generative statistical model is a Latent Dirichlet Allocation (LDA) model.

9. The method according to claim 6 , wherein calculating a similarity score includes calculating a cosine distance.

10. The method according to claim 6 , further comprising:

processing the spoken words of the plurality of tracks, wherein the processing includes: (i) white-space standardizing, (ii) lowercasing, (iii) removing stopwords, (iv) removing punctuation, (v) lemmatizing, (vi) removing character repetition based on a dictionary, or (vii) any combination of (i), (ii), (iii), (iv), (v), and (vi).

11. The method according to claim 6 , further comprising

receiving, for at least one of the plurality of tracks, a seed track lyrics vector; and

wherein calculating the similarity score for at least one of the plurality of tracks further includes calculating the similarity between the lyrics vector of the seed track and the lyrics vector of at least one of the plurality of tracks.

12. The method according to claim 11 , wherein calculating a similarity score further includes comparing a hybrid vector of at least one of the plurality of tracks to a hybrid vector of the seed track.

13. A system, comprising:

a computer-readable memory storing executable instructions; and

one or more processors in communication with the computer-readable memory, wherein, when the one or more processors execute the executable instructions, the one or more processors perform:

receiving a plurality of tracks at an information storage and retrieval platform via an electronic communication from a datastore of tracks, the plurality of tracks including spoken words;

extracting n topics summarizing the spoken words of the plurality of tracks, each topic consisting of a plurality of words found together within the spoken words of the plurality of tracks, where n is an integer;

generating, for each of the plurality of tracks, an n-dimensional lyrics vector using a generative statistical model based on the association of the spoken words of the track with the n topics, thereby generating a plurality of n-dimensional lyrics vectors;

receiving a seed track from among the plurality of tracks;

receiving, for at least one of the plurality of tracks, a seed track lyrics vector;

calculating a similarity score for each of the plurality of tracks to the seed track, based on their respective n-dimensional lyrics vectors, thereby generating a plurality of similarity scores, wherein calculating the similarity score for at least one of the plurality of tracks further includes calculating the similarity between the lyrics vector of the seed track and the lyrics vector of at least one of the plurality of tracks;

predicting, based on the n-dimensional lyrics vectors, a set of track tags for each of the plurality of tracks;

identifying one or more similar track tags from among the set of track tags for each of the plurality of tracks; and

generating a playlist of tracks based on the plurality of similarity scores and the one or more similar track tags.

14. The system of claim 13 , wherein extracting n topics summarizing the spoken words of the plurality of tracks comprises using a generative statistical model; and

the generative statistical model is a Latent Dirichlet Allocation (LDA) model.

15. The system of claim 13 , wherein calculating a similarity score includes calculating a cosine distance.

16. The system of claim 13 , wherein the one or more processors, when executing the executable instructions, further perform:

processing the spoken words of the plurality of tracks, wherein the processing includes: (i) white-space standardizing, (ii) lowercasing, (iii) removing stopwords, (iv) removing punctuation, (v) lemmatizing, (vi) removing character repetition based on a dictionary, or (vii) any combination of (i), (ii), (iii), (iv), (v), and (vi).

17. The system of claim 13 , wherein calculating a similarity score further includes comparing a hybrid vector of at least one of the plurality of tracks to a hybrid vector of the seed track.

18. A system, comprising:

a computer-readable memory storing executable instructions; and

one or more processors in communication with the computer-readable memory, wherein, when the one or more processors execute the executable instructions, the one or more processors perform:

receiving a plurality of tracks at an information storage and retrieval platform via an electronic communication from a datastore of tracks, the plurality of tracks including spoken words;

extracting n topics summarizing the spoken words of the plurality of tracks, each topic consisting of a plurality of words found together within the spoken words of the plurality of tracks, where n is an integer;

generating, for each of the plurality of tracks, an n-dimensional lyrics vector using a generative statistical model based on the association of the spoken words of the track with the n topics, thereby generating a plurality of n-dimensional lyrics vectors;

receiving a seed track from among the plurality of tracks;

calculating a similarity score for each of the plurality of tracks to the seed track, based on their respective n-dimensional lyrics vectors, thereby generating a plurality of similarity scores;

predicting, based on the n-dimensional lyrics vectors, a set of track tags for each of the plurality of tracks;

identifying one or more similar track tags from among the set of track tags for each of the plurality of tracks;

obtaining, for each track of the plurality of tracks, track metadata including: (i) at least one tag based on a playlist title, (ii) a title, (iii) an album identifier, (iv) an artist name, (v) a value representing popularity, (vi) a plurality of audio features, (vii) one or more genres, (viii) a sentiment score, or (ix) any combination of (i), (ii), (iii), (iv), (v), (vi), (vii) and (viii); and

generating a playlist of tracks based on the plurality of similarity scores and the one or more similar track tags.

19. The system according to claim 18 , wherein the one or more processors, when executing the executable instructions, further perform:

displaying the track metadata for at least one of the plurality of tracks.

20. The system according to claim 18 , wherein extracting n topics summarizing the spoken words of the plurality of tracks comprises using a generative statistical model; and

the generative statistical model is a Latent Dirichlet Allocation (LDA) model.

21. The system according to claim 18 , wherein calculating a similarity score includes calculating a cosine distance.

22. The system according to claim 18 , wherein the one or more processors, when executing the executable instructions, further perform:

processing the spoken words of the plurality of tracks, wherein the processing includes: (i) white-space standardizing, (ii) lowercasing, (iii) removing stopwords, (iv) removing punctuation, (v) lemmatizing, (vi) removing character repetition based on a dictionary, or (vii) any combination of (i), (ii), (iii), (iv), (v), and (vi).

23. The system according to claim 18 , wherein the one or more processors, when executing the executable instructions, further perform:

receiving, for at least one of the plurality of tracks, a seed track lyrics vector; and

wherein calculating the similarity score for at least one of the plurality of tracks further includes calculating the similarity between the lyrics vector of the seed track and the lyrics vector of at least one of the plurality of tracks.

24. The system according to claim 23 , wherein calculating a similarity score further includes comparing a hybrid vector of at least one of the plurality of tracks to a hybrid vector of the seed track.

25. A non-transitory computer-readable medium having stored thereon one or more sequences of instructions for causing one or more processors to perform:

receiving a plurality of tracks at an information storage and retrieval platform via an electronic communication from a datastore of tracks, the plurality of tracks including spoken words;

extracting n topics summarizing the spoken words of the plurality of tracks, each topic consisting of a plurality of words found together within the spoken words of the plurality of tracks, where n is an integer;

generating, for each of the plurality of tracks, an n-dimensional lyrics vector using a generative statistical model based on the association of the spoken words of the track with the n topics, thereby generating a plurality of n-dimensional lyrics vectors;

receiving a seed track from among the plurality of tracks;

receiving, for at least one of the plurality of tracks, a seed track lyrics vector;

calculating a similarity score for each of the plurality of tracks to the seed track, based on their respective n-dimensional lyrics vectors, thereby generating a plurality of similarity scores, wherein calculating the similarity score for at least one of the plurality of tracks further includes calculating the similarity between the lyrics vector of the seed track and the lyrics vector of at least one of the plurality of tracks;

predicting, based on the n-dimensional lyrics vectors, a set of track tags for each of the plurality of tracks;

identifying one or more similar track tags from among the set of track tags for each of the plurality of tracks; and

generating a playlist of tracks based on the plurality of similarity scores and the one or more similar track tags.

26. The non-transitory computer-readable medium of claim 25 , wherein extracting n topics summarizing the spoken words of the plurality of tracks comprises using a generative statistical model; and

the generative statistical model is a Latent Dirichlet Allocation (LDA) model.

27. The non-transitory computer-readable medium of claim 25 , wherein calculating a similarity score includes calculating a cosine distance.

28. The non-transitory computer-readable medium of claim 25 , having stored thereon one or more sequences of instructions for causing one or more processors to further perform:

processing the spoken words of the plurality of tracks, wherein the processing includes: (i) white-space standardizing, (ii) lowercasing, (iii) removing stopwords, (iv) removing punctuation, (v) lemmatizing, (vi) removing character repetition based on a dictionary, or (vii) any combination of (i), (ii), (iii), (iv), (v), and (vi).

29. The non-transitory computer-readable medium of claim 25 , wherein calculating a similarity score further includes comparing a hybrid vector of at least one of the plurality of tracks to a hybrid vector of the seed track.

30. A non-transitory computer-readable medium having stored thereon one or more sequences of instructions for causing one or more processors to perform:

receiving a plurality of tracks at an information storage and retrieval platform via an electronic communication from a datastore of tracks, the plurality of tracks including spoken words;

extracting n topics summarizing the spoken words of the plurality of tracks, each topic consisting of a plurality of words found together within the spoken words of the plurality of tracks, where n is an integer;

generating, for each of the plurality of tracks, an n-dimensional lyrics vector using a generative statistical model based on the association of the spoken words of the track with the n topics, thereby generating a plurality of n-dimensional lyrics vectors;

receiving a seed track from among the plurality of tracks;

calculating a similarity score for each of the plurality of tracks to the seed track, based on their respective n-dimensional lyrics vectors, thereby generating a plurality of similarity scores;

predicting, based on the n-dimensional lyrics vectors, a set of track tags for each of the plurality of tracks;

identifying one or more similar track tags from among the set of track tags for each of the plurality of tracks;

obtaining, for each track of the plurality of tracks, track metadata including: (i) at least one tag based on a playlist title, (ii) a title, (iii) an album identifier, (iv) an artist name, (v) a value representing popularity, (vi) a plurality of audio features, (vii) one or more genres, (viii) a sentiment score, or (ix) any combination of (i), (ii), (iii), (iv), (v), (vi), (vii) and (viii); and

generating a playlist of tracks based on the plurality of similarity scores and the one or more similar track tags.

31. The non-transitory computer-readable medium of claim 30 , having stored thereon one or more sequences of instructions for causing one or more processors to further perform:

displaying the track metadata for at least one of the plurality of tracks.

32. The non-transitory computer-readable medium of claim 30 , wherein extracting n topics summarizing the spoken words of the plurality of tracks comprises using a generative statistical model; and

the generative statistical model is a Latent Dirichlet Allocation (LDA) model.

33. The non-transitory computer-readable medium of claim 30 , wherein calculating a similarity score includes calculating a cosine distance.

34. The non-transitory computer-readable medium of claim 30 , having stored thereon one or more sequences of instructions for causing one or more processors to further perform:

processing the spoken words of the plurality of tracks, wherein the processing includes: (i) white-space standardizing, (ii) lowercasing, (iii) removing stopwords, (iv) removing punctuation, (v) lemmatizing, (vi) removing character repetition based on a dictionary, or (vii) any combination of (i), (ii), (iii), (iv), (v), and (vi).

35. The non-transitory computer-readable medium of claim 30 , having stored thereon one or more sequences of instructions for causing one or more processors to further perform:

receiving, for at least one of the plurality of tracks, a seed track lyrics vector; and

wherein calculating the similarity score for at least one of the plurality of tracks further includes calculating the similarity between the lyrics vector of the seed track and the lyrics vector of at least one of the plurality of tracks.

36. The non-transitory computer-readable medium of claim 35 , wherein calculating a similarity score further includes comparing a hybrid vector of at least one of the plurality of tracks to a hybrid vector of the seed track.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2023
From: NAZER, TAHORA H.; JEHAN, TRISTAN
To: SPOTIFY AB
Reel/Frame 062550/0873 →
Continuity (3)
Continuation 16111614 · Aug 24, 2018
Provisional Application 62552882 · Aug 31, 2017
Related Publication 20200394988A1 · Dec 17, 2020