IP Library › Granted Patent US 12,688,396
Granted Patent B2
US 12,688,396 · App. 17/320,439 · Granted Jul 21, 2026

Diversity aware media content recommendation

Inventors: Christian Hansen (Copenhagen, DK); Casper Hansen (Copenhagen, DK); Brian Christian Peter Brost (New York City, NY); Lucas Maystre (London, GB); Mounia Lalmas-Roelleke (Saffron Walden, GB); Rishabh Mehrotra (London, GB)
Assignee: Spotify AB
G06N3/042G06N3/044G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,396
App. No.
17/320,439
Filed
May 14, 2021
Granted
Jul 21, 2026
Kind
B2
Art Unit
2125
USPC
706/12
Abstract

A reinforcement learning ranker can take into account previously-recommended media content items to produce a ranked list of media content items to recommend next. The ranker finds a policy that gives the probability of sampling a media content item given a state. The policy is learned such that it maximizes a reward. A reward function associated with the media content item can be defined with respect to whether the user finds the media content item relevant (likelihood that the user will like the media content item) and a diversity score of the media content item.

Claims (35)

1 . A method for selecting a media content item, the method comprising:

obtaining first embedding data describing media content items from previous content consumption sessions of a user account, wherein the first embedding data is based on a weighting of session representations from the previous content consumption sessions;

obtaining second embedding data regarding media content items previously recommended during a current content consumption session of the user account;

obtaining third embedding data describing a potential media content item;

providing, to a reinforcement learning model, the first embedding data, the second embedding data, and the third embedding data, wherein the reinforcement learning model produces a score, wherein the score is based on: a relevance of features of the potential media content item with respect to the user account, and a diversity of the potential media content item with respect to the user account, and wherein the diversity is weighted by the relevance and a weighting factor; and

selecting, for the user account, the potential media content item based on the score, wherein the potential media content item is selected for playback to a device associated with the user account.

2 . The method of claim 1 , wherein the potential media content item is an audio track.

3 . The method of claim 1 , wherein the diversity is based on a dissimilarity of the potential media content item with respect to media content items from the previous content consumption sessions.

4 . A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to:

obtain first embedding data describing media content items from previous content consumption sessions of a user account, wherein the first embedding data is based on a weighting of session representations from the previous content consumption sessions;

obtain second embedding data regarding media content items previously recommended during a current content consumption session of the user account;

obtain third embedding data describing a potential media content item;

provide, to a reinforcement learning model, the first embedding data, the second embedding data, and the third embedding data, wherein the reinforcement learning model produces a score, wherein the score is based on: a relevance of features of the potential media content item with respect to the user account, and a diversity of the potential media content item with respect to the user account, and wherein the diversity is weighted by the relevance and a weighting factor; and

select, for the user account, the potential media content item based on the score, wherein the potential media content item is selected for playback to a device associated with the user account.

5 . A media-delivery system comprising:

one or more processors;

memory; and

program instructions, stored in the memory, that upon execution by the one or more processors cause the media-delivery system to perform operations comprising:

obtain first embedding data describing media content items from previous content consumption sessions of a user account, wherein the first embedding data is based on a weighting of session representations from the previous content consumption sessions;

obtain second embedding data regarding media content items previously recommended during a current content consumption session of the user account;

obtain third embedding data describing a potential media content item;

provide, to a reinforcement learning model, the first embedding data, the second embedding data, and the third embedding data, wherein the reinforcement learning model produces a score, wherein the score is based on: a relevance of features of the potential media content item with respect to the user account, and a diversity of the potential media content item with respect to the user account, and wherein the diversity is weighted by the relevance and a weighting factor;

select, for the user account, the potential media content item based on the score; and

transmit the potential media content item to a media-playback device for playback.

6 . The media-delivery system of claim 5 , wherein the diversity is based on a dissimilarity of the potential media content item with respect to media content items from the previous content consumption sessions.

7 . The method of claim 1 , wherein the reinforcement learning model applies a long short-term memory (LSTM) layer to the first embedding data and the second embedding data to produce an output, and wherein the reinforcement learning model applies a softmax function to a concatenation of the output and the third embedding data to produce the score.

8 . The method of claim 7 , wherein the LSTM layer comprises a stacked LSTM initialized based on a session meta feature relating to a combination of the first embedding data and data related to the user account.

9 . The non-transitory computer-readable medium of claim 4 , wherein the reinforcement learning model applies a long short-term memory (LSTM) layer to the first embedding data and the second embedding data to produce an output, and wherein the reinforcement learning model applies a softmax function to a concatenation of the output and the third embedding data to produce the score.

10 . The non-transitory computer-readable medium of claim 9 , wherein the LSTM layer comprises a stacked LSTM initialized based on a session meta feature relating to a combination of the first embedding data and data related to the user account.

11 . The media-delivery system of claim 5 , wherein the reinforcement learning model applies a long short-term memory (LSTM) layer to the first embedding data and the second embedding data to produce an output, and wherein the reinforcement learning model applies a softmax function to a concatenation of the output and the third embedding data to produce the score.

12 . The media-delivery system of claim 11 , wherein the LSTM layer comprises a stacked LSTM initialized based on a session meta feature relating to a combination of the first embedding data and data related to the user account.

13 . The method of claim 1 , further comprising:

transmitting the potential media content item to the device for the playback.

14 . The non-transitory computer-readable medium of claim 4 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to:

transmit the potential media content item to the device for the playback.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2024
From: HANSEN, CHRISTIAN; HANSEN, CASPER; BROST, BRIAN CHRISTIAN PETER; MAYSTRE, LUCAS; LALMAS-ROELLEKE, MOUNIA; MEHROTRA, RISHABH
To: SPOTIFY AB
Reel/Frame 068832/0917 →
Continuity (2)
Provisional Application 63025708 · May 15, 2020
Related Publication 20220012565A1 · Jan 13, 2022
References Cited (33)
US 10936653B2 · Levy · 2021 [cited by examiner]
US 20080201287A1 · Takeuchi · 2008 [cited by examiner]
US 20160299906A1 · Cartoon · 2016 [cited by applicant]
US 20180052921A1 · DeGlopper et al. · 2018 [cited by applicant]
US 20180341704A1 · Barkan · 2018 [cited by applicant]
US 20180349492A1 · Levy · 2018 [cited by examiner]
US 20190114687A1 · Krishnamurthy · 2019 [cited by examiner]
US 20190369948A1 · Gibson · 2019 [cited by applicant]
US 20210004682A1 · Gong · 2021 [cited by examiner]
US 20210081758A1 · Zadorojniy · 2021 [cited by examiner]
Aipe et al., “Sentiment-Aware Recommendation System for Healthcare using Social Media” (Year: 2019). [cited by examiner]
Bridge et al., “Sudden Death: A New Way to Compare Recommendation Diversification” (Year: 2019). [cited by examiner]
Gatzioura et al., “A Hybrid Recommender System for Improving Automatic Playlist Continuation” (Year: 2019). [cited by examiner]
Oh et al., “Novel Recommendation based on Personal Popularity Tendency” (Year: 2011). [cited by examiner]
Price et al., “Pull-Based Casting of Media Content Items” (Year: 2017). [cited by examiner]
Shih et al., “Automatic, Personalized, and Flexible Playlist Generation using Reinforcement Learning”, Sep. 12, 2018, arXiv preprint arXiv:1809.04214 (Year: 2018). [cited by examiner]
Price et al., “Pull-Based Casting of Media Content Items”, Dec. 12, 2017, Technical Disclosure Commons (Year: 2017). [cited by examiner]
Chen et al. “Top-K Off-Policy Correction for a REINFORCE Recommender System” WSDM, https://doi.org/10.1145/3289600.3290999, Feb. 2019, 9 pages. [cited by applicant]
Pei et al. “Value-aware Recommendation based on Reinforced Profit Maximization in E-commerce Systems,” arXiv:1902.00851v1, 2019, 8 pages. [cited by applicant]
Zhao et al. “Deep Reinforcement Learning for Page-wise Recommendations,” Association for Computing Machinery, 2018, 9 pages. [cited by applicant]
Zhao et al. “Recommendations with Negative Feedback via Pairwise Deep Reinforcement Learning,” Association for Computing Machinery, 2018, 9 pages. [cited by applicant]
Zheng et al. “DRN: A Deep Reinforcement Learning Framework for News Recommendation,” Track: Intelligent and Autonomous sytems on the web, 2018, 10 pages. [cited by applicant]
Wang et al. “Exploration in Interactive Personalized Music Recommendation: A Reinforcement Learning Approach,” https://bigbird.comp.nus.edu.sg/m2ap/wordpress/wp-content/uploads/2016/01/a7-wang.pdf, 2014, 22 pages. [cited by applicant]
Ie et al. “SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets,” Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI-19), 2019, 8 pages. [cited by applicant]
Anderson, et al., “Algorithmic Effects on the Diversity of Consumption on Spotify”, The World Wide Web Conference (2020), Taipei, Taiwan, 11 pages. [cited by applicant]
Clarke et al., “Novelty and Diversity in Information Retrieval Evaluation”, Proceedings of The 31 St Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2008), ACM, New … [cited by applicant]
Liebman et al., “DJ-MC: A reinforcement-learning agent for music playlist recommendation”, Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems (International Foundation for Auton… [cited by applicant]
Mehrotra et al., “Towards a fair marketplace: Counterfactual evaluation of the trade-off between relevance, fairness & satisfaction in recommendation systems”, Proceedings of the 27th ACM International Conference on Inf… [cited by applicant]
Waller and Anderson, “Generalists and Specialists: Using Community Embeddings to Quantify Activity Diversity in Online Platforms”, The World Wide Web Conference, ACM, 1954-1964 (2019), 36 pages. [cited by applicant]
Williams, Ronald, “Simple statistical gradient-following algorithms for connectionist reinforcement learning”, Machine Learning 8, 229-256 (1992). [cited by applicant]
Zhou, et al., “The Impact of You Tube Recommendation System on Video Views”, Proceedings of the 10th ACM SIGCOMM Conference on Internet Measurement, ACM, 404-410 (2010). [cited by applicant]
Sakai et al., Which Diversity Evaluation Measures are “Good”?, Session 7A: Relevance and Evaluation 1, SIGIR '19, Jul. 21-25, 2019, Paris, France, 10 pages (595-604). [cited by applicant]
Fleder and Hosanagar, Blockbuster Culture's Next Rise or Fall: The Impact of Recommender Systems on Sales Diversity, Management Science 55, 5, 697-712 (2009). [cited by applicant]