IP Library › Granted Patent US 12,259,950
Granted Patent B2
US 12,259,950 · App. 17/063,606 · Granted Mar 25, 2025

Model selection for production system via automated online experiments

Inventors: Zhenwen Dai (London, GB); Praveen Chandar Ravichandran (New York, NY); Ghazal Fazelnia (New York, NY); Benjamin Carterette (Wilmington, DE); Mounia Lalmas-Roelleke (Saffron Walden, GB)
Assignee: Spotify AB
G06F18/285G06F18/2113G06F18/2178G06N7/01G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,259,950
App. No.
17/063,606
Granted
Mar 25, 2025
Kind
B2
Abstract

Disclosed examples include an automated online experimentation mechanism that can perform model selection from a large pool of models with a relatively small number of online experiments. The probability distribution of the metric of interest that contains the model uncertainty is derived from a Bayesian surrogate model trained using historical logs. Disclosed techniques can be applied to identify a superior model by sequentially selecting and deploying a list of models from the candidate set that balance exploration-exploitation.

Claims (59)

1. A computer-implemented method comprising:

obtaining media content item analytics data, wherein the media content item analytics data include presentations of media content items to users and lengths of time the users engaged with the media content items that were presented;

iteratively repeating until an evaluation criterion is satisfied:

(i) generating a surrogate machine learning model of the media content item analytics data, wherein the surrogate machine learning model describes a distribution of the lengths of time,

(ii) based on samples from the distribution of the lengths of time, determining scores for a plurality of candidate machine learning models, wherein the plurality of candidate machine learning models each provide presentations of further media content items, and wherein the scores reflect how well the plurality of candidate machine learning models predict further lengths of time of engagement with the media content items that were presented,

(iii) based on the scores, selecting a highest-scoring machine learning model from the plurality of candidate machine learning models,

(iv) updating the media content item analytics data with additional media content item analytics data from online use of the highest scoring machine learning model;

in response to the evaluation criterion being satisfied, deploying a current highest-scoring machine learning model for online use; and

using the current highest-scoring machine learning model to select particular media content items for online streaming to a particular user.

2. The computer-implemented method of claim 1 , wherein the presentations of media content items to users comprises recommending the media content items to the users.

3. The computer-implemented method of claim 1 , wherein the media content items include audio content items, and wherein the lengths of time the users engaged with the media content items that were presented include amounts of time the audio content items were played out to the users.

4. The computer-implemented method of claim 3 , wherein the audio content items are songs or podcast episodes.

5. The computer-implemented method of claim 1 , wherein selecting the highest-scoring machine learning model is also based on an acquisition function that is characterizes an expected improvement over a previous machine learning model.

6. The computer-implemented method of claim 1 , wherein selecting the highest-scoring machine learning model is also based on an acquisition function that is characterized a probability of improvement over a previous machine learning model.

7. The computer-implemented method of claim 1 , wherein selecting the highest-scoring machine learning model is also based on an acquisition function that is characterized an upper confidence bound of the distribution.

8. The computer-implemented method of claim 1 , wherein the evaluation criterion being satisfied comprises at least one of:

the current highest-scoring machine learning model exhibiting an estimated performance above a first predetermined threshold;

a predetermined amount of time has elapsed;

a number of observations in the media content item analytics data is above a second predetermined threshold; or

a number of users associated with the media content item analytics data is above a third predetermined threshold.

9. A computer-implemented method of claim 1 , wherein the current highest-scoring machine learning model, when deployed for online use, provides recommendations of media content items to an application system, and wherein the application system provides the recommendations to user devices.

10. The computer-implemented method of claim 9 , wherein media playback engines on the user devices obtain and play out the media content items that were recommended.

11. A non-transitory computer-readable medium storing program instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations comprising:

obtaining media content item analytics data, wherein the media content item analytics data include presentations of media content items to users and lengths of time the users engaged with the media content items that were presented;

iteratively repeating until an evaluation criterion is satisfied:

(i) generating a surrogate machine learning model of the media content item analytics data, wherein the surrogate machine learning model describes a distribution of the lengths of time,

(ii) based on samples from the distribution of the lengths of time, determining scores for a plurality of candidate machine learning models, wherein the plurality of candidate machine learning models each provide presentations of further media content items, and wherein the scores reflect how well the plurality of candidate machine learning models predict further lengths of time of engagement with the media content items that were presented,

(iii) based on the scores, selecting a highest-scoring machine learning model from the plurality of candidate machine learning models,

(iv) updating the media content item analytics data with additional media content item analytics data from online use of the highest-scoring machine learning model;

in response to the evaluation criterion being satisfied, deploying a current highest-scoring machine learning model for online use; and

using the current highest-scoring machine learning model to select particular media content items for online streaming to a particular user.

12. The non-transitory computer-readable medium of claim 11 , wherein the presentations of media content items to users comprises recommending the media content items to the users.

13. The non-transitory computer-readable medium of claim 11 , wherein the media content items include audio content items, and wherein the lengths of time the users engaged with the media content items that were presented include amounts of time the audio content items were played out to the users.

14. The non-transitory computer-readable medium of claim 13 , wherein the audio content items are songs or podcast episodes.

15. The non-transitory computer-readable medium of claim 11 , wherein the evaluation criterion being satisfied comprises at least one of:

the current highest-scoring machine learning model exhibiting an estimated performance above a first predetermined threshold;

a predetermined amount of time has elapsed;

a number of observations in the media content item analytics data is above a second predetermined threshold; or

a number of users associated with the media content item analytics data is above a third predetermined threshold.

16. A computing system comprising:

one or more processors;

memory; and

program instructions, stored in the memory, that upon execution by the one or more processors cause the computing system to perform operations comprising:

obtaining media content item analytics data, wherein the media content item analytics data include presentations of media content items to users and lengths of time the users engaged with the media content items that were presented;

iteratively repeating until an evaluation criterion is satisfied:

(i) generating a surrogate machine learning model of the media content item analytics data, wherein the surrogate machine learning model describes a distribution of the lengths of time,

(ii) based on samples from the distribution of the lengths of time, determining scores for a plurality of candidate machine learning models, wherein the plurality of candidate machine learning models each provide presentations of further media content items, and wherein the scores reflect how well the plurality of candidate machine learning models predict further lengths of time of engagement with the media content items that were presented,

(iii) based on the scores, selecting a highest-scoring machine learning model from the plurality of candidate machine learning models,

(iv) updating the media content item analytics data with additional media content item analytics data from online use of the highest-scoring machine learning model;

in response to the evaluation criterion being satisfied, deploying a current highest-scoring machine learning model for online use; and

using the current highest-scoring machine learning model to select particular media content items for online streaming to a particular user.

17. The computing system of claim 16 , wherein the presentations of media content items to users comprises recommending the media content items to the users.

18. The computing system of claim 16 , wherein the media content items include audio content items, and wherein the lengths of time the users engaged with the media content items that were presented include amounts of time the audio content items were played out to the users.

19. The computing system of claim 18 , wherein the audio content items are songs or podcast episodes.

20. The computing system of claim 16 , wherein the evaluation criterion being satisfied comprises at least one of:

the current highest-scoring machine learning model exhibiting an estimated performance above a first predetermined threshold;

a predetermined amount of time has elapsed;

a number of observations in the media content item analytics data is above a second predetermined threshold; or

a number of users associated with the media content item analytics data is above a third predetermined threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2024
From: DAI, ZHENWEN; RAVICHANDRAN, PRAVEEN CHANDAR; FAZELNIA, GHAZAL; CARTERETTE, BENJAMIN; LALMAS-ROELLEKE, MOUNIA
To: SPOTIFY AB
Reel/Frame 066952/0651 →
Continuity (1)
Related Publication 20220108125A1 · Apr 7, 2022
References Cited (44)
US 20120191630A1 · Breckenridge · 2012 [cited by examiner]
US 20160110657A1 · Gibiansky et al. · 2016 [cited by applicant]
US 20170311035A1 · Lewis · 2017 [cited by examiner]
US 20190095756A1 · Agrawal et al. · 2019 [cited by applicant]
US 20200078685A1 · Aghdaie · 2020 [cited by examiner]
US 20220067087A1 · Reardon · 2022 [cited by examiner]
US 20220100772A1 · Kadarundalagi Raghura · 2022 [cited by examiner]
Malkomes, Gustavo, Charles Schaff, and Roman Garnett. “Bayesian optimization for automated model selection.” Advances in neural information processing systems 29 (2016). (Year: 2016). [cited by examiner]
Brochu, Eric, Vlad M. Cora, and Nando De Freitas. “A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning.” arXiv preprint arXiv… [cited by examiner]
Zhou, Weilin, and Frederic Precioso. “Adaptive bayesian linear regression for automated machine learning.” arXiv preprint arXiv:1904.00577 (2019). (Year: 2019). [cited by examiner]
“AutoAI Overview”, IBM Cloud Pak for Data (2020). Available online at: https://dataplatform.cloud.ibm.com/docs/content/wsj/analyze-data/autoai-overview.html <accessed Aug. 12, 2020>. [cited by applicant]
H. Akaike, “A new look at the statistical model identification,” IEEE Transactions on Automatic Control, vol. 19, No. 6, pp. 716-723 (1974). [cited by applicant]
G. Schwarz, “Estimating the dimension of a model,” Annals of Statistics, vol. 6, pp. 461-464 (1978). [cited by applicant]
J. Snoek et al., “Practical bayesian optimization of machine learning algorithms,” in Advances in neural information processing systems, pp. 2951-2959 (2012). [cited by applicant]
A. Klein et al., “Meta-surrogate benchmarking for hyperparameter optimization,” in Advances in Neural Information Processing Systems 32, pp. 6270-6280 (2019). [cited by applicant]
J. R. Lloyd et al., “Automatic construction and natural-language description of nonparametric regression models,” in Proceedings of the Twenty-Eighth AAAI Conference on Artificial Intelligence, p. 1242-1250 (2014). [cited by applicant]
G. Malkomes et al., “Bayesian optimization for automated model selection,” in Advances in Neural Information Processing Systems 29, pp. 2900-2908 (2016). [cited by applicant]
H. Kim and Y. W. Teh, “Scaling up the automatic statistician: Scalable structure discovery using gaussian processes,” in Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics… [cited by applicant]
X. Lu et al., “Structured variationally auto-encoded optimization,” in Proceedings of the 35th International Conference on Machine Learning, pp. 3267-3275 (2018). [cited by applicant]
T. Elsken et al., “Neural architecture search: A survey,” Journal of Machine Learning Research, vol. 20, No. 55, pp. 1-21 (2019). [cited by applicant]
H. Chai et al., “Automated model selection with Bayesian quadrature,” in Proceedings of the 36th International Conference on Machine Learning, pp. 931-940 (2019). [cited by applicant]
V. Muthukumar et al, “Best of many worlds: Robust model selection for online supervised learning,” in Proceedings of Machine Learning Research, pp. 3177-3186 (2019). [cited by applicant]
D. Precup et al., “Eligibility traces for off-policy policy evaluation,” in ICML, pp. 759-766 (2000). [cited by applicant]
M. Dudik et al., “Doubly robust policy evaluation and optimization,” Statistical Science, vol. 29, pp. 485-511 (2014). [cited by applicant]
M. Farajtabar et al., “More robust doubly robust off-policy evaluation,” in Proceedings of the 35th International Conference on Machine Learning, pp. 1447-1456 (2018). [cited by applicant]
Y. Liu et al., “Representation balancing mdps for off-policy policy evaluation,” in Advances in Neural Information Processing Systems 31, pp. 2644-2653, (2018). [cited by applicant]
N. Vlassis et al., “On the design of estimators for bandit off-policy evaluation,” in Proceedings of the 36th International Conference on Machine Learning (2019). [cited by applicant]
A. Irpan et al., “Off-policy evaluation via off-policy classification,” in Advances in Neural Information Processing Systems 32, pp. 5437-5448 (2019). [cited by applicant]
M. Ghavamzadeh and Y. Engel, “Bayesian policy gradient algorithms,” in Advances in neural information processing systems, pp. 457-464 (2007). [cited by applicant]
M. Ghavamzadeh et al., “Bayesian policy gradient and actor-critic algorithms,” The Journal of Machine Learning Research, vol. 17, No. 1, pp. 2319-2371 (2016). [cited by applicant]
G. Lee et al., “Bayesian policy optimization for model uncertainty,” in International Conference on Learning Representations (2019). [cited by applicant]
B. Letham and E. Bakshy, “Bayesian optimization for policy search via online-offline experimentation,” Journal of Machine Learning Research, vol. 20, No. 145, pp. 1-30 (2019). [cited by applicant]
D. Russo, “Simple bayesian algorithms for best arm identification,” in 29th Annual Conference on Learning Theory, pp. 1417-1418 (2016). [cited by applicant]
K. Chaloner and I. Verdinelli, “Bayesian experimental design: A review,” Statistical Science, pp. 273-304 (1995). [cited by applicant]
J. M. Hernández-Lobato et al., “Predictive entropy search for efficient global optimization of black-box functions,” in Advances in neural information processing systems, pp. 918-926 (2014). [cited by applicant]
A. Foster et al., “Variational bayesian optimal experimental design,” in Advances in Neural Information Processing Systems, pp. 14036-14047 (2019). [cited by applicant]
J. Vanlier et al., “A bayesian approach to targeted experiment design,” Bioinformatics, vol. 28, No. 8, pp. 1136-1142 (2012). [cited by applicant]
D. Golovin et al., “Near-optimal bayesian active learning with noisy observations,” in Advances in Neural Information Processing Systems, pp. 766-774 (2010). [cited by applicant]
B. Shababo et al., “Bayesian inference and online experimental design for mapping neural microcircuits,” in Advances In Neural Information Processing Systems, pp. 1304-1312 (2013). [cited by applicant]
Sato, Masa-aki, “Online Model Selection Based on the Variational Bayes”, Neural Computation 13, 1649-1681 (2001). [cited by applicant]
Bishop, Christopher M., “Pattern Recognition and Machine Leaning”, Eds. Jordan et al., Springer, 2006, 758 pages. [cited by applicant]
Hennig, Philipp et al., “Entropy Search for Information-Efficient Global Optimization”, Journal of Machine Learning Research 13 (2012), 1809-1837. [cited by applicant]
Harper, F. Maxwell et al., “The MovieLens Datasets: History and Context”, ACM Transactions on Interactive Intelligent Systems, vol. 5, No. 4, Article 19, Dec. 2015, 19 pages. [cited by applicant]
Hug, Nicolas, “Surprise: A Python library for recommender systems”, Journal of Open Source Software 5(52), 2174, 3 pages. [cited by applicant]
Cited By (1)
US 12,479,090