IP Library Granted Patent US 12,462,290
Granted Patent B2
US 12,462,290 · App. 18/530,907 · Granted Nov 4, 2025

Utilizing additive decomposition for universal off-policy evaluation of digital content slate recommendations

Inventors: Shreyas Chaudhari (Amherst, MA); Nikolaos Vlassis (San Jose, CA); Georgios Theocharous (San Jose, CA); David Arbour (San Jose, CA)
Assignee: Adobe Inc.
G06Q30/0631
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,290
App. No.
18/530,907
Granted
Nov 4, 2025
Kind
B2
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods for performing off-policy evaluations of slate recommendation policies through additive decomposition. In particular, in one or more embodiments, the disclosed systems receive historical data corresponding to digital slate recommendations performed by a first slate recommendation policy, with each slate recommendation comprising a plurality of digital slot recommendations. Additionally, in some embodiments, the disclosed systems generate a second slate action using a second slate recommendation policy conditioned on user context. Further, in some embodiments, the disclosed systems generate a plurality of importance weights by summing a plurality of slot-level density ratios generated by comparing the slate actions of the second slate recommendation policy to the slate actions of the first slate recommendation policy. In some embodiments, the disclosed systems apply the plurality of importance weights to generate a predicted reward distribution for evaluation of the second slate recommendation policy.

Claims (49)

1 . A computer-implemented method comprising:

receiving historical slate data comprising observed rewards from selecting slate actions for a plurality of digital slots of a digital slate utilizing a first slate recommendation policy;

generating, for a second slate recommendation policy, a plurality of importance weights from the historical slate data by summing slot-level density ratios between the first slate recommendation policy and the second slate recommendation policy for the slate actions; and

generating a predicted reward distribution for the second slate recommendation policy by applying the plurality of importance weights to the historical slate data for the first slate recommendation policy.

2 . The computer-implemented method of claim 1 , wherein generating the plurality of importance weights comprises:

determining, for a slate action, a first slot-level density ratio between the first slate recommendation policy and the second slate recommendation policy for a first slot of the digital slate;

determining, for the slate action, a second slot-level density ratio between the first slate recommendation policy and the second slate recommendation policy for a second slot of the digital slate; and

summing the first slot-level density ratio and the second slot-level density ratio to determine an importance weight for the slate action.

3 . The computer-implemented method of claim 2 , wherein generating the plurality of importance weights comprises summing, for an additional slate action, slot-level density ratios for the plurality of digital slots to generate an additional importance weight.

4 . The computer-implemented method of claim 1 , wherein generating the slot-level density ratios comprises:

determining a first slot-level probability of selecting a first slot-level action utilizing the first slate recommendation policy; and

determining a second slot-level probability of selecting the first slot-level action utilizing the second slate recommendation policy.

5 . The computer-implemented method of claim 4 , wherein generating the slot-level density ratios comprises generating a first slot-level density ratio from the first slot-level probability and the second slot-level probability.

6 . The computer-implemented method of claim 4 , wherein receiving the historical slate data further comprises receiving a client device context analyzed by the first slate recommendation policy in selecting the slate actions, and further comprising:

determining, from the historical slate data, a client device context embedding from a plurality of client device context embeddings utilized to select the first slot-level action; and

determining the second slot-level probability of selecting the first slot-level action utilizing the second slate recommendation policy in light of the client device context embedding.

7 . The computer-implemented method of claim 6 , further comprising determining the first slot-level probability of selecting the first slot-level action utilizing the first slate recommendation policy in light of the client device context.

8 . The computer-implemented method of claim 7 , wherein generating the predicted reward distribution for the second slate recommendation policy comprises generating a cumulative distribution function by applying the plurality of importance weights to the observed rewards from the historical slate data.

9 . A system comprising:

one or more memory devices comprising historical slate data comprising observed rewards from selecting slate actions for a plurality of digital slots of a digital slate utilizing a first digital policy; and

one or more processors configured to cause the system to:

generate, for a second digital policy, a plurality of importance weights from the historical slate data by:

summing, for a first slate action, a first plurality of slot-level density ratios for the plurality of digital slots to generate a first importance weight; and

summing, for a second slate action, a second plurality of slot-level density ratios for the plurality of digital slots to generate a second importance weight; and

generate a predicted reward distribution for the second digital policy by applying the plurality of importance weights to the historical slate data.

10 . The system of claim 9 , wherein the one or more processors are configured to cause the system to generate the first plurality of slot-level density ratios by:

determining, for the first slate action, a first slot-level probability of selecting a first slot-level action for a first slot utilizing the first digital policy;

determining, for the first slate action, a second slot-level probability of selecting the first slot-level action for the first slot utilizing the second digital policy; and

generating a first slot-level density ratio from the first slot-level probability and the second slot-level probability.

11 . The system of claim 10 , wherein the one or more processors are configured to cause the system to generate the first plurality of slot-level density ratios by:

determining, for the first slate action, a third slot-level probability of selecting a second slot-level action for a second slot utilizing the first digital policy; and

determining, for the first slate action, a fourth slot-level probability of selecting the second slot-level action for the second slot utilizing the second digital policy.

12 . The system of claim 11 , wherein the one or more processors are configured to cause the system to generate the first plurality of slot-level density ratios by generating a second slot-level density ratio from the third slot-level probability and the fourth slot-level probability.

13 . The system of claim 12 , wherein the one or more processors are configured to cause the system to generate the first importance weight for the first slate action by summing the first slot-level density ratio and the second slot-level density ratio.

14 . The system of claim 10 , wherein the historical slate data comprises client device context data analyzed by the first digital policy in selecting the slate actions and wherein the one or more processors are further configured to cause the system to determine the first slot-level probability of selecting the first slot-level action utilizing the first digital policy in light of a first client device context from the client device context data.

15 . The system of claim 9 , wherein the one or more processors are configured to cause the system to generate the predicted reward distribution for the second digital policy by generating a cumulative distribution function from the plurality of importance weights and the observed rewards from the historical slate data.

16 . A non-transitory computer readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

receiving historical slate data comprising observed rewards from selecting slate actions for a plurality of slots of a digital slate utilizing a first digital policy;

generating, for a second digital policy, a plurality of importance weights from the historical slate data corresponding to the first digital policy by:

determining, for a slate action, a first slot-level density ratio between the first digital policy and the second digital policy for a first slot of the digital slate;

determining, for the slate action, a second slot-level density ratio between the first digital policy and the second digital policy for a second slot of the digital slate; and

summing the first slot-level density ratio and the second slot-level density ratio to determine an importance weight for the slate action; and

generating a predicted reward distribution for the second digital policy by applying the plurality of importance weights to the historical slate data.

17 . The non-transitory computer readable medium of claim 16 , wherein generating the plurality of importance weights comprises summing, for an additional slate action, a third slot-level density ratio for the first slot of the digital slate and a fourth slot-level density ratio for the second slot of the digital slate to generate an additional importance weight.

18 . The non-transitory computer readable medium of claim 16 , wherein determining the first slot-level density ratio between the first digital policy and the second digital policy for the first slot of the digital slate comprises determining a first slot-level probability of selecting a first slot-level action utilizing the first digital policy in light of a client device context.

19 . The non-transitory computer readable medium of claim 18 , wherein determining the first slot-level density ratio between the first digital policy and the second digital policy for the first slot of the digital slate comprises:

determining a second slot-level probability of selecting the first slot-level action utilizing the second digital policy in light of the client device context; and

determining the first slot-level density ratio by summing the first slot-level probability and the second slot-level probability.

20 . The non-transitory computer readable medium of claim 16 , wherein generating the predicted reward distribution for the second digital policy comprises generating at least one of a cumulative distribution function or a probability density function from the plurality of importance weights and the observed rewards.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2023
From: CHAUDHARI, SHREYAS; VLASSIS, NIKOLAOS; ARBOUR, DAVID
To: ADOBE INC.
Reel/Frame 065800/0965 →
AGREEMENT Recorded Dec 7, 2023
From: THEOCHAROUS, GEORGIOS
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 065820/0689 →
CHANGE OF NAME Recorded Dec 7, 2023
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 065834/0513 →
Continuity (1)
Related Publication 20250191047A1 · Jun 12, 2025
References Cited (35)
Audrey Huang, Liu Leqi, Zachary Lipton, and Kamyar Azizzadenesheli. 2021. Off-policy risk assessment in contextual bandits. Advances in Neural Information Processing Systems 34 (2021), 23714-23726. [cited by applicant]
Badrul Sarwar, George Karypis, Joseph Konstan, and John Riedl. 2000. Analysis of recommendation algorithms for e-commerce. In Proceedings of the 2nd ACM Conference on Electronic Commerce. 158-167. [cited by applicant]
Carlos A Gomez-Uribe and Neil Hunt. 2015. The netflix recommender system: Algorithms, business value, and innovation. ACM Transactions on Management Information Systems (TMIS) 6, 4 (2015), 1-19. [cited by applicant]
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender. 2005. Learning to rank using gradient descent. In Proceedings of the 22nd international conference on Machine learning… [cited by applicant]
Daniel G Horvitz and Donovan J Thompson. 1952. A generalization of sampling without replacement from a finite universe. Journal of the American statistical Association 47, 260 (1952), 663-685. [cited by applicant]
Daphne Koller and Nir Friedman. 2009. Probabilistic graphical models: principles and techniques. MIT press. [cited by applicant]
Dror G Feitelson, Eitan Frachtenberg, and Kent L Beck. 2013. Development and deployment at facebook. IEEE Internet Computing 17, 4 (2013), 8-17. [cited by applicant]
Ehtsham Elahi and Ashok Chandrashekar. 2020. Learning representations of hierarchical slates in collaborative filtering. In Proceedings of the 14th ACM Conference on Recommender Systems. 703-707. [cited by applicant]
F Maxwell Harper and Joseph A Konstan. 2015. The movielens datasets: History and context. Acm transactions on interactive intelligent systems (tiis) 5, 4 (2015), 1-19. [cited by applicant]
Guy Shani and Asela Gunawardana. 2011. Evaluating recommendation systems. In Recommender systems handbook. Springer, 257-297. [cited by applicant]
Harald Steck. 2019. Embarrassingly shallow autoencoders for sparse data. In The World Wide Web Conference. 3251-3257. [cited by applicant]
Haruka Kiyohara, Yuta Saito, Tatsuya Matsuhiro, Yusuke Narita, Nobuyuki Shimizu, and Yasuo Yamamoto. 2022. Doubly robust off-policy evaluation for ranking policies under the cascade behavior model. In Proceedings of the… [cited by applicant]
Jason M Altschuler, Victor-Emmanuel Brunel, and Alan Malek. 2019. Best Arm Identification for Contaminated Bandits. J. Mach. Learn. Res. 20, 91 (2019), 1-39. [cited by applicant]
Jesús Bobadilla, Fernando Ortega, Antonio Hernando, and Abraham Gutiérrez. 2013. Recommender systems survey. Knowledge-based systems 46 (2013), 109-132. [cited by applicant]
Jie Lu, Dianshuang Wu, Mingsong Mao, Wei Wang, and Guangquan Zhang. 2015. Recommender system application developments: a survey. Decision support systems 74 (2015), 12-32. [cited by applicant]
Julia L Wirch and Mary R Hardy. 2001. Distortion risk measures: Coherence and stochastic dominance. In International congress on insurance: Mathematics and economics. 15-17. [cited by applicant]
Kalervo Järvelin and Jaana Kekäläinen. 2017. IR evaluation methods for retrieving highly relevant documents. In ACM SIGIR Forum, vol. 51. ACM New York, NY, USA, 243-250. [cited by applicant]
Michael A Stephens. 1974. EDF statistics for goodness of fit and some comparisons. Journal of the American statistical Association 69, 347 (1974), 730-737. [cited by applicant]
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li. 2014. Doubly robust policy evaluation and optimization. Statist. Sci. 29, 4 (2014), 485-511. [cited by applicant]
Miroslav Dudík, John Langford, and Lihong Li. 2011. Doubly robust policy evaluation and learning. arXiv preprint arXiv: 1103.4601 (2011). [cited by applicant]
Nathan Kallus and Masatoshi Uehara. 2019. Intrinsically efficient, stable, and bounded off-policy evaluation for reinforcement learning. Advances in neural information processing systems 32 (2019). [cited by applicant]
Nicolo Cesa-Bianchi and Gábor Lugosi. 2012. Combinatorial bandits. J. Comput. System Sci. 78, 5 (2012), 1404-1422. [cited by applicant]
Nikos Vlassis, Ashok Chandrashekar, Fernando Amat, and Nathan Kallus. 2021. Control variates for slate off-policy evaluation. Advances in Neural Information Processing Systems 34 (2021). [cited by applicant]
Philip S Thomas. 2015. Safe reinforcement learning. (2015). [cited by applicant]
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh. 2015. High-confidence off-policy evaluation. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 29. [cited by applicant]
Pranab K Sen and Julio M Singer. 1994. Large sample methods in statistics: an introduction with applications. vol. 25. CRC press. [cited by applicant]
R Tyrrell Rockafellar, Stanislav Uryasev, et al. 2000. Optimization of conditional value-at-risk. Journal of risk 2 (2000), 21-42. [cited by applicant]
Ramtin Keramati, Christoph Dann, Alex Tamkin, and Emma Brunskill. 2020. Being optimistic to be conservative: Quickly learning a cvar policy. In Proceedings of the AAAI conference on artificial intelligence, vol. 34. 443… [cited by applicant]
Richard S Sutton and Andrew G Barto. 2018. Reinforcement learning: An introduction. MIT press. [cited by applicant]
Ron Kohavi and Roger Longbotham. 2017. Online Controlled Experiments and A/B Testing. Encyclopedia of machine learning and data mining 7, 8 (2017), 922-929. [cited by applicant]
Shuai Li, Yasin Abbasi-Yadkori, Branislav Kveton, Shan Muthukrishnan, Vishwa Vinay, and Zheng Wen. 2018. Offline evaluation of ranking policies with click models. In Proceedings of the 24th ACM SIGKDD International Conf… [cited by applicant]
Yash Chandak, Georgios Theocharous, James Kostas, Scott Jordan, and Philip Thomas. 2019. Learning action representations for reinforcement learning. In International conference on machine learning. PMLR, 941-950. [cited by applicant]
Yash Chandak, Scott Niekum, Bruno da Silva, Erik Learned-Miller, Emma Brunskill, and Philip S Thomas. 2021. Universal off-policy evaluation. Advances in Neural Information Processing Systems 34 (2021). [cited by applicant]
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik. 2017. Optimal and adaptive off-policy evaluation in contextual bandits. In International Conference on Machine Learning. PMLR, 3589-3597. [cited by applicant]
Yuta Saito, Aihara Shunsuke, Matsutani Megumi, and Narita Yusuke. 2020. Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation. arXiv preprint arXiv:2008.07146 (2020). [cited by applicant]