IP Library › Granted Patent US 12,586,114
Granted Patent B2
US 12,586,114 · App. 17/367,134 · Granted Mar 24, 2026

Generating digital recommendations utilizing collaborative filtering, reinforcement learning, and inclusive sets of negative feedback

Inventors: Saayan Mitra (San Jose, CA); Xiang Chen (Palo Alto, CA); Vahid Azizi (Piscataway, PA)
Assignee: Adobe Inc.
G06Q30/0631G06F18/2178G06N3/044G06N3/088G06Q30/0202
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,114
App. No.
17/367,134
Granted
Mar 24, 2026
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer readable media that utilize collaborative filtering and a reinforcement learning model having an actor-critic framework to provide digital content items across client devices. In particular, in one or more embodiments, the disclosed systems monitor interactions of a client device with one or more digital content items to generate item embeddings (e.g., utilizing a collaborative filtering model). The disclosed systems further utilize a reinforcement learning model to generate a recommendation (e.g., determine one or more additional digital content items to provide to the client device) based on the user interactions. In some implementations, the disclosed systems utilize the reinforcement learning model to analyze every negative and positive interaction observed when generating the recommendation. Further, the disclosed systems utilize the reinforcement learning model to analyze item embeddings, which encode the relationships among the digital content items, when generating the recommendation.

Claims (83)

1 . A non-transitory computer-readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

generating, for a plurality of digital content items, a set of item embeddings that encode interactions across client devices associated with the plurality of digital content items;

monitoring user interactions of a client device with one or more digital content items from the plurality of digital content items during an interaction session;

determining, utilizing the set of item embeddings, a negative interaction map and a positive interaction map from the user interactions of the client device during the interaction session by:

determining the positive interaction map or the negative interaction map using, from the set of item embeddings, item embeddings that correspond to the user interactions of the client device during the interaction session; and

determining at least one of the positive interaction map or the negative interaction map using, from the set of item embeddings, additional item embeddings selected based on distances within an item embedding space between the additional item embeddings and the item embeddings that correspond to the user interactions, the additional item embeddings corresponding to digital content items with which the client device did not interact during the interaction session;

determining, utilizing a reinforcement learning model, one or more additional digital content items from the plurality of digital content items to provide for display by:

generating, utilizing a first convolutional gated recurrent unit neural network layer of the reinforcement learning model, a negative state for the client device based on the negative interaction map;

generating, utilizing a second convolutional gated recurrent unit neural network layer of the reinforcement learning model, a positive state for the client device based on the positive interaction map; and

generating a recommendation for the one or more additional digital content items based on the set of item embeddings, the negative state, and the positive state;

determining, using one or more rectified linear interaction neural network layers of the reinforcement learning model, a value function indicating a measure of quality of the one or more additional digital content items; and

modifying parameters of the reinforcement learning model using the value function.

2 . The non-transitory computer-readable medium of claim 1 , wherein:

determining the positive interaction map or the negative interaction map using the item embeddings that correspond to the user interactions comprises determining, for the negative interaction map and from the set of item embeddings, one or more negative item embeddings by determining an item embedding for each negative interaction from the user interactions; and

determining at least one of the positive interaction map or the negative interaction map using the additional item embeddings selected based on the distances within the item embedding space comprises determining, for the negative interaction map and from the set of item embeddings, one or more additional negative item embeddings based on a proximity to the one or more negative item embeddings within the item embedding space.

3 . The non-transitory computer-readable medium of claim 1 , wherein modifying the parameters of the reinforcement learning model using the value function comprises:

modifying, using the value function, a first set of parameters for the first convolutional gated recurrent unit neural network layer; and

modifying, using the value function, a second set of parameters for the second convolutional gated recurrent unit neural network layer.

4 . The non-transitory computer-readable medium of claim 3 , wherein:

modifying, using the value function, the first set of parameters for the first convolutional gated recurrent unit neural network layer comprises back propagating the value function to the first convolutional gated recurrent unit neural network layer; and

modifying, using the value function, the second set of parameters for the second convolutional gated recurrent unit neural network layer comprises back propagating the value function to the second convolutional gated recurrent unit neural network layer.

5 . The non-transitory computer-readable medium of claim 1 , wherein determining, utilizing the reinforcement learning model, the one or more additional digital content items based on the negative state, the positive state, and the set of item embeddings comprises:

generating a first similarity metric between the positive state and the set of item embeddings;

generating a second similarity metric between the negative state and the set of item embeddings; and

determining the one or more additional digital content items utilizing the first similarity metric and the second similarity metric.

6 . The non-transitory computer-readable medium of claim 1 , wherein generating the set of item embeddings for the plurality of digital content items comprises generating the set of item embeddings via collaborative filtering or graph embedding to encode the interactions across the client devices associated with the plurality of digital content items.

7 . The non-transitory computer-readable medium of claim 1 , wherein:

determining the positive interaction map or the negative interaction map using the item embeddings that correspond to the user interactions comprises determining, for the negative interaction map and from the set of item embeddings, one or more negative item embeddings by determining an item embedding for each negative interaction from the user interactions; and

determining at least one of the positive interaction map or the negative interaction map using the additional item embeddings selected based on the distances within the item embedding space comprises determining, for the positive interaction map and from the set of item embeddings, one or more positive item embeddings based on a distance from the one or more negative item embeddings within the item embedding space.

8 . The non-transitory computer-readable medium of claim 1 , wherein determining the negative interaction map from the user interactions of the client device during the interaction session comprises determining the negative interaction map without sampling a subset of negative interactions from the user interactions.

9 . The non-transitory computer-readable medium of claim 1 , wherein:

generating, utilizing the first convolutional gated recurrent unit neural network layer, the negative state for the client device based on the negative interaction map comprises generating, using the first convolutional gated recurrent unit neural network layer having a set of shared weights, the negative state for the client device based on the negative interaction map and a previous negative state corresponding to a previous interaction session of the client device; and

generating, utilizing the second convolutional gated recurrent unit neural network layer, the positive state for the client device based on the positive interaction map comprises generating, using the second convolutional gated recurrent unit neural network layer having the set of shared weights, the positive state for the client device based on the positive interaction map and a previous positive state corresponding to the previous interaction session of the client device.

10 . The non-transitory computer-readable medium of claim 9 , wherein determining at least one of the positive interaction map or the negative interaction map using the additional item embeddings corresponding to the digital content items with which the client device did not interact during the interaction session comprises using the additional item embeddings to determine the positive interaction map and the negative interaction map.

11 . The non-transitory computer-readable medium of claim 1 , wherein:

determining the positive interaction map or the negative interaction map using the item embeddings that correspond to the user interactions comprises determining, for the positive interaction map and from the set of item embeddings, one or more positive item embeddings by determining an item embedding for each positive interaction from the user interactions; and

determining at least one of the positive interaction map or the negative interaction map using the additional item embeddings selected based on the distances within the item embedding space comprises determining, for the positive interaction map and from the set of item embeddings, one or more additional positive item embeddings based on a proximity to the one or more positive item embeddings within the item embedding space.

12 . A computer-implemented method comprising:

monitoring user interactions of a client device with one or more digital content items from a plurality of digital content items during an interaction session;

determining, utilizing a set of item embeddings that encode interactions with the plurality of digital content items across client devices, a negative interaction map and a positive interaction map from the user interactions of the client device during the interaction session by:

determining the positive interaction map or the negative interaction map using, from the set of item embeddings, item embeddings that correspond to the user interactions of the client device during the interaction session; and

determining at least one of the positive interaction map or the negative interaction map using, from the set of item embeddings, additional item embeddings selected based on distances within an item embedding space between the additional item embeddings and the item embeddings that correspond to the user interactions, the additional item embeddings corresponding to digital content items with which the client device did not interact during the interaction session;

generating, utilizing a first convolutional gated recurrent unit neural network layer of a reinforcement learning model, a negative state for the client device based on the negative interaction map;

generating, utilizing a second convolutional gated recurrent unit neural network layer of the reinforcement learning model, a positive state for the client device based on the positive interaction map;

determining, utilizing the reinforcement learning model, one or more additional digital content items from the plurality of digital content items based on the set of item embeddings, the negative state, and the positive state;

providing the one or more additional digital content items for display via the client device;

determining, using one or more rectified linear interaction neural network layers of the reinforcement learning model, a value function indicating a measure of quality of the one or more additional digital content items provided for display; and

modifying parameters of the reinforcement learning model using the value function.

13 . The computer-implemented method of claim 12 , further comprising:

determining a previous negative state for the client device and a previous positive state for the client device,

wherein generating, utilizing the first convolutional gated recurrent unit neural network layer, the negative state for the client device based on the negative interaction map comprises generating, utilizing the first convolutional gated recurrent unit neural network layer, the negative state based on the negative interaction map and the previous negative state; and

wherein generating, utilizing the second convolutional gated recurrent unit neural network layer, the positive state for the client device based on the positive interaction map comprises generating, utilizing the second convolutional gated recurrent unit neural network layer, the positive state based on the positive interaction map and the previous positive state.

14 . The computer-implemented method of claim 12 , further comprising generating the set of item embeddings utilizing a factorization machine neural network.

15 . The computer-implemented method of claim 12 , wherein:

determining the positive interaction map or the negative interaction map using the item embeddings that correspond to the user interactions comprises determining, for the positive interaction map and from the set of item embeddings, one or more positive item embeddings by determining an item embedding for each positive interaction from the user interactions; and

determining at least one of the positive interaction map or the negative interaction map using the additional item embeddings selected based on the distances within the item embedding space comprises determining, for the negative interaction map and from the set of item embeddings, one or more negative item embeddings based on a distance from the one or more positive item embeddings within an item embedding space.

16 . A system comprising:

one or more memory devices; and

one or more server devices configured to cause the system to:

generate interaction maps from user interactions of a client device with one or more digital content items from a plurality of digital content items during an interaction session by:

determining a positive interaction map or a negative interaction map using item embeddings that correspond to the user interactions of the client device during the interaction session; and

determining at least one of the positive interaction map or the negative interaction map using additional item embeddings selected based on distances within an item embedding space between the additional item embeddings and the item embeddings that correspond to the user interactions, the additional item embeddings corresponding to digital content items with which the client device did not interact during the interaction session;

determine, utilizing an actor model of a reinforcement learning model and based on the interaction maps, one or more additional digital content items from the plurality of digital content items to provide for display via the client device by:

generating, utilizing a first convolutional gated recurrent unit neural network layer of the actor model, a negative state for the client device based on the interaction maps; and

generating, utilizing a second convolutional gated recurrent unit neural network layer of the actor model, a positive state for the client device based on the interaction maps;

determining to provide the one or more additional digital content items for display via the client device based on the positive state and the negative state; and

generate, utilizing a critic model of the reinforcement learning model, a value function to modify parameters of the actor model by:

determining a first state-action vector utilizing the positive state and a second state-action vector utilizing the negative state; and

generating, utilizing one or more rectified linear interaction neural network layers of the critic model, the value function based on the first state-action vector and the second state-action vector.

17 . The system of claim 16 , wherein generating, utilizing the one or more rectified linear interaction neural network layers, the value function based on the first state-action vector and the second state-action vector comprises:

determining, utilizing a first rectified linear interaction neural network layer, a first set of feature values based on the first state-action vector;

determining, utilizing a second rectified linear interaction neural network layer, a second set of feature values based on the second state-action vector; and

generating, utilizing a third rectified linear interaction neural network layer, the value function based on the first set of feature values and the second set of feature values.

18 . The system of claim 16 , wherein the one or more server devices are further configured to cause the system to:

monitor additional user interactions of the client device with the one or more additional digital content items during an additional interaction session; and

determine, utilizing the actor model having the modified parameters, one or more other digital content items from the plurality of digital content items to provide for display based on the additional user interactions.

19 . The system of claim 16 , wherein determining the first state-action vector utilizing the positive state and the second state-action vector utilizing the negative state comprises:

determining the first state-action vector utilizing the positive state and the one or more additional digital content items; and

determining the second state-action vector utilizing the negative state and the one or more additional digital content items.

20 . The system of claim 19 , wherein:

the one or more server devices are further configured to cause the system to generate an action vector utilizing item embeddings corresponding to the one or more additional digital content items;

determining the first state-action vector utilizing the positive state and the one or more additional digital content items comprises determining the first state-action vector by combining the action vector and the positive state; and

determining the second state-action vector utilizing the negative state and the one or more additional digital content items comprises determining the second state-action vector by combining the action vector and the negative state.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2021
From: MITRA, SAAYAN; CHEN, XIANG; AZIZI, VAHID
To: ADOBE INC.
Reel/Frame 056748/0225 →
Continuity (1)
Related Publication 20230022396A1 · Jan 26, 2023
References Cited (71)
US 10929743B2 · Liu · 2021 [cited by examiner]
US 20140122502A1 · Kalmes · 2014 [cited by examiner]
US 20170140263A1 · Kaiser · 2017 [cited by examiner]
US 20180089553A1 · Liu et al. · 2018 [cited by applicant]
US 20180240030A1 · Shen · 2018 [cited by examiner]
US 20210082471A1 · Detroja · 2021 [cited by examiner]
US 20220270155A1 · Volkovs · 2022 [cited by examiner]
Zhao, Xiangyu, et al. “Deep reinforcement learning for page-wise recommendations.” (Year: 2018). [cited by examiner]
Han, Jianhua, et al. “Optimizing ranking algorithm in recommender system via deep reinforcement learning.” (Year: 2019). [cited by examiner]
He, Xiangnan, et al. “Fast matrix factorization for online recommendation with implicit feedback.” (Year: 2016). [cited by examiner]
Wang, Xiang, et al. “Reinforced negative sampling over knowledge graph for recommendation.” (Year: 2020). [cited by examiner]
Wang, Wei, and Longbing Cao. “Interactive sequential basket recommendation by learning basket couplings and positive/negative feedback.” (Year: 2021). [cited by examiner]
Liu, Feng, et al. “Deep reinforcement learning based recommendation with explicit user-item interactions modeling.” (Year: 2018). [cited by examiner]
Wang, et al. Learning performance prediction via convolutional GRU and explainable neural networks in e-learning environments (Year: 2019). [cited by examiner]
Zhao, Xiangyu, et al. “Recommendations with negative feedback via pairwise deep reinforcement learning.” (Year: 2018). [cited by examiner]
Wang, et al., “Interactive sequential basket recommendation by learning basket couplings and positive/negative feedback” (Year: 2021). [cited by examiner]
Rendle, Steffen, et al. “BPR: Bayesian personalized ranking from implicit feedback” (Year: 2012). [cited by examiner]
Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, and Anil Anthony Bharath. 2017. Deep reinforcement learning: A brief survey. IEEE Signal Processing Magazine 34, 6 (2017), 26-38. [cited by applicant]
Nicolas Ballas, Li Yao, Chris Pal, and Aaron Courville. 2015. Delving deeper into convolutional networks for learning video representations. arXiv preprint arXiv:1511.06432 (2015). [cited by applicant]
Yoshua Bengio, Patrice Simard, and Paolo Frasconi. 1994. Learning long-term dependencies with gradient descent is difficult. IEEE transactions on neural networks 5, 2 (1994), 157-166. [cited by applicant]
José Bento, Stratis Ioannidis, S Muthukrishnan, and Jinyun Yan. 2013. A time and space efficient algorithm for Contextual Linear Bandits. In Joint European Conference on Machine Learning and Knowledge Discovery in Datab… [cited by applicant]
Xiaocong Chen, Chaoran Huang, Lina Yao, Xianzhi Wang, Wei Liu, and Wenjie Zhang. 2020. Knowledge-guided Deep Reinforcement Learning for Interactive Recommendation. arXiv:2004.08068 [cs.IR]. [cited by applicant]
Xu Chen, Hongteng Xu, Yongfeng Zhang, Jiaxi Tang, Yixin Cao, Zheng Qin, and Hongyuan Zha. 2018. Sequential Recommendation with User Memory Networks. 108-116. [cited by applicant]
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al. 2016. Wide & deep learning for recommender systems. In Proceedings o… [cited by applicant]
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire. 2011. Contextual bandits with linear payoff functions. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. JMLR Works… [cited by applicant]
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555 (2014). [cited by applicant]
Mukund Deshpande and George Karypis. 2004. Item-based top-n recommendation algorithms. ACM Transactions on Information Systems (TOIS) 22, 1 (2004), 143-177. [cited by applicant]
Georges E Dupret and Benjamin Piwowarski. 2008. A user browsing model to predict search engine click data from past observations.. In Proceedings of the 31st annual international ACM SIGIR conference on Research and dev… [cited by applicant]
Scott Fujimoto, Herke van Hoof, and David Meger. 2018. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning. PMLR, 1582-1591. [cited by applicant]
Kun Gai, Xiaoqiang Zhu, Han Li, Kai Liu, and Zhe Wang. 2017. Learning piece-wise linear models from large scale data for ad click prediction. arXiv preprint arXiv:1704.05194 (2017). [cited by applicant]
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins. 1999. Learning to forget: Continual prediction with LSTM. (1999). [cited by applicant]
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction. arXiv preprint arXiv:1703.04247 (2017). [cited by applicant]
F. Maxwell Harper and Joseph A. Konstan. 2015. The MovieLens Datasets: History and Context. In ACM Transactions on Interactive Intelligent Systems (TiiS), vol. 5. [cited by applicant]
Ruining He and Julian McAuley. 2016. Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering. In proceedings of the 25th international conference on world wide web. 507-517. [cited by applicant]
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. Neural collaborative filtering. In Proceedings of the 26th international conference on world wide web. 173-182. [cited by applicant]
Xiangnan He, Hanwang Zhang, Min-Yen Kan, and Tat-Seng Chua. 2016. Fast matrix factorization for online recommendation with implicit feedback. In Proceedings of the 39th International ACM SIGIR conference on Research and… [cited by applicant]
Yujing Hu, Qing Da, Anxiang Zeng, Yang Yu, and Yinghui Xu. 2018. Reinforcement learning to rank in e-commerce search engine: Formalization, analysis, and application. In Proceedings of the 24th ACM SIGKDD International … [cited by applicant]
Yuchin Juan, Yong Zhuang, Wei-Sheng Chin, and Chih-Jen Lin. 2016. Field-aware factorization machines for CTR prediction. In Proceedings of the 10th ACM conference on recommender systems. 43-50. [cited by applicant]
Vijay R Konda and John N Tsitsiklis. 2003. Onactor-critic algorithms. SIAM journal on Control and Optimization 42, 4 (2003), 1143-1166. [cited by applicant]
Joseph A Konstan, Bradley N Miller, David Maltz, Jonathan L Herlocker, Lee R Gordon, and John Riedl. 1997. Grouplens: Applying collaborative filtering to usenet news. Commun. ACM 40, 3 (1997), 77-87. [cited by applicant]
Yehuda Koren, Robert Bell, and Chris Volinsky. 2009. Matrix factorization techniques for recommender systems. Computer 42, 8 (2009), 30-37. [cited by applicant]
Lihong Li, Wei Chu, John Langford, and Robert E Schapire. 2010. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th international conference on World wide web. 661-670. [cited by applicant]
Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xdeepfm: Combining explicit and implicit feature interactions for recommender systems. In Proceedings of the 24th ACM SIGKDD… [cited by applicant]
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 … [cited by applicant]
Greg Linden, Brent Smith, and Jeremy York. 2003. Amazon.com recommendations: Item-to-item collaborative filtering. IEEE Internet computing 7, 1 (2003), 76-80. [cited by applicant]
Feng Liu, Ruiming Tang, Xutao Li, Yunming Ye, Haokun Chen, Huifeng Guo, and Yuzhou Zhang. 2018. Deep Reinforcement Learning based Recommendation with Explicit User-Item Interactions Modeling. CoRR abs/1810.12027 (2018).… [cited by applicant]
Hao Ma, Irwin King, and Michael R Lyu. 2007. Effective missing data prediction for collaborative filtering. In Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information… [cited by applicant]
Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton Van Den Hengel. 2015. Image-based recommendations on styles and substitutes. In Proceedings of the 38th international ACM SIGIR conference on research and deve… [cited by applicant]
H Brendan McMahan, Gary Holt, David Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al. 2013. Ad click prediction: a view from the trenches. In Proceedings… [cited by applicant]
Raymond J Mooney and Loriene Roy. 2000. Content-based book recommending using learning for text categorization. In Proceedings of the fifth ACM conference on Digital libraries. 195-204. [cited by applicant]
Steffen Rendle. 2010. Factorization machines. In 2010 IEEE International Conference on Data Mining. IEEE, 995-1000. [cited by applicant]
Steffen Rendle. 2012. Factorization machines with libfm. ACM Transactions on Intelligent Systems and Technology (TIST) 3, 3 (2012), 1-22. [cited by applicant]
Badrul Sarwar, George Karypis, Joseph Konstan, and John Riedl. 2001. Item-based collaborative filtering recommendation algorithms. In Proceedings of the 10th international conference on World Wide Web. 285-295. [cited by applicant]
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. 2015. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438 (2015). [cited by applicant]
Hinrich Schütze, Christopher D Manning, and Prabhakar Raghavan. 2008. Introduction to information retrieval. Vol. 39. Cambridge University Press Cambridge. [cited by applicant]
Guy Shani, David Heckerman, Ronen I Brafman, and Craig Boutilier. 2005. An MDP-based recommender system. Journal of Machine Learning Research 6, 9 (2005). [cited by applicant]
Yue Shi, Alexandros Karatzoglou, Linas Baltrunas, Martha Larson, Alan Hanjalic, and Nuria Oliver. 2012. Tfmap: optimizing map for top-n context-aware recommendation. In Proceedings of the 35th international ACM SIGIR co… [cited by applicant]
Xiaoyuan Su and Taghi M Khoshgoftaar. 2009. A survey of collaborative filtering techniques. Advances in artificial intelligence 2009 (2009). [cited by applicant]
Richard S Sutton and Andrew G Barto. 2018. Reinforcement learning: An introduction. MIT press. [cited by applicant]
Nima Taghipour and Ahmad Kardan. 2008. A hybrid web recommender system based on q-learning. In Proceedings of the 2008 ACM symposium on Applied computing. 1164-1168. [cited by applicant]
Huazheng Wang, Qingyun Wu, and Hongning Wang. 2016. Learning hidden features for contextual bandits. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management. 1633-1642. [cited by applicant]
Jun Wang, Arjen P De Vries, and Marcel JT Reinders. 2006. Unifying userbased and item-based collaborative filtering approaches by similarity fusion. In Proceedings of the 29th annual international ACM SIGIR conference o… [cited by applicant]
Pengfei Wang, Yu Fan, Long Xia, Wayne Zhao, Shaozhang Niu, and Jimmy Huang. 2020. KERL: A Knowledge-Guided Reinforcement Learning Model for Sequential Recommendation. 209-218. https://doi.org/10.1145/3397271.3401134. [cited by applicant]
Xinxi Wang, Yi Wang, David Hsu, and Ye Wang. 2014. Exploration in interactive personalized music recommendation: a reinforcement learning approach. ACM Transactions on Multimedia Computing, Communications, and Applicati… [cited by applicant]
YiningWang, LiweiWang, Yuanzhi Li, Di He, and Tie-Yan Liu. 2013. A theoretical analysis of NDCG type ranking measures. In Conference on Learning Theory. 25-54. [cited by applicant]
Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay. 2019. Deep Learning Based Recommender System: A Survey and New Perspectives. ACM Comput. Surv. 52, 1, Article 5 (Feb. 2019), 38 pages. [cited by applicant]
Weinan Zhang, Tianqi Chen, Jun Wang, and Yong Yu. 2013. Optimizing top-n collaborative filtering via dynamic negative item sampling. In Proceedings of the 36th international ACM SIGIR conference on Research and developm… [cited by applicant]
Xiangyu Zhao, Long Xia, Liang Zhang, Zhuoye Ding, Dawei Yin, and Jiliang Tang. 2018. Deep Reinforcement Learning for Page-wise Recommendations. RecSys (2018). [cited by applicant]
Xiangyu Zhao, Liang Zhang, Zhuoye Ding, Long Xia, Jiliang Tang, and Dawei Yin. 2018. Recommendations with negative feedback via pairwise deep reinforcement learning. In Proceedings of the 24th ACM SIGKDD International C… [cited by applicant]
Guanjie Zheng, Fuzheng Zhang, Zihan Zheng, Yang Xiang, Nicholas Yuan, Xing Xie, and Zhenhui Li. 2018. DRN: A Deep Reinforcement Learning Framework for News Recommendation. WWW '18: Proceedings of the 2018 World Wide Web… [cited by applicant]
Office Action as received in Chinese Application No. 202210700134.9 dated Dec. 24, 2025. [cited by applicant]