Dynamically Personalized Product Recommendation Engine Using Stochastic and Adversarial Bandits
A method for recommending products to a user includes providing a user profile with product related data. At least one bandit is generated to model product related recommendations. The bandit model(s) are passed to a recommendation module that provides recommendations to the user based on the bandit model and expected payoff. User interactions in response to the recommendation can be evaluated to adjust further recommendations.
1 . A method for recommending products to a user, the method comprising the steps of:
providing a user profile with product related data;
generating at least one bandit to model product related recommendations;
passing the bandit model to a recommendation module that provides recommendations to the user based on the bandit model and expected payoff; and
evaluating user interactions in response to the recommendation to adjust further recommendations.
2 . The method of claim 1 , wherein the user profile data is derived at least partially from at least one of product related user data and traffic-based link data.
3 . The method of claim 1 , wherein the bandit is an adversarial bandit.
4 . The method of claim 1 , wherein the bandit is an adaptive adversarial bandit.
5 . The method of claim 1 , wherein the bandit is an stationary adversarial bandit.
6 . The method of claim 1 , wherein the bandit is a federation bandit.
7 . The method of claim 1 , wherein the bandit is a tuning bandit.
8 . The method of claim 1 , wherein the bandit uses a reward functions based on reciprocal rank.
9 . The method of claim 1 , wherein the bandit uses a reward functions based on similarity score.
10 . The method of claim 1 , wherein the recommendation module provides dynamic personalization.
11 . A method for dynamically recommending products to a user, the method comprising the steps of:
receiving a request for a personal recommendation;
weighting a bandit payoff;
assembling bandit recommendations;
providing recommendations to the user; and
evaluating further user interactions in response to the provided recommendation to adjust weighting of the bandit payoff.