IP Library Granted Patent US 10,706,454
Granted Patent B2
US 10,706,454 · App. 15/943,807 · Granted Jul 7, 2020

Method, medium, and system for training and utilizing item-level importance sampling models

Inventors: Shuai Li (Sha Tin, HK); Zheng Wen (Fremont, CA); Yasin Abbasi Yadkori (San Francisco, CA); Vishwa Vinay (Karnataka, IN); Branislav Kveton (San Jose, CA)
Assignee: ADOBE INC.
G06Q30/0631G06N20/00G06Q30/0253G06Q30/0633
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,706,454
App. No.
15/943,807
Granted
Jul 7, 2020
Kind
B2
Abstract

The present disclosure is directed toward systems, methods, and computer readable media for training and utilizing an item-level importance sampling model to evaluate and execute digital content selection policies. For example, systems described herein include training and utilizing an item-level importance sampling model that accurately and efficiently predicts a performance value that indicates a probability that a target user will interact with ranked lists of digital content items provided in accordance with a target digital content selection policy. Specifically, systems described herein can perform an offline evaluation of a target policy in light of historical user interactions corresponding to a training digital content selection policy to determine item-level importance weights that account for differences in digital content item distributions between the training policy and the target policy. In addition, the systems described herein can apply the item-level importance weights to training data to train item-level importance sampling model.

Claims (51)

1. In a digital medium environment for selecting and providing digital item lists to client devices in accordance with content selection policies, a computer-implemented method of training a policy model for offline evaluation and execution of a content selection policy, the method comprising:

identifying digital interactions by users with respect to digital content items, the digital content items selected and presented as part of training item lists to computing devices of the users in accordance with a training digital content selection policy;

training an item-level importance sampling model to predict a performance value indicating a probability that a user will interact with one or more items from a target item list selected in accordance with a target digital content selection policy by:

for a first digital content item of the digital content items, determining an item-level importance weight based on the training digital content selection policy and the target digital content selection policy; and

applying the item-level importance weight to a first interaction from the identified digital interactions by a first computing device of a first user with the first digital content item from a first item list to train the item-level importance sampling model;

training the item-level importance sampling model to predict a second performance value indicating a second probability that a user will interact with one or more items from a second target item list selected in accordance with a second target digital content selection policy; and

comparing the performance value associated with the target digital content selection policy with the second performance value associated with the second target digital content selection policy.

2. The method of claim 1 , further comprising based on comparing the performance value associated with the target digital content selection policy with the second performance value associated with the second target digital content selection policy, executing the target digital content selection policy by generating and providing an item list to a client device.

3. The method of claim 1 , wherein executing the target digital content selection policy comprises conducting an online A/B hypothesis test between the target digital content selection policy and a second target digital content selection policy.

4. The method of claim 1 , wherein the training item lists generated in accordance with the training digital content selection policy differ from target item lists generated in accordance with the target digital content selection policy.

5. A non-transitory computer readable medium storing instructions thereon that, when executed by at least one processor, cause a computer system to:

identify digital interactions by users with respect to digital content items, the digital content items selected and presented as part of training item lists to computing devices of the users in accordance with a training digital content selection policy; and

train an item-level importance sampling model to predict a performance value indicating a probability that a user will interact with one or more items from a target item list selected in accordance with a target digital content selection policy by:

for a first digital content item of the digital content items, determining an item-level importance weight based on the training digital content selection policy and the target digital content selection policy; and

applying the item-level importance weight to a first interaction from the identified digital interactions by a first computing device of a first user with the first digital content item from a first item list to train the item-level importance sampling model;

train the item-level importance sampling model to predict a second performance value indicating a second probability that a user will interact with one or more items from a second target item list selected in accordance with a second target digital content selection policy; and

compare the performance value associated with the target digital content selection policy with the second performance value associated with the second target digital content selection policy.

6. The non-transitory computer readable medium of claim 5 , wherein the instructions further cause the computer system to select the target digital content selection policy for execution by comparing the performance value associated with the target digital content selection policy with the second performance value associated with the second target digital content selection policy.

7. The non-transitory computer readable medium of claim 6 , wherein the instructions cause the computer system to execute the target digital content selection policy by:

selecting a digital content item list utilizing the target digital content selection policy; and

providing the digital content item list for display to a client device.

8. The non-transitory computer readable medium of claim 6 , wherein executing the target digital content selection policy comprises conducting an online A/B hypothesis test between the target digital content selection policy and another digital content selection policy.

9. The non-transitory computer readable medium of claim 5 , wherein determining the item-level importance weight comprises:

determining a first training item-level selection probability for the first digital content item based on the training digital content selection policy and a first target item-level selection probability for the first digital content item based on the target digital content selection policy; and

determining the item-level importance weight based on the first training item-level selection probability and the first target item-level selection probability.

10. The non-transitory computer readable medium of claim 5 , wherein training item lists generated in accordance the training digital content selection policy differ from target item lists generated in accordance with the target digital content selection policy.

11. The non-transitory computer readable medium of claim 5 , wherein training the item-level importance sampling model further comprises:

for a second digital content item of the digital content items, determining an additional item-level importance weight based on the training digital content selection policy and the target digital content selection policy; and

applying the additional item-level importance weight to a second interaction from the identified digital interactions with a second digital content item from the digital content items.

12. The non-transitory computer readable medium of claim 11 , wherein training the item-level importance sampling model further comprises:

comparing the additional item-level importance weight to a clipping threshold; and

based on the comparison between the additional item-level importance weight and the clipping threshold, applying the clipping threshold to the second interaction with the second digital content item to further train the item-level importance sampling model.

13. The non-transitory computer readable medium of claim 5 , wherein training the item-level importance sampling model comprises training an item-position importance sampling model that generates the performance value based on both a position of the first digital content item within the first item list and content of the first digital content item.

14. The non-transitory computer readable medium of claim 5 , wherein training the item-level importance sampling model comprises training a content-based importance sampling model that generates the performance value based on content of the first digital content item and independent of position of the first digital content item within the first item list.

15. A system comprising:

at least one processor;

a non-transitory computer readable medium comprising instructions that, when executed by the at least one processor, cause the system to:

identify digital interactions by users with respect to digital content items, the digital content items selected and presented as part of training item lists to computing devices of the users in accordance with a training digital content selection policy;

train an item-level importance sampling model to generate a performance value that indicates a probability that a user will interact with one or more items from a target item list selected in accordance with a target digital content selection policy by:

for a first digital content item of the digital content items, determining an item-level importance weight based on the training digital content selection policy and the target digital content selection policy; and

applying the item-level importance weight to a first interaction from the digital interactions by a first computing device of a first user with the first digital content item from a first item list train the item-level importance sampling model;

train the item-level importance sampling model to predict a second performance value indicating a second probability that a user will interact with one or more items from a second target item list selected in accordance with a second target digital content selection policy; and

compare the performance value associated with the target digital content selection policy with the second performance value associated with the second target digital content selection policy.

16. The system of claim 15 , wherein the instructions cause the system to:

select the target digital content selection policy for execution by comparing the performance value associated with the target digital content selection policy with the second performance value associated with the second target digital content selection policy.

17. The system of claim 15 , further comprising instructions that, when executed by the at least one processor, cause the system to train the item-level importance sampling model by training an item-position importance sampling model that generates a performance value indicating a probability of interaction with the first digital content item based on both a position of the first digital content item within the first item list and content of the first digital content item.

18. The system of claim 15 , further comprising instructions that, when executed by the at least one processor, cause the system to train the item-level importance sampling model by training a content-based importance sampling model that generates a performance value indicating a probability of interaction with the first digital content item based on content of the first digital content item and independent of position of the first digital content item within the first item list.

19. The system of claim 15 , wherein the instructions cause the system to estimate the training digital content selection policy based on the training item lists presented to computing devices of training users.

20. The system of claim 15 ,

wherein the training digital content selection policy comprises a distribution of training item lists in accordance with a plurality of contexts; and

wherein the instructions further cause the system to train the item-level importance sampling model to predict a probability that a particular user will interact with a set of items from a target digital content item list selected in accordance with the target digital content selection policy based on a context associated with the particular user.

Assignments (2)
CHANGE OF NAME Recorded Nov 30, 2018
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 047688/0635 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2018
From: LI, SHUAI; WEN, ZHENG; YADKORI, YASIN ABBASI; VINAY, VISHWA; KVETON, BRANISLAV
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 045421/0380 →