IP Library Patent Application 14080038
Patent Application
App. No. 14/080,038

TESTING A MARKETING STRATEGY OFFLINE USING AN APPROXIMATE SIMULATOR

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/080,038
Abstract

In various example embodiments, a system and method for testing marketing strategies and approximate simulators offline for lifetime value marketing. In example embodiments, real world data, simulated data, and one or more policies that resulted in the simulated data are obtained. Errors between the real world data and the simulated data are determined. Using the determined errors, bounds are determined. Simulators are ranked based on the determined bounds, whereby a lower bound indicates a first simulator providing simulated data closer to the real world data then a second simulator having a higher bound.

Claims (48)

1 . A method for testing policies and simulators offline for lifetime value marketing, the method comprising:

obtaining real world data indicating a number of actual interactions of a user, simulated data indicating a number of simulated interactions, and one or more policies that resulted in the simulated data, the simulators and the one or more policies used to predict a series of information to provide to the user to maximize the number of actual interactions by the user;

determining errors between the real world data and the simulated data, the errors being used to bound a lifetime difference between the number of actual interactions and the number of simulated interactions;

determining, using a hardware processor, bounds using the determined errors; and

ranking the simulators based on the determined bounds, a lower bound indicating a first simulator providing simulated data closer to the real world data than a second simulator having a higher bound.

2 . The method of claim 1 , wherein the ranking comprises ranking the simulators from lowest bounds to highest bounds, each simulator recommending the series of information to present to the user to maximize the simulated number of interactions.

3 . The method of claim 1 , further comprising:

presenting the ranking of the simulator to a user; and

allowing the user to select one of the simulators for future use, each simulator recommending the series of information to present to the user to maximize the simulated number of interactions.

4 . The method of claim 1 , further comprising ranking the one or more polices based on the determined bounds, a lower bound indicating a first policy providing simulated data closer to the real world data than a second policy having a higher bound, each policy indicating what information to show and how often to show the information in order to maximize the simulated number of interactions.

5 . The method of claim 4 , wherein the ranking comprises ranking the one or more policies from lowest bounds to highest bounds.

6 . The method of claim 4 , further comprising:

presenting the ranking of the policies to a user; and

allowing the user to selecting one of the policies for future use.

7 . The method of claim 1 wherein the bounds are based on at least a selection of a type of error from the group consisting of:

a difference between a true reward function and an estimated reward, δ 1 ;

a smoothness of the reward functions, α and δ 2 ;

a difference between true dynamics and estimated dynamics, ε 1 ; and

a smoothness of dynamics, ε 2 and β.

8 . The method of claim 1 , further comprising applying the one or more policies to one or more simulators to obtain the simulated data.

9 . A non-transitory machine-readable medium in communication with at least one processor, the non-transitory machine-readable medium storing instructions which, when executed by the at least one processor of a machine, causes the machine to perform operations comprising:

obtaining real world data indicating a number of actual interactions of a user, simulated data indicating a number of simulated interactions, and one or more policies that resulted in the simulated data, the simulators and the one or more policies used to predict a series of information to provide to the user to maximize the number of actual interactions by the user;

determining errors between the real world data and the simulated data, the errors being used to bound a lifetime difference between the number of actual interactions and the number of simulated interactions;

determining bounds using the determined errors; and

ranking simulators based on the determined bounds, a lower bound indicating a first simulator providing simulated data closer to the real world data than a second simulator having a higher bound.

10 . The non-transitory machine-readable medium of claim 9 , wherein the ranking comprises ranking the simulators from lowest bounds to highest bounds, each simulator recommending the series of information to present to the user to maximize the simulated number of interactions.

11 . The non-transitory machine-readable medium of claim 9 , further comprising:

presenting the ranking of the simulator to a user; and

allowing the user to select one of the simulators for future use, each simulator recommending the series of information to present to the user to maximize the simulated number of interactions.

12 . The non-transitory machine-readable medium of claim 9 , further comprising ranking the one or more polices based on the determined bounds, a lower bound indicating a first policy providing simulated data closer to the real world data than a second policy having a higher bound, each policy indicating what information to show and how often to show the information in order to maximize the simulated number of interactions.

13 . The non-transitory machine-readable medium of claim 12 , wherein the ranking comprises ranking the one or more policies from lowest bounds to highest bounds.

14 . The non-transitory machine-readable medium of claim 12 , further comprising:

presenting the ranking of the policies to a user; and

allowing the user to selecting one of the policies for future use.

15 . The non-transitory machine-readable medium of claim 9 wherein the bounds are based on at least a selection of a type of error from the group consisting of:

a difference between a true reward function and an estimated reward, δ 1 ;

a smoothness of the reward functions, α and δ 2 ;

a difference between true dynamics and estimated dynamics, ε 1 ; and

a smoothness of dynamics, ε 2 and β.

16 . The non-transitory machine-readable medium of claim 9 , further comprising applying the one or more policies to one or more simulators to obtain the simulated data.

17 . A system comprising:

A hardware processor of a machine;

a communication module to obtain real world data indicating a number of actual interactions of a user, simulated data indicating a number of simulated interactions, and one or more policies that resulted in the simulated data, the simulators and the one or more policies used to predict a series of information to provide to the user to maximize the number of actual interactions by the user;

a bounding module to determine errors between the real world data and the simulated data, and to determine, using the hardware processor, bounds using the determined errors, the errors being used to bound a lifetime difference between the number of actual interactions and the number of simulated interactions; and

an analysis module to rank simulators based on the determined bounds, a lower bound indicating a first simulator providing simulated data closer to the real world data than a second simulator having a higher bound.

18 . The system of claim 17 , wherein the analysis module ranks the simulators by ranking the simulators from lowest bounds to highest bounds, each simulator recommending the series of information to present to the user to maximize the simulated number of interactions.

19 . The system of claim 17 , wherein the analysis module is further to rank the one or more polices based on the determined bounds, a lower bound indicating a first policy providing simulated data closer to the real world data then a second policy having a higher bound, each policy indicating what information to show and how often to show the information in order to maximize the simulated number of interactions.

20 . The system of claim 19 , wherein the analysis module ranks the one or more policies from lowest bounds to highest bounds.

Assignments (2)
CHANGE OF NAME Recorded Nov 29, 2018
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 047687/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2013
From: HALLAK, ASSAF; THEOCHAROUS, GEORGIOS
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 031603/0003 →