IP Library Patent Application 14551975
Patent Application
App. No. 14/551,975

Searching for Safe Policies to Deploy

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/551,975
Abstract

Risk quantification, policy search, and automated safe policy deployment techniques are described. In one or more implementations, techniques are utilized to determine safety of a policy, such as to express a level of confidence that a new policy will exhibit an increased measure of performance (e.g., interactions or conversions) over a currently deployed policy. In order to make this determination, reinforcement learning and concentration inequalities are utilized, which generate and bound confidence values regarding the measurement of performance of the policy and thus provide a statistical guarantee of this performance. These techniques are usable to quantify risk in deployment of a policy, select a policy for deployment based on estimated performance and a confidence level in this estimate (e.g., which may include use of a policy space to reduce an amount of data processed), used to create a new policy through iteration in which parameters of a policy are iteratively adjusted and an effect of those adjustments are evaluated, and so forth.

Claims (31)

1 . In a digital medium environment for identifying and deploying potential digital advertising campaigns, where campaigns can be altered, removed, or replaced on demand, a method for optimizing campaign selection in the digital medium environment, the method comprising:

controlling replacement of one or more deployed polices of a content provider that are used to select advertisements with at least one of a plurality of policies, the controlling including:

searching a plurality of policies to locate the at least one said policy that is deemed safe to replace the one or more deployed policies, the at least one said policy deemed safe if a measure of performance of the at least one said policy is greater than a threshold measure of performance and within a defined level of confidence as indicated by one or more statistical guarantees computed through use of reinforcement learning and concentration inequalities on deployment data generated by the one or more deployed policies; and

responsive to the location of the at least one said policy that is deemed safe to replace the one or more other policies, causing the replacement of the one or more other policies with the at least one said policy.

2 . A method as described in claim 1 , wherein:

each of the plurality of policies is expressed using a high-dimensional vector; and

the searching includes computing a direction in a policy space that is expected to point towards a safe region.

3 . A method as described in claim 2 , wherein the searching is constrained to line searches of the high-dimensional vectors of the plurality of policies that correspond to the direction.

4 . A method as described in claim 3 , wherein the searching further comprises determining that the at least one said policy, of the plurality of policies having high-dimensional vectors corresponding to the direction, exhibits a highest level of the measure of performance, one to another.

5 . A method as described in claim 2 , wherein the direction is a generalized natural policy gradient.

6 . A method as described in claim 1 , wherein the one or more statistical guarantees are configured as performance bounds defined by the concentration inequality on the likely performance of the at least one said policy.

7 . A method as described in claim 1 , wherein the threshold is based at least in part on measured performance of the one or more deployed policies and a set margin.

8 . A method as described in claim 7 , wherein the threshold is set such that the estimated values of the at least one said policy exhibit an improvement in the measurement of performance over the one or more deployed policies.

9 . A method as described in claim 1 , wherein deployment data does not describe deployment of the at least one said policy.

10 . A method as described in claim 1 , wherein received deployment data also describes deployment of the at least one said policy.

11 . A system comprising:

one or more computing devices configured to perform operations including selecting at least one of a plurality to policies to replace one or more deployed policies of a content provider that are used to select advertisements to be included with content, the selecting including:

accessing a plurality of high-dimensional vectors that express respective ones of the plurality of policies;

computing a direction in a policy space of the plurality of policies that is expected to point towards a region that is expected to be safe as including the policies that have a measure of performance that is greater than a threshold measure of performance and within a defined level of confidence; and

selecting the at least one said policy of the plurality of policies having high-dimensional vectors that correspond to the direction and that exhibits a highest level of the measure of performance.

12 . A system as described in claim 11 , wherein the selecting includes searching the plurality of policies as constrained to line searches of the high-dimensional vectors of the plurality of policies that correspond to the direction.

13 . A system as described in claim 11 , wherein the direction is a generalized natural policy gradient.

14 . A system as described in claim 11 , wherein the measure of performance is computed through use of reinforcement learning and concentration inequalities on deployment data generated by the one or more deployed policies.

15 . A content provider comprising one or more computing devices configured to perform operations including:

deploying a policy to select advertisements to be included with content based on one or more characteristics associated with a request for the content; and

replacing the deployed policy with another policy that is selected from a plurality of policies by computing a direction in a policy space of a plurality of high-dimensional vectors of the plurality of policies that is expected to point towards a region that is expected to be safe as including the policies that have a measure of performance that is greater than a threshold measure of performance and within a defined level of confidence.

16 . A system as described in claim 15 , wherein the selecting includes searching the plurality of policies as constrained to line searches of the high-dimensional vectors of the plurality of policies that correspond to the direction.

17 . A system as described in claim 15 , wherein the direction is a generalized natural policy gradient.

18 . A system as described in claim 15 , wherein the measure of performance is computed through use of reinforcement learning and concentration inequalities on deployment data generated by the deployed policy.

19 . A system as described in claim 18 , wherein deployment data does not describe deployment of the other policy.

20 . A system as described in claim 18 , wherein the concentration inequality is configured to enforce performance bounds on the measurement of performance as part of the reinforcement learning.

Assignments (2)
CHANGE OF NAME Recorded Jan 18, 2019
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 048097/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2014
From: THOMAS, PHILIP S.; THEOCHAROUS, GEORGIOS; GHAVAMZADEH, MOHAMMAD
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 034289/0868 →