System and method for dynamically determining resource-hold-time recommendations based on estimated causal effects
A method can include upon receiving, from a policy update engine, one or more hold-time recommendations, selectively determining, based on one or more selection rules, one or more selected hold-time values of the one or more hold-time recommendations. The method further can include implementing the one or more selected hold-time values, as determined. The method additionally can include after implementing the one or more selected hold-time values, determining one or more effects associated with the one or more selected hold-time values. The method also can include transmitting the one or more selected hold-time values and the one or more effects to the policy update engine for retraining. Other embodiments are disclosed.
1 . A system comprising:
one or more processors; and
one or more non-transitory computer-readable media storing computing instructions configured to, when run on the one or more processors, cause the one or more processors to perform:
training a causal inference machine learning (ML) model to determine treatment effects related to user engagement with a website associated with hold-time treatment levels applied to resources accessed through the website;
training a policy update engine to determine one or more hold-time recommendations using the treatment effects and the hold-time treatment levels;
wherein the policy update engine comprises a reinforcement learning model trained by policy iteration;
upon receiving, from the policy update engine, one or more hold-time recommendations, selectively determining, based on one or more selection rules, one or more selected hold-time values of the one or more hold-time recommendations, wherein:
the causal inference ML model determines a respective treatment effect related to user engagement with the website associated with a respective hold-time treatment level for each grouping of one or more experimental groupings, wherein each grouping of the one or more experimental groupings comprises one or more respective grouping treatment units of treatment observation units and one or more respective grouping control units of control observation units assigned to the each grouping based on a respective threshold and a respective similarity level between each respective pair of the one or more respective grouping treatment units and the one or more respective grouping control units; and
the policy update engine determines the one or more hold-time recommendations using the respective treatment effect associated with the respective hold-time treatment level for each grouping, determined by the causal inference ML model; and
implementing the one or more selected hold-time values, as determined, on the resources accessed through the website;
after implementing the one or more selected hold-time values, determining one or more effects on user engagement with the website associated with the one or more selected hold-time values;
transmitting the one or more selected hold-time values and the one or more effects to the policy update engine and retraining the policy update engine on the one or more selected hold-time values and the one or more effects as feedback;
re-determining, by the retrained policy update engine, the one or more hold-time recommendations; and
retraining the causal inference ML model using the one or more hold-time recommendations, as re-determined by the policy update engine.
2 . The system in claim 1 , wherein:
the causal inference ML model is configured to determine the respective treatment effect associated with the respective hold-time treatment level for the each grouping of the one or more experimental groupings by:
classifying the treatment observation units in a treatment population and the control observation units in a control population into the one or more experimental groupings; and
determining, by one or more respective machine learning models of the causal inference ML model for the each grouping of the one or more experimental groupings, the respective treatment effect for the each grouping based on one or more respective causal inference values associated with the respective hold-time treatment level for (a) each respective treatment unit of the one or more respective grouping treatment units of the treatment observation units and (b) a respective matched control unit of the one or more respective grouping control units of the control observation units for the each grouping.
3 . The system in claim 2 , wherein classifying the treatment observation units and the control observation units into the one or more experimental groupings further comprises:
before determining the respective matched control unit for the each respective treatment unit, determining, by a matching model, the respective similarity level between the each respective treatment unit and the respective matched control unit based on at least one of:
recursive partitioning based on one or more respective features for the each respective treatment unit and the respective matched control unit;
a respective cosine distance between respective feature embeddings for the each respective treatment unit and the respective matched control unit; or
propensity score matching based on the one or more respective features for the each respective treatment unit and the respective matched control unit.
4 . The system in claim 3 , wherein determining the respective treatment effect for the each grouping further comprises training the causal inference ML model by training at least one of the one or more respective machine learning models or the matching model.
5 . The system in claim 2 , wherein determining the respective treatment effect for the each grouping further comprises, after the one or more respective grouping treatment units of the treatment observation units and the one or more respective grouping control units of the control observation units are assigned to the each grouping:
training a respective treatment model of the one or more respective machine learning models for a respective treatment group of each grouping of the one or more experimental groupings based on the one or more respective grouping treatment units of the respective treatment group to determine a respective treatment causal inference value associated with a treatment level; and
training a respective control model of the one or more respective machine learning models for a respective control group of each grouping of the one or more experimental groupings based on one or more respective grouping control units of the respective control group to determine a respective control causal inference value associated with a non-treatment level.
6 . The system in claim 5 , wherein retraining the causal inference ML model comprises retraining at least one of the respective treatment model and the respective control model based at least in part on the one or more hold-time recommendations, as re-determined by the policy update engine.
7 . The system in claim 2 , wherein determining the respective treatment effect for the each grouping further comprises:
determining, by the one or more respective machine learning models, a respective treatment causal inference value of the one or more respective causal inference values for each respective treatment unit of the one or more respective grouping treatment units for the each grouping;
determining, by the one or more respective machine learning models, a respective control causal inference value of the one or more respective causal inference values for the respective matched control unit of the one or more respective grouping control units for the each grouping; and
determining, as the respective treatment effect for the each grouping, an average value of the respective treatment causal inference value for each respective treatment unit of the one or more respective grouping treatment units and the respective control causal inference value for the respective matched control unit of the one or more respective grouping control units.
8 . The system in claim 1 , wherein the computing instructions are further configured, when run on the one or more processors, to cause the one or more processors to perform one or more of:
re-determining, by the causal inference ML model, the respective treatment effect associated with the one or more hold-time recommendations, as re-determined by the policy update engine, for the each grouping of the one or more experimental groupings for the policy update engine to iteratively re-determine the one or more hold-time recommendations.
9 . The system in claim 1 , wherein determining the one or more hold-time recommendations comprises:
evaluating the respective hold-time treatment level for the each grouping of the one or more experimental groupings based on a respective estimated reward determined by a state-value function with the respective treatment effect associated with the respective hold-time treatment level for the each grouping of the one or more experimental groupings; and
updating the respective hold-time treatment level by a greedy function with the respective estimated reward.
10 . A method being implemented via execution of computing instructions configured to run at one or more processors and stored at one or more non-transitory computer-readable media, the method comprising:
training a causal inference machine learning (ML) model to determine treatment effects related to user engagement with a website associated with hold-time treatment levels applied to resources accessed through the website;
training a policy update engine to determine one or more hold-time recommendations using the treatment effects and the hold-time treatment levels;
wherein the policy update engine comprises a reinforcement learning model trained by policy iteration;
upon receiving, from the policy update engine, one or more hold-time recommendations, selectively determining, based on one or more selection rules, one or more selected hold-time values of the one or more hold-time recommendations, wherein:
the causal inference ML model determines a respective treatment effect related to user engagement with the website associated with a respective hold-time treatment level for each grouping of one or more experimental groupings, wherein each grouping of the one or more experimental groupings comprises one or more respective grouping treatment units of treatment observation units and one or more respective grouping control units of control observation units assigned to the each grouping based on a respective threshold and a respective similarity level between each respective pair of the one or more respective grouping treatment units and the one or more respective grouping control units; and
the policy update engine determines the one or more hold-time recommendations using the respective treatment effect associated with the respective hold-time treatment level for each grouping, determined by the causal inference ML model; and
implementing the one or more selected hold-time values, as determined, on the resources accessed through the website;
after implementing the one or more selected hold-time values, determining one or more effects on user engagement with the website associated with the one or more selected hold-time values;
transmitting the one or more selected hold-time values and the one or more effects to the policy update engine and retraining the policy update engine on the one or more selected hold-time values and the one or more effects as feedback;
re-determining, by the retrained policy update engine, the one or more hold-time recommendations; and
retraining the causal inference ML model using the one or more hold-time recommendations, as re-determined by the policy update engine.
11 . The method in claim 10 , wherein:
the causal inference ML model is configured to determine the respective treatment effect associated with the respective hold-time treatment level for the each grouping of the one or more experimental groupings by:
classifying the treatment observation units in a treatment population and the control observation units in a control population into the one or more experimental groupings; and
determining, by one or more respective machine learning models of the causal inference ML model for the each grouping of the one or more experimental groupings, the respective treatment effect for the each grouping based on one or more respective causal inference values associated with the respective hold-time treatment level for (a) each respective treatment unit of the one or more respective grouping treatment units of the treatment observation units and (b) a respective matched control unit of the one or more respective grouping control units of the control observation units for the each grouping.
12 . The method in claim 11 , wherein classifying the treatment observation units and the control observation units into the one or more experimental groupings further comprises:
before determining the respective matched control unit for the each respective treatment unit, determining, by a matching model, the respective similarity level between the each respective treatment unit and the respective matched control unit based on at least one of:
recursive partitioning based on one or more respective features for the each respective treatment unit and the respective matched control unit;
a respective cosine distance between respective feature embeddings for the each respective treatment unit and the respective matched control unit; or
propensity score matching based on the one or more respective features for the each respective treatment unit and the respective matched control unit.
13 . The method in claim 12 , wherein determining the respective treatment effect for the each grouping further comprises training the causal inference ML model by training at least one of the one or more respective machine learning models or the matching model.
14 . The method in claim 11 , wherein determining the respective treatment effect for the each grouping further comprises, after the one or more respective grouping treatment units of the treatment observation units and the one or more respective grouping control units of the control observation units are assigned to the each grouping:
training a respective treatment model of the one or more respective machine learning models for a respective treatment group of each grouping of the one or more experimental groupings based on the one or more respective grouping treatment units of the respective treatment group to determine a respective treatment causal inference value associated with a treatment level; and
training a respective control model of the one or more respective machine learning models for a respective control group of each grouping of the one or more experimental groupings based on one or more respective grouping control units of the respective control group to determine a respective control causal inference value associated with a non-treatment level.
15 . The method in claim 10 , wherein the computing instructions are further configured, when run on the one or more processors, to cause the one or more processors to perform one or more of:
re-determining, by the causal inference ML model, the respective treatment effect associated with the one or more hold-time recommendations, as re-determined by the policy update engine, for the each grouping of the one or more experimental groupings for the policy update engine to iteratively re-determine the one or more hold-time recommendations.
16 . The method in claim 11 , wherein determining the respective treatment effect for the each grouping further comprises:
determining, by the one or more respective machine learning models, a respective treatment causal inference value of the one or more respective causal inference values for each respective treatment unit of the one or more respective grouping treatment units for the each grouping;
determining, by the one or more respective machine learning models, a respective control causal inference value of the one or more respective causal inference values for the respective matched control unit of the one or more respective grouping control units for the each grouping; and
determining, as the respective treatment effect for the each grouping, an average value of the respective treatment causal inference value for each respective treatment unit of the one or more respective grouping treatment units and the respective control causal inference value for the respective matched control unit of the one or more respective grouping control units.
17 . The method in claim 10 , wherein determining the one or more hold-time recommendations comprises:
evaluating the respective hold-time treatment level for the each grouping of the one or more experimental groupings based on a respective estimated reward determined by a state-value function with the respective treatment effect associated with the respective hold-time treatment level for the each grouping of the one or more experimental groupings; and
updating the respective hold-time treatment level by a greedy function with the respective estimated reward.