IP Library › Granted Patent US 12,632,753
Granted Patent B2
US 12,632,753 · App. 18/051,324 · Granted May 19, 2026

Efficient real-world experimentation using causal inference models

Inventors: Adam Evan Foster (Cambridge, GB); Cheng Zhang (Cambridge, GB); Desislava Rosenova Ivanova (Oxford, GB); Joel Nicholas Jennings (Cambridge, GB)
Assignee: Microsoft Technology Licensing, LLC.
G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,753
App. No.
18/051,324
Granted
May 19, 2026
Kind
B2
Abstract

An experiment design is determined for a plurality of physical experiment entities, based on a training loss that is dependent on a critic function and an action parameter individually associated with each physical experiment entity, with the aim of increasing (e.g., optimizing) information gain with respect to a plurality of test entities. The training loss encodes a predicted information gain between a predicted experiment outcome and a predicted test quantity. The predicted experiment outcome associated therewith is sampled from a joint probability distribution based on an entity context. A numerical output is computed using the critic function applied to the predicted experiment outcome and the predicted test quantity.

Claims (90)

1 . A computer-implemented method comprising:

receiving, for a first physical experiment entity of a plurality of physical experiment entities, a first context that individually characterizes the first physical experiment entity;

receiving, for a first physical test entity of a plurality of physical test entities, a second context that individually characterizes the first physical test entity;

computing, for the first physical experiment entity, an initial value of an action parameter that is individually associated with the first physical experiment entity;

computing, in a sequence of training iterations, an updated value of the action parameter, the sequence comprising:

an initial training iteration based on the initial value of the action parameter, and

a subsequent training iteration, for updating the action parameter, based on a value of the action parameter computed in a preceding training iteration of the sequence;

wherein a training iteration of the sequence comprises:

sampling, for the first physical experiment entity, a first predicted experiment outcome from a joint probability distribution, the first predicted experiment outcome being associated with the first physical experiment entity;

sampling, for the first physical test entity, a first predicted test quantity from the joint probability distribution, the first predicted test quantity being associated with the first physical test entity, wherein the joint probability distribution is based on the first context, the second context, and the action parameter,

computing a numerical output using a critic function applied to the first predicted experiment outcome and the first predicted test quantity;

updating, using the numerical output, the action parameter based on a training loss that:

is dependent on the critic function and the action parameter, and

encodes a predicted information gain between the first predicted experiment outcome and the first predicted test quantity;

outputting, for the first physical experiment entity, an indication of a first real-world experiment action, as defined by the updated value of the action parameter associated with the first physical experiment entity as computed in a final training iteration in the sequence of training iterations;

determining a real-world test action individually associated with the first physical test entity based on an outcome of the first real-world experiment action performed on the first physical experiment entity; and

performing the real-world test action on the first physical test entity.

2 . The computer-implemented method of claim 1

wherein the joint probability distribution comprises a Bayesian model of a joint probability distribution over predicted experiment outcomes and predicted test quantities, the Bayesian model being parameterised by a world parameter vector, such that different world models are obtained by sampling different values of the world parameter vector, the joint probability distribution being conditional on the world parameter vector, the action parameter of the first physical experiment entity, the first context, and the second context.

3 . The computer-implemented method of claim 1 , comprising evaluating a real-world test quantity associated with the first physical test entity based on an outcome of the performing of the first real-world experiment action on the first physical experiment entity.

4 . The computer-implemented method of claim 1 , wherein:

the critic function is parameterised by a critic parameter and the computer-implemented method includes computing an initial value of the critic parameter;

the initial training iteration is based on the initial value of the critic parameter; and

the training iterations comprises computing, based on the training loss, an updated value of the action parameter, such that the critic parameter and the action parameter are jointly optimized across the sequence of training iterations.

5 . The computer-implemented method of claim 4 , wherein the critic parameter and the action parameter are jointly optimized across the sequence of training iterations via gradient-based optimization of the training loss.

6 . The computer-implemented method of claim 5 , wherein the sequence of training iterations is extended until a training limit, convergence threshold, or maximum number of iterations has been reached.

7 . The computer-implemented method of claim 5 , wherein at least one of the action parameter and the critic parameter is non-scalar.

8 . The computer-implemented method of claim 5 , wherein the critic parameter comprises a neural network weight.

9 . The computer-implemented method of claim 1 , wherein:

in the training iteration, a world state is sampled from a world-state distribution; and

the first predicted experiment outcome associated with the first physical experiment entity and the first predicted test quantity are sampled based on the sampled world state.

10 . The computer-implemented method of claim 9 , wherein:

in the training iteration, multiple world states are sampled from the world-state distribution; and

for the first physical experiment entity and for the first physical test entity, predicted experiment outcomes and predicted test quantities are sampled for the multiple world states.

11 . The computer-implemented method of claim 1 , further comprising:

receiving, for a second physical experiment entity of the plurality of physical experiment entities, a third context that individually characterizes the second physical experiment entity;

receiving, for a second physical test entity of the plurality of physical test entities, a fourth context that individually characterizes the second physical test entity;

computing, for the second physical experiment entity, an initial value of a second action parameter that is individually associated with the second physical experiment entity;

computing, in a second sequence of training iterations, an updated value of the second action parameter, the second sequence comprising:

an initial training iteration based on the initial value of the second action parameter, and

a subsequent training iteration, for updating the second action parameter, based on a value of the second action parameter computed in a preceding training iteration of the second sequence;

wherein a training iteration of the second sequence comprises:

sampling, for the second physical experiment entity, a second predicted experiment outcome from the joint probability distribution, the second predicted experiment outcome being associated with the second physical experiment entity;

sampling, for the second physical test entity, a second predicted test quantity from the joint probability distribution, the second predicted test quantity being associated with the second physical test entity, wherein the joint probability distribution is based on the third context, the fourth context, and the second action parameter;

computing a second numerical output using the critic function applied to the second predicted experiment outcome and the second predicted test quantity; and

updating, using the second numerical output, the second action parameter based on a second training loss that is dependent on the critic function and the second action parameter and that encodes a predicted information gain between the second predicted experiment outcome and the second predicted test quantity;

outputting, for the second physical experiment entity, an indication of a second real-world experiment action, as defined by the updated value of the second action parameter associated with the second physical experiment entity as computed in a final training iteration of the sequence of training iterations; and

performing a second real-world test action individually associated with the second physical test entity on the second physical test entity, wherein the second real-world test action is based on the first and second real-world experiment actions performed on the first and second physical experiment entities, respectively.

12 . The computer-implemented method of claim 1 , wherein the method is used to design a clinical trial, and the first physical experiment entity and the first physical test entity are living beings.

13 . The computer-implemented method of claim 1 , wherein the first physical experiment entity and the first physical test entity are particular physical configurations of a machine.

14 . The computer-implemented method of claim 1 , wherein the first predicted experiment outcome comprises a predicted reward associated with the first physical experiment entity, and the first predicted test quantity comprises a predicted maximum reward associated with the first physical test entity.

15 . The computer-implemented method of claim 1 , further comprising: causing a control signal to be transmitted to an actuator or device configured to implement the first real-world experiment action.

16 . A computer system comprising:

a memory embodying computer-readable instructions; and

a processor coupled to the memory and configured to execute the computer-readable instructions, the computer-readable instructions configured to cause the processor to:

receive, for a physical test entity of a plurality of physical test entities, a first context that individually characterizes the physical test entity;

receive, for a physical experiment entity of a plurality of physical experiment entities, a second context that individually characterizes the physical experiment entity;

determine, for the physical experiment entity, a first action parameter that is individually associated with the physical experiment entity;

sample a first predicted experiment outcome associated with the physical experiment entity and a predicted test quantity associated with the physical test entity from a first joint probability distribution based on: the first context, the second context, and the first action parameter of the physical experiment entity; sample a first predicted experiment outcome associated with the physical experiment entity and a first predicted test quantity associated with the physical test entity from a first joint probability distribution based on: the first context, the second context, and the first action parameter of the physical experiment entity;

compute a first numerical output using a critic function applied to the first predicted experiment outcome associated with the physical experiment entity and the first predicted test quantity associated with the physical test entity;

determine, for the physical experiment entity, a second action parameter individually associated with the physical experiment entity based on a training loss applied to the first numerical output computed using the critic function and the first action parameter of the physical experiment entity, the training loss applied to the first numerical output and the first action parameter encoding a predicted information gain between the first predicted experiment outcome associated with the physical experiment entity and the first predicted test quantity associated with the physical test entity;

sample a second predicted experiment outcome associated with the physical experiment entity and a second predicted test quantity associated with the physical test entity from a second joint probability distribution based on: the first context, the second context, and the second action parameter of the physical experiment entity;

compute a second numerical output using the critic function applied to the second predicted experiment outcome associated with the physical experiment entity and the second predicted test quantity associated with the physical test entity;

determine, for the physical experiment entity, a third action parameter individually associated with the physical experiment entity based on the training loss applied to the second numerical output computed using the critic function and the second action parameter of the physical experiment entity, the training loss applied to the second numerical output and the second action parameter encoding a predicted information gain between the second predicted experiment outcome associated with the physical experiment entity and the second predicted test quantity associated with the physical test entity;

output an experiment design based on the third action parameter, wherein a real-world experiment action according to the experiment design is performed on the physical experiment entity to obtain a real-world experiment result;

determine a real-world test action individually associated with the physical test entity by updating a probability distribution of outcomes using the real-world experiment result and applying the updated probability distribution to the first context; and

cause the real-world test action to be performed on the physical test entity.

17 . The computer system of claim 16 , wherein the computer-readable instructions are configured to cause the processor to:

sample a third predicted experiment outcome associated the physical experiment entity and a third predicted test quantity associated with the physical test entity from a third joint probability distribution based on: the first context, the second context, and the third action parameter of the physical experiment entity;

compute a third numerical output using the critic function applied to the third predicted experiment outcome associated with the physical experiment entity and the third predicted test quantity associated with the physical test entity; and

determine, for the physical experiment entity, a fourth action parameter individually associated with the physical experiment entity based on the training loss applied to the third numerical output computed using the critic function and the third action parameter of the physical experiment entity, the training loss applied to the third numerical output and the third action parameter encoding a predicted information gain between the third predicted experiment outcome associated with the physical experiment entity and the third predicted test quantity associated with the physical test entity; and

wherein the experiment design is additionally based on the fourth action parameter.

18 . The computer system of claim 16 , wherein the computer-readable instructions are configured to cause the processor to output the experiment design via a graphical user interface associated with the computer system.

19 . The computer system of claim 16 , wherein:

the critic function is parameterised by a first critic parameter, wherein the first numerical output is computed using the first critic parameter; and

a second critic parameter is computed based on the training loss applied to the first numerical output and the second action parameter of the physical experiment entity, wherein the second numerical output is computed using the critic function applied to the second predicted experiment outcome, the second predicted test quantity, and the second critic parameter.

20 . Computer-readable storage media embodying computer-readable instructions configured, when executed on a computer processor, to cause the computer processor to carry out operations comprising:

computing an initial value of a critic parameter;

computing, for a physical experiment entity of a plurality of experimental entities, an initial value of an action parameter individually associated with the physical experiment entity;

computing, in a sequence of training iterations, an updated value of the critic parameter and an updated value of the action parameter, the sequence including:

an initial training iteration based on the initial value of the critic parameter and the initial value of the action parameter, and

a subsequent training iteration for updating the critic parameter and the action parameter, based on a value of the action parameter computed in a preceding one of the training iterations,

wherein a training iteration of the sequence comprises:

sampling, for the physical experiment entity, a predicted experiment outcome from a joint probability distribution, the predicted experiment outcome being associated with the physical experiment entity;

sampling, for a physical test entity, a predicted test quantity associated with physical test entity, wherein the joint probability distribution is based on: a context of the physical experiment entity, a context of the physical test entity, and the action parameter,

computing a numerical output using a critic function parameterised by the critic parameter and applied to the predicted experiment outcome and the predicted test quantity, and

updating, using the numerical output, the critic parameter and the action parameter based on a training loss that is dependent on the critic function and the action parameter and that encodes a predicted information gain between the predicted experiment outcome and the predicted test quantity; updating, using the numerical output, the critic parameter and the action parameter based on a training loss that is dependent on the critic function and the action parameter and that encodes a predicted information gain between the predicted experiment outcomes and the predicted test quantity;

outputting, for the physical experiment entity, an indication of a real-world experiment action defined by the updated value of the action parameter;

determining a real-world test action individually associated with the physical test entity based on an outcome of the real-world experiment action performed on the physical experiment entity; and

causing the real-world test action to be performed on the physical test entity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2022
From: FOSTER, ADAM EVAN; ZHANG, CHENG; IVANOVA, DESISLAVA ROSENOVA; JENNINGS, JOEL NICHOLAS
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061606/0205 →
Continuity (2)
Provisional Application 63347986 · Jun 1, 2022
Related Publication 20230394339A1 · Dec 7, 2023
References Cited (54)
US 10318674B2 · Morgan · 2019 [cited by examiner]
US 11049043B2 · Dalli et al. · 2021 [cited by applicant]
US 11126493B2 · Guha et al. · 2021 [cited by applicant]
US 11354566B1 · Chen et al. · 2022 [cited by applicant]
US 20180060466A1 · Morgan · 2018 [cited by examiner]
US 20200117492A1 · Natarajan et al. · 2020 [cited by applicant]
US 20210056353A1 · Vahdat et al. · 2021 [cited by applicant]
US 20210142190A1 · Isahagian et al. · 2021 [cited by applicant]
US 20230111115A1 · Ramarao · 2023 [cited by examiner]
CA 3110395A1 · 2020 [cited by examiner]
WO WO2022129610A1 · 2022 [cited by examiner]
Foster, et al., “Deep Adaptive Design: Amortizing Sequential Bayesian Experimental Design”, In Repository of arXiv:2103.02438v2, Jun. 11, 2021, 28 Pages. [cited by applicant]
Ivanova, et al., “CO-BED: Information-Theoretic Contextual Optimization via Bayesian Experimental Design”, In Repository of arXiv:2302.14015v1, Feb. 27, 2023, 19 Pages. [cited by applicant]
Ivanova, et al., “Efficient Real-world Testing of Causal Decision Making via Bayesian Experimental Design for Contextual Optimisation”, In Repository of arXiv:2207.05250v1, Jul. 12, 2022, pp. 1-16. [cited by applicant]
Ivanova, et al., “Implicit Deep Adaptive Design: Policy-Based Experimental Design without Likelihoods”, In Repository of arXiv:2111.02329v1, Nov. 3, 2021, pp. 1-33. [cited by applicant]
Kleinegesse, et al., “Gradient-based Bayesian Experimental Design for Implicit Models using Mutual Information Lower Bounds”, In Repository of arXiv:2105.04379v1, May 10, 2021, pp. 1-52. [cited by applicant]
Lim, et al., “Policy-Based Bayesian Experimental Design for Non-Differentiable Implicit Models”, In Repository of arXiv:2203.04272v1, Mar. 8, 2022, 15 Pages. [cited by applicant]
“International Search Report and Written Opinion issued in PCT Application No. PCT/US23/021606”, Mailed Date: Aug. 22, 2023, 16 Pages. [cited by applicant]
Agrawal, et al., “Thompson Sampling for Contextual Bandits with Linear Payoffs”, In Proceedings of 30th International Conference on Machine Learning, May 26, 2013, 9 Pages. [cited by applicant]
Auer, Peter, “Using Confidence Bounds for Exploitation-Exploration Trade-Offs”, In Journal of Machine Learning Research, vol. 3, Nov. 2002, pp. 397-422. [cited by applicant]
Char, et al., “Offline Contextual Bayesian Optimization”, In Proceedings of 33rd Conference on Neural Information Processing Systems, Dec. 8, 2019, 12 Pages. [cited by applicant]
Chu, et al., “Contextual Bandits with Linear Payoff Functions”, In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, Jun. 14, 2011, pp. 208-214. [cited by applicant]
Fauvel, et al., “Contextual Bayesian Optimization with Binary Outputs”, In Repository of arXiv:2111.03447v1, Nov. 5, 2021, 20 Pages. [cited by applicant]
Filippi, et al., “Parametric Bandits: The Generalized Linear Case”, In Proceedings of 24th Annual Conference on Neural Information Processing Systems, Dec. 6, 2010, 9 Pages. [cited by applicant]
Frazier, et al., “The Knowledge-Gradient Policy for Correlated Normal Beliefs”, In INFORMS Journal on Computing, vol. 21, Issue 4, Nov. 2009, pp. 599-613. [cited by applicant]
Geffner, et al., “Deep End-to-end Causal Inference”, In Repository of arXiv:2202.02195v2, Jun. 20, 2022, 31 Pages. [cited by applicant]
Ginsbourger, et al., “Bayesian Adaptive Reconstruction of Profile Optima and Optimizers”, In SIAM/ASA Journal on Uncertainty Quantification, vol. 2, Issue 1, Sep. 23, 2014, pp. 490-510. [cited by applicant]
Han, et al., “Sequential Batch Learning in Finite-Action Linear Contextual Bandit”, In Repository of arXiv:2004.06321v1, Apr. 14, 2020, 34 Pages. [cited by applicant]
Kolyshkina, et al., “The CRISP-ML Approach to Handling Causality and Interpretability Issues in Machine Learning”, In Proceedings of IEEE International Conference on Big Data, Dec. 15, 2021, pp. 2306-2312. [cited by applicant]
Krause, et al., “Contextual Gaussian Process Bandit Optimization”, In Proceedings of 25th Annual Conference on Neural Information Processing Systems, Dec. 12, 2011, 9 Pages. [cited by applicant]
Krishnamurthy, et al., “Contextual Bandits with Continuous Actions: Smoothing, Zooming, and Adapting”, In Proceedings of 32nd Annual Conference on Learning Theory, Jun. 25, 2019, 3 pages. [cited by applicant]
Majzoubi, et al., “Efficient Contextual Bandits with Continuous Actions”, In Proceedings of 34th Conference on Neural Information Processing Systems, Dec. 6, 2020, 12 Pages. [cited by applicant]
Metzen, Jan H. , “Active Contextual Entropy Search”, In Repository of arXiv:1511.04211v2, Nov. 16, 2015, 6 Pages. [cited by applicant]
Metzen, et al., “Bayesian Optimization for Contextual Policy Search”, In Proceedings of the Second Machine Learning in Planning and Control of Robot Motion Workshop, 2015, 2 Pages. [cited by applicant]
Oord, et al., “Representation Learning with Contrastive Predictive Coding”, In Repository of arXiv:1807.03748v1, Jul. 10, 2018, 13 Pages. [cited by applicant]
Pearce, et al., “Continuous Multi-Task Bayesian Optimisation with Correlation”, In European Journal of Operational Research, vol. 270, Issue 3, Nov. 1, 2018, 31 Pages. [cited by applicant]
Pearce, et al., “Practical Bayesian Optimization of Objectives with Conditioning Variables”, In Repository of arXiv:2002.09996v2, Nov. 2, 2020, 22 Pages. [cited by applicant]
Si, et al., “Distributional Robust Batch Contextual Bandits”, In Repository of arXiv:2006.05630v1, Jun. 10, 2020, 40 Pages. [cited by applicant]
Sutton, et al., “Reinforcement Learning: An Introduction”, In Publication of MIT Press, Nov. 13, 2018, 548 Pages. [cited by applicant]
Swaminathan, et al., “Counterfactual Risk Minimization: Learning from Logged Bandit Feedback”, In Proceedings of 32nd International Conference on Machine Learning, Jul. 6, 2015, 10 Pages. [cited by applicant]
Turek, Matt, “Explainable Artificial Intelligence (XAI)”, Retrieved from: https://www.darpa.mil/program/explainable-artificial-intelligence, Retrieved on: Jul. 13, 2022, 4 Pages. [cited by applicant]
Wang, et al., “Max-Value Entropy Search for Efficient Bayesian Optimization”, In Proceedings of 34th International Conference on Machine Learning, Aug. 6, 2017, 9 Pages. [cited by applicant]
Zhou, et al., “Neural Contextual Bandits with Ucb-Based Exploration”, In Proceedings of 37th International Conference on Machine Learning, Jul. 13, 2020, 11 Pages. [cited by applicant]
Cranmer, et al., “The Frontier of Simulation-based Inference”, In Journal of the National Academy of Sciences, vol. 117, Issue 48, Dec. 1, 2020, pp. 30055-30062. [cited by applicant]
Foster, et al., “A Unified Stochastic Gradient Approach to Designing Bayesian-Optimal Experiments”, In Proceedings of 23rd International Conference on Artificial Intelligence and Statistics, Aug. 26, 2020, 10 Pages. [cited by applicant]
Green, et al., “Bayesian Computation: A Summary of the Current State, and Samples Backwards and Forwards”, In Journal of Statistics and Computing, vol. 25, Issue 4, Jun. 11, 2015, pp. 835-862. [cited by applicant]
Jang, et al., “Categorical Reparameterization with Gumbel-Softmax”, In Repository of arXiv:1611.01144v1, Nov. 3, 2016, 13 Pages. [cited by applicant]
Kleinegesse, et al., “Efficient Bayesian Experimental Design for Implicit Models”, In Proceedings of 22nd International Conference on Artificial Intelligence and Statistics, Apr. 16, 2019, 10 Pages. [cited by applicant]
Kleinegesse, et al., “Bayesian Experimental Design for Implicit Models by Mutual Information Neural Estimation”, In Proceedings of 37th International Conference on Machine Learning, Jul. 13, 2020, 11 Pages. [cited by applicant]
Lindley, Dennis V. , “On a Measure of the Information Provided by an Experiment”, In Journal of the Annals of Mathematical Statistics, vol. 27, Issue 4, Dec. 1956, pp. 986-1005. [cited by applicant]
Maddison, et al., “The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”, In Repository of arXiv:1611.00712v1, Nov. 2, 2016, 17 Pages. [cited by applicant]
Mohamed, et al., “Monte Carlo Gradient Estimation in Machine Learning”, In Journal of Machine Learning Research, vol. 21, Issue 132, Jul. 2020, 62 Pages. [cited by applicant]
Rainforth, et al., “On Nesting Monte Carlo Estimators”, In Proceedings of 35th International Conference on Machine Learning, Jul. 10, 2018, 10 Pages. [cited by applicant]
Zanette, et al., “Design of Experiments for Stochastic Contextual Linear Bandits”, In Proceedings of 35th Conference on Neural Information Processing Systems, Dec. 6, 2021, 12 Pages. [cited by applicant]