Post-hoc loss-calibration for Bayesian neural networks
A computing device and computer-implemented method for post-hoc correction of a decision generated by a machine learning model. The computing device accesses a trained first machine learning (ML) model, a dataset, and a utility function. The computing device trains a second ML model based on performing post-hoc correction of a first set of decisions generated by the first ML model on the dataset. The training includes processing the first set of decisions with respect to a second set of decisions made by the second ML model on the dataset. The training further includes configuring, based on the processing, the second ML model with parameters from a set of parameters optimizing a loss-objective function that concurrently maximizes utility of the second set of decisions according to the utility function and a log-likelihood on the dataset. After training, the second ML model is outputted as a loss-calibrated ML model.
1 . A method of using a computing device for post-hoc correction of a decision generated by a machine learning model, the method comprising:
accessing, by the computing device, a trained first machine learning (ML) model, a first dataset, and a utility function over a prescribed set of actions;
training, by the computing device, a second ML model by iteratively:
causing the first dataset to be input into the first ML model,
causing the first dataset to be input into the second ML model,
receiving a first posterior predictive distribution for the first dataset from the first ML model based at least in part on a learned posterior distribution,
receiving a second posterior predictive distribution for the first dataset from the second ML model based at least in part on the learned posterior distribution,
causing a distance between the first posterior predictive distribution and the second posterior predictive distribution to be reduced by modifying weights of the second ML model, and
in response to the distance between the first posterior predictive distribution and the second posterior predictive distribution falling below a threshold, designating the second posterior predictive distribution as a loss-calibrated posterior predictive distribution; and
outputting, by the computing device, the second ML model as a loss-calibrated ML model.
2 . The method of claim 1 , further comprising:
causing a second dataset to be input into the loss-calibrated ML model; and
receiving a set of decisions and/or predictions for the second dataset, wherein the set of decisions and/or predictions are generated by the loss-calibrated model.
3 . The method of claim 1 , wherein the first ML model and the second ML model are both trained using Bayesian inference, wherein a same subset of the first dataset is input into the first ML model and the second ML model during respective iterations of training the second ML model.
4 . The method of claim 1 , wherein the first dataset includes calibration data without labels and is independent of any data used to train the first ML model.
5 . The method of claim 4 , further comprising:
replacing the first posterior predictive distribution with an amortized approximation of the first posterior predictive distribution.
6 . The method of claim 1 , wherein modifying the weights of the second ML model includes:
processing the first posterior predictive distribution with respect to the second posterior predictive distribution, by determining a Kullback-Leibler (KL) divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution; and
determining a set of weights for the second ML model that maximizes utility of the second posterior predictive distribution according to the utility function and minimizes the KL divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution.
7 . The method of claim 1 , further comprising:
configuring the second ML model with parameters that optimize a loss-objective function by:
in a coordinate ascent fashion, alternating between fixing a decision in the second posterior predictive distribution and taking a gradient step in a direction maximizing the utility function with respect to a given parameter, and then fixing the parameter maximizing the decision, wherein the decision is maximized by enumerating an expected utility of all decisions and selecting a decision with a highest utility according to the utility function.
8 . An information processing system for post-hoc correction of a decision generated by a machine learning model, the information processing system comprising:
a memory having program instructions embodied therewith;
a processor communicatively coupled to the memory, wherein the program instructions are executable by the processor to cause the processor to perform a method comprising:
accessing a trained first machine learning (ML) model, a first dataset, and a utility function over a prescribed set of actions;
training a second ML model by iteratively:
causing the first dataset to be input into the first ML model,
causing the first dataset to be input into the second ML model,
receiving a first posterior predictive distribution for the first dataset from the first ML model based at least in part on a learned posterior distribution,
receiving a second posterior predictive distribution for the first dataset from the second ML model based at least in part on the learned posterior distribution,
causing a distance between the first posterior predictive distribution and the
second posterior predictive distribution to be reduced by modifying weights of the second ML model, and
in response to the distance between the first posterior predictive distribution and the second posterior predictive distribution falling below a threshold, designating the second posterior predictive distribution as a loss-calibrated posterior predictive distribution; and
outputting the second ML model as a loss-calibrated ML model.
9 . The information processing system of claim 8 , wherein the method executed by the processor further comprises:
processing the first posterior predictive distribution with respect to the second posterior predictive distribution by determining a Kullback-Leibler (KL) divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution; and
configuring the second ML model with parameters by determining a set of weights for the second ML model that maximizes utility of the second posterior predictive distribution according to the utility function and minimizes the KL divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution.
10 . The information processing system of claim 8 , wherein the method executed by the processor further comprises:
configuring the second ML model with parameters that optimize a loss-objective function by:
in a coordinate ascent fashion, alternating between fixing a decision in the second posterior predictive distribution and taking a gradient step in a direction maximizing the utility function with respect to a given parameter, and then fixing the parameter maximizing the decision, wherein the decision is maximized by enumerating an expected utility of all decisions and selecting a decision with a highest utility.
11 . A computer program product for post-hoc correction of a decision generated by a machine learning model, the computer program product comprising a computer-readable storage medium having program instructions embodied thereon, the program instructions executable by a computing device to cause the computing device to:
access a trained first machine learning (ML) model, a first dataset, and a utility function over a prescribed set of actions;
train a second ML model by iteratively:
causing the first dataset to be input into the first ML model,
causing the first dataset to be input into the second ML model,
receiving a first posterior predictive distribution for the first dataset from the first ML model based at least in part on a learned posterior distribution,
receiving a second posterior predictive distribution for the first dataset from the second ML model based at least in part on the learned posterior distribution,
causing a distance between the first posterior predictive distribution and the second posterior predictive distribution to be reduced by modifying weights of the second ML model, and
in response to the distance between the first posterior predictive distribution and the second posterior predictive distribution falling below a threshold, designating the second posterior predictive distribution as a loss-calibrated posterior predictive distribution; and
output the second ML model as a loss-calibrated ML model.
12 . The computer program product of claim 11 , wherein the first ML model and the second ML model are both trained using Bayesian inference, wherein a same subset of the first dataset is input into the first ML model and the second ML model during respective iterations of training the second ML model.
13 . The computer program product of claim 11 , wherein modifying the weights of the second ML model includes:
processing the first posterior predictive distribution with respect to the second posterior predictive distribution, by determining a Kullback-Leibler (KL) divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution; and
configuring the second ML model with parameters from a set of parameters by determining a set of weights for the second ML model that maximizes utility of the second posterior predictive distribution according to the utility function and minimizes the KL divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution.
14 . The computer program product of claim 11 , wherein the program instructions cause the computing device to configure the second ML model with parameters that optimize a loss-objective function that concurrently maximizes utility of the second set of decisions according to the utility function and a log-likelihood on the first dataset by:
in a coordinate ascent fashion, alternating between fixing a decision in the second posterior predictive distribution and taking a gradient step in a direction maximizing the utility function with respect to a given parameter in a set of parameters, and then fixing the parameter maximizing the decision, wherein the decision is maximized by enumerating an expected utility of all decisions and selecting a decision with a highest utility.
15 . The computer program product of claim 11 , wherein the program instructions cause the computing device to:
replace the first posterior predictive distribution with an amortized approximation of the first posterior predictive distribution; and
configure the second ML model with parameters that optimize a loss-objective function that concurrently maximizes utility of the second set of decisions according to the utility function and a log-likelihood on the first dataset by:
in a coordinate ascent fashion, alternating between fixing a decision in the second posterior predictive distribution and taking a gradient step in a direction maximizing the utility function with respect to a given parameter in a set of parameters, and then fixing the parameter maximizing the decision, wherein the decision is maximized by enumerating an expected utility of all decisions and selecting a decision with a highest utility,
wherein modifying the weights of the second ML model includes:
processing the first posterior predictive distribution with respect to the second posterior predictive distribution, by determining a Kullback-Leibler (KL) divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution, and
configuring the second ML model with parameters from a set of parameters by determining a set of weights for the second ML model that maximizes utility of the second posterior predictive distribution according to the utility function and minimizes the KL divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution,
wherein the first ML model and the second ML model are both trained using Bayesian inference,
wherein a same subset of the first dataset is input into the first ML model and the second ML model during respective iterations of training the second ML model.