IP Library › Granted Patent US 12,737,613
Granted Patent B2
US 12,737,613 · App. 17/345,099 · Granted Sep 15, 2026

Post-hoc loss-calibration for Bayesian neural networks

Inventors: Meet Prakash Vadera (Northampton, MA); Uri Kartoun (Cambridge, MA); Soumya Ghosh (Boston, MA); Kenney Ng (Arlington, MA)
Assignee: International Business Machines Corporation
G06N3/08G06N3/045G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,613
App. No.
17/345,099
Granted
Sep 15, 2026
Kind
B2
Abstract

A computing device and computer-implemented method for post-hoc correction of a decision generated by a machine learning model. The computing device accesses a trained first machine learning (ML) model, a dataset, and a utility function. The computing device trains a second ML model based on performing post-hoc correction of a first set of decisions generated by the first ML model on the dataset. The training includes processing the first set of decisions with respect to a second set of decisions made by the second ML model on the dataset. The training further includes configuring, based on the processing, the second ML model with parameters from a set of parameters optimizing a loss-objective function that concurrently maximizes utility of the second set of decisions according to the utility function and a log-likelihood on the dataset. After training, the second ML model is outputted as a loss-calibrated ML model.

Claims (67)

1 . A method of using a computing device for post-hoc correction of a decision generated by a machine learning model, the method comprising:

accessing, by the computing device, a trained first machine learning (ML) model, a first dataset, and a utility function over a prescribed set of actions;

training, by the computing device, a second ML model by iteratively:

causing the first dataset to be input into the first ML model,

causing the first dataset to be input into the second ML model,

receiving a first posterior predictive distribution for the first dataset from the first ML model based at least in part on a learned posterior distribution,

receiving a second posterior predictive distribution for the first dataset from the second ML model based at least in part on the learned posterior distribution,

causing a distance between the first posterior predictive distribution and the second posterior predictive distribution to be reduced by modifying weights of the second ML model, and

in response to the distance between the first posterior predictive distribution and the second posterior predictive distribution falling below a threshold, designating the second posterior predictive distribution as a loss-calibrated posterior predictive distribution; and

outputting, by the computing device, the second ML model as a loss-calibrated ML model.

2 . The method of claim 1 , further comprising:

causing a second dataset to be input into the loss-calibrated ML model; and

receiving a set of decisions and/or predictions for the second dataset, wherein the set of decisions and/or predictions are generated by the loss-calibrated model.

3 . The method of claim 1 , wherein the first ML model and the second ML model are both trained using Bayesian inference, wherein a same subset of the first dataset is input into the first ML model and the second ML model during respective iterations of training the second ML model.

4 . The method of claim 1 , wherein the first dataset includes calibration data without labels and is independent of any data used to train the first ML model.

5 . The method of claim 4 , further comprising:

replacing the first posterior predictive distribution with an amortized approximation of the first posterior predictive distribution.

6 . The method of claim 1 , wherein modifying the weights of the second ML model includes:

processing the first posterior predictive distribution with respect to the second posterior predictive distribution, by determining a Kullback-Leibler (KL) divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution; and

determining a set of weights for the second ML model that maximizes utility of the second posterior predictive distribution according to the utility function and minimizes the KL divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution.

7 . The method of claim 1 , further comprising:

configuring the second ML model with parameters that optimize a loss-objective function by:

in a coordinate ascent fashion, alternating between fixing a decision in the second posterior predictive distribution and taking a gradient step in a direction maximizing the utility function with respect to a given parameter, and then fixing the parameter maximizing the decision, wherein the decision is maximized by enumerating an expected utility of all decisions and selecting a decision with a highest utility according to the utility function.

8 . An information processing system for post-hoc correction of a decision generated by a machine learning model, the information processing system comprising:

a memory having program instructions embodied therewith;

a processor communicatively coupled to the memory, wherein the program instructions are executable by the processor to cause the processor to perform a method comprising:

accessing a trained first machine learning (ML) model, a first dataset, and a utility function over a prescribed set of actions;

training a second ML model by iteratively:

causing the first dataset to be input into the first ML model,

causing the first dataset to be input into the second ML model,

receiving a first posterior predictive distribution for the first dataset from the first ML model based at least in part on a learned posterior distribution,

receiving a second posterior predictive distribution for the first dataset from the second ML model based at least in part on the learned posterior distribution,

causing a distance between the first posterior predictive distribution and the

second posterior predictive distribution to be reduced by modifying weights of the second ML model, and

in response to the distance between the first posterior predictive distribution and the second posterior predictive distribution falling below a threshold, designating the second posterior predictive distribution as a loss-calibrated posterior predictive distribution; and

outputting the second ML model as a loss-calibrated ML model.

9 . The information processing system of claim 8 , wherein the method executed by the processor further comprises:

processing the first posterior predictive distribution with respect to the second posterior predictive distribution by determining a Kullback-Leibler (KL) divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution; and

configuring the second ML model with parameters by determining a set of weights for the second ML model that maximizes utility of the second posterior predictive distribution according to the utility function and minimizes the KL divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution.

10 . The information processing system of claim 8 , wherein the method executed by the processor further comprises:

configuring the second ML model with parameters that optimize a loss-objective function by:

in a coordinate ascent fashion, alternating between fixing a decision in the second posterior predictive distribution and taking a gradient step in a direction maximizing the utility function with respect to a given parameter, and then fixing the parameter maximizing the decision, wherein the decision is maximized by enumerating an expected utility of all decisions and selecting a decision with a highest utility.

11 . A computer program product for post-hoc correction of a decision generated by a machine learning model, the computer program product comprising a computer-readable storage medium having program instructions embodied thereon, the program instructions executable by a computing device to cause the computing device to:

access a trained first machine learning (ML) model, a first dataset, and a utility function over a prescribed set of actions;

train a second ML model by iteratively:

causing the first dataset to be input into the first ML model,

causing the first dataset to be input into the second ML model,

receiving a first posterior predictive distribution for the first dataset from the first ML model based at least in part on a learned posterior distribution,

receiving a second posterior predictive distribution for the first dataset from the second ML model based at least in part on the learned posterior distribution,

causing a distance between the first posterior predictive distribution and the second posterior predictive distribution to be reduced by modifying weights of the second ML model, and

in response to the distance between the first posterior predictive distribution and the second posterior predictive distribution falling below a threshold, designating the second posterior predictive distribution as a loss-calibrated posterior predictive distribution; and

output the second ML model as a loss-calibrated ML model.

12 . The computer program product of claim 11 , wherein the first ML model and the second ML model are both trained using Bayesian inference, wherein a same subset of the first dataset is input into the first ML model and the second ML model during respective iterations of training the second ML model.

13 . The computer program product of claim 11 , wherein modifying the weights of the second ML model includes:

processing the first posterior predictive distribution with respect to the second posterior predictive distribution, by determining a Kullback-Leibler (KL) divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution; and

configuring the second ML model with parameters from a set of parameters by determining a set of weights for the second ML model that maximizes utility of the second posterior predictive distribution according to the utility function and minimizes the KL divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution.

14 . The computer program product of claim 11 , wherein the program instructions cause the computing device to configure the second ML model with parameters that optimize a loss-objective function that concurrently maximizes utility of the second set of decisions according to the utility function and a log-likelihood on the first dataset by:

in a coordinate ascent fashion, alternating between fixing a decision in the second posterior predictive distribution and taking a gradient step in a direction maximizing the utility function with respect to a given parameter in a set of parameters, and then fixing the parameter maximizing the decision, wherein the decision is maximized by enumerating an expected utility of all decisions and selecting a decision with a highest utility.

15 . The computer program product of claim 11 , wherein the program instructions cause the computing device to:

replace the first posterior predictive distribution with an amortized approximation of the first posterior predictive distribution; and

configure the second ML model with parameters that optimize a loss-objective function that concurrently maximizes utility of the second set of decisions according to the utility function and a log-likelihood on the first dataset by:

in a coordinate ascent fashion, alternating between fixing a decision in the second posterior predictive distribution and taking a gradient step in a direction maximizing the utility function with respect to a given parameter in a set of parameters, and then fixing the parameter maximizing the decision, wherein the decision is maximized by enumerating an expected utility of all decisions and selecting a decision with a highest utility,

wherein modifying the weights of the second ML model includes:

processing the first posterior predictive distribution with respect to the second posterior predictive distribution, by determining a Kullback-Leibler (KL) divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution, and

configuring the second ML model with parameters from a set of parameters by determining a set of weights for the second ML model that maximizes utility of the second posterior predictive distribution according to the utility function and minimizes the KL divergence between the first posterior predictive distribution with respect to the second posterior predictive distribution,

wherein the first ML model and the second ML model are both trained using Bayesian inference,

wherein a same subset of the first dataset is input into the first ML model and the second ML model during respective iterations of training the second ML model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2021
From: VADERA, MEET PRAKASH; KARTOUN, URI; GHOSH, SOUMYA; NG, KENNEY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 056508/0985 →
Continuity (1)
Related Publication 20220405570A1 · Dec 22, 2022
References Cited (22)
US 8688603B1 · Kurup · 2014 [cited by applicant]
US 8909564B1 · Kaplow · 2014 [cited by applicant]
US 10438114B1 · Blundell · 2019 [cited by examiner]
US 11113632B1 · Evans · 2021 [cited by examiner]
US 11468262B2 · Cheng · 2022 [cited by examiner]
US 20170053646A1 · Watanabe · 2017 [cited by examiner]
US 20170300605A1 · Ardis · 2017 [cited by examiner]
US 20190098039A1 · Gates · 2019 [cited by examiner]
US 20200134493A1 · Bhide · 2020 [cited by applicant]
US 20200210893A1 · Harada · 2020 [cited by examiner]
US 20200334569A1 · Moghadam · 2020 [cited by applicant]
US 20200356905A1 · Luk · 2020 [cited by applicant]
US 20200380310A1 · Weider · 2020 [cited by applicant]
US 20210034976A1 · Arik · 2021 [cited by examiner]
US 20210202070A1 · Bannae · 2021 [cited by examiner]
US 20210326664A1 · Dasgupta · 2021 [cited by examiner]
WO 2020109074A1 · 2020 [cited by applicant]
Cui et al., “Accelerating Monte Carlo Bayesian Prediction via Approximating Predictive Uncertainty Over the Simplex,” Dec. 23, 2020, IEEE Transactions on Neural Networks and Learning Systems, vol. 33, No. 4, pp. 1492-15… [cited by examiner]
Kuśmierczyk et al., “Correcting Predictions for Approximate Bayesian Inference”, Apr. 3, 2020, Proceedings of the AAAI Conference on Artificial Intelligence, 34(04), 4511-4518. https://doi.org/10.1609/aaai.v34i04.5879 (… [cited by examiner]
Cobb et al., “Loss-Calibrated Approximate Inference in Bayesian Neural Networks”, arXiv:1805.03901v1 [stat.ML], May 10, 2018, 12 pages. [cited by applicant]
Kusmierczyk et al., “Correcting Predictions for Approximate Bayesian Inference”, arXiv:1909.04919v1 [stat.ML], Sep. 11, 2019, 12 pages. [cited by applicant]
Lacoste-Julien et al., “Approximate inference for the loss-calibrated Bayesian”, Appearing in Proceedings of the 14th International Conference on Artifcial Intelligence and Statistics (AISTATS), 2011, 09 pages. [cited by applicant]