Identifying and correcting label bias in machine learning
The present disclosure is directed to systems and methods for identifying and correcting label bias in machine learning via intelligent re-weighting of training examples. In particular, aspects of the present disclosure leverage a problem formulation which assumes the existence of underlying, unknown, and unbiased labels which are overwritten by an agent who intends to provide accurate labels but may have biases towards certain groups. Despite the fact that a biased training dataset provides only observations of the biased labels, the systems and methods described herein can nevertheless correct the bias by re-weighting the data points without changing the labels.
1 . A computer-implemented method to reduce bias in a machine-learned classification model configured without post-processing bias calibration of model outputs, the method comprising:
obtaining, by one or more computing devices, a training dataset comprising a plurality of training examples, each training example comprising an example input and a respective example label applied to the example input, wherein the example labels of the training dataset exhibit a bias against one or more subgroups of the example inputs;
initializing, by the one or more computing devices, a plurality of weights that are respectively associated with the plurality of training examples;
for each of one or more training iterations:
determining, by the one or more computing devices, one or more constraint violation values for the machine-learned classification model on the training dataset relative to one or more fairness constraints applied to the one or more subgroups of the example inputs, wherein the one or more fairness constraints comprise a disparate impact constraint;
updating, by the one or more computing devices, one or more re-weighting control values respectively associated with the one or more fairness constraints based at least in part on the one or more constraint violation values;
modifying, by the one or more computing devices, at least one of the plurality of weights associated with the plurality of training examples based at least in part on the one or more re-weighting control values to form a plurality of modified weights; and
re-training, by the one or more computing devices, the machine-learned classification model using the training dataset weighted according to the plurality of modified weights;
wherein both a true positive re-weighting control value and a false positive re-weighting control value are associated with at least one of the one or more fairness constraints.
2 . The computer-implemented method of claim 1 , wherein a single re-weighting control value is associated with at least one of the one or more fairness constraints.
3 . The computer-implemented method of claim 1 , wherein the one or more fairness constraints comprise one or more of: a demographic parity constraint, or an equal opportunity constraint.
4 . The computer-implemented method of claim 1 , wherein the one or more fairness constraints comprise an equalized odds constraint.
5 . The computer-implemented method of claim 1 , wherein modifying, by the one or more computing devices, at least one of the plurality of weights associated with the plurality of training examples based at least in part on one or more re-weighting control values to form the plurality of modified weights comprises:
determining, by the one or more computing devices, for each of plurality of weights, an intermediate weight value equal to an exponential raised to a sum of the re-weighting control values for which the corresponding example input is included in the corresponding subgroup; and
normalizing, by the one or more computing devices, the intermediate weight values for the plurality of weights to form the plurality of modified weights.
6 . The computer-implemented method of claim 1 , wherein updating, by the one or more computing devices, the one or more re-weighting control values comprises subtracting, from the one or more re-weighting control values, the one or more constraint violation values multiplied by a step size.
7 . The computer-implemented method of claim 1 wherein the one or more re-weighting control values comprise Lagrange multipliers.
8 . The computer-implemented method of claim 1 , wherein modifying, by the one or more computing devices, at least one of the plurality of weights associated with the plurality of training examples based at least in part on one or more re-weighting control values to form a plurality of modified weights has, when a positive prediction rate of the machine-learned classification model with respect to a first subgroup of the example inputs is below a target value, a first effect of increasing the weight associated with training examples in which the corresponding example input is included in the first subgroup and the corresponding example label is a positive label and a second effect of decreasing the weight associated with training examples in which the corresponding example input is included in the first subgroup and the corresponding example label is a negative label.
9 . The computer-implemented method of claim 1 , wherein the machine-learned classification model comprises an artificial neural network.
10 . The computer-implemented method of claim 1 , wherein the machine-learned classification model comprises a logistic regression classifier model.
11 . Non-transitory computer-readable media storing a machine-learned classification model trained according to the method of claim 1 .
12 . The computer-implemented method of claim 1 , further comprising:
receiving one or more sensor inputs characterizing a state of an environment;
providing data associated with the one or more sensor inputs as input to the machine-learned classification model;
generating, with the machine-learned classification model, one or more classifications associated with the state of the environment based on the data associated with the one or more sensor inputs; and
generating a control signal for a machine based at least in part on the one or more classifications.
13 . The computer-implemented method of claim 1 , further comprising:
receiving input data;
providing the input data as input to the machine-learned classification model;
generating, with the machine-learned classification model, one or more classifications based on the input data; and
processing one or more images based at least in part on the one or more classifications.
14 . A computer system comprising:
one or more processors; and
one or more non-transitory computer readable media that collectively store:
a machine-learned classification model configured without post-processing bias calibration of model outputs; and
instructions that, when executed by the one or more processors, cause the computer system to perform operations, the operations comprising:
initializing, by the one or more processors, a plurality of weights that are respectively associated with a plurality of training examples wherein the plurality of training examples comprise a training dataset;
for each of one or more training iterations:
determining, by the one or more processors, one or more constraint violation values for the machine-learned classification model on the training dataset relative to one or more fairness constraints applied to one or more subgroups of the example inputs, wherein the one or more fairness constraints comprise a disparate impact constraint;
updating, by the one or more processors, one or more re-weighting control values respectively associated with the one or more fairness constraints based at least in part on the one or more constraint violation values;
modifying, by the one or more processors, at least one of the plurality of weights associated with the plurality of training examples based at least in part on the one or more re-weighting control values to form a plurality of modified weights; and
re-training, by the one or more processors, the machine-learned classification model using the training dataset weighted according to the plurality of modified weights;
wherein both a true positive re-weighting control value and a false positive re-weighting control value are associated with at least one of the one or more fairness constraints.
15 . One or more non-transitory computer readable media that collectively store a machine-learned classification model configured without post-processing bias calibration of model outputs and instructions that, when executed by one or more processors, cause a computer system to perform operations, the operations comprising:
obtaining, by one or more computing devices, a training dataset comprising a plurality of training examples, each training example comprising an example input and a respective example label applied to the example input, wherein the example labels of the training dataset exhibit a bias against one or more subgroups of the example inputs;
initializing, by the one or more computing devices, a plurality of weights that are respectively associated with the plurality of training examples;
for each of one or more training iterations:
determining, by the one or more computing devices, one or more constraint violation values for the machine-learned classification model on the training dataset relative to one or more fairness constraints applied to the one or more subgroups of the example inputs, wherein the one or more fairness constraints comprise a disparate impact constraint;
updating, by the one or more computing devices, one or more re-weighting control values respectively associated with the one or more fairness constraints based at least in part on the one or more constraint violation values;
modifying, by the one or more computing devices, at least one of the plurality of weights associated with the plurality of training examples based at least in part on the one or more re-weighting control values to form a plurality of modified weights; and
re-training, by the one or more computing devices, the machine-learned classification model using the training dataset weighted according to the plurality of modified weights;
wherein both a true positive re-weighting control value and a false positive re-weighting control value are associated with at least one of the one or more fairness constraints.
16 . The one or more non-transitory computer readable media of claim 15 , wherein a single re-weighting control value is associated with at least one of the one or more fairness constraints.
17 . The one or more non-transitory computer readable media of claim 15 , wherein the one or more fairness constraints comprise one or more of: a demographic parity constraint, or an equal opportunity constraint.
18 . The one or more non-transitory computer readable media of claim 15 , wherein the one or more fairness constraints comprise an equalized odds constraint.
19 . The one or more non-transitory computer readable media of claim 15 , wherein modifying, by the one or more computing devices, at least one of the plurality of weights associated with the plurality of training examples based at least in part on one or more re-weighting control values to form the plurality of modified weights comprises:
determining, by the one or more computing devices, for each of plurality of weights, an intermediate weight value equal to an exponential raised to a sum of the re-weighting control values for which the corresponding example input is included in the corresponding subgroup; and
normalizing, by the one or more computing devices, the intermediate weight values for the plurality of weights to form the plurality of modified weights.
20 . The one or more non-transitory computer readable media of claim 15 , wherein updating, by the one or more computing devices, the one or more re-weighting control values comprises subtracting, from the one or more re-weighting control values, the one or more constraint violation values multiplied by a step size.