System and method for machine learning fairness testing
Systems and methods for diagnosing and testing fairness of machine learning models based on detecting individual violations of group definitions of fairness, via adversarial attacks that aim to perturb model inputs to generate individual violations. The systems and methods employ auxiliary machine learning models using a local surrogate for identifying group membership and assess fairness by measuring the transferability of attacks from this model. The systems and methods generate fairness indicator values indicative of discrimination risk due to the target predictions generated by the machine learning model, by comparing gradients of the machine learning model to gradients of an auxiliary machine learning model.
1 . A system for estimating fairness of a target machine learning model by using an auxiliary machine learning model for identifying a protected group membership from an input variable of the target machine learning model, the auxiliary machine learning model configured to generate adversarial attacks from perturbed model inputs, the system comprising:
one or more processors operating in conjunction with computer memory, the one or more processors configured to:
instantiate the auxiliary machine learning model configured for inferring the protected group membership from the input variable of the target machine learning model;
receive data representative of the input variable of the target machine learning model, the target machine learning model configured to generate target predictions based on the input variable;
generate, at a given point X 0 , a first vector indicative of a gradient of the target predictions of the target machine learning model with respect to the input variable;
generate, in parallel to generating the first vector, a second vector as a locally linearized surrogate of a gradient of the input variable with respect to the protected group membership, the second vector computed in a numeric local neighborhood around the given point X 0 via a Moore-Penrose pseudoinverse of a first-order Taylor expansion utilizing a gradient of the auxiliary target machine learning model evaluated at the given point X 0 ;
generate a fairness indicator value based on a projection value obtained by projecting the first vector onto the second vector and dividing by a L2 norm of the second vector;
generate downstream computing instructions representative of the fairness indicator value as compared to a predetermined fairness threshold to determine whether the target machine learning model exhibits individual-level unfairness with respect to the protected group membership, the downstream computing instructions indicating a modification instruction when the fairness indicator value is less than the predetermined fairness threshold; and
transmit, to a downstream computing subsystem, the downstream computing instructions, the downstream computing subsystem configured to automatically select, from a plurality of candidate target machine learning models, an alternative machine learning model in accordance with the modification instruction.
2 . The system of claim 1 , wherein the downstream computing subsystem randomizes model weights of the target machine learning model and retrains the target machine learning model using an alternative data set.
3 . The system of claim 1 , wherein the gradient of the target machine learning model and the gradient of the auxiliary target machine learning model are evaluated using a Fast Gradient Method approximation.
4 . The system of claim 1 , wherein the target machine learning model is a first supervised learning model, and the auxiliary machine learning model is a second supervised learning model trained at least partially based on known values of the protected group membership.
5 . The system of claim 1 , wherein the second vector is representative of an output of a sign function of the gradient of the auxiliary machine learning model evaluated at the value of the input variable.
6 . The system of claim 1 , wherein the second vector is indicative of a modified gradient when the gradient of the auxiliary machine learning model is associated with out-of-distribution predictions of the target machine learning model.
7 . The system of claim 1 , wherein a plurality of fairness indicator values is generated, each corresponding to different protected attributes of the one or more attributes, and the downstream instructions are indicative of whether an aggregated measure of the plurality of fairness indicators exceeds a predefined aggregate fairness threshold.
8 . The system of claim 7 , wherein each of the plurality of fairness indicator values is indicative of a covariance between the gradient of the target machine learning model and the gradient of the auxiliary machine learning model.
9 . The system of claim 7 , wherein the aggregated measure is an L-p norm.
10 . The system of claim 1 , wherein the input variable includes an observable attribute correlated with at least one of the one or more protected attributes via an unobserved latent variable.
11 . The method of claim 1 , wherein the downstream computing subsystem randomizes model weights of the target machine learning model and retrains the target machine learning model using an alternative data set.
12 . A method for estimating fairness of a target machine learning model by using an auxiliary machine learning model for identifying a protected group membership from an input variable of the target machine learning model, the auxiliary machine learning model configured to generate adversarial attacks from perturbed model inputs, the method comprising:
instantiating the auxiliary machine learning model configured for inferring the protected group membership from the input variable of the target machine learning model;
receiving data representative of the input variable of the target machine learning model, the target machine learning model configured to generate target predictions based on the input variable;
generating, at a given point X 0 , a first vector indicative of a gradient of the target predictions of the target machine learning model with respect to the input variable;
generating, in parallel to generating the first vector, a second vector as a locally linearized surrogate of a gradient of the input variable with respect to the protected group membership, the second vector computed in a numeric local neighborhood around the given point X 0 via a Moore-Penrose pseudoinverse of a first-order Taylor expansion utilizing a gradient of the auxiliary machine learning model evaluated at the given point X 0 ;
generating a fairness indicator value based on a projection value obtained by projecting the first vector onto the second vector and dividing by a L2 norm of the second vector;
generating downstream computing instructions representative of the fairness indicator value as compared to a predetermined fairness threshold to determine whether the target machine learning model exhibits individual-level unfairness with respect to the protected group membership, the downstream computing instructions indicating a modification instruction when the fairness indicator value is less than the predetermined fairness threshold; and
transmitting, to a downstream computing subsystem, the downstream computing instructions, the downstream computing subsystem configured to automatically select, from a plurality of candidate target machine learning models, an alternative machine learning model in accordance with the modification instruction.
13 . The method of claim 12 , wherein the gradient of the target machine learning model and the gradient of the auxiliary machine learning model are evaluated using a Fast Gradient Method approximation.
14 . The method of claim 12 , wherein the target machine learning model is a first supervised learning model, and the auxiliary machine learning model is a second supervised learning model trained at least partially based on known values of the protected group membership.
15 . The method of claim 12 , wherein the second vector is representative of an output of a sign function of the gradient of the auxiliary machine learning model evaluated at the value of the input variable.
16 . The method of claim 12 , wherein the second vector is indicative of a modified gradient when the gradient of the auxiliary machine learning model is associated with out-of-distribution predictions of the target machine learning model.
17 . The method of claim 12 , wherein a plurality of fairness indicator values is generated, each corresponding to different protected attributes of the one or more attributes, and the downstream instructions are indicative of whether an aggregated measure of the plurality of fairness indicators exceeds a predefined aggregate fairness threshold.
18 . The method of claim 17 , wherein each of the plurality of fairness indicator values is indicative of a covariance between the gradient of the target machine learning model and the gradient of the auxiliary machine learning model.
19 . The method of claim 17 , wherein the aggregated measure is an L-p norm.
20 . A non-transitory computer readable medium storing machine interpretable instructions, which when executed by a processor, cause the processor to perform a method for estimating fairness of a target machine learning model by using an auxiliary machine learning model for identifying a protected group membership from an input variable of the target machine learning model, the auxiliary machine learning model configured to generate adversarial attacks from perturbed model inputs, the method comprising:
instantiating the auxiliary machine learning model configured for inferring the protected group membership from the input variable of the target machine learning model;
receiving data representative of the input variable of the target machine learning model, the target machine learning model configured to generate target predictions based on the input variable;
generating, at a given point X 0 , a first vector indicative of a gradient of the target predictions of the target machine learning model with respect to the input variable;
generating, in parallel to generating the first vector, a second vector as a locally-linearized surrogate of a gradient of the input variable with respect to the protected group membership, the second vector computed in a numeric local neighborhood around the given point X 0 via a Moore-Penrose pseudoinverse of a first-order Taylor expansion utilizing of a gradient of the auxiliary machine learning model evaluated at the given point X 0 ;
generating a fairness indicator value based on a projection value obtained by projecting the first vector onto the second vector and dividing by a L2 norm of the second vector;
generating downstream computing instructions representative of the fairness indicator value as compared to a predetermined fairness threshold to determine whether the target machine learning model exhibits individual-level unfairness with respect to the protected group membership, the downstream computing instructions indicating a modification instruction when the fairness indicator value is less than the predetermined fairness threshold; and
transmitting, to a downstream computing subsystem, the downstream computing instructions, the downstream computing subsystem configured to automatically select, from a plurality of candidate target machine learning models, an alternative machine learning model in accordance with the modification instruction.