Latent feature based model bias mitigation in artificial intelligence systems
To eliminating bias from artificial intelligent (AI) systems, a list of class identifiers and features derived from class identifiers represented in training data fed to an AI system are identified for purpose of training a predictive model. Correlation analysis of input features is conducted from a list of raw variables, r, in a dataset and a plurality of derived features, x, with one or more class identifiers in the list of class identifiers and features derived from these class identifiers. A first list of input features is identified, one or more input features are in the first list belonging to and correlated with the one or more class identifiers or features derived from class identifiers. A second list of sets of input features is created to identify a set of combinations of input features that are not allowed to interact based on identifying biased latent features.
1 . A computer-implemented method for execution by one or more processors in a special purpose computing machine to eliminate bias from artificial intelligent (AI) systems, wherein the execution of the method comprises:
identifying a plurality of features derived from one or more class identifiers represented in training data fed to an AI system for purpose of training a predictive model in the AI system, the predictive model having:
one or more input layers,
one or more output layers,
one or more hidden layers connecting the one or more input layers to the one or more output layers, and
at least one edge connecting two layers in the predictive model, the edge representing an interaction between features in the two layers and being associated with a weight, which is adjustable to train the predictive model towards less bias;
conducting correlation analysis of input features from a list of raw variables, r, in a dataset and a plurality of derived features, x, associated with the one or more class identifiers;
identifying a first list of features, one or more features in the first list correlated with the one or more class identifiers according to the correlation analysis;
creating a second list including sets of input features associated with at least one latent feature in a hidden layer of the predictive model, the second list identifying combinations of input features that are not allowed to interact due to learned nonlinearities that result in bias in the hidden layer;
training the predictive model using the first list and the second list to eliminate the bias from the predictive model by removing from the predictive model one or more interactions between the identified set of combinations of input features that are not allowed to interact due to the learned nonlinearities,
wherein for one or more hidden layers in an interpretable neural network model, interpretable latent features in the hidden layers are extracted to investigate whether a first latent feature from among the latent features in the hidden layers contains a bias,
wherein the first latent feature is determined to be biased, in response to determining that the first latent feature results in a discriminatory distribution against a protected class of individuals identified by the one or more class identifiers,
wherein for a protected class, the latent feature output is binned into N bins, such that N is a universal constant specified per latent feature, and
wherein a two-way table is generated with counts, C ij , where a latent feature,
LF
j
k
is binned into the N bins, and a protected class PC m , has P class values, and a cell value, Cij represents an instance of the i th class value in the j th bin.
2 . The method of claim 1 , wherein an expected value Eij is given by:
E
i
j
=
(
∑
j
C
ij
)
*
(
∑
i
C
ij
)
(
∑
i
,
j
C
ij
)
.
3 . The method of claim 2 , wherein chi-square statistics is given by:
X
=
∑
i
,
j
(
C
ij
-
E
ij
)
2
E
ij
.
4 . The method of claim 3 , wherein a P-value for the chi-square statistics is computed to determine a statistical significance of difference in a chi-square distribution with df degrees of freedom.
5 . The method of claim 1 , wherein determination of a biased latent feature towards a class value results in determining the combination of input features contributing to the latent feature and the combination of features being added to the second list of sets of input features.
6 . The method of claim 5 , wherein the biased latent feature is approximated with a sparse set of multiple latent features to explode the latent feature into a set of lower complexity latent features and nonlinearities, the sparse set of lower complexity latent features being investigated for bias to determine which lower complexity latent features are identified as being biased, wherein the identified latent features are added to the second list of sets of input features.
7 . A system comprising:
at least one programmable processor; and
a non-transitory machine-readable medium storing instructions that, when executed by the at least one programmable processor, cause the at least one programmable processor to perform operations comprising:
identifying a plurality of features derived from one or more class identifiers represented in training data fed to an AI system for purpose of training a predictive model in the AI system, the predictive model having:
one or more input layers,
one or more output layers,
one or more hidden layers connecting the one or more input layers to the one or more output layers, and
at least one edge connecting two layers in the predictive model, the edge representing an interaction between features in the two layers and being associated with a weight, which is adjustable to train the predictive model towards less bias;
conducting correlation analysis of input features from a list of raw variables, r, in a dataset and a plurality of derived features, x, associated with the one or more class identifiers;
identifying a first list of features, one or more features in the first list correlated with the one or more class identifiers according to the correlation analysis;
creating a second list including sets of input features associated with at least one latent feature in a hidden layer of the predictive model, the second list identifying combinations of input features that are not allowed to interact due to learned nonlinearities that result in bias in the hidden layer;
training the predictive model using the first list and the second list to eliminate the bias from the predictive model by removing from the predictive model one or more interactions between the identified set of combinations of input features that are not allowed to interact due to the learned nonlinearities,
wherein for one or more hidden layers in an interpretable neural network model, interpretable latent features in the hidden layers are extracted to investigate whether a first latent feature from among the latent features in the hidden layers contains a bias,
wherein the first latent feature is determined to be biased, in response to determining that the first latent feature results in a discriminatory distribution against a protected class of individuals identified by the one or more class identifiers,
wherein for a protected class, the latent feature output is binned into N bins, such that N is a universal constant specified per latent feature, and
wherein a two-way table is generated with counts, C ij , where a latent feature,
LF
j
k
is binned into the N bins, and a protected class PC m , has P class values, and a cell value, Cij represents an instance of the i th class value in the j th bin.
8 . The system of claim 7 , wherein an expected value Eij is given by:
E
ij
=
(
∑
j
C
ij
)
*
(
∑
i
C
ij
)
(
∑
i
,
j
C
ij
)
.
9 . The system of claim 8 , wherein chi-square statistics is given by:
X
=
∑
i
,
j
(
C
ij
-
E
ij
)
2
E
ij
.
10 . A computer program product comprising a non-transitory machine-readable medium storing instructions that, when executed by at least one programmable processor, cause the at least one programmable processor to perform operations comprising:
identifying a plurality of features derived from one or more class identifiers represented in training data fed to an AI system for purpose of training a predictive model in the AI system, the predictive model having:
one or more input layers,
one or more output layers,
one or more hidden layers connecting the one or more input layers to the one or more output layers, and
at least one edge connecting two layers in the predictive model, the edge representing an interaction between features in the two layers and being associated with a weight, which is adjustable to train the predictive model towards less bias;
conducting correlation analysis of input features from a list of raw variables, r, in a dataset and a plurality of derived features, x, associated with the one or more class identifiers;
identifying a first list of features, one or more features in the first list correlated with the one or more class identifiers according to the correlation analysis;
creating a second list including sets of input features associated with at least one latent feature in a hidden layer of the predictive model, the second list identifying combinations of input features that are not allowed to interact due to learned nonlinearities that result in bias in the hidden layer;
training the predictive model using the first list and the second list to eliminate the bias from the predictive model by removing from the predictive model one or more interactions between the identified set of combinations of input features that are not allowed to interact due to the learned nonlinearities,
wherein for one or more hidden layers in an interpretable neural network model, interpretable latent features in the hidden layers are extracted to investigate whether a first latent feature from among the latent features in the hidden layers contains a bias,
wherein the first latent feature is determined to be biased, in response to determining that the first latent feature results in a discriminatory distribution against a protected class of individuals identified by the one or more class identifiers,
wherein for a protected class, the latent feature output is binned into N bins, such that N is a universal constant specified per latent feature, and
wherein a two-way table is generated with counts, C ij , where a latent feature,
LF
j
k
is binned into the N bins, and a protected class PC m , has P class values, and a cell value, Cij represents an instance of the i th class value in the j th bin.