Compensating for vulnerabilities in machine learning algorithms
A method performed by a processing system including at least one processor includes obtaining an output of a machine learning algorithm, identifying a vulnerability in the output of the machine learning algorithm, wherein the vulnerability relates to a bias in the output, integrating auxiliary data from an auxiliary data source of a plurality of auxiliary data sources into the machine learning algorithm to try to compensate for the vulnerability, determining whether the integrating has compensated for the vulnerability, and generating a runtime output using the machine learning algorithm when the processing system determines that the integrating has compensated for the vulnerability.
1 . A method comprising:
obtaining, by a processing system including at least one processor, an output of a machine learning model;
identifying, by the processing system, a vulnerability in the output of the machine learning model, wherein the vulnerability causes the machine learning model to perpetuate a bias against a subgroup of a human population in the output, and wherein the vulnerability is identified by detecting, in a set of training data used to train the machine learning model, at least one of: an item of training data which is known to reflect the bias, an input feature which is known to be susceptible to the bias, or a combination of features and feature values which is known to be susceptible to the bias;
augmenting, by the processing system, the training data with auxiliary data that is curated to address the vulnerability to produce augmented training data, wherein the auxiliary data includes outputs of another machine learning model for which another vulnerability that caused the another machine learning model to perpetuate the bias against the subgroup of the human population has been minimized;
retraining, by the processing system, the machine learning model using the augmented training data to produce an updated version of the machine learning model in which the bias is minimized, wherein the retraining assigns a greater weight to the auxiliary data than to other data in the augmented training data to minimize an influence of the bias in the augmented training data;
determining, by the processing system, whether the retraining has minimized the bias in the updated version of the machine learning model by a desired amount; and
generating, by the processing system, a runtime output using the updated version of the machine learning model when the processing system determines that the retraining has minimized the bias by the desired amount.
2 . The method of claim 1 , wherein the output comprises at least one of: generated multimedia content, a list of prioritized samples, or attributes and values that are considered high-value by the machine learning model or based on a domain knowledge.
3 . The method of claim 1 , wherein the bias is a result of a misrepresentation of at least one aspect of a sample provided to the machine learning model.
4 . The method of claim 3 , wherein the misrepresentation is due to at least one of: a bias of a human who labeled the sample, a bias in a sampling process used to generate the sample, or a systemic bias.
5 . The method of claim 1 , wherein the augmenting further comprises weighting the auxiliary data more heavily than training data originally used to train the machine learning model during the retraining.
6 . The method of claim 1 , wherein the auxiliary data comprises a collection of ground-truth data that is used to enhance a degree of confidence in the output of the machine learning model.
7 . The method of claim 1 , wherein the auxiliary data is curated to target a different potential vulnerability of a plurality of potential vulnerabilities.
8 . The method of claim 1 , wherein the augmenting further comprises utilizing the auxiliary data as an alarm mechanism to detect when the output of the machine learning model differs by more than a threshold from items of the auxiliary data that have input features similar to the output of the machine learning model.
9 . The method of claim 1 , wherein the determining comprises re-running the machine learning model and checking an output of the re-running for the vulnerability.
10 . The method of claim 1 , further comprising:
repeating, by the processing system subsequent to the determining but prior to the generating, the augmenting and the re-training when the processing system determines that the augmenting has not minimized the bias.
11 . The method of claim 1 , wherein the auxiliary data comprises a newly discovered auxiliary data, and wherein the auxiliary data is stored for use in minimizing a bias in a future machine learning model.
12 . The method of claim 1 , further comprising, subsequent to the identifying but prior to the augmenting:
prioritizing, by the processing system, the vulnerability relative to a plurality of vulnerabilities that were identified in the output of the machine learning model, wherein the prioritizing indicates which vulnerabilities of the plurality of vulnerabilities are most important to compensate for.
13 . The method of claim 12 , wherein an importance of a respective vulnerability of the plurality of vulnerabilities is defined in terms of an impact on a performance of the machine learning model.
14 . The method of claim 12 , wherein an importance of a respective vulnerability of the plurality of vulnerabilities is defined in terms of a fairness consideration.
15 . The method of claim 6 , wherein the ground-truth data is independently verified by a party who is separate from a party who is performing at least one of: analyzing the machine learning model, using the machine learning model, providing the auxiliary data, or maintaining the auxiliary data.
16 . The method of claim 1 , wherein the machine learning model is a neural network.
17 . The method of claim 1 , wherein the machine learning model is deployed to moderate speech in a social media space, and the auxiliary data comprises a set of keywords that is considered acceptable in a context of the social media space, but expected to trigger blocking by machine learning models in other spaces.
18 . The method of claim 1 , wherein the auxiliary data contains data that was unknown to the machine learning model at a time at which the machine learning model was initially trained using the training data.
19 . A non-transitory computer-readable medium storing instructions which, when executed by a processing system including at least one processor, cause the processing system to perform operations, the operations comprising:
obtaining an output of a machine learning model;
identifying a vulnerability in the output of the machine learning model, wherein the vulnerability causes the machine learning model to perpetuate a bias against a subgroup of a human population in the output, and wherein the vulnerability is identified by detecting, in a set of training data used to train the machine learning model, at least one of: an item of training data which is known to reflect the bias, an input feature which is known to be susceptible to bias, or a combination of features and feature values which is known to be susceptible to bias;
augmenting the training data with auxiliary data that is curated to address the vulnerability to produce augmented training data, wherein the auxiliary data includes outputs of another machine learning model for which another vulnerability that caused the another machine learning model to perpetuate the bias against the subgroup of the human population has been minimized;
retraining the machine learning model using the augmented training data to produce an updated version of the machine learning model in which the bias is minimized, wherein the retraining assigns a greater weight to the auxiliary data than to other data in the augmented training data;
determining whether the retraining has minimized the bias in the updated version of the machine learning model by a desired amount; and
generating a runtime output using the updated version of the machine learning model when the processing system determines that the retraining has minimized the bias by the desired amount.
20 . A device comprising:
a processing system including at least one processor; and
a non-transitory computer-readable medium storing instructions which, when executed by the processing system, cause the processing system to perform operations, the operations comprising:
obtaining an output of a machine learning model;
identifying a vulnerability in the output of the machine learning model, wherein the vulnerability causes the machine learning model to perpetuate a bias against a subgroup of a human population in the output, and wherein the vulnerability is identified by detecting, in a set of training data used to train the machine learning model, at least one of: an item of training data which is known to reflect the bias, an input feature which is known to be susceptible to bias, or a combination of features and feature values which is known to be susceptible to bias;
augmenting the training data with auxiliary data that is curated to address the vulnerability to produce augmented training data, wherein the auxiliary data includes outputs of another machine learning model for which another vulnerability that caused the another machine learning model to perpetuate the bias against the subgroup of the human population has been minimized;
retraining the machine learning model using the augmented training data to produce an updated version of the machine learning model in which the bias is minimized, wherein the retraining assigns a greater weight to the auxiliary data than to other data in the augmented training data;
determining whether the retraining has minimized the bias in the updated version of the machine learning model by a desired amount; and
generating a runtime output using the updated version of the machine learning model when the processing system determines that the retraining has minimized the bias by the desired amount.