Removing undesirable inferences from a machine learning model
A method and system for removing undesirable inferences from a machine learning model include a search component configured to receive a rejected explanation of model output provided by the machine learning model, identify data samples to unlearn by selecting training samples from training data that were used to train the machine learning model, the selected training samples being associated with explanations that are similar to the rejected explanation according to a calculated similarity measure, and pass the data samples to unlearn to a machine unlearning unit.
1 . A system for removing undesirable inferences from a machine learning model, the system comprising:
a machine learning model stored on a computer;
a machine unlearning unit executed by the computer;
a search component, the search component being executed by the computer and configured to:
receive a rejected explanation of model output provided by the machine learning model, wherein the rejected explanation is an explanation of a recommendation produced by the machine learning model;
identify data samples to unlearn by selecting training samples from training data that were used to train the machine learning model, the selected training samples being associated with training explanations that are similar to the rejected explanation according to a calculated similarity measure; and
pass the data samples to unlearn to the machine unlearning unit;
wherein either:
the explanations are rule-based, and wherein the search component is configured to identify the training explanations that are associated with training samples as being similar to the rejected explanation in response to one or more of i) premises, ii) structure, and iii) decision values of the rule-based explanations being the same or similar; or
the explanations comprises feature explanations, wherein each feature is assigned an importance weight, and wherein the search component is configured to calculate the explanation similarity measure as either i) a distance metric between the rejected explanation and the training explanations, the distance metric comprising one or more of a Euclidean distance, a City-Block distance, a Mahalanobis distance, a Huber distance, an elastic distance measure, or a Jaccard similarity coefficient;
wherein the machine unlearning unit is configured to remove undesirable inferences from the machine learning model by either:
i) assigning, to the data samples to unlearn, weightings for a training loss function before retraining the machine learning model using an updated set of the training data; or
ii) retraining the machine learning model using an updated set of the training data from which the data samples to unlearn have been removed; or
iii) altering the data samples to unlearn by finding perturbations to the data samples that result in an explanation derived from the perturbed data samples matching a non-rejected explanation;
wherein the machine learning model is configured to receive, as the training data, one or more signals from sensors at monitored and operating industrial plant equipment and to produce recommendations to an end user to take an action, the action involving controlling the operation of the industrial equipment.
2 . The system of claim 1 , wherein the search component is configured to identify the data samples to unlearn by:
obtaining a training explanation associated with a respective candidate training sample selected from the training data;
calculating an explanation similarity measure between the rejected explanation and the training explanation;
comparing the calculated explanation similarity measure to an explanation similarity threshold; and
selecting the training sample for inclusion in the data samples to unlearn in response to the calculated explanation similarity measure satisfying the explanation similarity threshold.
3 . The system of claim 2 , wherein the search component is further configured to calculate a sample similarity measure between the rejected sample and the training sample, to compare the calculated sample similarity measure to a sample similarity threshold, and to identify as data samples to unlearn only those selected training samples whose calculated sample similarity measures also satisfy the sample similarity threshold.
4 . The system of claim 2 , wherein the search component is further configured to repeat the obtaining, calculating, comparing, and selecting for one or more further training samples in the training data.
5 . The system of claim 1 , wherein the weightings assigned to the data samples to unlearn are based on the similarity measures calculated for those data samples.
6 . The system of claim 1 , wherein the weightings are assigned to some but not all of a plurality of signals in the data samples to unlearn.
7 . The system of claim 1 , wherein the machine unlearning unit is configured incrementally to change the weightings based on the Implicit Function Theorem applied to optimality conditions of a machine model training loss minimization problem.
8 . A method for removing undesirable inferences from a machine learning model, the method comprising:
receiving a rejected explanation of model output provided by the machine learning model, wherein the rejected explanation is an explanation of a recommendation produced by the machine learning model;
identifying data samples to unlearn by selecting training samples from training data that were used to train the machine learning model, the selected training samples being associated with explanations that are similar to the rejected explanation according to a calculated similarity measure; and
passing the data samples to unlearn to a machine unlearning unit;
by a machine learning unit, removing undesirable inferences from the machine learning model by either:
i) assigning, to the data samples to unlearn, weightings for a training loss function before retraining the machine learning model using an updated set of the training data;
ii) retaining the machine learning model using an updated set of the training data from which the data samples to unlearn have been removed, or
iii) altering the data samples to unlearn by finding perturbations to the data samples that result in an explanation derived from the perturbed data samples matching a non-rejected explanation;
wherein the machine learning model is configured to receive, as the training data, one or more signals from sensors at monitored and operating industrial plant equipment and to produce recommendations to an end user to take an action, the action involving controlling the operation of the industrial equipment.
9 . The method of claim 8 , wherein the method is carried out by a search component, the method further comprising:
obtaining a training explanation associated with a respective candidate training sample selected from the training data;
calculating an explanation similarity measure between the rejected explanation and the training explanation;
comparing the calculated explanation similarity measure to an explanation similarity threshold; and
selecting the training sample for inclusion in the data samples to unlearn in response to the calculated explanation similarity measure satisfying the explanation similarity threshold.
10 . The method of claim 9 , wherein the search component is further configured to calculate a sample similarity measure between the rejected sample and the training sample, to compare the calculated sample similarity measure to a sample similarity threshold, and to identify as data samples to unlearn only those selected training samples whose calculated sample similarity measures also satisfy the sample similarity threshold.
11 . The method of claim 9 , wherein the search component is further configured to repeat the obtaining, calculating, comparing, and selecting for one or more further training samples in the training data.
12 . The method of claim 9 , wherein the search component is configured to calculate the explanation similarity measure as either (i) a distance metric between the rejected explanation and the training explanation, the distance metric comprising one or more of a Euclidian distance, a City-Block distance, a Mahalanobis distance, a Huber distance, or an elastic distance measure such as Dynamic Time Warping or a Levenshtein distance, or (ii) a Jaccard similarity coefficient.
13 . The method of claim 8 , wherein the explanations are rule-based, and wherein the search component is configured to identify the explanations that are associated with training samples as being similar to the rejected explanation in response to one or more of (i) premises, (ii) structure, and (iii) decision values of the rule-based explanations being the same or similar.
14 . The method of claim 8 , wherein the explanations comprise prototypes, and wherein the search component is configured to identify the explanations that are associated with training samples as being similar to the rejected explanation in response to the same prototype being produced.