IP Library › Granted Patent US 12,493,796
Granted Patent B2
US 12,493,796 · App. 17/124,018 · Granted Dec 9, 2025

Using generative adversarial networks to construct realistic counterfactual explanations for machine learning models

Inventors: Karoon Rashedi Nia (Vancouver, CA); Tayler Hetherington (Vancouver, CA); Zahra Zohrevand (Vancouver, CA); Yasha Pushak (Vancouver, CA); Sanjay Jinturkar (Santa Clara, CA); Nipun Agarwal (Saratoga, CA)
Assignee: Oracle International Corporation
G06N3/088G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,796
App. No.
17/124,018
Granted
Dec 9, 2025
Kind
B2
Abstract

Herein are counterfactual explanations of machine learning (ML) inferencing provided by generative adversarial networks (GANs) that ensure realistic counterfactuals and use latent spaces to optimize perturbations. In an embodiment, a first computer trains a generator model in a GAN. A same or second computer hosts a classifier model that inferences an original label for original feature values respectively for many features. Runtime ML explainability (MLX) occurs on the first or second or a third computer as follows. The generator model from the GAN generates a sequence of revised feature values that are based on noise. The noise is iteratively optimized based on a distance between the original feature values and current revised feature values in the sequence of revised feature values. The classifier model inferences a current label respectively for each counterfactual in the sequence of revised feature values. Satisfactory discovered counterfactuals are promoted as explanations of behavior of the classifier model.

Claims (40)

1 . A method comprising:

inferring, by a trained classifier, an original classification from at least three classes for an original plurality of feature values respectively for a plurality of features;

generating, by a generator model from a generative adversarial network (GAN), a first revised plurality of feature values that are based on noise;

detecting that the trained classifier infers the original classification from the first revised plurality of feature values;

optimizing the noise, including minimizing a loss, including increasing the loss in response to said detecting, wherein the loss is based on a distance between: i) the first revised plurality of feature values and ii) the original plurality of feature values;

generating, by the generator model after said optimizing, a second revised plurality of feature values that are based on the noise;

inferring, by the trained classifier, a second classification for the second revised plurality of feature values, wherein the second classification is different from the original classification; and

generating and displaying a counterfactual explanation, for the original classification, that comprises the second revised plurality of feature values;

wherein the method is performed by one or more computers.

2 . The method of claim 1 wherein said inferring for each revised plurality of feature values comprises repeatedly inferring until a threshold amount of a sequence of classifications have a desired class that is not a class of the original classification.

3 . The method of claim 1 wherein said optimizing the noise is further based on at least one selected from a group consisting of a partial derivative of the noise and a partial derivative of said distance.

4 . The method of claim 1 wherein said optimizing the noise entails none of: randomness, training the generator model, and adjusting the generator model.

5 . The method of claim 1 wherein said optimizing the noise entails at least one selected from a group consisting of: gradient descent without a machine learning model, backpropagation without an artificial neural network, and adjustment of a plurality of numbers that initially were randomly generated as said noise.

6 . The method of claim 5 wherein said adjustment of the plurality of numbers comprises replacement, by random generation, of a number of the plurality of numbers upon occurrence of a resampling condition that is based on at least one of: said number, a mean, and a standard deviation.

7 . The method of claim 5 further comprising initially randomly generating said plurality of numbers in said noise based on a normal distribution having at least one statistic selected from a group consisting of: zero as a mean and one as a standard deviation.

8 . The method of claim 1 wherein said generating said first revised plurality of feature values comprises the generator model from the GAN inferring based on a noise feature vector that consists of said noise.

9 . The method of claim 8 wherein said generator model inferring comprises the generator model from the GAN inferring based on said noise feature vector that does not contain a same amount of members as the plurality of features.

10 . The method of claim 1 without using an autoencoder.

11 . The method of claim 1 further comprising training the generator model in the GAN based on a training corpus that does not contain said original plurality of feature values.

12 . The method of claim 1 further comprising:

training the generator model in the GAN;

training a discriminator model in the GAN;

deciding, by said discriminator model, to exclude said first revised plurality of feature values.

13 . The method of claim 12 further comprising in response to said deciding to exclude said first revised plurality of feature values, performing one selected from a group consisting of:

said optimizing the noise, and

replacement, by random generation, of at least one number in the noise.

14 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

inferring, by a trained classifier, an original classification from at least three classes for an original plurality of feature values respectively for a plurality of features;

generating, by a generator model from a generative adversarial network (GAN), a first revised plurality of feature values that are based on noise;

detecting that the trained classifier infers the original classification from the first revised plurality of feature values;

optimizing the noise, including minimizing a loss, including increasing the loss in response to said detecting, wherein the loss is based on a distance between: i) the first revised plurality of feature values and ii) the original plurality of feature values;

generating, by the generator model after said optimizing, a second revised plurality of feature values that are based on the noise;

inferring, by the trained classifier, a second classification for the second revised plurality of feature values, wherein the second classification is different from the original classification; and

generating and displaying a counterfactual explanation, for the original classification, that comprises the second revised plurality of feature values.

15 . The one or more non-transitory computer-readable media of claim 14 wherein said inferring for each revised plurality of feature values comprises repeatedly inferring until a threshold amount of a sequence of classifications have a desired class that is not a class of the original classification.

16 . The one or more non-transitory computer-readable media of claim 14 wherein said optimizing the noise is further based on at least one selected from a group consisting of a partial derivative of the noise and a partial derivative of said distance.

17 . The one or more non-transitory computer-readable media of claim 14 wherein said optimizing the noise entails none of: randomness, training the generator model, and adjusting the generator model.

18 . The one or more non-transitory computer-readable media of claim 14 wherein said optimizing the noise entails at least one selected from a group consisting of: gradient descent without a machine learning model, backpropagation without an artificial neural network, and adjustment of a plurality of numbers that initially were randomly generated as said noise.

19 . The one or more non-transitory computer-readable media of claim 14 wherein said generating said first revised plurality of feature values comprises the generator model from the GAN inferring based on a noise feature vector that consists of said noise.

20 . The one or more non-transitory computer-readable media of claim 14 wherein when executed the instructions do not use an autoencoder.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2021
From: NIA, KAROON RASHEDI; HETHERINGTON, TAYLER; ZOHREVAND, ZAHRA; PUSHAK, YASHA; JINTURKAR, SANJAY; AGARWAL, NIPUN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 054896/0521 →
Continuity (1)
Related Publication 20220188645A1 · Jun 16, 2022
References Cited (74)
US 20190130212A1 · Cheng · 2019 [cited by examiner]
US 20190197357A1 · Anderson et al. · 2019 [cited by applicant]
US 20190197368A1 · Madani · 2019 [cited by examiner]
US 20190385019A1 · Bazrafkan · 2019 [cited by examiner]
US 20200098139A1 · Kaplanyan · 2020 [cited by examiner]
US 20200110982A1 · Gou · 2020 [cited by applicant]
US 20200226212A1 · Tan · 2020 [cited by applicant]
US 20200302318A1 · Hetherington · 2020 [cited by applicant]
US 20220129791A1 · Nia et al. · 2022 [cited by applicant]
Linardatos et al., “Explainable AI: A Review of Machine Learning Interpretability Methods”, Entropy, 23(1), 2021, 45 pages. [cited by applicant]
Kurakin et al., “Adversarial Examples In The Physical World”, ICLR 2017, Feb. 11, 2017, 14 pages. [cited by applicant]
Goodfellow et al., “Explaining and Harnessing Adversarial Examples”, ICLR 2015, Mar. 20, 2015, 11 pages. [cited by applicant]
Saito et al., Improving LIME Robustness with Smarter Locality Sampling, arXiv.org, available at https://arxiv.org/pdf/2006.12302v2.pdf (revised Jul. 24, 2020). [cited by applicant]
Molnar et al., Limitations of Interpretable Machine Learning Methods, available at https://slds-lmu.github.io/iml_methods_limitations/ (Oct. 5, 2020). [cited by applicant]
Laugel, Local post-hoc interpretability for black-box classifiers. Machine Learning [cs.LG]. Sorbonne Universite, 2020 (Year: 2020). [cited by applicant]
Laugel et al., Defining Locality for Surrogates in Post-hoc Interpretablity, arXiv.org, available at https://arxiv.org/pdf/1806.07498.pdf (Jun. 19, 2018). [cited by applicant]
Lahiri A, Edakunni NU. Accurate and Intuitive Contextual Explanations using Linear Model Trees. arXiv preprint arXiv:2009.05322. Sep. 11, 2020, originally presented at KDD Workshop on ML in Finance 2020, Aug. 24, 2020 (… [cited by applicant]
Guidotti R, Monreale A, Ruggieri S, Pedreschi D, Turini F, Giannotti F. Local rule-based explanations of black box decision systems. arXiv preprint arXiv:1805.10820. May 28, 2018 (Year: 2018). [cited by applicant]
Deep Hyperspherical Learning, arXiv.org, available at https://arxiv.org/pdf/1711.03189.pdf (Jan. 30, 2018). [cited by applicant]
Letham et al., “Interpretable Classifiers Using Rules and Bayesian Analysis: Building a Better Stroke Prediction Model”, The Annals of Applied Statistics, dated 2015, vol. 9, No. 3, 23 pages. [cited by applicant]
Agrawal et al., “Fast Discovery of Association Rules”, dated 1996, 22 pages. [cited by applicant]
Baehrens et al., “How to Explain Individual Classification Decisions”, Journal of Machine Learning Research 11 dated 2010, 29 pages. [cited by applicant]
Dhurandhar et al., “Explanations based on the Missing: Towards Contrastive Explanations with Pertinent Negatives”, dated 2018, 12 pages. [cited by applicant]
Dougherty et al., “Supervised and Unsupervised Discretization of Continous Features”, dated 1995, 9 pages. [cited by applicant]
Egan et al., “Generalized Latent Variable Recovery for Generative Adversarial Networks”, dated Oct. 19, 2018, 9 pages. [cited by applicant]
Fayyad, “Multi-Interval Discretization of Continuous-Valued Attributed for Classification Learning”, dated 1993, 6 pages. [cited by applicant]
Goodfellow et al., “Explaining and Harnessing Adversarial Examples”, Published as a conference paper at ICLR dated 2015, 11 pages. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets”, dated 2014, 9 pages. [cited by applicant]
Holte, Robert, “Very Simple Classification Rules Perform Well on Most Commonly Used Datasets”, 1993 Kluwer Academic Publishers, Boston. Manufactured in The Netherlands, 28 pages. [cited by applicant]
Jin et al., “Data discretization unification”, Regular Paper, Springer-Verlag London Limited 2008, 29 pages. [cited by applicant]
Lakkaraju et al., “Interpretable & Explorable Approximations of Black Box Models”, KDD, dated 2017, 5 pages. [cited by applicant]
“Decision Trees”, dated Jan. 26, 2005, 10 pages. [cited by applicant]
Laugel et al., “Comparison-based Inverse Classification for Interpretability in Machine Learning”., (IPMU 2018), dated Jun. 2018, Cadix, Spain, 13 pages. [cited by applicant]
Xu et al., “Modeling Tabular Data using Conditional GAN”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), dated 2019, Vancouver, Canada, 11 pages. [cited by applicant]
Letham et al., “Interpretable classifiers using rules and Bayesian analysis: Building a better stroke prediction model”, vol. 9, No. 3, dated 2015, 2 pages. [cited by applicant]
Lin et al., “Experiencing SAX: a novel symbolic representation of time series”, Data Min Knowl Disc (2007), 38 pages. [cited by applicant]
Lipton et al., “Precise Recovery of Latent Vectors From Generative Adversarial Networks”, Workshop track—ICLR dated Feb. 17, 2017, 4 pages. [cited by applicant]
Pedregosa et al., “Scikit-learn: Machine Learning in Python”, Journal of Machine Learning Research 12, dated 2011, 6 pages. [cited by applicant]
Quinlan et al., “Induction of Decision Trees”, dated 1986 Kluwer Academic Publishers, Boston—Manufactured in The Netherlands, 26 pages. [cited by applicant]
Ramírez-Gallego et al., “Data discretization: taxonomy and big data challenge”, WIREs Data Mining Knowl Discov 2015, 17 pages. [cited by applicant]
Ribeiro et al., “Why Should I Trust You?” Explaining the Predictions of Any Classifier, Publication rights licensed to ACM, dated 2016, 10 pages. [cited by applicant]
Robnik-Sikonja et al., “Explaining Classifications for Individual Instances”, IEEE Transactions on Knowledge and Data Engineering, 20:589-600, dated 2008, 24 pages. [cited by applicant]
Samek et al., “EXPLAINABLEARTIFICIALINTELLIGENCE: Understanding, VISUALIZINGAND Interpreting Deep Learning Models”, ITU Journal: ICT Discoveries, Special Issue No. 1, Oct. 13, 2017, 10 pages. [cited by applicant]
Utgoff, Paul, “Incremental Induction of Decision Trees”, 1989 Kluwer Academic Publishers, Boston. Manufactured in The Netherlands, 26 pages. [cited by applicant]
Van Looveren et al., “Interpretable Counterfactual Explanations Guided by Prototypes”, dated Feb. 18, 2020, 17 pages. [cited by applicant]
Wachter et al., “Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR”, dated 2017, 52 pages. [cited by applicant]
Lakkaraju et al., “Interpretable Decision Sets: A Joint Framework for Description and Prediction”, KDD, PMC dated Nov. 14, 2016, 24 pages. [cited by applicant]
Zhao et al., “Data-driven risk-averse stochastic optimization with Wasserstein metric”, Operations Research Letters, dated 2018, 6 pages. [cited by applicant]
Vanschoren et al., “OpenML: Networked Science in Machine Learning”, dated Aug. 1, 2014, 12 pages. [cited by applicant]
Sweke et al., “On the Quantum versus Classical Learnability of Discrete Distributions”, dated Jul. 28, 2020, 27 pages. [cited by applicant]
Strumbelj et al., “Explaining Predictions Models and Individual Predictions with Contributions”, Knowledge and Information Systems 41.3, dated Dec. 2014, pp. 647-665. [cited by applicant]
Strumbelj et al., “An Efficient Explanation of Individual Classifications Using Game Theory”, Journal of Machine Learning Research 11, dated 2010, 18 pages. [cited by applicant]
Shapley, Lloyd S. “A Value for N-person Games”, Contributions to the Theory of Games 2.28, dated Aug. 21, 1951, 19 pages. [cited by applicant]
Roth, Alvin, “The Shapley Value”, Essays in honor of Lloyd S. Shapley, Cambridge University Press, dated 1988, 338 pages. [cited by applicant]
Ribeiro et al., “Why Should I Trust You?” Explaining the Predictions of Any Classifier, KDD dated 2016 San Francisco, CA, USA, 10 pages. [cited by applicant]
Plumb et al., “Model Agnostic Supervised Local Explanations”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), dated 2018, Montréal, Canada, 10 pages. [cited by applicant]
Lundberg et al., “A Unified Approach to Interpreting Model Predictions”, 31st Conference on Neural Information Processing Systems (NIPS 2017), dated 2017, Long Beach, CA, USA, 10 pages. [cited by applicant]
Laugel et al., “Defining Locality for Surrogates in Post-hoc Interpretablity”, dated Jun. 19, 2018 ICML Workshop on Human Interpretability in Machine Learning (WHI 2018), 7 pages. [cited by applicant]
Heusel et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 12 pages. [cited by applicant]
Github.com, “slundberg/shap”, https://github.com/slundberg/shap/blob/master/shap/explainers/_permutation.py, updated on Nov. 17, 2020, 5 pages. [cited by applicant]
Dua, D. and Graff, C., “UCI Machine Learning Repository” [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science, dated 2019, 2 pages. [cited by applicant]
Bloniarz et al., “Supervised Neighborhoods for Distributed Nonparametric Regression”, Proceedings of the 19th International Conference on Artificial Intelligence and Statistics dated 2016, 10 pages. [cited by applicant]
Yeh RA et al., “Semantic image inpainting with deep generative models”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 5485-5493). (Year: 2017). [cited by applicant]
Srivastava A. et al., “VEEGAN—Reducing Mode Collapse in GANs using Implicit Variational Learning”, Advances in Neural Information Processing Systems (Year: 2017). [cited by applicant]
Liu Y. et al., “Collaborative sampling in generative adversarial networks”, In Proceedings of the AAAI Conference on Artificial Intelligence 2020 (vol. 34, No. 04, pp. 4948-4956) (Apr. 3, 2020). [cited by applicant]
Jung AB, “Learning to avoid errors in gans by manipulating input spaces”, ARXIV Preprint ARXIV:1707.00768 (Jul. 3, 2017). [cited by applicant]
Hoshen Y. et al., “Non-adversarial image synthesis with generative latent nearest neighbors”, Inproceedings of the IEEE/Cvf Conference on Computer Vision and Pattern Recognition (pp. 5811-5819) (Year: 2019). [cited by applicant]
Che T. et al., “Your gan is secretly an energy-based model and you should use discriminator driven latent sampling. Advances in Neural Information Processing Systems” 12275-87 (Year: 2020). [cited by applicant]
Bhattarai B. et al., “Sampling strategies for gan synthetic data”, INICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 2303-2307) (May 4, 2020). [cited by applicant]
Arvanitidis G. et al., “Latent space oddity: on the curvature of deep generative models”, ARXIV Preprint ARXIV:1710.11379 (Oct. 31, 2017). [cited by applicant]
Vlassopoulos G. et al., “Explaining predictions by approximating the local decision boundary”, arXiv preprint arXiv:2006.07985. (Jun. 14, 2020). [cited by applicant]
Ren, S. et al., “Generating natural language adversarial examples through probability weighted word saliency”, In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 1085-1097) (… [cited by applicant]
Greco, S. “Explaining black-box models in the context of Natural Language Processing” (Doctoral dissertation, Politecnico diTorino). (Year: 2019). [cited by applicant]
Alvarez-Melis, D. et al., “A causal framework for explaining the predictions of black-box sequence-to-sequence models”, ARXIV Preprint ARXIV:1707.01943 (Nov. 14, 2017). [cited by applicant]