IP Library › Granted Patent US 12,670,439
Granted Patent B2
US 12,670,439 · App. 18/048,341 · Granted Jun 30, 2026

Generating locally invariant explanations for machine learning

Inventors: Amit Dhurandhar (Yorktown Heights, NY); Karthikeyan Natesan Ramamurthy (Pleasantville, NY); Kartik Ahuja (Montreal, CA); Vijay Arya (Gurgaon, IN)
Assignees: International Business Machines Corporation; Université de Montréal
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,439
App. No.
18/048,341
Filed
Oct 20, 2022
Granted
Jun 30, 2026
Kind
B2
Art Unit
2144
USPC
706/12
Abstract

Techniques for generating explanations for machine learning (ML) are disclosed. These techniques include identifying an ML model, an output from the ML model, and a plurality of constraints, and generating a plurality of neighborhoods relating to the ML model based on the plurality of constraints. The techniques further include generating a predictor for each of the plurality of neighborhoods using the ML model and the plurality of constraints, constructing a combined predictor based on combining each of the respective predictors for the plurality of neighborhoods, and creating one or more explanations relating to the ML model and the output from the ML model using the combined predictor.

Claims (56)

1 . A method comprising:

identifying a machine learning (ML) model and a plurality of constraints, wherein the plurality of constraints comprises:

both: (i) a least absolute shrinkage and selection operator (LASSO) constraint and (ii) an l ∞ constraint;

wherein the l ∞ constraint provides a constraint for each predictor relating to only the neighborhood for which the predictor is generated, and

wherein the l ∞ constraint forces a matching sign relating to two generated predictors relating to two neighborhoods of the plurality of neighborhoods;

generating a plurality of neighborhoods relating to the ML model and an output from the ML model, based on the plurality of constraints;

generating a predictor for each of the plurality of neighborhoods using the ML model and the plurality of constraints, based on the output from the ML model, wherein generating the predictor comprises fitting a respective predictor to each neighborhood of the plurality of neighborhoods;

constructing a combined predictor based on combining each of the respective predictors for the plurality of neighborhoods;

creating one or more explanations relating to the ML model using the combined predictor; and

reporting the one or more explanations to explain the output from the ML model.

2 . The method of claim 1 , wherein generating the predictor for each of the plurality of neighborhoods using the input ML model and the plurality of constraints comprises:

generating a first predictor for a first neighborhood of the plurality of neighborhoods based on the plurality of constraints; and

generating a second predictor for a second neighborhood of the plurality of neighborhoods based on both an output relating to the first predictor and the plurality of constraints.

3 . The method of claim 2 , wherein constructing the combined predictor based on combining each of the respective predictors for the plurality of neighborhoods comprises:

summing a first output relating to the first predictor and a second output relating to the second predictor.

4 . The method of claim 1 , wherein the plurality of constraints comprises at least one of: (i) a least absolute shrinkage and selection operator (LASSO) constraint or (ii) an l ∞ constraint.

5 . The method of claim 1 , further comprising:

receiving an input number of neighborhoods,

wherein generating the plurality of neighborhoods comprises generating the input number of neighborhoods based on the receiving the input number of neighborhoods.

6 . The method of claim 1 , wherein generating the plurality of neighborhoods comprises using random perturbation to generate the plurality of neighborhoods.

7 . The method of claim 1 , wherein generating the plurality of neighborhoods comprises at least one of: (i) generating or (ii) selecting realistic neighborhoods.

8 . The method of claim 1 , wherein the ML model comprises a neural network, the output from the ML model comprises an inference by the neural network, and the one or more explanations relate to generating the inference using the neural network.

9 . A system, comprising:

a processor; and

a memory having instructions stored thereon which, when executed on the processor, performs operations comprising:

identifying a machine learning (ML) model, and a plurality of constraints, wherein the plurality of constraints comprises:

both: (i) a least absolute shrinkage and selection operator (LASSO) constraint and (ii) an l ∞ constraint;

wherein the l ∞ constraint provides a constraint for each predictor relating to only the neighborhood for which the predictor is generated, and

wherein the l ∞ constraint forces a matching sign relating to two generated predictors relating to two neighborhoods of the plurality of neighborhoods;

generating a plurality of neighborhoods relating to the ML model and an output from the ML model, based on the plurality of constraints;

generating a predictor for each of the plurality of neighborhoods using the ML model and the plurality of constraints, based on the output from the ML model, wherein generating the predictor comprises fitting a respective predictor to each neighborhood of the plurality of neighborhoods;

constructing a combined predictor based on combining each of the respective predictors for the plurality of neighborhoods;

creating one or more explanations relating to the ML model using the combined predictor; and

reporting the one or more explanations to explain the output from the ML model.

10 . The system of claim 9 , wherein generating the predictor for each of the plurality of neighborhoods using the input ML model and the plurality of constraints comprises:

generating a first predictor for a first neighborhood of the plurality of neighborhoods based on the plurality of constraints; and

generating a second predictor for a second neighborhood of the plurality of neighborhoods based on both an output relating to the first predictor and the plurality of constraints.

11 . The system of claim 9 , further comprising:

receiving an input number of neighborhoods,

wherein generating the plurality of neighborhoods comprises generating the input number of neighborhoods based on the receiving the input number of neighborhoods.

12 . A computer program product comprising:

a computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code executable by one or more computer processors to perform operations comprising:

identifying a machine learning (ML) model, an output from the ML model, and a plurality of constraints, wherein the plurality of constraints comprises:

both: (i) a least absolute shrinkage and selection operator (LASSO) constraint and (ii) an l ∞ constraint,

wherein the l ∞ constraint provides a constraint for each predictor relating to only the neighborhood for which the predictor is generated, and

wherein the l ∞ constraint forces a matching sign relating to two generated predictors relating to two neighborhoods of the plurality of neighborhoods;

generating a plurality of neighborhoods relating to the ML model, based on the plurality of constraints;

generating a predictor for each of the plurality of neighborhoods using the ML model and the plurality of constraints;

constructing a combined predictor based on combining each of the respective predictors for the plurality of neighborhoods; and

creating one or more explanations relating to the ML model using the combined predictor.

13 . The computer program product of claim 12 , wherein generating the predictor for each of the plurality of neighborhoods using the input ML model and the plurality of constraints comprises:

generating a first predictor for a first neighborhood of the plurality of neighborhoods based on the plurality of constraints; and

generating a second predictor for a second neighborhood of the plurality of neighborhoods based on both an output relating to the first predictor and the plurality of constraints.

14 . The computer program product of claim 12 , further comprising:

receiving an input number of neighborhoods,

wherein generating the plurality of neighborhoods comprises generating the input number of neighborhoods based on the receiving the input number of neighborhoods.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2022
From: AHUJA, KARTIK
To: UNIVERSITE DE MONTREAL
Reel/Frame 061655/0945 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2022
From: DHURANDHAR, AMIT; NATESAN RAMAMURTHY, KARTHIKEYAN; ARYA, VIJAY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061487/0609 →
Continuity (2)
Related Publication 20240135239A1 · Apr 25, 2024
Related Publication 20240232687A9 · Jul 11, 2024
References Cited (61)
US 20170277855A1 · De La Torre · 2017 [cited by examiner]
US 20180188301A1 · McBrearty · 2018 [cited by examiner]
US 20200356891A1 · Saito · 2020 [cited by applicant]
US 20210012897A1 · Katuwal et al. · 2021 [cited by applicant]
US 20210241115A1 · Ibrahim et al. · 2021 [cited by applicant]
US 20230025712A1 · Qin · 2023 [cited by examiner]
US 20230143235A1 · Dong · 2023 [cited by examiner]
US 20230309133A1 · Rafael · 2023 [cited by examiner]
US 20230410488A1 · Hamamoto · 2023 [cited by examiner]
Hansen, Jakob Vogdrup. “Combining predictors: comparison of five meta machine learning methods.” Information Sciences 119.1-2 (1999): 91-105 (Year: 1999). [cited by examiner]
Golshanrad, Paria, et al. “MEGA: Predicting the best classifier combination using meta-learning and a genetic algorithm.” Intelligent Data Analysis 25.6 (2021): 1547-1563 (Year: 2021). [cited by examiner]
Xu, et al. “Explainable AI: A Brief Survey on History, Research Areas, Approaches and Challenges”—Scientific Figure on ResearchGate. https://www.researchgate.net/figure/Two-categories-of-Explainable-AI-work-transparency… [cited by applicant]
V. Arya et al., “One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques,” [Submitted on Sep. 6, 2019 (v1), last revised Sep. 14, 2019 (this version, v2)], <https://arxiv.org/abs/1909.03… [cited by applicant]
X. Zhao et al., “BayLIME: Bayesian Local Interpretable Model-Agnostic Explanations,” [Submitted on Dec. 5, 2020 (v1) last revised May 29, 2021 (this version, v5)], https://arxiv.org/abs/2012.03058. [cited by applicant]
K. Ahuja et al., “Invariant risk minimization games,” Proc. 37th International Conference on Machine Learning, Online, PMLR 119, pp. 145-155, Nov. 2020. [cited by applicant]
K. Ramamurthy et al., “Analogies and Feature Attributions for Model Agnostic Explanation of Similarity Learners,” [Submitted on Feb. 2, 2022], <https://arxiv.org/abs/2202.01153>. [cited by applicant]
K. Ramamurthy et al., “Model Agnostic Multilevel Explanations,” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, 33, pp. 5968-5979, 2020. <https://proceedings.neurips.cc/paper/2020/has… [cited by applicant]
Grace Period Disclosure: Locally Invariant Explanations: Towards Stable and Unidirectional Explanations through Local Invariant Learning, Amit Dhurandhar, Karthikeyen Ramamurthy, Karthik Ahuja, Vijay Arya, Jan. 28, 2022. [cited by applicant]
Martin Arjovsky et al., “Invariant Risk Minimization,” arXiv:1907.02893v3, dated Mar. 27, 2020, pp. 1-31. [cited by applicant]
“Medical Expenditure Panel Survey (MEPS)”, Agency for Healthcare Research and Quality, Aug. 8, 2019, 2 pages, doi: https://meps.ahrq.gov/mepsweb/. [cited by applicant]
A. A. Shrotri et al., “Constraint-Driven Explanations for Black-Box ML Models”, Proceedings of the AAAI Conference on Artificial Intelligence, Jun. 2022, 11 pages, vol. 36, doi: 10.1609/aaai.v36i8.20805. [cited by applicant]
A. Dhurandhar et al., “Enhancing Simple Models by Exploiting What They Already Know”, Proceedings of the 37th International Conference on Machine Learning, 2020, 10 pages, vol. 119. [cited by applicant]
A. Dhurandhar et al., “Explanations based on the Missing: Towards Contrastive Explanations with Pertinent Negatives”, Advances in Neural Information Processing Systems, Feb. 2018, 22 pages, DOI: 10.48550/arXiv.1802.0762… [cited by applicant]
A. Dhurandhar et al., “Improving Simple Models with Confidence Profiles”, arXiv, Jul. 19, 2018, 17 pages, doi: https://doi.org/10.48550/arXiv.1807.07506. [cited by applicant]
A. Ghorbani et al., “Interpretation of neural networks is fragile”, arXiv, Oct. 29, 2019, 21 pages, doi: https://doi.org/10.48550/arXiv.1710.10547. [cited by applicant]
A. H. Karimi et al., “Algorithmic Recourse: from Counterfactual Explanations to Interventions”, arXiv, Feb. 14, 2020, 10 pages, doi: https://doi.org/10.48550/arXiv.2002.06278. [cited by applicant]
A. Kumar et al., “Variational Inference of Disentangled Latent Concepts from Unlabeled Observations”, arXiv, Nov. 2, 2017, 16 pages, doi: https://doi.org/10.48550/arXiv.1711.00848. [cited by applicant]
B. Kim et al., “Examples are not enough, learn to criticize! Criticism for Interpretability”, Part of Advances in Neural Information Processing Systems 29 (NIPS 2016), 2016, 9 pages. [cited by applicant]
B. Pang et al., “Thumbs up? Sentiment Classification Using Machine Learning Techniques”, arXiv, Jun. 2002, 9 pages, doi: 10.3115/1118693.1118704. [cited by applicant]
Bareinboim et al., “Local Characterizations of Causal Bayesian Networks”, in Croitoru, M., Rudolph, S., Wilson, N., Howse, J., Corby, O. (eds) Graph Structures for Knowledge Representation and Reasoning. Lecture Notes i… [cited by applicant]
C. Aggarwal et al., “Data Clustering Algorithms and Applications”, CRC Press, Aug. 2013, 20 pages. [cited by applicant]
C. Frye et al., “Asymmetric shapley values: incorporating causal knowledge into model-agnostic explainability”, arXiv, NeurIPS, Oct. 14, 2019, 14 pages, doi: https://doi.org/10.48550/arXiv.1910.06358. [cited by applicant]
C. J. Anders et al., “Fairwashing explanations with off-manifold detergent”, Proceedings of the 37th International Conference on Machine Learning, Jul. 2020, 10 pages, vol. 119. [cited by applicant]
C. K. Yeh et al., “On the (In)fidelity and Sensitivity for Explanations”, Advances in Neural Information Processing Systems, Jan. 27, 2019, 25 pages, doi: https://doi.org/10.48550/arXiv.1901.09392. [cited by applicant]
C. Molnar “Interpretable machine learning”, 2019, 251 pages, doi: https://originalstatic.aminer.cn/misc/pdf/Molnar-interpretable-machine-learning_compressed.pdf. [cited by applicant]
C. Rudin et al., “Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead”, arXiv, Nov. 26, 2018, 20 pages, doi: https://doi.org/10.48550/arXiv.1811.10154. [cited by applicant]
D. Dheeru et al., “UCI Machine Learning Repository”, Jul. 13, 2017, 2 pages, doi: https://archive.ics.uci.edu/ml. [cited by applicant]
D. Gunning et al., “Explainable artificial intelligence (xai)”, Darpa/120, Nov. 2017, 36 pages. [cited by applicant]
E. Creager et al., “Environment Inference for Invariant Learning”, Proceedings of the 38th International Conference on Machine Learning, 2021, 8 pages, vol. 139. [cited by applicant]
G. Debreu, “A Social Equilibrium Existence Theorem”, Proceedings of the National Academy of Sciences of the United States of America, 1952, pp. 886-893, vol. 38, https://doi.org/10.1073/pnas.38.10.886. [cited by applicant]
G. Hinton et al., “Distilling the Knowledge in a Neural Network”, arXiv, Mar. 9, 2015, 9 pages, doi: https://doi.org/10.48550/arXiv.1503.02531. [cited by applicant]
G. Plump et al., “Model Agnostic Supervised Local Explanations”, arXiv, Jul. 9, 2018, 10 pages, doi: https://doi.org/10.48550/arXiv.1807.02910. [cited by applicant]
H. Lakkaraju et al., “Robust and Stable Black Box Explanations”, arXiv, Nov. 12, 2020, 14 pages, doi: https://doi.org/10.48550/arXiv.2011.06169. [cited by applicant]
H. Xiao et al., “Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms”, arXiv, Aug. 25, 2017, 6 pages, doi: https://doi.org/10.48550/arXiv.1708.07747. [cited by applicant]
H. Zhang et al., “The Limitations of Adversarial Training and the Blind-Spot Attack”, arXiv, Jan. 15, 2019, 16 pages, doi: https://doi.org/10.48550/arXiv.1901.04684. [cited by applicant]
Judea Pearl et al., “Casuality”, Cambridge University Press, 2009, 487 pages, doi: https://doi.org/10.1017/CBO9780511803161. [cited by applicant]
K. Ahuja et al., “Linear Regression Games: Convergence Guarantees to Approximate Out-of-Distribution Solutions”, arXiv, Oct. 28, 2020, 38 pages, doi: https://doi.org/10.48550/arXiv.2010.15234. [cited by applicant]
K. Gurumoorthy et al., “Efficient Data Representation by Selecting Prototypes with Importance Weights”, ArXiv, Jul. 5, 2017, 10 pages, doi: https://doi.org/10.48550/arXiv.1707.01212. [cited by applicant]
K. Simonyan et al., “Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps”, arXiv, Dec. 20, 2013, 8 pages, doi: https://doi.org/10.48550/arXiv.1312.6034. [cited by applicant]
L. Hancox-Li, “Robustness in machine learning explanations: does it matter?”, In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. ACM, Jan. 2020, 8 pages, doi: http://dx.doi.org/10.1145/… [cited by applicant]
M. Sundararajan et al., “Axiomatic Attribution for Deep Networks”, Proceedings of the 34th International Conference on Machine Learning, Mar. 2017, 10 pages, vol. 70, doi: https://doi.org/10.48550/arXiv.1703.01365. [cited by applicant]
M. T. Ribeiro et al., “Why Should I Trust You?': Explaining the Predictions of Any Classifier”, arXiv, Feb. 16, 2016, 10 pages, doi: https://doi.org/10.48550/arXiv.1602.04938. [cited by applicant]
P. K. Dutta “Strategies and Games: Theory and Practice”, MIT Press, 1999, 384 pages, doi: https://arteref.com/wp-content/uploads/2019/12/Prajit-K.-Dutta-Strategies-and-Games_-Theory-and-Practice-The-MIT-Press-1999.pdf. [cited by applicant]
P. N. Yannella et al., “Analysis: Article 29 Working Party Guidelines on Automated Decision Making Under GDPR”, Posted in Article 29 Working Party, Compliance, Cross-Border Data Flow, Data Protection, EU Regulation, Gen… [cited by applicant]
S. Feng et al., “Pathologies of Neural Models Make Interpretations Difficult”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Oct. 2018, pp. 3719-3728, doi: https://arxiv.org/abs… [cited by applicant]
S. Lapuschkin et al., “The LRP Toolbox for Artificial Neural Networks”, Journal of Machine Learning Research, 2016, 5 pages, vol. 17, No. 114. [cited by applicant]
S. M. Lundberg et al., “A Unified Approach to Interpreting Model Predictions”, Advances in Neural Information Processing Systems, Dec. 2017, 10 pages, doi: 10.48550/arXiv.1705.07874. [cited by applicant]
T. Botari et al., “MeLIME: Meaningful Local Explanation for Machine Learning Models”, arXiv, Sep. 12, 2020, 21 pages, doi: https://doi.org/10.48550/arXiv.2009.05818. [cited by applicant]
T. Heskes et al., “Causal Shapley Values: Exploiting Causal Knowledge to Explain Individual Predictions of Complex Models”, Advances in Neural Information Processing Systems, Nov. 3, 2020, 21 pages, doi: https://doi.org… [cited by applicant]
Tim Miller, “Explanation in artificial intelligence: Insights from the social sciences”, Artificial Intelligence, Feb. 2019, pp. 1-38, vol. 267, doi: https://doi.org/10.1016/j.artint.2018.07.007. [cited by applicant]
Y. Roth et al., “Updating our approach to misleading information”, X Blog, May 11, 2020, 3 pages, doi: https://blog. twitter.com/en_us/topics/product/2020/updating-our-approach-to-misleading-information.html. [cited by applicant]