IP Library › Granted Patent US 12,271,789
Granted Patent B2
US 12,271,789 · App. 17/197,535 · Granted Apr 8, 2025

Interpretable model changes

Inventors: Elizabeth Daly (Dublin, IE); Rahul Nair (Dublin, IE); Oznur Alkan (Dublin, IE); Massimiliano Mattetti (Dublin, IE); Dennis Wei (Sunnyvale, CA); Yunfeng Zhang (Chappaqua, NY)
Assignee: International Business Machines Corporation
G06N20/00G06N5/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,271,789
App. No.
17/197,535
Granted
Apr 8, 2025
Kind
B2
Abstract

In a method for interpreting output of a machine learning model, a processor receives a first interpretable rule set. A processor may also receive a second interpretable rule set generated from a dataset and model-predicted labels classifying the dataset. A processor may also generate a difference metric and mapping between the first interpretable rule set and the second interpretable rule set.

Claims (47)

1. A computer-implemented method for generating a difference metric for a machine learning model, comprising:

receiving, by one or more processors, a first interpretable rule set comprising first reasons for categorizing data;

generating model-predicted labels from a dataset using the machine learning model;

generating a second interpretable rule set comprising second reasons for categorizing data, by using an objective function to map literals between the dataset and the model-predicted labels while imposing grounding constraints that introduce a penalty term in the objective function to make the first interpretable rules set more likely to persist in the second interpretable rules set;

generating a difference metric and mapping between the first interpretable rule set and the second interpretable rule set; and

updating the machine learning model with additional data based on the difference metric being within a range.

2. The method of claim 1 , wherein the difference metric comprises a selection from a group consisting of a comparison view highlighting the differences between the first interpretable rule set and the second interpretable rule set, an edit distance between the first interpretable rule set and the second interpretable rule set, and a set of descriptions about the changes between the first interpretable rule set and the second interpretable rule set.

3. The method of claim 1 , wherein the first interpretable rule set comprises user-designated rules.

4. The method of claim 1 , wherein the first interpretable rule set comprises rules based on ground truth labels.

5. The method of claim 1 , wherein the first interpretable rule set comprises rules generated from the dataset and first model-predicted labels classifying the dataset.

6. The method of claim 1 , wherein the labels comprise binarized features, wherein the binarized features are binarized by a selection from a group consisting of: thresholding and one-hot encoding.

7. A computer program product comprising:

one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions comprising:

program instructions to receive a first interpretable rule set comprising first reasons for categorizing data;

program instructions to generate model-predicted labels from a dataset using the machine learning model;

program instructions to generate a second interpretable rule set comprising second reasons for categorizing data, by using an objective function to map literals between the dataset and the model-predicted labels while imposing grounding constraints that introduce a penalty term in the objective function to make the first interpretable rules set more likely to persist in the second interpretable rules set;

program instructions to generate a difference metric and mapping between the first interpretable rule set and the second interpretable rule set; and

program instructions to update the machine learning model with additional data based on the difference metric being within a range.

8. The computer program product of claim 7 , wherein the difference metric comprises a selection from a group consisting of a comparison view highlighting the differences between the first interpretable rule set and the second interpretable rule set, an edit distance between the first interpretable rule set and the second interpretable rule set, and a set of descriptions about the changes between the first interpretable rule set and the second interpretable rule set.

9. The computer program product of claim 7 , wherein the first interpretable rule set comprises user-designated rules.

10. The computer program product of claim 7 , wherein the first interpretable rule set comprises rules based on ground truth labels.

11. The computer program product of claim 7 , wherein the first interpretable rule set comprises rules generated from the dataset and first model-predicted labels classifying the dataset.

12. The computer program product of claim 7 , wherein the labels comprise binarized features, wherein the features are binarized by a selection from a group consisting of: thresholding and one-hot encoding.

13. A computer system comprising:

one or more computer processors, one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising:

program instructions to receive a first interpretable rule set comprising first reasons for categorizing data;

program instructions to generate model-predicted labels from a dataset using the machine learning model;

program instructions to generate a second interpretable rule set comprising second reasons for categorizing data, by using an objective function to map literals between the dataset and the model-predicted labels while imposing grounding constraints that introduce a penalty term in the objective function to make the first interpretable rules set more likely to persist in the second interpretable rules set;

program instructions to generate a difference metric and mapping between the first interpretable rule set and the second interpretable rule set; and

program instructions to update the machine learning model with additional data based on the difference metric being within a range.

14. The computer system of claim 13 , wherein the difference metric comprises a selection from a group consisting of a comparison view highlighting the differences between the first interpretable rule set and the second interpretable rule set, an edit distance between the first interpretable rule set and the second interpretable rule set, and a set of descriptions about the changes between the first interpretable rule set and the second interpretable rule set.

15. The computer system of claim 13 , wherein the first interpretable rule set comprises user-designated rules.

16. The computer system of claim 13 , wherein the first interpretable rule set comprises rules based on ground truth labels.

17. The computer system of claim 13 , wherein the first interpretable rule set comprises rules generated from the dataset and first model-predicted labels classifying the dataset.

18. The computer system of claim 13 , wherein the labels comprise binarized features, wherein the features are binarized by a selection from a group consisting of: thresholding and one-hot encoding.

19. A computer-implemented method for generating a difference metric between a first machine learning model and a second machine learning model, comprising:

receiving, by one or more processors, a first interpretable rule set comprising first reasons for categorizing data generated from a dataset and first model-predicted labels from the first machine learning model classifying the dataset;

receiving a second interpretable rule set comprising second reasons for categorizing data generated by an objective function operating on the dataset and second model-predicted labels from the second machine learning model classifying the dataset while imposing grounding constraints that introduce a penalty term in the objective function to make the first interpretable rules set more likely to persist in the second interpretable rules set;

generating a difference metric and mapping between the first interpretable rule set and the second interpretable rule set; and

updating the second machine learning model with additional data based on the difference metric being within a range.

20. A computer-implemented method for generating a difference metric between a first machine learning model and a second machine learning model, comprising:

generating first model-predicted labels from a dataset using the first machine learning model;

generating a first interpretable rule set comprising first reasons for categorizing data by comparing the dataset with the first model-predicted labels;

generating second model-predicted labels from the dataset using the second machine learning model;

generating a second interpretable rule set comprising second reasons for categorizing data using an objective function comparing the dataset with the second model-predicted labels while imposing grounding constraints that introduce a penalty term in the objective function to make the first interpretable rules set more likely to persist in the second interpretable rules set;

generating a difference metric and mapping between the first interpretable rule set and the second interpretable rule set; and

updating the second machine learning model with additional data based on the difference metric being within a range.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2021
From: DALY, ELIZABETH; NAIR, RAHUL; ALKAN, OZNUR; MATTETTI, MASSIMILIANO; WEI, DENNIS; ZHANG, YUNFENG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055561/0745 →
Continuity (1)
Related Publication 20220292391A1 · Sep 15, 2022
References Cited (33)
US 8533222B2 · Breckenridge et al. · 2013 [cited by applicant]
US 8887286B2 · Dupont et al. · 2014 [cited by applicant]
US 10324951B1 · Miller et al. · 2019 [cited by applicant]
US 10482183B1 · Vargas et al. · 2019 [cited by applicant]
US 10510022B1 · Tharrington, Jr. et al. · 2019 [cited by applicant]
US 10552002B1 · Maclean et al. · 2020 [cited by applicant]
US 10558554B2 · Bhandarkar et al. · 2020 [cited by applicant]
US 20090091443A1 · Chen et al. · 2009 [cited by applicant]
US 20160071027A1 · Brand et al. · 2016 [cited by applicant]
US 20160371601A1 · Grove et al. · 2016 [cited by applicant]
US 20190056918A1 · Langdon · 2019 [cited by applicant]
US 20190156216A1 · Gupta et al. · 2019 [cited by applicant]
US 20200090038A1 · Teredesai et al. · 2020 [cited by applicant]
US 20200097439A1 · Sinay et al. · 2020 [cited by applicant]
US 20200184494A1 · Joseph et al. · 2020 [cited by applicant]
US 20200250556A1 · Nourian et al. · 2020 [cited by applicant]
US 20200279140A1 · Pai et al. · 2020 [cited by applicant]
WO 2017201107A1 · 2017 [cited by applicant]
Wang et al., “Bayesian Rule Sets for Interpretable Classification,” 2016 IEEE 16th International Conference on Data Mining (ICDM), Barcelona, Spain, 2016, pp. 1269-1274, doi: 10.1109/ICDM.2016.0171. (Year: 2016). [cited by examiner]
Fletcher et al., Measuring the Similarity between Rule Lists, 2016, In Proc. of the 14th Australasian Data Mining Conference (AusDM) At: Canberra, Australia, Dec. 6-8, vol. 170 (Year: 2016). [cited by examiner]
Zhang et al., Diverse Rule Sets, 2020 In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD '20). Association for Computing Machinery, New York, Ny, USA, 1532-1541. htt… [cited by examiner]
Letham et al., “Interpretable classifiers using rules and Bayesian analysis: Building a better stroke prediction model”, 2019, Annals of Applied Statistics 2015, vol. 9, No. 3, 1350-1371, https://doi.org/10.48550/arXiv.… [cited by examiner]
Lakkaraju et al., Interpretable Decision Sets: A Joint Framework for Description and Prediction, 2016, In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '16). Ass… [cited by examiner]
Bu et al., “Model Change Detection with Application to Machine Learning,”, 2019, ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, pp. 5341-5346, doi: 10.1… [cited by examiner]
Arya et al., “One Explanation Does Not Fit All: A Toolkit And Taxonomy Of AI Explainability Techniques”, Cornell University, arXivLabs, Sep. 6, 2019, 18 pages, <https://arxiv.org/abs/1909.03012>. [cited by applicant]
Baena-Garcia et al., “Early Drift Detection Method”, Proceedings of the Fourth International Workshop on Knowledge Discovery From Data Streams, Sep. 18-22, 2006, Berlin, Germany, 10 pages, <http://www.machine-learning.e… [cited by applicant]
Bansal et al., “Updates in Human-AI Teams: Understanding and Addressing the Performance/Compatibility Tradeoff”, Proceedings of the Thirty-third AAAI Conference on Artificial Intelligence, Jan. 27-Feb. 1, 2019, Honolulu… [cited by applicant]
Dash et al., “Boolean Decision Rules via Column Generation”, Proceedings of the 32nd Conference on Neural Information Processing Systems, Dec. 2-8, 2018, Montréal, Canada, 11 pages. [cited by applicant]
Dries et al., “Adaptive Concept Drift Detection”, Statistical Analysis and Data Mining: The ASA Data Science Journal, vol. 2, Issue 5-6, Nov. 18, 2009, pp. 311-317. [cited by applicant]
Firat et al., “Constructing Classification Trees Using Column Generation”, Cornell University, arXivLabs, Jul. 11, 2019, 29 pages, <https://arxiv.org/abs/1810.06684v1>. [cited by applicant]
Margot, Vincent, “A Rigorous Method To Compare Interpretability Of Rule-Based Algorithms”, Cornell University, arXivLabs, Apr. 7, 2020, 9 pages, <https://arxiv.org/abs/2004.01570>. [cited by applicant]
Rajapaksha et al., “LoRMIKA: Local Rule-Based Model Interpretability With k-Optimal Associations”, Information Sciences, vol. 540, Nov. 2020, pp. 221-241. [cited by applicant]
Sutton et al., “Data Diff: Interpretability, Executable Summaries of Changes In Distributions For Data Wrangling”, Proceedings of the 24th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), Aug. 19-23, … [cited by applicant]