IP Library › Granted Patent US 12,566,983
Granted Patent B2
US 12,566,983 · App. 17/121,871 · Granted Mar 3, 2026

Machine learning classifiers prediction confidence and explanation

Inventors: Qingzi Liao (White Plains, NY); Yunfeng Zhang (Chappaqua, NY); Rachel Katherine Emma Bellamy (Bedford, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N5/045G06F18/2148G06F18/2415G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,983
App. No.
17/121,871
Granted
Mar 3, 2026
Kind
B2
Abstract

A method, a computer system, and a computer program product for generating explanations for different confidence levels of machine learning classifiers is provided. Embodiments of the present invention may include obtaining a dataset. Embodiments of the present invention may include training a first classifier using the dataset to generate probabilities. Embodiments of the present invention may include generating confidence scores using the first classifier. Embodiments of the present invention may include defining targeted confidence zones by transforming the generated probabilities into the confidence scores. Embodiments of the present invention may include training a second classifier to derive explanations. Embodiments of the present invention may include providing the explanations as an output.

Claims (42)

1 . A method comprising:

obtaining a structured dataset comprising numerical or categorical features;

training a first machine learning classifier using the structured dataset to generate prediction probabilities, each prediction probability indicating a likelihood of correctness for a corresponding data point;

transforming the prediction probabilities into confidence scores using the first classifier to generate a confidence level for each data point;

defining targeted confidence zones based on the confidence scores, wherein each confidence zone corresponds to a defined confidence range and is domain-specific or task-dependent;

training a second classifier on records labeled with the targeted confidence zones to derive zone-specific explanations of why the first classifier is confident or not confident for records assigned to each targeted confidence zone, wherein the second classifier comprises a rule-deduction model configured to generate symbolic rules based on features corresponding to the targeted confidence zones;

identifying explanation features for a current input based on the confidence zone assigned by the first classifier; and

outputting the explanations for the current input as symbolic rules that specify feature conditions associated with the assigned targeted confidence zone.

2 . The method of claim 1 , wherein the first machine learning classifier is trained to classify the dataset into one or more classes using the structured features of the dataset.

3 . The method of claim 1 , wherein the second machine learning classifier generates explanations of prediction confidence levels that indicate feature conditions under which the first classifier produces high-confidence or low-confidence predictions.

4 . The method of claim 1 , wherein the targeted confidence zones and the dataset are used to train the second machine learning classifier such that the second classifier learns symbolic rules specific to each confidence zone.

5 . The method of claim 1 , wherein the explanations are derived by partitioning the targeted confidence zones into low, medium and high confidence levels.

6 . The method of claim 1 , wherein the explanations are derived by using the targeted confidence zones and data to identify features.

7 . The method of claim 1 , wherein the output is provided in the form of symbolic rules learned for the targeted confidence zone.

8 . A computer-implemented method comprising:

obtaining an unstructured text dataset comprising textual features;

training a first machine learning classifier using the unstructured text dataset to generate prediction probabilities;

generating confidence scores based on the prediction probabilities using the first classifier;

segmenting the confidence scores into targeted confidence zones comprising at least a low-confidence zone and a high-confidence zone, wherein each zone is defined based on a task-specific threshold;

training a second classifier, different from the first classifier, using zone-labeled text data to derive zone-specific symbolic rules that explain why the first classifier is confident or not confident for text data points assigned to the targeted confidence zones, wherein the second classifier comprises a rule-based model configured to generate symbolic rules based on textual features corresponding to the targeted confidence zones;

identifying keywords associated with the targeted confidence zone for a current user-submitted input;

filtering the identified keywords based on semantic proximity to a previous user input; and

presenting the filtered zone-specific symbolic rules corresponding to the confidence zone assigned to the input.

9 . The computer-implemented method of claim 8 , wherein the first machine learning classifier is trained to classify the unstructured text dataset into one or more classes based on textual features.

10 . The computer-implemented method of claim 8 , wherein the second machine learning classifier generates explanations of prediction confidence levels that indicate feature conditions associated with the low-confidence zone or the high-confidence zone.

11 . The computer-implemented method of claim 8 , wherein the targeted confidence zones and the dataset are used to train the second machine learning classifier to learn symbolic rules specific to each confidence zone.

12 . The computer-implemented method of claim 8 , wherein the explanations are derived by partitioning the targeted confidence zones into low, medium and high confidence levels.

13 . The computer-implemented method of claim 8 , wherein the explanations are derived by using the targeted confidence zones and text data to identify features.

14 . The computer-implemented method of claim 8 , wherein the output is provided in a form of rules that state textual feature conditions for the assigned targeted confidence zone.

15 . A computer-implemented method comprising:

obtaining an image-based dataset comprising pixel values or image regions;

training a first machine learning classifier using the image-based dataset to generate prediction probabilities for image regions, each prediction probability indicating a likelihood of correctness for a corresponding image region;

generating confidence scores based on the prediction probabilities using the first classifier;

segmenting the confidence scores into targeted confidence zones comprising at least a low-confidence zone and a high-confidence zone, wherein each zone is defined based on a task-specific threshold;

training a second machine learning classifier, different from the first machine learning classifier, using zone-labeled image regions to derive zone-specific symbolic rules that explain why the first machine learning classifier is confident or not confident for data points assigned to the targeted confidence zones, wherein the second machine learning classifier comprises a rule-based model configured to generate symbolic rules based on pixel-level or region-level features corresponding to the targeted confidence zones;

identifying symbolic rules associated with the targeted confidence zone for a current input; and

presenting to the user the symbolic rules corresponding to the targeted confidence zone assigned to the image region.

16 . The computer-implemented method of claim 15 , wherein the first machine learning classifier is trained to classify image regions into one or more classes based on pixel-level or region-level features.

17 . The computer-implemented method of claim 15 , wherein the second machine learning classifier generates explanations of prediction confidence levels that indicate pixel-level feature conditions associated with the corresponding targeted confidence zone.

18 . The computer-implemented method of claim 15 , wherein the targeted confidence zones and the dataset are used to train the second machine learning classifier to learn symbolic rules specific to each confidence zone.

19 . The computer-implemented method of claim 15 , wherein the explanations are derived by partitioning the targeted confidence zones into low, medium and high confidence level.

20 . The computer-implemented method of claim 15 , wherein the explanations are derived by using the targeted confidence zones and image-based data to identify pixel-level or region-level features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2020
From: LIAO, QINGZI; ZHANG, YUNFENG; BELLAMY, RACHEL KATHERINE EMMA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054646/0556 →
Continuity (1)
Related Publication 20220188674A1 · Jun 16, 2022
References Cited (41)
US 8676731B1 · Sathyanarayana · 2014 [cited by applicant]
US 8868472B1 · Lin · 2014 [cited by examiner]
US RE45770E · Flinn · 2015 [cited by applicant]
US 9600779B2 · Hoover · 2017 [cited by examiner]
US 9679261B1 · Hoover · 2017 [cited by examiner]
US 9691027B1 · Sawant · 2017 [cited by applicant]
US 10824959B1 · Chatterjee · 2020 [cited by examiner]
US 11694093B2 · Perez · 2023 [cited by examiner]
US 20140229164A1 · Martens · 2014 [cited by examiner]
US 20160155069A1 · Hoover · 2016 [cited by examiner]
US 20160259842A1 · Knight · 2016 [cited by examiner]
US 20180285771A1 · Lee · 2018 [cited by examiner]
US 20190130226A1 · Guo · 2019 [cited by examiner]
US 20190197431A1 · Gopalakrishnan · 2019 [cited by examiner]
US 20190354805A1 · Hind · 2019 [cited by examiner]
US 20200020000A1 · Guy · 2020 [cited by examiner]
US 20200111572A1 · Bergen · 2020 [cited by examiner]
US 20210089957A1 · Ermans · 2021 [cited by examiner]
US 20230316204A1 · Liu · 2023 [cited by examiner]
WO WO2020185101A1 · 2020 [cited by examiner]
WO WO2020185101A9 · 2021 [cited by examiner]
Shah et al. “Evaluating Explanations of Convolutional Neural Network Image Classifications” (Year: 2020). [cited by examiner]
Soko el at. “LIMEtree Consistent and Faithful Multi-class Explanations” (Year: 2020). [cited by examiner]
Sarwar, B., Karypis, G., Konstan, J., & Riedl, J. (2001). Item-based Collaborative Filtering Recommendation Algorithms. Proceedings of ACM World Wide Web Conference, 1. (Year: 2001). [cited by examiner]
Kolter, J., & Maloof, M. (2004). Learning to detect malicious executables in the wild. In Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 470-478). Association fo… [cited by examiner]
Terrance DeVries, & Graham W. Taylor. (2018). Learning Confidence for Out-of-Distribution Detection in Neural Networks. (Year: 2018). [cited by examiner]
W. Zhang, H. Wang, K. Ren and J. Song, “Chinese sentence based lexical similarity measure for artificial intelligence chatbot,” 2016 8th International Conference on Electronics, Computers and Artificial Intelligence (EC… [cited by examiner]
Mehmet Yigit Yildirim, Mert Ozer, & Hasan Davulcu. (2019). Leveraging Uncertainty in Deep Learning for Selective Classification. (Year: 2019). [cited by examiner]
Kirill Bykov, Marina M.-C. Höhne, Klaus-Robert Müller, Shinichi Nakajima, & Marius Kloft. (2020). How Much Can I Trust You?—Quantifying Uncertainties in Explaining Neural Networks. (Year: 2020). [cited by examiner]
Augasta et al., “Reverse Engineering the Neural Networks for Rule Extraction in Classification Problems”, Article, Neural Processing Letters, Apr. 2012, https://doi.org/10.1007/s11063-011-9207-8, 4 pages. [cited by applicant]
Carvalho et al., “Machine Learning Interpretability: A Survey on Methods and Metrics”, Electronics 2019, 8, 832, www.mdpi.com/journal/electronics, pp. 1-34. [cited by applicant]
Cohen, “Fast Effective Rule Induction”, ScienceDirect, Machine Learning Proceedings 1995, Proceedings of the Twelfth International Conference on Machine Learning, Tahoe City, California, Jul. 9-12, 1995, 2 pages. [cited by applicant]
Deutch et al., “Explaining White-box Classifications to Data Scientists (technical report)”, 2018, pp. 1-28. [cited by applicant]
Jiang et al., “To Trust or Not to Trust a Classifier”, 32nd Conference on Neural Information Processing Systems (NIPS 2018), Montreal, Canada, 25 pages. [cited by applicant]
Kailkhura et al., “Reliable and explainable machine-learning methods for accelerated material discovery”, npj Computational Materials (2019) 5:108 ; https://doi.org/10.1038/s41524-019-0248-2, pp. 1-9. [cited by applicant]
Petkovic et al., Improving the explainability of Random Forest classifier-user centered approachAuthor Manuscript, , HHS Public Access, Pac Symp Biocomput. Author manuscript; available in PMC Jan. 1, 2018, pp. 1-16. [cited by applicant]
Phillips et al., “Interpretable Active Learning”, Journal of Machine Learning Research 81:1-13, 2018, Conference on Fairness, Accountability, and Transparency, 13 pages. [cited by applicant]
Ribeiro et al., “Anchors: High-Precision Model-Agnostic Explanations”, Copyright 2018, Association for the Advancement of Artificial Intelligence, 9 pages. [cited by applicant]
Ribeiro et al., “Why Should I Trust You?” Explaining the Predictions of Any Classifier, KDD 2016, San Francisco, CA, USA, 10 pages. [cited by applicant]
Teso, “Why Should I Trust Interactive Learners?” Explaining Interactive Queries of Classifiers to Users, May 22, 2018, DeepAI, 28 pages. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, Recommendations of the National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]