IP Library › Granted Patent US 12,242,367
Granted Patent B2
US 12,242,367 · App. 17/744,690 · Granted Mar 4, 2025

Feature importance based model optimization

Inventors: Jing Xu (Xi'an, CN); Xue Ying Zhang (Xi'an, CN); Si Er Han (Xi'an, CN); Jing James Xu (Xi'an, CN); Xiao Ming Ma (Xi'an, CN); Jun Wang (Xi'an, CN); Wen Pei Yu (Xi'an, CN)
Assignee: International Business Machines Corporation
G06F11/3447G06F18/23
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,367
App. No.
17/744,690
Granted
Mar 4, 2025
Kind
B2
Abstract

Disclosed are a computer-implemented method, a system and a computer program product for model exploration. Model feature importance of each model of a plurality of models can be obtained, the plurality of models can be grouped into a plurality of model clusters based on the model feature importance of each model, and the model feature importance can be presented by box-plot or confidence interval.

Claims (57)

1. A computer-implemented method for

exploring predictive models, the method comprising:

computing feature importance for each of a plurality of features for each of a plurality of predictive models;

clustering predictive models among the plurality of predictive models to form a plurality of model clusters based on a similarity of feature importance of the computed feature importance for each of the plurality of features for each of the plurality of predictive models;

deriving a cluster feature importance for each model cluster of the plurality of model clusters;

computing a feature importance for a data case using a preliminary model;

selecting a model cluster for the data case from the plurality of model clusters based on similarity between the feature importance for the data case and the cluster feature importance of each model cluster of the plurality of model clusters; and

selecting a predictive model within the selected model cluster to score the data case based on at least one criteria thereby optimizing an evaluation of the plurality of predictive models among a massive set of predictive models.

2. The computer-implemented method of claim 1 ,

wherein the preliminary model is selected based on at least one of a plurality of model attributes from the plurality of predictive models to be explored.

3. The computer-implemented method of claim 2 , wherein the plurality of model attributes comprise model performance, model size, model accuracy, and model complexity.

4. The computer-implemented method of claim 1 ,

further comprising:

performing a model interpretation algorithm on the preliminary model to deduce importance values for features of the data case to compute the feature importance for the data case using the preliminary model.

5. The computer-implemented method of claim 1 ,

wherein the feature importance is computed for each of the plurality of predictive models using sensitivity analysis.

6. The computer-implemented method of claim 1 ,

wherein the computed feature importance for each of the plurality of features for each of the plurality of predictive models is displayed via a box-plot or a confidence interval to show a distribution of feature importance among the plurality of predictive models.

7. The computer-implemented method of claim 1 ,

wherein the at least one criteria comprises one or more of the following in the group consisting of a model type, a model size, a model accuracy, and a model performance.

8. The computer-implemented method of claim 1 , further comprising:

clustering predictive models among the plurality of predictive models to form the plurality of model clusters using a k-means based algorithm.

9. The computer-implemented method of claim 1 ,

wherein the similarity between the feature importance for the data case and the cluster feature importance of each model cluster of the plurality of model clusters is computed via a measured distance between vectors.

10. A computer program product for exploring predictive models, the computer program product comprising one or more computer readable storage mediums having program code embodied therewith, the program code comprising programming instructions for:

computing feature importance for each of a plurality of features for each of a plurality of predictive models;

clustering predictive models among the plurality of predictive models to form a plurality of model clusters based on a similarity of feature importance of the computed feature importance for each of the plurality of features for each of the plurality of predictive models;

deriving a cluster feature importance for each model cluster of the plurality of model clusters;

computing a feature importance for a data case using a preliminary model;

selecting a model cluster for the data case from the plurality of model clusters based on similarity between the feature importance for the data case and the cluster feature importance of each model cluster of the plurality of model clusters; and

selecting a predictive model within the selected model cluster to score the data case based on at least one criteria thereby optimizing an evaluation of the plurality of predictive models among a massive set of predictive models.

11. The computer program product of claim 10 , wherein the preliminary model is selected based on at least one of a plurality of model attributes from the plurality of predictive models to be explored.

12. The computer program product of claim 11 , wherein the plurality of model attributes comprise model performance, model size, model accuracy, and model complexity.

13. The computer program product of claim 10 , wherein the program code further comprises the programming instructions for:

performing a model interpretation algorithm on the preliminary model to deduce importance values for features of the data case to compute the feature importance for the data case using the preliminary model.

14. The computer program product of claim 10 , wherein the feature importance is computed for each of the plurality of predictive models using sensitivity analysis.

15. The computer program product of claim 10 , wherein the computed feature importance for each of the plurality of features for each of the plurality of predictive models is displayed via a box-plot or a confidence interval to show a distribution of feature importance among the plurality of predictive models.

16. The computer program product of claim 10 , wherein the at least one criteria comprises one or more of the following in the group consisting of a model type, a model size, a model accuracy, and a model performance.

17. The computer program product of claim 10 , wherein the program code further comprises the programming instructions for:

clustering predictive models among the plurality of predictive models to form the plurality of model clusters using a k-means based algorithm.

18. The computer program product of claim 10 , wherein the similarity between the feature importance for the data case and the cluster feature importance of each model cluster of the plurality of model clusters is computed via a measured distance between vectors.

19. A system, comprising:

a memory for storing a computer program for exploring predictive models; and

a processor connected to the memory, wherein the processor is configured to execute program instructions of the computer program comprising:

computing feature importance for each of a plurality of features for each of a plurality of predictive models;

clustering predictive models among the plurality of predictive models to form a plurality of model clusters based on a similarity of feature importance of the computed feature importance for each of the plurality of features for each of the plurality of predictive models;

deriving a cluster feature importance for each model cluster of the plurality of model clusters;

computing a feature importance for a data case using a preliminary model;

selecting a model cluster for the data case from the plurality of model clusters based on similarity between the feature importance for the data case and the cluster feature importance of each model cluster of the plurality of model clusters; and

selecting a predictive model within the selected model cluster to score the data case based on at least one criteria thereby optimizing an evaluation of the plurality of predictive models among a massive set of predictive models.

20. The system of claim 19 , wherein the preliminary model is selected based on at least one of a plurality of model attributes from the plurality of predictive models to be explored.

21. The system of claim 20 , wherein the plurality of model attributes comprise model performance, model size, model accuracy, and model complexity.

22. The system of claim 19 , wherein the program instructions of the computer program further comprise:

performing a model interpretation algorithm on the preliminary model to deduce importance values for features of the data case to compute the feature importance for the data case using the preliminary model.

23. The system of claim 19 , wherein the feature importance is computed for each of the plurality of predictive models using sensitivity analysis.

24. The system of claim 19 , wherein the computed feature importance for each of the plurality of features for each of the plurality of predictive models is displayed via a box-plot or a confidence interval to show a distribution of feature importance among the plurality of predictive models.

25. The system of claim 19 , wherein the at least one criteria comprises one or more of the following in the group consisting of a model type, a model size, a model accuracy, and a model performance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2022
From: XU, JING; ZHANG, XUE YING; HAN, SI ER; XU, JING JAMES; MA, XIAO MING; WANG, JUN; YU, WEN PEI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059914/0149 →
Continuity (1)
Related Publication 20230367689A1 · Nov 16, 2023
References Cited (18)
US 20140136165A1 · Sarma · 2014 [cited by examiner]
US 20150286704A1 · Shyr · 2015 [cited by examiner]
CN 108268460A · 2018 [cited by applicant]
CN 106897199B · 2020 [cited by applicant]
CN 111931009A · 2020 [cited by applicant]
IN 201821013348A · 2019 [cited by applicant]
WO 2019193570A1 · 2019 [cited by applicant]
WO 2020008392A2 · 2020 [cited by applicant]
Fang et al., “Better Model Selection with a new Definition of Feature Importance”, arXiv: 2009.07708v1 [stat.ML] Sep. 16, 2020, 8 pps <https://arxiv.org/pdf/2009.07708.pdf> (hereinafter Fang). (Year: 2009). [cited by examiner]
Interpret model predictions using Permutation Feature Importance, Oct. 21, 2021, 7 pps., Microsoft Docs, https://docs.microsoft.com/en-us/machine-learning/how-to-guides/explain-machine learning-model-permutation-feature… [cited by examiner]
By Shah, “Using Feature Importance Rank Ensembling (Fire) for Advanced Feature Selection”, DataRobot, 10pps., 2021 DataRobot, Inc., https://www.datarobot.com/blog/using-feature-importance-rank-ensembling-fire-for-advanc… [cited by examiner]
“Automated Model Nuggets”, SPSS Modeler 18.2.1, 2 pps., IBM documentation, © Copyright IBM Corporation 1994, 2019, <https://www.ibm.com/docs/en/spss-modeler/18.2.1?topic=nodes-automated-model-nuggets >. [cited by applicant]
“Interpret model predictions using Permutation Feature Importance”, Oct. 21, 2021, 7 pps., Microsoft Docs, <https://docs.microsoft.com/en-us/dotnet/machine-learning/how-to-guides/explain-machine-learning-model-permutati… [cited by applicant]
Chahal et al., “Prowl: Towards Predicting the Runtime of Batch Workloads”, WOSP-C Workshop, CPE'18 Companion, Apr. 9-13, 2018, Berlin, Germany, 2 pps. [cited by applicant]
Fang et al., “Better Model Selection with a new Definition of Feature Importance”, arXiv:2009.07708v1 [stat.ML] Sep. 16, 2020, 8 pps., <https://arxiv.org/pdf/2009.07708.pdf>. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, U.S. Department of Commerce, NIST Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Shah, “Using Feature Importance Rank Ensembling (Fire) for Advanced Feature Selection”, DataRobot, 10 pps., © 2021 DataRobot, Inc., <https://www.datarobot.com/blog/using-feature-importance-rank-ensembling-fire-for-advan… [cited by applicant]
U.S. Appl. No. 17/443,831, filed Jul. 28, 2021. [cited by applicant]