IP Library › Granted Patent US 11,742,081
Granted Patent B2
US 11,742,081 · App. 16/863,452 · Granted Aug 29, 2023

Data model processing in machine learning employing feature selection using sub-population analysis

Inventors: Uri Kartoun (Cambridge, MA); Kristen Severson (Somerville, MA); Kenney Ng (Arlington, MA); Paul D. Myers (Bloomfield Hills, MI); Wangzhi Dai (Cambridge, MA); Collin M. Stultz (Newton, MA)
Assignees: International Business Machines Corporation; Massachusetts Institute of Technology
G16H40/67G06F17/18G06F18/23G16H50/50G16H50/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,742,081
App. No.
16/863,452
Granted
Aug 29, 2023
Kind
B2
Abstract

A computer system selects features of a dataset for predictive modeling. A first set of features that are relevant to outcome are selected from a dataset comprising a plurality of cases and controls. A subset of cases and controls having similar values for the first set of features is identified. The subset is analyzed to select a set of additional features relevant to outcome. A first and second predictive model are evaluated to determine that the second predictive model more accurately predicts outcome, wherein the first predictive model is based on the first set of features and the second predictive model is based on the first set of features and the additional features. The second predictive model is utilized to predict outcomes. Embodiments of the present invention further include a method and program product for selecting features of a dataset for predictive modeling in substantially the same manner described above.

Claims (65)

1. A computer-implemented method in a data processing system comprising at least one processor and at least one memory, the at least one memory comprising instructions executed by the at least one processor to cause the at least one processor to select features of a dataset for predictive modeling, the computer-implemented method comprising:

selecting, from a dataset comprising a plurality of cases and controls, a first set of features that are relevant to outcome;

identifying a subset of cases and controls having similar values for the first set of features;

analyzing the subset of cases and controls to select a second set of additional features that are relevant to outcome;

evaluating performance of a first predictive model against a second predictive model to determine that the second predictive model more accurately predicts outcome, wherein the first predictive model is based on the first set of features and the second predictive model is based on the first set of features and the second set of additional features; and

utilizing the second predictive model to predict outcomes.

2. The computer-implemented method of claim 1 , wherein the subset of cases and controls is identified by:

identifying a plurality of case-control matchings by applying propensity score matching with different combinations of caliper values and case-control ratio values to the dataset to match cases with controls;

comparing, for each case-control matching, values of the cases to values of the controls to determine a similarity value, wherein the compared values comprise values for the first set of features; and

selecting a case-control matching based on a ranking of similarity values of the plurality of case-control matchings.

3. The computer-implemented method of claim 2 , wherein the similarity value for each case-control matching is determined by:

dividing cases and controls of the case-control matching into a training set and a testing set;

training a predictive model using the training set;

testing the predictive model using the testing set to identify true positives and false positives; and

calculating the similarity value by determining an area under a receiver operating characteristic curve, wherein the receiver operating characteristic curve is based on the identified true positives and false positives.

4. The computer-implemented method of claim 1 , wherein evaluating performance of the first predictive model against the second predictive model comprises:

testing each of the first predictive model and the second predictive model using a same model testing set to identify true positives and false positives; and

comparing a first area under a first receiver operating curve to a second area under a second receiver operating curve, wherein the first receiver operating curve is based on identified true and false positives of the first predictive model and wherein the second receiver operating curve is based on identified true and false positives of the second predictive model.

5. The computer-implemented method of claim 1 , wherein the second set of additional features are identified based on a statistical significance of each feature satisfying a threshold significance value to predict the outcome.

6. The computer-implemented method of claim 5 , wherein the statistical significance is determined using one or more of: a chi square test, a t-test, and a non-parametric test.

7. The computer-implemented method of claim 1 , wherein the dataset comprises clinical data and wherein the outcome comprises a medical outcome.

8. A computer system for selecting features of a dataset for predictive modeling, the computer system comprising:

one or more computer processors;

one or more computer readable storage media;

program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising instructions to:

select, from a dataset comprising a plurality of cases and controls, a first set of features that are relevant to outcome;

identify a subset of cases and controls having similar values for the first set of features;

analyze the subset of cases and controls to select a second set of additional features that are relevant to outcome;

evaluate performance of a first predictive model against a second predictive model to determine that the second predictive model more accurately predicts outcome, wherein the first predictive model is based on the first set of features and the second predictive model is based on the first set of features and the second set of additional features; and

utilize the second predictive model to predict outcomes.

9. The computer system of claim 8 , wherein the program instructions to identify the subset of cases and controls comprise instructions to:

identify a plurality of case-control matchings by applying propensity score matching with different combinations of caliper values and case-control ratio values to the dataset to match cases with controls;

compare, for each case-control matching, values of the cases to values of the controls to determine a similarity value, wherein the compared values comprise values for the first set of features; and

select a case-control matching based on a ranking of similarity values of the plurality of case-control matchings.

10. The computer system of claim 9 , wherein the similarity value for each case-control matching is determined by:

dividing cases and controls of the case-control matching into a training set and a testing set;

training a predictive model using the training set;

testing the predictive model using the testing set to identify true positives and false positives; and

calculating the similarity value by determining an area under a receiver operating characteristic curve, wherein the receiver operating characteristic curve is based on the identified true positives and false positives.

11. The computer system of claim 8 , wherein the program instructions to evaluate performance of the first predictive model against the second predictive model comprise instructions to:

test each of the first predictive model and the second predictive model using a same model testing set to identify true positives and false positives; and

compare a first area under a first receiver operating curve to a second area under a second receiver operating curve, wherein the first receiver operating curve is based on identified true and false positives of the first predictive model and wherein the second receiver operating curve is based on identified true and false positives of the second predictive model.

12. The computer system of claim 8 , wherein the second set of additional features are identified based on a statistical significance of each feature satisfying a threshold significance value to predict the outcome.

13. The computer system of claim 12 , wherein the statistical significance is determined using one or more of: a chi square test, a t-test, and a non-parametric test.

14. The computer system of claim 8 , wherein the dataset comprises clinical data and wherein the outcome comprises a medical outcome.

15. A computer program product for selecting features of a dataset for predictive modeling, the computer program product comprising one or more computer readable storage media collectively having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to:

select, from a dataset comprising a plurality of cases and controls, a first set of features that are relevant to outcome;

identify a subset of cases and controls having similar values for the first set of features;

analyze the subset of cases and controls to select a second set of additional features that are relevant to outcome;

evaluate performance of a first predictive model against a second predictive model to determine that the second predictive model more accurately predicts outcome, wherein the first predictive model is based on the first set of features and the second predictive model is based on the first set of features and the second set of additional features; and

utilize the second predictive model to predict outcomes.

16. The computer program product of claim 15 , wherein the program instructions to identify the subset of cases and controls cause the computer to:

identify a plurality of case-control matchings by applying propensity score matching with different combinations of caliper values and case-control ratio values to the dataset to match cases with controls;

compare, for each case-control matching, values of the cases to values of the controls to determine a similarity value, wherein the compared values comprise values for the first set of features; and

selecting a case-control matching based on a ranking of similarity values of the plurality of case-control matchings.

17. The computer program product of claim 16 , wherein the similarity value for each case-control matching is determined by:

dividing cases and controls of the case-control matching into a training set and a testing set;

training a predictive model using the training set;

testing the predictive model using the testing set to identify true positives and false positives; and

calculating the similarity value by determining an area under a receiver operating characteristic curve, wherein the receiver operating characteristic curve is based on the identified true positives and false positives.

18. The computer program product of claim 15 , wherein the program instructions to evaluate performance of the first predictive model against the second predictive model cause the computer to:

test each of the first predictive model and the second predictive model using a same model testing set to identify true positives and false positives; and

compare a first area under a first receiver operating curve to a second area under a second receiver operating curve, wherein the first receiver operating curve is based on identified true and false positives of the first predictive model and wherein the second receiver operating curve is based on identified true and false positives of the second predictive model.

19. The computer program product of claim 15 , wherein the second set of additional features are identified based on a statistical significance of each feature satisfying a threshold significance value to predict the outcome.

20. The computer program product of claim 19 , wherein the statistical significance is determined using one or more of: a chi square test, a t-test, and a non-parametric test.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2020
From: MYERS, PAUL D.; DAI, WANGZHI; STULTZ, COLLIN M.
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 052546/0572 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2020
From: KARTOUN, URI; SEVERSON, KRISTEN; NG, KENNEY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052546/0614 →
Continuity (1)
Related Publication 20210343421A1 · Nov 4, 2021
Cited By (1)
US 12,730,945