IP Library › Granted Patent US 11,068,743
Granted Patent B2
US 11,068,743 · App. 15/844,833 · Granted Jul 20, 2021

Feature selection impact analysis for statistical models

Inventors: Cagri Ozcaglar (Sunnyvale, CA); Vijay K. Dialani (Fremont, CA); Sara S. Gerrard (Stanford, CA); Sahin C. Geyik (Redwood City, CA); Anish R. Nair (Fremont, CA)
Assignee: Microsoft Technology Licensing, LLC
G06K9/6231G06K9/6212G06K9/6219G06K9/6262
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,068,743
App. No.
15/844,833
Granted
Jul 20, 2021
Kind
B2
Abstract

The disclosed embodiments provide a system for processing data. During operation, the system obtains a set of feature additions and an evaluation metric for assessing the performance of a statistical model. Next, the system automatically builds treatment versions of the statistical model using a set of baseline features for the statistical model and feature combinations generated using the feature additions. The system then uses a hypothesis test and a fixed set of feature values to compare a baseline value of the evaluation metric for a baseline version of the statistical model that is built using the set of baseline features with additional values of the evaluation metric for the treatment versions. Finally, the system outputs a result of the hypothesis test for use in assessing an impact of the feature combinations on a performance of the statistical model.

Claims (64)

1. A method, comprising:

obtaining a set of feature additions and an evaluation metric for assessing a performance of a statistical model;

automatically building, by one or more computer systems, treatment versions of the statistical model using a set of baseline features for the statistical model and feature combinations generated using the set of feature additions;

using a hypothesis test and a fixed set of feature values to compare, by the one or more computer systems, a baseline value of the evaluation metric for a baseline version of the statistical model that is built using the set of baseline features with additional values of the evaluation metric for the treatment versions; and

outputting a result of the hypothesis test for use in assessing an impact of the feature combinations on a performance of the statistical model.

2. The method of claim 1 , further comprising:

when the result comprises a treatment version of the statistical model with an improvement in performance over the baseline version, automatically adding to the set of baseline features a feature combination used to build the treatment version.

3. The method of claim 1 , wherein obtaining the set of feature additions for the statistical model comprises:

obtaining a selection of one or more feature additions from a user.

4. The method of claim 1 , wherein obtaining the set of feature additions for the statistical model comprises:

using a feature selection method to generate the set of feature additions.

5. The method of claim 4 , wherein the feature selection method comprises at least one of:

a feature-label correlation;

a feature-feature correlation; and

a variance.

6. The method of claim 1 , wherein automatically building the treatment versions of the statistical model comprises:

building the treatment versions of the statistical model from all possible combinations generated from the set of feature additions.

7. The method of claim 1 , wherein automatically building the treatment versions of the statistical model comprises:

using a fixed set of training data to build the treatment versions and the baseline version of the statistical model.

8. The method of claim 1 , wherein using the hypothesis test and the fixed set of feature values to compare the baseline value of the evaluation metric with the additional values of the evaluation metric comprises:

using the baseline version and the fixed set of feature values for the baseline features to generate the baseline value of the evaluation metric;

using the treatment versions and the fixed set of feature values for the baseline features and the feature combinations to generate the additional values of the evaluation metric;

using the hypothesis test to compare the additional values with the baseline value; and

determining a statistical significance associated with differences between the additional values and the baseline value.

9. The method of claim 8 , wherein the result of the hypothesis test comprises a treatment version of the statistical model with a statistically significant improvement in performance over the baseline version.

10. The method of claim 1 , further comprising:

obtaining a model type of the statistical model with the set of feature additions and the evaluation metric.

11. The method of claim 1 , wherein obtaining the set of feature additions for the statistical model comprises:

obtaining the set of feature additions as new features added to a feature repository.

12. The method of claim 1 , wherein the hypothesis test comprises a paired t-test.

13. A system, comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the system to:

obtain a set of feature additions and an evaluation metric for assessing a performance of a statistical model;

automatically build treatment versions of the statistical model using a set of baseline features for the statistical model and feature combinations generated using the set of feature additions;

use a hypothesis test and a fixed set of feature values to compare a baseline value of the evaluation metric for a baseline version of the statistical model that is built using the set of baseline features with additional values of the evaluation metric for the treatment versions; and

output a result of the hypothesis test for use in assessing an impact of the feature combinations on a performance of the statistical model.

14. The system of claim 13 , wherein the memory further stores instructions that, when executed by the one or more processors, cause the system to:

automatically add to the set of baseline features a feature combination used to build the treatment version when the result comprises a treatment version of the statistical model with an improvement in performance over the baseline version.

15. The system of claim 13 , wherein obtaining the set of feature additions for the statistical model comprises at least one of:

obtaining a selection of one or more feature additions from a user;

using a feature selection method to generate the set of feature additions; and

obtaining the set of feature additions as new features added to a feature repository.

16. The system of claim 15 , wherein the feature selection method comprises at least one of:

a feature-label correlation;

a feature-feature correlation; and

a variance.

17. The system of claim 13 , wherein using the hypothesis test and the fixed set of feature values to compare the baseline value of the evaluation metric with the additional values of the evaluation metric comprises:

using the baseline version and the fixed set of feature values for the baseline features to generate the baseline value of the evaluation metric;

using the treatment versions and the fixed set of feature values for the baseline features and the feature combinations to generate the additional values of the evaluation metric;

using the hypothesis test to compare the additional values with the baseline value; and

determining a statistical significance associated with differences between the additional values and the baseline value.

18. The system of claim 13 , wherein the memory further stores instructions that, when executed by the one or more processors, cause the system to:

obtain the set of feature additions as new features added to a feature repository.

19. A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method, the method comprising:

obtaining a set of feature additions and an evaluation metric for assessing a performance of a statistical model;

automatically building treatment versions of the statistical model using a set of baseline features for the statistical model and feature combinations generated using the set of feature additions;

using a hypothesis test and a fixed set of feature values to compare a baseline value of the evaluation metric for a baseline version of the statistical model that is built using the set of baseline features with additional values of the evaluation metric for the treatment versions; and

outputting a result of the hypothesis test for use in assessing an impact of the feature combinations on a performance of the statistical model.

20. The non-transitory computer-readable storage medium of claim 19 , wherein using the hypothesis test and the fixed set of feature values to compare the baseline value of the evaluation metric with the additional values of the evaluation metric comprises:

using the baseline version and the fixed set of feature values for the baseline features to generate the baseline value of the evaluation metric;

using the treatment versions and the fixed set of feature values for the baseline features and the feature combinations to generate the additional values of the evaluation metric;

using the hypothesis test to compare the additional values with the baseline value; and

determining a statistical significance associated with differences between the additional values and the baseline value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2018
From: OZCAGLAR, CAGRI; DIALANI, VIJAY K.; GERRARD, SARA S.; GEYIK, SAHIN C.; NAIR, ANISH R.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 044584/0806 →
Continuity (1)
Related Publication 20190188531A1 · Jun 20, 2019
Cited By (1)
US 12,217,145