IP Library › Granted Patent US 11,494,699
Granted Patent B2
US 11,494,699 · App. 16/868,145 · Granted Nov 8, 2022

Automatic machine learning feature backward stripping

Inventor: Jacques Doan Huu (Montigny le Bretonneux, FR)
Assignee: SAP SE
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,494,699
App. No.
16/868,145
Granted
Nov 8, 2022
Kind
B2
Abstract

Features are used to train one or more ML models in a modelling layer. In a feature selection layer, each generated ML model is analyzed to determine, for each input feature, a degree of importance of the feature on the results generated by the ML model. Features with low importance are identified and the information is propagated backward to the data source and feature engineering layers. In response, the data source and feature engineering layers refrain from gathering or generating the unimportant features. Based on a confidence measure of the determination that each feature is important or unimportant, a number of periods between reevaluation of the feature importance is determined. After the number of periods has elapsed, a removed feature is restored to the pipeline.

Claims (73)

1. A method comprising:

generating, by one or more processors, a first set of machine learning (ML) models each using a first set of features;

determining, for each feature of the first set of features and each ML model of the first set of ML models, an importance measure based on a relation between the feature and output of the ML model;

based on the importance measures, defining a second set of features comprising a strict subset of the first set of features; and

generating, by the one or more processors, a second set of ML models using the second set of features.

2. The method of claim 1 , wherein the defining of the second set of features comprises:

determining, for each feature of the first set of features, based on the importance measures for the feature and a first predetermined threshold, a percentage of the first set of ML models for which the feature is not important;

determining, for the one or more removed features, based on a percentage of ML models for which the feature is not important and a second predetermined threshold, that the feature is not likely to be important to future ML models; and

based on the determination that the one or more removed features are not likely to be important to future ML models, removing the one or more removed features from the first set of features to define the second set of features.

3. The method of claim 2 , wherein the second predetermined threshold is 50%.

4. The method of claim 1 , further comprising:

determining, for each feature of the second set of features and each ML model of the second set of ML models, an importance measure based on a relation between the feature and output of the ML model;

based on the importance measures, defining a third set of features comprising the second set of features with one or more features removed; and

generating, by the one or more processors, a third set of ML models using the third set of features.

5. The method of claim 1 , further comprising:

determining, based on a percentage of ML models for which a first removed feature of the one or more removed features is important, a confidence measure for the determination that the first removed feature is not likely to be important to future ML models;

defining a third set of features comprising the second set of features and a first removed feature of the one or more removed features, based on the confidence measure, a third predetermined threshold, and a number of retrains elapsed since the removal of the first removed feature from the second set of features; and

generating, by the one or more processors, a third set of ML models using the third set of features.

6. The method of claim 1 , wherein the first set of ML models are generated at a predetermined rate.

7. The method of claim 1 , further comprising:

prior to generating the first set of ML models, accessing a first plurality of features over a network;

based on the first plurality of features, generating a second feature of the first set of features;

creating a feature dependency graph that relates the second feature to the first plurality of features; and

based on the feature dependency graph and the second feature being one of the one or more removed features, refraining from accessing the first plurality of features for the generation of the second set of ML models.

8. A system comprising:

a memory that stores instructions; and

one or more processors configured by the instructions to perform operations comprising:

generating a first set of machine learning (ML) models each using a first set of features;

determining, for each feature of the first set of features and each ML model of the first set of ML models, an importance measure based on a relation between the feature and output of the ML model;

based on the importance measures, defining a second set of features comprising a strict subset of the first set of features; and

generating a second set of ML models using the second set of features.

9. The system of claim 8 , wherein the defining of the second set of features comprises:

determining, for each feature of the first set of features, based on the performance measures for the feature and a first predetermined threshold, a percentage of the first set of ML models for which the feature is not important;

determining, for the one or more removed features, based on the percentage of ML models for which the feature is not important and a second predetermined threshold, that the feature is not likely to be important to future ML models; and

based on the determination that the one or more removed features are not likely to be important to future ML models, removing the one or more removed features from the first set of features to define the second set of features.

10. The system of claim 9 , wherein the second predetermined threshold is 50%.

11. The system of claim 8 , wherein the operations further comprise:

determining, for each feature of the second set of features and each ML model of the second set of ML models, an importance measure based on a relation between the feature and output of the ML model;

based on the importance measures, defining a third set of features comprising the second set of features with one or more features removed; and

generating, by the one or more processors, a third set of ML models using the third set of features.

12. The system of claim 8 , wherein the operations further comprise:

determining, based on a percentage of ML models for which a first removed feature of the one or more removed features is important, a confidence measure for the determination that the first removed feature is not likely to be important to future ML models;

defining a third set of features comprising the second set of features and the first removed feature, based on the confidence measure, a third predetermined threshold, and a number of retrains elapsed since the removal of the first removed feature from the second set of features; and

generating, by the one or more processors, a third ML model using the third set of features.

13. The system of claim 8 , wherein the first set of ML models are generated at a predetermined rate.

14. The system of claim 8 , wherein the operations further comprise:

prior to generating the first set of ML models, accessing a first plurality of features over a network;

based on the first plurality of features, generating a second feature of the first set of features;

creating a feature dependency graph that relates the second feature to the first plurality of features; and

based on the feature dependency graph and the second feature being one of the one or more removed features, refraining from accessing the first plurality of features for the generation of the second set of ML models.

15. A non-transitory machine-readable medium that stores instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

generating a first set of machine learning (ML) models each using a first set of features;

determining, for each feature of the first set of features and each ML model of the first set of ML models, an importance measure based on a relation between the feature and output of the ML model;

based on the importance measures, defining a second set of features comprising a strict subset of the first set of features; and

generating a second set of ML models using the second set of features.

16. The non-transitory machine-readable medium of claim 15 , wherein the defining of the second set of features comprises:

determining, for each feature of the first set of features, based on the importance measures for the feature and a first predetermined threshold, a percentage of the first set of ML models for which the feature is not important;

determining, for the one or more removed features, based on a percentage of ML models for which the feature is not important and a second predetermined threshold, that the feature is not likely to be important to future ML models; and

based on the determination that the one or more removed features are not likely to be important to future ML models, removing the one or more removed features from the first set of features to define the second set of features.

17. The non-transitory machine-readable medium of claim 16 , wherein the second predetermined threshold is 50%.

18. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise:

determining, for each feature of the second set of features and each ML model of the second set of ML models, an importance measure based on a relation between the feature and output of the ML model;

based on the importance measures, defining a third set of features comprising the second set of features with one or more features removed; and

generating, by the one or more processors, a third set of ML models using the third set of features.

19. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise:

determining, based on a percentage of ML models for which a first removed feature of the one or more removed features is important, a confidence measure for the determination that the first removed feature is not likely to be important to future ML models;

defining a third set of features comprising the second set of features and the first removed feature, based on the confidence measure, a third predetermined threshold, and a number of retrains elapsed since the removal of the first removed feature from the second set of features; and

generating, by the one or more processors, a third set of ML models using the third set of features.

20. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise:

prior to generating the first set of ML models, accessing a first plurality of features over a network;

based on the first plurality of features, generating a second feature of the first set of features;

creating a feature dependency graph that relates the second feature to the first plurality of features; and

based on the feature dependency graph and the second feature being one of the one or more removed features, refraining from accessing the first plurality of features for the generation of the second set of ML models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2020
From: HUU, JACQUES DOAN
To: SAP SE
Reel/Frame 052590/0110 →
Continuity (1)
Related Publication 20210350273A1 · Nov 11, 2021
Cited By (1)
US 12,711,423