IP Library › Granted Patent US 12,437,237
Granted Patent B2
US 12,437,237 · App. 17/859,978 · Granted Oct 7, 2025

Sequential synthesis and selection for feature engineering

Inventor: Michael Langford (Plano, TX)
Assignee: Capital One Services, LLC
G06N20/00G06F18/211G06F18/2155
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,237
App. No.
17/859,978
Granted
Oct 7, 2025
Kind
B2
Abstract

Systems and methods, as described herein, relate to sequential synthesis and selection for feature engineering. A dataset may be associated with a label defining a machine-learning target attribute and a received operation that can be applied to at least one of the existing features of the dataset. One or more potential features may be generated by applying the operation to one or more existing features. For each of the one or more potential features, a feature importance algorithm may be applied to the respective feature along with the one or more existing features, generating a respective feature importance value. Respective feature importance values may be generated for each of the one or more existing features based on applying the feature importance algorithm and used to sort the potential features. A level of correlation to each of the one or more existing features may be determined to make sure it is under a threshold level to avoid new features heavily correlated to existing ones.

Claims (95)

1. A computer-implemented method, executing on a computing device, comprising:

receiving a dataset with one or more existing features, wherein the received dataset is associated with a dataset label defining a machine-learning target attribute;

receiving an operation that can be applied to at least one of the one or more existing features of the dataset;

generating one or more potential features by applying the operation to the at least one of the one or more existing features;

for each of the one or more potential features:

applying a feature importance algorithm to the respective feature along with the one or more existing features, and

generating a respective feature importance value for the respective feature based on applying the feature importance algorithm;

generating respective feature importance values for each of the one or more existing features based on applying the feature importance algorithm; and

updating, based on applying the feature importance algorithm, the respective feature importance value for the respective feature by:

summing the respective feature importance values for the one or more existing features, and

dividing by the respective feature importance value:

sorting the generated one or more potential features by their respective feature importance values;

receiving a threshold level of correlation;

iterating through the sorted, generated one or more potential features until a total number of new features are added or until no more potential features are left by, for a given potential feature:

determining one or more correlations between the given potential feature and each of the one or more existing features; and

adding, based on determining that each of the one or more correlations is under the threshold level of correlation, the given potential feature to the dataset as a new feature; and

training, using the dataset, a machine learning model.

2. The method of claim 1 , further comprising:

skipping adding a second feature of the sorted, generated one or more potential features to the dataset based on determining that a correlation between the second feature and at least one of the one or more existing features is at or over the threshold level of correlation.

3. The method of claim 1 , wherein for each of the one or more potential features:

a plurality of respective feature importance values for the respective feature are obtained based on applying the feature importance algorithm a plurality of times,

a plurality of respective feature importance values for each of the one or more existing features are obtained based on applying the feature importance algorithm the plurality of times,

the respective feature importance value for the respective feature is obtained by averaging the plurality of respective feature importance values for the respective feature, and

the respective feature importance values for the one or more existing features is obtained by averaging the plurality of respective feature importance values for each respective one or more existing features.

4. The method of claim 1 , wherein an added first feature of the dataset is used as an input to the trained machine-learning model.

5. The method of claim 1 , further comprising:

receiving a number of runs to perform; and

wherein for each of the one or more potential features:

the number of runs to perform of respective feature importance values for the respective feature are obtained based on applying the feature importance algorithm the number of runs to perform times,

the number of runs to perform of respective feature importance values for each of the one or more existing features are obtained based on applying the feature importance algorithm the number of runs to perform times,

the respective feature importance value for the respective feature is obtained by averaging the number of runs to perform of respective feature importance values for the respective feature, and

the respective feature importance values for the one or more existing features is obtained by averaging the number of runs to perform of respective feature importance values for each respective one or more existing features.

6. The method of claim 1 , further comprising:

receiving an indication of the total number of new features to add.

7. The method of claim 6 further comprising:

sending an alert based on determining the number of new features added consequent to iterating through the sorted, generated one or more potential features is less than the total number of new features.

8. A computing system comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the computing system to:

receive a dataset with one or more existing features, wherein the received dataset is associated with a dataset label defining a machine-learning target attribute;

receive an operation that can be applied to at least one of the one or more existing features of the dataset;

generate one or more potential features by applying the operation to the at least one of the one or more existing features;

for each of the one or more potential features:

apply a feature importance algorithm to the respective feature along with the one or more existing features,

generate a respective feature importance value for the respective feature based on applying the feature importance algorithm, and

generate respective feature importance values for each of the one or more existing features based on applying the feature importance algorithm;

sort the generated one or more potential features by their respective feature importance values;

receive a threshold level of correlation;

iterate through the sorted, generated one or more potential features until a total number of new features are added or until no more potential features are left by, for a given potential feature:

determining one or more correlations between the given potential feature and each of the one or more existing features; and

adding, based on determining that each of the one or more correlations is under the threshold level of correlation, the given potential feature to the dataset as a new feature; and

train, using the dataset, a machine learning model, wherein an added first feature of the dataset is used as an input to the trained machine-learning model.

9. The computing system of claim 8 , wherein the memory is further storing instructions that, when executed by the one or more processors, cause the computing system to:

skip adding a second feature of the sorted, generated one or more potential features to the dataset based on determining that a correlation between the second feature and at least one of the one or more existing features is at or over the threshold level of correlation.

10. The computing system of claim 8 , wherein for each of the one or more potential features:

a plurality of respective feature importance values for the respective feature are obtained based on applying the feature importance algorithm a plurality of times,

a plurality of respective feature importance values for each of the one or more existing features are obtained based on applying the feature importance algorithm the plurality of times,

the respective feature importance value for the respective feature is obtained by averaging the plurality of respective feature importance values for the respective feature, and

the respective feature importance values for the one or more existing features is obtained by averaging the plurality of respective feature importance values for each respective one or more existing features.

11. The computing system of claim 8 , wherein the memory is further storing instructions that, when executed by the one or more processors, cause the computing system to update the respective feature importance value for the respective feature based on applying the feature importance algorithm by summing the respective feature importance values for the one or more existing features and dividing by the respective feature importance value.

12. The computing system of claim 8 , wherein the memory is further storing instructions that, when executed by the one or more processors, cause the computing system to:

receive a number of runs to perform; and

wherein for each of the one or more potential features:

the number of runs to perform of respective feature importance values for the respective feature are obtained based on applying the feature importance algorithm the number of runs to perform times,

the number of runs to perform of respective feature importance values for each of the one or more existing features are obtained based on applying the feature importance algorithm the number of runs to perform times,

the respective feature importance value for the respective feature is obtained by averaging the number of runs to perform of respective feature importance values for the respective feature, and

the respective feature importance values for the one or more existing features is obtained by averaging the number of runs to perform of respective feature importance values for each respective one or more existing features.

13. The computing system of claim 8 , wherein the memory is further storing instructions that, when executed by the one or more processors, cause the computing system to:

receive the total number of new features to add.

14. The computing system of claim 13 , wherein the memory is further storing instructions that, when executed by the one or more processors, cause the computing system to:

send an alert based on determining the number of new features added consequent to iterating through the sorted, generated one or more potential features is less than the total number of new features.

15. A non-transitory computer-readable storage medium comprising instructions that, when executed, cause a computing system to:

receive a dataset with one or more existing features, wherein the received dataset is associated with a dataset label defining a machine-learning target attribute;

receive an operation that can be applied to at least one of the one or more existing features of the dataset;

generate one or more potential features by applying the operation to the at least one of the one or more existing features;

for each of the one or more potential features:

apply a feature importance algorithm to the respective feature along with the one or more existing features;

generate a respective feature importance value for the respective feature based on applying the feature importance algorithm; and

generate respective feature importance values for each of the one or more existing features based on applying the feature importance algorithm, and;

update, based on applying the feature importance algorithm, the respective feature importance value for the respective feature by:

summing the respective feature importance values for the one or more existing features, and

dividing by the respective feature importance value:

sort the generated one or more potential features by their respective feature importance values;

receive a threshold level of correlation;

iterate through the sorted, generated one or more potential features until a total number of new features are added or until no more potential features are left by, for a given potential feature:

determine one or more correlations between the given potential feature and each of the one or more existing features;

adding, based on determining that each of the one or more correlations is under the threshold level of correlation, the given potential feature to the dataset as a new feature;

skip adding a second feature of the sorted, generated one or more potential features to the dataset based on determining that a correlation between the second feature and at least one of the one or more existing features is at or over the threshold level of correlation; and

train, using the dataset, a machine learning model.

16. The non-transitory computer-readable storage medium of claim 15 , wherein for each of the one or more potential features:

a plurality of respective feature importance values for the respective feature are obtained based on applying the feature importance algorithm a plurality of times,

a plurality of respective feature importance values for each of the one or more existing features are obtained based on applying the feature importance algorithm the plurality of times,

the respective feature importance value for the respective feature is obtained by averaging the plurality of respective feature importance values for the respective feature, and

the respective feature importance values for the one or more existing features is obtained by averaging the plurality of respective feature importance values for each respective one or more existing features.

17. The non-transitory computer-readable storage medium of claim 15 , wherein an added first feature of the dataset is used as an input to the machine-learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2022
From: LANGFORD, MICHAEL
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 060458/0186 →
Continuity (1)
Related Publication 20240013089A1 · Jan 11, 2024
References Cited (7)
US 12306909B1 · Matar · 2025 [cited by examiner]
US 20100257501A1 · Ouali · 2010 [cited by examiner]
US 20170017900A1 · Maor · 2017 [cited by examiner]
US 20190318248A1 · Moreira-Matias · 2019 [cited by examiner]
US 20200311611A1 · Kennedy · 2020 [cited by examiner]
US 20210133630A1 · Dalli · 2021 [cited by examiner]
US 20220091818A1 · Gu · 2022 [cited by examiner]