IP Library Granted Patent US 12,730,945
Granted Patent B2
US 12,730,945 · App. 17/736,613 · Granted Sep 8, 2026

Feature selection method and system for regression analysis / model construction

Inventors: Vladimir Sevastyanov (Fort Worth, TX); Abhay Dabholkar (Allen, TX); Rakhi Gupta (Frisco, TX); James H. Pratt (Round Rock, TX); Nikhlesh Agrawal (McKinney, TX)
Assignee: AT&T Intellectual Property I, L.P.
G06F30/20G06F2111/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,945
App. No.
17/736,613
Granted
Sep 8, 2026
Kind
B2
Abstract

Aspects of the subject disclosure may include, for example, dividing a feature range of a feature into a plurality of subsets that span the feature range, calculating an average target variable value for each subset of the plurality of subsets, resulting in a plurality of average target variable values, and estimating a measure of feature significance with respect to a target variable by determining a difference between a maximum average target variable value and a minimum average target variable value in the plurality of average target variable values. Other embodiments are disclosed.

Claims (41)

1 . A method, comprising:

selecting, by a processing system including a processor, a feature from a plurality of features in a data set for generating a machine learning model, wherein:

the plurality of features are determined to potentially impact a target variable of the machine learning model, and

the data set comprises feature values of the plurality of features associated with target variable values of the target variable;

dividing, by the processing system, a feature range of the feature into a plurality of subsets that span the feature range, wherein the feature range comprises a value range of the feature values for the feature from the data set;

calculating, by the processing system, an average target variable value from the target variable values associated with the feature values for each subset of the plurality of subsets, resulting in a plurality of average target variable values;

estimating, by the processing system, a measure of feature significance with respect to the target variable by determining a difference between a maximum average target variable value and a minimum average target variable value in the plurality of average target variable values; and

generating the machine learning model using at least the feature based on the measure of feature significance indicating an impact on the target variable.

2 . The method of claim 1 , wherein the average target variable value comprises an arithmetic mean value.

3 . The method of claim 1 , wherein the average target variable value comprises a median value.

4 . The method of claim 1 , wherein the average target variable value comprises a mode value.

5 . The method of claim 1 , wherein the average target variable value comprises a mid-range value.

6 . The method of claim 1 , wherein the feature comprises a numeric feature, wherein the method further comprises determining, by the processing system, the feature range for the numeric feature based on a difference between maximum and minimum values of the numeric feature, wherein the dividing the feature range comprises splitting the feature range into a set of equal sub-intervals, and wherein the calculating the average target variable value for each subset of the plurality of subsets comprises calculating an average target variable value for each of the equal sub-intervals.

7 . The method of claim 1 , wherein the feature comprises a categorical feature, wherein the dividing the feature range comprises determining a list of unique categorical values, and wherein the calculating the average target variable value for each subset of the plurality of subsets comprises calculating an average target variable value for each of the unique categorical values.

8 . The method of claim 1 , wherein the feature comprises one or more of an integer feature or a logical feature.

9 . The method of claim 1 , wherein the machine learning model comprises a regression model.

10 . A device, comprising:

a processing system including a processor; and

a memory that stores executable instructions that, when executed by the processing system, facilitate performance of operations, the operations comprising:

determining a type of each feature of a plurality of features associated with a target variable in a data set for generating a machine learning model, wherein:

the plurality of features are determined to potentially impact a target variable of the machine learning model, and

the data set comprises feature values of the plurality of features associated with target variable values of the target variable;

for each feature of the plurality of features, estimating a respective significance value for that feature, relative to the target variable, based on the type of that feature and the feature values and the target variable values of that feature, resulting in a plurality of respective significant values;

performing filtering of the plurality of features based on the plurality of respective significant values; and

generating the machine learning model using a set of features that were filtered from the plurality of features.

11 . The device of claim 10 , wherein the machine learning model comprises a regression model or predictive model.

12 . The device of claim 10 , wherein the performing the filtering comprises sorting the plurality of respective significant values to identify the set of features from the plurality of features that each has a determined significant impact on the target variable.

13 . The device of claim 10 , wherein the performing the filtering comprises comparing the plurality of respective significant values with a threshold.

14 . The device of claim 10 , wherein the estimating the respective significance value involves calculation of arithmetic mean values.

15 . The device of claim 10 , wherein the estimating the respective significance value involves calculation of median values.

16 . The device of claim 10 , wherein the estimating the respective significance value involves calculation of mode values.

17 . The device of claim 10 , wherein the estimating the respective significance value involves calculation of mid-range values.

18 . A non-transitory machine-readable medium, comprising executable instructions that, when executed by a processing system including a processor, facilitate performance of operations, the operations comprising:

identifying a plurality of subsets for a feature in a data set, wherein the data set comprises:

a plurality of features that are determined to potentially impact a target variable of the machine learning model, and

feature values of the plurality of features associated with target variable values of the target variable;

determining an average target variable value from the target variable values associated with the feature for each subset of the plurality of subsets, resulting in a plurality of average target variable values;

calculating a difference between a maximum average target variable value and a minimum average target variable value in the plurality of average target variable values, wherein the difference corresponds to a measure of feature significance with respect to the target variable; and

based on the measure of feature significance satisfying a threshold, utilizing the feature to construct a regression or predictive model.

19 . The non-transitory machine-readable medium of claim 18 , wherein the average target variable value comprises an arithmetic mean value, a median value, a mode value, or a mid-range value.

20 . The non-transitory machine-readable medium of claim 18 , wherein the feature comprises a numeric feature, a categorical feature, an integer feature, or a logical feature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2022
From: SEVASTYANOV, VLADIMIR; DABHOLKAR, ABHAY; GUPTA, RAKHI; PRATT, JAMES H.; AGRAWAL, NIKHLESH
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 059998/0981 →
Continuity (1)
Related Publication 20230359781A1 · Nov 9, 2023
References Cited (32)
US 7970718B2 · Guyon · 2011 [cited by examiner]
US 8190537B1 · Singh · 2012 [cited by examiner]
US 9189750B1 · Narsky · 2015 [cited by examiner]
US 9471871B2 · Bruillard · 2016 [cited by examiner]
US 10452793B2 · Joshi · 2019 [cited by examiner]
US 11257000B2 · Mathew · 2022 [cited by examiner]
US 11429899B2 · Kartoun · 2022 [cited by examiner]
US 11444926B1 · Gama · 2022 [cited by examiner]
US 11500932B2 · Kulkarni · 2022 [cited by examiner]
US 11742081B2 · Kartoun · 2023 [cited by examiner]
US 11899650B2 · Havel · 2024 [cited by examiner]
US 12112261B2 · Partee · 2024 [cited by examiner]
US 20110106735A1 · Weston · 2011 [cited by examiner]
US 20120143799A1 · Wilson · 2012 [cited by examiner]
US 20150242759A1 · Bruillard · 2015 [cited by examiner]
US 20160026917A1 · Weisberg · 2016 [cited by examiner]
US 20160171134A1 · Mirabella · 2016 [cited by examiner]
US 20160224705A1 · Joshi · 2016 [cited by examiner]
US 20180128619A1 · Herberth · 2018 [cited by examiner]
US 20180137219A1 · Goldfarb · 2018 [cited by examiner]
US 20210182357A1 · Partee · 2021 [cited by examiner]
US 20210231834A1 · Prochnow · 2021 [cited by examiner]
US 20210365498A1 · Kulkarni · 2021 [cited by examiner]
US 20220351004A1 · Kanter · 2022 [cited by examiner]
US 20220374765A1 · Wu · 2022 [cited by examiner]
US 20230076592A1 · Pratt · 2023 [cited by examiner]
US 20230117925A1 · Havel · 2023 [cited by examiner]
US 20230205847A1 · Ackerman · 2023 [cited by examiner]
US 20230359781A1 · Sevastyanov · 2023 [cited by examiner]
Ting et al. (Efficient Learning and Feature Selection in High-Dimensional Regression, Neural Computation 22, 831-886 (2010)) (Year: 2010). [cited by examiner]
Rendall et al. (Wide spectrum feature selection (WiSe) for regression model building, Computers and Chemical Engineering 121 (2019) 99-110) (Year: 2019). [cited by examiner]
Guyon, Isabelle et al., “An Introduction to Variable and Feature Selection”, Journal of Machine Learning Research, 2003, 26 pages. [cited by applicant]