IP Library › Granted Patent US 12,639,586
Granted Patent B2
US 12,639,586 · App. 17/733,420 · Granted May 26, 2026

Scoring correlated independent variables for elimination from a dataset

Inventor: Mridul Kumar Nath (Bangalore, IN)
Assignee: ORACLE FINANCIAL SERVICES SOFTWARE LIMITED
G06N5/022G06N5/041
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,586
App. No.
17/733,420
Granted
May 26, 2026
Kind
B2
Abstract

Techniques are disclosed as an optimization data system for eliminating correlated independent variables programmatically from data with ranked exclusion scores. The system can obtain an initial dataset comprising variables, determine a set of correlation values by analyzing linear correlation between the variables, generate a correlation matrix using at least in part the set of correlation values and corresponding variables from the initial data, calculate exclusion scores for the variables in the correlation matrix that exhibit multicollinearity, and update the initial dataset by removing at least one variable with the highest exclusion score from the variables to generate an updated dataset comprising optimized variables. The steps for correlation and elimination of variables are iterated until an updated dataset without any correlation is obtained and then a machine learning model may be trained using the updated dataset.

Claims (55)

1 . A method comprising:

obtaining an initial dataset, wherein the initial dataset comprises a plurality of independent variables;

determining a plurality of correlation values by analyzing linear correlation between at least two independent variables in the plurality of independent variables in the initial dataset;

generating a correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset;

calculating redundancy values for the independent variables in the correlation matrix, wherein the redundancy value for a corresponding independent variable depends on a number of times the corresponding independent variable appears in the correlation matrix;

calculating prediction strengths for the independent variables in the correlation matrix, wherein the prediction strength is related to an explanatory power of the corresponding independent variable;

calculating exclusion scores for the independent variables in the correlation matrix that exhibit multicollinearity based on redundancy values for the independent variables in the correlation matrix and prediction strengths for the independent variables in the correlation matrix;

updating the initial dataset by removing at least one independent variable with a highest exclusion score from the plurality of independent variables to generate an updated training dataset comprising an optimized plurality of independent variables, wherein the plurality of correlation values are determined, the correlation matrix is generated, and the initial dataset is updated iteratively until multicollinearity is no longer exhibited between the independent variables;

training a machine learning model using the updated training dataset, wherein the machine learning model comprises one or more algorithms to make an inference based on relationships between the optimized plurality of independent variables, the training comprising iteratively inputting the updated training dataset to the machine learning model to determine a set of optimal hyperparameters that minimizes a cost function; and

generating the trained machine learning model having the set of optimal hyperparameters,

wherein the trained machine learning model is configured to, based on a set of independent variables provided as an input, output a prediction for at least one dependent variable.

2 . The method of claim 1 , wherein the initial dataset comprises historical transaction data, loan records, credit records, a customer's profile, or any combination thereof.

3 . The method of claim 1 , wherein generating the correlation matrix using at least in part the plurality of correlation values and corresponding variables in the plurality of independent variables in the initial dataset, includes a process to:

generate a threshold value used to compare with each correlation value in the plurality of correlation values; and

incorporate the correlation values that are greater than the threshold value with corresponding independent variables into the correlation matrix.

4 . The method of claim 1 , wherein a pre-set criteria is used to determine whether a significant linear correlation exists between at least two variables in the plurality of independent variables via a threshold value, and wherein the independent variables determined to exhibit the significant linear correlation are identified as the independent variables in the correlation matrix that exhibit multicollinearity.

5 . The method of claim 1 , wherein the training further comprises adapting the machine learning model to minimize a difference between a final value for a dependent variable and ground truth information.

6 . The method of claim 1 , wherein the trained machine learning model is deployed in a real-world environment.

7 . A computing system comprising:

a processor; and

a memory including instructions that, when executed with the processor, cause the computing system to, at least:

obtain an initial dataset, wherein the initial dataset comprises a plurality of independent variables;

determine a plurality of correlation values by analyzing linear correlation between at least two independent variables in the plurality of independent variables in the initial dataset;

generate a correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset;

calculate redundancy values for the independent variables in the correlation matrix, wherein the redundancy value for a corresponding independent variable depends on a number of times the corresponding independent variable appears in the correlation matrix;

calculate prediction strengths for the independent variables in the correlation matrix, wherein the prediction strength is related to an explanatory power of the corresponding independent variable;

calculate exclusion scores for the independent variables in the correlation matrix that exhibit multicollinearity based on redundancy values for the independent variables in the correlation matrix and prediction strengths for the independent variables in the correlation matrix;

update the initial dataset by removing at least one independent variable with a highest exclusion score from the plurality of independent variables to generate an updated training dataset comprising an optimized plurality of independent variables, wherein the plurality of correlation values are determined, the correlation matrix is generated, and the initial dataset is updated iteratively until multicollinearity is no longer exhibited between the independent variables;

train a machine learning model using the updated training dataset, wherein the machine learning model comprises one or more algorithms to make an inference based on relationships between the optimized plurality of independent variables, the training including iteratively inputting the updated training dataset to the machine learning model to determine a set of optimal hyperparameters that minimizes a cost function; and

generate the trained machine learning model having the set of optimal hyperparameters,

wherein the trained machine learning model is configured to, based on a set of independent variables provided as an input, output a prediction for at least one dependent variable.

8 . The computing system of claim 7 , wherein the initial dataset comprises historical transaction data, loan records, credit records, a customer's profile, or any combination thereof.

9 . The computing system of claim 7 , wherein generating the correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset, includes a process to:

generate a threshold value used to compare with each correlation value in the plurality of correlation values; and

incorporate the correlation values that are greater than the threshold value with corresponding independent variables into the correlation matrix.

10 . The computing system of claim 7 , wherein a pre-set criteria is used to determine whether a significant linear correlation exists between at least two variables in the plurality of independent variables via a threshold value, and wherein the independent variables determined to exhibit the significant linear correlation are identified as the independent variables in the correlation matrix that exhibit multicollinearity.

11 . The computing system of claim 7 , wherein the training further includes adapting the machine learning model to minimize a difference between a final value for a dependent variable and ground truth information.

12 . The computing system of claim 11 , wherein the trained machine learning model is deployed in a real-world environment.

13 . A non-transitory computer readable medium storing specific computer-executable instructions that, when executed by a processor, cause a computer system to at least:

obtain an initial dataset, wherein the initial dataset comprises a plurality of independent variables;

determine a plurality of correlation values by analyzing linear correlation between at least two independent variables in the plurality of independent variables in the initial dataset;

generate a correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset;

calculate redundancy values for the independent variables in the correlation matrix, wherein the redundancy value for a corresponding independent variable depends on a number of times the corresponding independent variable appears in the correlation matrix;

calculate prediction strengths for the independent variables in the correlation matrix, wherein the prediction strength is related to an explanatory power of the corresponding independent variable;

calculate exclusion scores for the independent variables in the correlation matrix that exhibit multicollinearity based on redundancy values for the independent variables in the correlation matrix and prediction strengths for the independent variables in the correlation matrix;

update the initial dataset by removing at least one independent variable with a highest exclusion score from the plurality of independent variables to generate an updated training dataset comprising an optimized plurality of independent variables, wherein the plurality of correlation values are determined, the correlation matrix is generated, and the initial dataset is updated iteratively until multicollinearity is no longer exhibited between the independent variables;

train a machine learning model using the updated training dataset, wherein the machine learning model comprises one or more algorithms to make an inference based on relationships between the optimized plurality of independent variables, the training including iteratively inputting the updated training dataset to the machine learning model to determine a set of optimal hyperparameters that minimizes a cost function; and

generate the trained machine learning model having the set of optimal hyperparameters,

wherein the trained machine learning model is configured to, based on a set of independent variables provided as an input, output a prediction for at least one dependent variable.

14 . The non-transitory computer readable medium of claim 13 , wherein generating the correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset, includes a process to:

generate a threshold value used to compare with each correlation value in the plurality of correlation values; and

incorporate the correlation values that are greater than the threshold value with corresponding independent variables into the correlation matrix.

15 . The non-transitory computer readable medium of claim 13 , a pre-set criteria is used to determine whether a significant linear correlation exists between at least two variables in the plurality of independent variables via a threshold value, and wherein the independent variables determined to exhibit the significant linear correlation are identified as the independent variables in the correlation matrix that exhibit multicollinearity.

16 . The non-transitory computer readable medium of claim 13 , wherein the training further includes adapting the machine learning model to minimize a difference between a final value for a dependent variable and ground truth information.

17 . The non-transitory computer readable medium of claim 13 , wherein the trained machine learning model is deployed in a real-world environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2022
From: NATH, MRIDUL KUMAR
To: ORACLE FINANCIAL SERVICES SOFTWARE LIMITED
Reel/Frame 059717/0978 →
Continuity (1)
Related Publication 20230351211A1 · Nov 2, 2023
References Cited (11)
Vu, Dao Hoang, Kashem M. Muttaqi, and Ashish P. Agalgaonkar. “A variance inflation factor and backward elimination based robust regression model for forecasting monthly electricity demand using climatic variables.” Appl… [cited by examiner]
Ballabio, Davide, et al. “A novel variable reduction method adapted from space-filling designs.” Chemometrics and Intelligent Laboratory Systems 136 (2014): 147-154. (Year: 2014) (Year: 2014). [cited by examiner]
Tanioka, Kensuke, Yuki Furotani, and Satoru Hiwa. “Thresholding approach for low-rank correlation matrix based on mm algorithm.” Entropy 24.5 (2022): 579. (Year: 2022) (Year: 2022). [cited by examiner]
Oh, Taeseob, et al. “Machine learning-based diagnosis and risk factor analysis of cardiocerebrovascular disease based on KNHANES.” Scientific reports 12.1 (2022): 2250. (Year: 2022) (Year: 2022). [cited by examiner]
Chan, Jireh Yi-Le, et al. “Mitigating the multicollinearity problem and its machine learning approach: a review.” Mathematics 10.8 ( 2022): 1283. (Year: 2022). [cited by examiner]
Nie, Guangli, et al. “Credit card churn forecasting by logistic regression and decision tree.” Expert Systems with Applications 38. 12 (2011): 15273-15285. (Year: 2011) (Year: 2011). [cited by examiner]
Peng, Hanchuan, Fuhui Long, and Chris Ding. “Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy.” IEEE Transactions on pattern analysis and machine intelligence 2… [cited by examiner]
James, Gareth, et al. An introduction to statistical learning. vol. 112. New York: springer, 2013. (Year: 2013) (Year: 2013). [cited by examiner]
Adnan et al., A Comparative Study on Some Methods for Handling Multicollinearity Problems, Matematka, vol. 22, No. 2, Jan. 2006, pp. 109-119. [cited by applicant]
Garg et al., Comparison of Regression Analysis, Artificial Neural Network and Genetic Programming in Handling the Multicollinearity Problem, Proceedings of International Conference on Modelling, Identification and Contr… [cited by applicant]
Paul, Multicollinearity: Causes, Effects and Remedies, Mathematics, Available Online at: https://www.semanticscholar.org/paper/MULTICOLLINEARITY%3ACAUSES%2CEFFECTSANDREMEDIESPaul/300f806f39ea84cfbe999cb52903726a2e8f0458… [cited by applicant]