Scoring correlated independent variables for elimination from a dataset
View Patent ↗Techniques are disclosed as an optimization data system for eliminating correlated independent variables programmatically from data with ranked exclusion scores. The system can obtain an initial dataset comprising variables, determine a set of correlation values by analyzing linear correlation between the variables, generate a correlation matrix using at least in part the set of correlation values and corresponding variables from the initial data, calculate exclusion scores for the variables in the correlation matrix that exhibit multicollinearity, and update the initial dataset by removing at least one variable with the highest exclusion score from the variables to generate an updated dataset comprising optimized variables. The steps for correlation and elimination of variables are iterated until an updated dataset without any correlation is obtained and then a machine learning model may be trained using the updated dataset.
1 . A method comprising:
obtaining an initial dataset, wherein the initial dataset comprises a plurality of independent variables;
determining a plurality of correlation values by analyzing linear correlation between at least two independent variables in the plurality of independent variables in the initial dataset;
generating a correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset;
calculating redundancy values for the independent variables in the correlation matrix, wherein the redundancy value for a corresponding independent variable depends on a number of times the corresponding independent variable appears in the correlation matrix;
calculating prediction strengths for the independent variables in the correlation matrix, wherein the prediction strength is related to an explanatory power of the corresponding independent variable;
calculating exclusion scores for the independent variables in the correlation matrix that exhibit multicollinearity based on redundancy values for the independent variables in the correlation matrix and prediction strengths for the independent variables in the correlation matrix;
updating the initial dataset by removing at least one independent variable with a highest exclusion score from the plurality of independent variables to generate an updated training dataset comprising an optimized plurality of independent variables, wherein the plurality of correlation values are determined, the correlation matrix is generated, and the initial dataset is updated iteratively until multicollinearity is no longer exhibited between the independent variables;
training a machine learning model using the updated training dataset, wherein the machine learning model comprises one or more algorithms to make an inference based on relationships between the optimized plurality of independent variables, the training comprising iteratively inputting the updated training dataset to the machine learning model to determine a set of optimal hyperparameters that minimizes a cost function; and
generating the trained machine learning model having the set of optimal hyperparameters,
wherein the trained machine learning model is configured to, based on a set of independent variables provided as an input, output a prediction for at least one dependent variable.
2 . The method of claim 1 , wherein the initial dataset comprises historical transaction data, loan records, credit records, a customer's profile, or any combination thereof.
3 . The method of claim 1 , wherein generating the correlation matrix using at least in part the plurality of correlation values and corresponding variables in the plurality of independent variables in the initial dataset, includes a process to:
generate a threshold value used to compare with each correlation value in the plurality of correlation values; and
incorporate the correlation values that are greater than the threshold value with corresponding independent variables into the correlation matrix.
4 . The method of claim 1 , wherein a pre-set criteria is used to determine whether a significant linear correlation exists between at least two variables in the plurality of independent variables via a threshold value, and wherein the independent variables determined to exhibit the significant linear correlation are identified as the independent variables in the correlation matrix that exhibit multicollinearity.
5 . The method of claim 1 , wherein the training further comprises adapting the machine learning model to minimize a difference between a final value for a dependent variable and ground truth information.
6 . The method of claim 1 , wherein the trained machine learning model is deployed in a real-world environment.
7 . A computing system comprising:
a processor; and
a memory including instructions that, when executed with the processor, cause the computing system to, at least:
obtain an initial dataset, wherein the initial dataset comprises a plurality of independent variables;
determine a plurality of correlation values by analyzing linear correlation between at least two independent variables in the plurality of independent variables in the initial dataset;
generate a correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset;
calculate redundancy values for the independent variables in the correlation matrix, wherein the redundancy value for a corresponding independent variable depends on a number of times the corresponding independent variable appears in the correlation matrix;
calculate prediction strengths for the independent variables in the correlation matrix, wherein the prediction strength is related to an explanatory power of the corresponding independent variable;
calculate exclusion scores for the independent variables in the correlation matrix that exhibit multicollinearity based on redundancy values for the independent variables in the correlation matrix and prediction strengths for the independent variables in the correlation matrix;
update the initial dataset by removing at least one independent variable with a highest exclusion score from the plurality of independent variables to generate an updated training dataset comprising an optimized plurality of independent variables, wherein the plurality of correlation values are determined, the correlation matrix is generated, and the initial dataset is updated iteratively until multicollinearity is no longer exhibited between the independent variables;
train a machine learning model using the updated training dataset, wherein the machine learning model comprises one or more algorithms to make an inference based on relationships between the optimized plurality of independent variables, the training including iteratively inputting the updated training dataset to the machine learning model to determine a set of optimal hyperparameters that minimizes a cost function; and
generate the trained machine learning model having the set of optimal hyperparameters,
wherein the trained machine learning model is configured to, based on a set of independent variables provided as an input, output a prediction for at least one dependent variable.
8 . The computing system of claim 7 , wherein the initial dataset comprises historical transaction data, loan records, credit records, a customer's profile, or any combination thereof.
9 . The computing system of claim 7 , wherein generating the correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset, includes a process to:
generate a threshold value used to compare with each correlation value in the plurality of correlation values; and
incorporate the correlation values that are greater than the threshold value with corresponding independent variables into the correlation matrix.
10 . The computing system of claim 7 , wherein a pre-set criteria is used to determine whether a significant linear correlation exists between at least two variables in the plurality of independent variables via a threshold value, and wherein the independent variables determined to exhibit the significant linear correlation are identified as the independent variables in the correlation matrix that exhibit multicollinearity.
11 . The computing system of claim 7 , wherein the training further includes adapting the machine learning model to minimize a difference between a final value for a dependent variable and ground truth information.
12 . The computing system of claim 11 , wherein the trained machine learning model is deployed in a real-world environment.
13 . A non-transitory computer readable medium storing specific computer-executable instructions that, when executed by a processor, cause a computer system to at least:
obtain an initial dataset, wherein the initial dataset comprises a plurality of independent variables;
determine a plurality of correlation values by analyzing linear correlation between at least two independent variables in the plurality of independent variables in the initial dataset;
generate a correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset;
calculate redundancy values for the independent variables in the correlation matrix, wherein the redundancy value for a corresponding independent variable depends on a number of times the corresponding independent variable appears in the correlation matrix;
calculate prediction strengths for the independent variables in the correlation matrix, wherein the prediction strength is related to an explanatory power of the corresponding independent variable;
calculate exclusion scores for the independent variables in the correlation matrix that exhibit multicollinearity based on redundancy values for the independent variables in the correlation matrix and prediction strengths for the independent variables in the correlation matrix;
update the initial dataset by removing at least one independent variable with a highest exclusion score from the plurality of independent variables to generate an updated training dataset comprising an optimized plurality of independent variables, wherein the plurality of correlation values are determined, the correlation matrix is generated, and the initial dataset is updated iteratively until multicollinearity is no longer exhibited between the independent variables;
train a machine learning model using the updated training dataset, wherein the machine learning model comprises one or more algorithms to make an inference based on relationships between the optimized plurality of independent variables, the training including iteratively inputting the updated training dataset to the machine learning model to determine a set of optimal hyperparameters that minimizes a cost function; and
generate the trained machine learning model having the set of optimal hyperparameters,
wherein the trained machine learning model is configured to, based on a set of independent variables provided as an input, output a prediction for at least one dependent variable.
14 . The non-transitory computer readable medium of claim 13 , wherein generating the correlation matrix using at least in part the plurality of correlation values and corresponding independent variables in the plurality of independent variables in the initial dataset, includes a process to:
generate a threshold value used to compare with each correlation value in the plurality of correlation values; and
incorporate the correlation values that are greater than the threshold value with corresponding independent variables into the correlation matrix.
15 . The non-transitory computer readable medium of claim 13 , a pre-set criteria is used to determine whether a significant linear correlation exists between at least two variables in the plurality of independent variables via a threshold value, and wherein the independent variables determined to exhibit the significant linear correlation are identified as the independent variables in the correlation matrix that exhibit multicollinearity.
16 . The non-transitory computer readable medium of claim 13 , wherein the training further includes adapting the machine learning model to minimize a difference between a final value for a dependent variable and ground truth information.
17 . The non-transitory computer readable medium of claim 13 , wherein the trained machine learning model is deployed in a real-world environment.