Early warning for issue validation failure using machine learning model
A computing system receives issue data for a plurality of issues from a risk management system. The issue data for each respective issue of the plurality of issues includes an issue identifier. The computing system generates issue validation risk scores for the plurality of issues. An issue validation risk score for the respective issue indicates a predicted probability that a resolution of the respective issue will fail validation within a fixed period of time. The computing system generates a warning message for at least the respective issue of the plurality of issues based on the issue validation risk score for the respective issue being greater than a threshold. The computing system revises parameters of the at least one machine learning model based on validation failure outcomes associated with the plurality of issues.
1 . A method comprising:
receiving, by a computing system, issue data for a plurality of issues each related to at least one of governance or compliance associated with a policy from a risk management system, wherein the issue data for each respective issue of the plurality of issues includes an issue identifier;
generating, by the computing system using at least one machine learning model and based on the issue data, issue validation risk scores for the plurality of issues, wherein an issue validation risk score for the respective issue indicates a predicted probability that a resolution of the respective issue to satisfy the policy will fail validation within a fixed period of time;
generating a warning message for at least the respective issue of the plurality of issues based on the issue validation risk score for the respective issue being greater than a threshold;
tagging, by the computing system, the respective issue of the warning message for use as a feature input for training the at least one machine learning model;
updating, by the computing system, training data to include the feature input associated with the tagged respective issue, wherein the training data includes monitored validation failure outcomes associated with one or more issues of the plurality of issues for which warning messages were generated, and wherein the training data is used to train the at least one machine learning model;
determining, by the computing system, one or more features corresponding to a level of modified organizational attention applied to the tagged respective issue in response to the warning message; and
revising, by the computing system and as part of training the at least one machine learning model using the updated training data that includes the monitored validation failure outcomes and the feature input associated with the tagged respective issue, parameters of the at least one machine learning model, wherein revising the parameters of the at least one machine learning model includes:
retuning hyperparameters of the at least one machine learning model using the monitored validation failure outcomes and the feature input associated with the tagged respective issue included in the updated training data and based on the one or more features corresponding to the level of modified organizational attention applied to the tagged respective issue.
2 . The method of claim 1 , further comprising generating data representative of a dashboard for display via a display device, wherein the dashboard indicates one or more issue identifiers of one or more issues of the plurality of issues along with the issue validation risk scores associated with the one or more issues, and wherein the dashboard further indicates a ranking of the one or more issues based on the issue validation risk scores associated with the one or more issues.
3 . The method of claim 1 , further comprising generating, by the computing system using at least one second machine learning model, second issue validation risk scores for the plurality of issues, wherein a second issue validation risk score for the respective issue indicates a predicted probability that the resolution of the respective issue will fail validation within a second fixed period of time, wherein the second fixed period of time is longer than the fixed period of time.
4 . The method of claim 3 , wherein generating the second issue validation risk scores for the plurality of issues comprises generating the second issue validation risk scores for the plurality of issues in parallel with generating the issue validation risk scores for the plurality of issues.
5 . The method of claim 1 , wherein the at least one machine learning model is a random forest model that includes an ensemble of predictive elements, wherein each predictive element determines whether the resolution of the respective issue will fail validation within the fixed period of time based on one or more features of the issue data for the respective issue, and wherein the issue validation risk score for the respective issue comprises a percentage of the predictive elements that determined that the resolution of the respective issue will fail validation within the fixed period of time.
6 . The method of claim 1 , further comprising preprocessing, by the computing system, the issue data for the plurality of issues, wherein preprocessing includes creating one or more variables from raw data of the issue data for the plurality of issues, wherein the one or more variables include one or more of numeric, categorical, or text variables.
7 . The method of claim 1 , further comprising processing, by the computing system, one or more variables created from the issue data for the plurality of issues into one or more features used as input to the at least one machine learning model, wherein processing includes performing feature extraction on the one or more variables.
8 . The method of claim 7 , wherein performing feature extraction includes generating text features as term frequency-inverse document frequency (TFIDF) vectors from text variables created from the issue data for the plurality of issues, wherein the TFIDF vectors comprise numeric presentations of the text variables.
9 . The method of claim 7 , wherein performing feature extraction includes generating numeric features by imputing and scaling numeric variables created from the issue data for the plurality of issues.
10 . The method of claim 7 , wherein performing feature extraction includes generating categorical features by encoding categorical variables created from the issue data for the plurality of issues.
11 . The method of claim 1 , further comprising monitoring performance of the at least one machine learning model over time, wherein monitoring performance includes monitoring validation failure outcomes for the one or more issues of the plurality of issues for which warning messages were generated to determine the monitored validation failure outcomes, and wherein the monitored validation failure outcomes further comprise one or more of failed validations or passed validations.
12 . A computing system comprising:
one or more storage devices; and
processing circuitry in communication with the one or more storage devices, the processing circuitry configured to:
receive issue data for a plurality of issues each related to at least one of governance or compliance associated with a policy from a risk management system, wherein the issue data for each respective issue of the plurality of issues includes an issue identifier;
generate, using at least one machine learning model and based on the issue data, issue validation risk scores for the plurality of issues, wherein an issue validation risk score for the respective issue indicates a predicted probability that a resolution of the respective issue to satisfy the policy will fail validation within a fixed period of time;
generate a warning message for at least the respective issue of the plurality of issues based on the issue validation risk score for the respective issue being greater than a threshold;
tag the respective issue of the warning message for use as a feature input for training the at least one machine learning model;
update training data to include the feature input associated with the tagged respective issue, wherein the training data includes monitored validation failure outcomes associated with one or more issues of the plurality of issues for which warning messages were generated, and wherein the training data is used to train the at least one machine learning model;
determine one or more features corresponding to a level of modified organizational attention applied to the tagged respective issue in response to the warning message; and
revise, as part of training the at least one machine learning model using the updated training data that includes the monitored validation failure outcomes and the feature input associated with the tagged respective issue, parameters of the at least one machine learning model, wherein to revise the parameters of the at least one machine learning model, the processing circuitry is configured to:
retune hyperparameters of the at least one machine learning model using the monitored validation failure outcomes and the feature input associated with the tagged respective issue included in the updated training data and based on the one or more features corresponding to the level of modified organizational attention applied to the tagged respective issue.
13 . The computing system of claim 12 , wherein the processing circuitry is further configured to generate data representative of a dashboard for display via a display device, wherein the dashboard indicates one or more issue identifiers of one or more issues of the plurality of issues along with the issue validation risk scores associated with the one or more issues, and wherein the dashboard further indicates a ranking of the one or more issues based on the issue validation risk scores associated with the one or more issues.
14 . The computing system of claim 12 , wherein the processing circuitry is further configured to
generate, using at least one second machine learning model, second issue validation risk scores for the plurality of issues, wherein a second issue validation risk score for the respective issue indicates a predicted probability that the resolution of the respective issue will fail validation within a second fixed period of time, wherein the second fixed period of time is longer than the fixed period of time.
15 . The computing system of claim 14 , wherein to generate the second issue validation risk scores for the plurality of issues, the processing circuitry is configured to generate the second issue validation risk scores for the plurality of issues in parallel with generation of the issue validation risk scores for the plurality of issues.
16 . The computing system of claim 12 , wherein the at least one machine learning model is a random forest model that includes an ensemble of predictive elements, wherein each predictive element determines whether the resolution of the respective issue will fail validation within the fixed period of time based on one or more features of the issue data for the respective issue, and wherein the issue validation risk score for the respective issue comprises a percentage of the predictive elements that determined that the resolution of the respective issue will fail validation within the fixed period of time.
17 . The computing system of claim 12 , wherein the processing circuitry is further configured to preprocess the issue data for the plurality of issues, wherein to preprocess the issue data, the processing circuitry is configured to create one or more variables from raw data of the issue data for the plurality of issues, wherein the one or more variables include one or more of numeric, categorical, or text variables.
18 . The computing system of claim 12 , wherein the processing circuitry is further configured to process one or more variables created from the issue data for the plurality of issues into one or more features used as input to the at least one machine learning model, wherein to process the one or more variables, the processing circuitry is configured to perform feature extraction on the one or more variables.
19 . A non-transitory computer-readable storage medium comprising instructions that, when executed, cause processing circuitry to:
receive issue data for a plurality of issues each related to at least one of governance or compliance associated with a policy from a risk management system, wherein the issue data for each respective issue of the plurality of issues includes an issue identifier;
generate, using at least one machine learning model and based on the issue data, issue validation risk scores for the plurality of issues, wherein an issue validation risk score for the respective issue indicates a predicted probability that a resolution of the respective issue to satisfy the policy will fail validation within a fixed period of time;
generate a warning message for at least the respective issue of the plurality of issues based on the issue validation risk score for the respective issue being greater than a threshold;
tag the respective issue of the warning message for use as a feature input for training the at least one machine learning model;
update training data to include the feature input associated with the tagged respective issue, wherein the training data includes monitored validation failure outcomes associated with one or more issues of the plurality of issues for which warning messages were generated, and wherein the training data is used to train the at least one machine learning model;
determine one or more features corresponding to a level of modified organizational attention applied to the tagged respective issue in response to the warning message; and
revise, as part of training the at least one machine learning model using the updated training data that includes the monitored validation failure outcomes and the feature input associated with the tagged respective issue, parameters of the at least one machine learning model, wherein to revise the parameters of the at least one machine learning model, the processing circuitry is configured to:
retune hyperparameters of the at least one machine learning model using the monitored validation failure outcomes and the feature input associated with the tagged respective issue included in the updated training data and based on the one or more features corresponding to the level of modified organizational attention applied to the tagged respective issue.