IP Library Granted Patent US 12,205,690
Granted Patent B1
US 12,205,690 · App. 18/119,654 · Granted Jan 21, 2025

Systems and methods for excluded risk factor predictive modeling

Inventors: Marc Maier (Springfield, MA); Shanshan Li (Springfield, MA); Hayley Carlotto (Springfield, MA)
Assignee: Massachusetts Mutual Life Insurance Company
G16H10/60G06N20/00G16H10/20G16H20/10G16H50/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,690
App. No.
18/119,654
Granted
Jan 21, 2025
Kind
B1
Abstract

A suite of fluidless predictive machine learning models includes a fluidless mortality module, smoking propensity model, and prescription fills model. The fluidless machine learning models are trained against a corpus of historical underwriting applications of a sponsoring enterprise, including clinical data of historical applicants. A data appended procedure supplements historical applications data with public records and credit risks. Various features of this data are engineered for improved predictive characteristics. Fluidless models are trained by application of a random forest ensemble including survival, regression and classification models. The trained models produce high-resolution, individual mortality scores. A fluidless underwriting protocol runs these predictive models to assess mortality risk and other risk attributes of a fluidless application that excludes clinical data to determine whether to present an accelerated underwriting offer. If any of the fluidless predictive models determines a high risk target, the applicant is required to submit clinical data.

Claims (47)

1. A method for processing an electronic application without clinical data, the method comprising:

receiving, by a server, information for an electronic application from a user device, wherein the information for the electronic application excludes all clinical data for an applicant;

upon receiving the information for the electronic application from the user display device, retrieving, by the server, public data identified with the applicant of the electronic application from one or more third party sources;

executing, by the server, a first predictive machine learning model configured to determine a first risk rank representative of a mortality risk for the electronic application and to classify the electronic application into one of a first high risk group and a first low risk group based upon the first risk rank,

wherein the first predictive machine learning model is trained by inputting into the first predictive machine learning model a plurality of historical application records supplemented with public data identified with an applicant of the respective historical application record, wherein the first predictive machine learning model is iteratively trained using updated public data to select variables of the public data with highest importance based on mortality risk;

executing, by the server, a second predictive machine learning model configured to determine a second risk rank representative of a smoking propensity of the electronic application and to classify the electronic application into one of a second high risk group and a second low risk group based upon the second risk rank;

executing, by the server, a third predictive machine learning model configured to determine a third risk rank representative of prescription drug data of the electronic application, and to classify the electronic application into one of a third high risk group and a third low risk group based upon the third risk rank; and

responsive to the server classifying the electronic application into all of the first low risk group, the second low risk group, and the third low risk group, generating, by the server, a user interface that displays information associated with an accelerated application offer.

2. The method of claim 1 , wherein the first predictive machine learning model is a regression random forest model.

3. The method of claim 1 , wherein the public data identified with the applicant of the electronic application, and the public data identified with the respective applicant of each historical application record, comprise public records and credit risk data.

4. The method of claim 3 , wherein the first predictive machine learning model is configured to fit a survival random forest model to the public records and the credit risk data record to select actuarially important attributes as covariates of first predictive machine learning model.

5. The method of claim 1 , wherein the first predictive machine learning model is configured to fit a survival model to the public data to select attributes of the public data with highest importance based on mortality risk as variables of the first predictive machine learning model.

6. The method of claim 5 , wherein the survival model is configured to approximate a survival function that describes probability that an event occurs later than some given time.

7. The method of claim 1 , wherein the second predictive machine learning model comprises a random forest classification model configured to estimate a propensity of the applicant of the respective historical application record to be a smoker.

8. The method of claim 1 , wherein the third predictive machine learning model comprises a prescription drug data random forest classification model configured to determine disqualifying medical risks based on information derived from prescription drug fills for the applicant of the electronic application.

9. The method of claim 1 , further comprising, when the server classifies the electronic application into at least one of the first high risk group, the second high risk group, and the third high risk group, generating and presenting a user interface that displays an instruction to submit the clinical data for the applicant of the electronic application.

10. A method for processing an electronic application without clinical data, the method comprising:

receiving, by a server, information for an electronic application from a user device, wherein the information for the electronic application excludes all clinical data for an applicant;

upon receiving the information for the electronic application from the user display device,

retrieving, by the server, public data identified with the applicant of the electronic application from one or more third party sources;

executing, by the server, a first predictive machine learning model configured to determine a first risk rank representative of a mortality risk for the electronic application and to classify the electronic application into one of a first high risk group and a first low risk group based upon the first risk rank, wherein the first predictive machine learning model is trained by inputting into the first predictive machine learning model a plurality of historical application records supplemented with public data identified with an applicant of the respective historical application record, wherein the first predictive machine learning model is iteratively trained using updated public data to select variables of the public data with highest importance based on mortality risk;

executing, by the server, a second predictive machine learning model configured to determine a second risk rank representative of a smoking propensity of the electronic application and to classify the electronic application into one of a second high risk group and a second low risk group based upon the second risk rank, wherein the second predictive machine learning model applies random forest classification to estimate the smoking propensity of the electronic application and to determine a smoking/non-smoking binary target;

executing, by the server, a third predictive machine learning model configured to determine a third risk rank representative of prescription drug data of the electronic application, and to classify the electronic application into one of a third high risk group and a third low risk group based upon the third risk rank; and

responsive to the server classifying the electronic application into all of the first low risk group, the second low risk group, and the third low risk group, generating, by the server, a user interface that displays information associated with an accelerated application offer.

11. The method of claim 10 , further comprising, when the server classifies the electronic application into at least one of the first high risk group and the second high risk group, of generating and presenting a user interface that displays an instruction to submit the clinical data for the applicant of the electronic application.

12. The method of claim 10 , wherein the first predictive machine learning model is configured to fit a survival random forest model to the public records and the credit risk data record to select the actuarially important attributes of the public records and credit risk data as covariates of the first predictive machine learning model.

13. The processor-based method of claim 10 , wherein the first predictive machine learning model is a regression random forest model.

14. The processor-based method of claim 10 , further comprising, upon receiving the information for the electronic application from a user device, of

executing the third predictive machine learning module to determine the third risk rank representative of a third user risk attribute, and to classify the electronic application into one of the third high risk group and the third low risk group based upon the third risk rank; and

when the processor classifies the electronic application into all of the first low risk group, the second low risk group, and the third low risk group, generating and presenting the user interface that displays the accelerated application offer.

15. The processor-based method of claim 14 , wherein the third predictive machine learning model comprises a prescription drug data model configured to determine disqualifying medical risks derived from prescription drug fills for the applicant of the electronic application.

16. The processor-based method of claim 10 , further comprising:

imputing missing values of public records and credit risk data to supplement the historical application records; and

inputting the supplemented historical application records into a random forest model ensemble.

17. The processor-based method of claim 10 , further comprising:

transforming variables of public records and credit risk data to supplement the historical application records via one or more of constructing count indicator variables, calculating ratios between countervailing quantitative variables, and determining temporal rates and temporal extents of time-dependent variables; and

inputting the supplemented historical application records into a random forest model ensemble.

18. A system comprising:

an analytical engine server comprising:

a first module configured for receiving information for an electronic application from a user device that excludes all clinical data for an applicant of the electronic application, and for retrieving public data identified with the applicant of the received electronic application from one or more third party source;

a second module configured for executing a predictive machine learning module to determine a mortality risk rank for the electronic application and classify the electronic application into a first low risk group or a first high risk group,

wherein the first predictive machine learning model is trained by inputting into the first predictive machine learning model a plurality of historical application records supplemented with public data identified with an applicant of the respective historical application record, wherein the first predictive machine learning model is iteratively trained using updated public data to select variables of the public data with highest importance based on mortality risk;

a third module configured for executing a smoking propensity predictive model, wherein the smoking propensity model is configured to estimate a propensity of the applicant of the electronic application to be a smoker and determine a smoking/non-smoking binary target;

a fourth module configured for executing a prescription drug data predictive model configured to determine disqualifying medical risks based on information derived from prescription drug fills for the applicant of the electronic application; and

a fifth module configured for generating and presenting a user interface that displays information associated with an accelerated application offer when the analytical engine server classifies the electronic application into the first low risk group, determines the non-smoking binary target, and does not determine the disqualifying medical risk.

19. The system of claim 18 , wherein the fifth module is further configured for generating and presenting a user interface that displays an instruction to submit the clinical data for the applicant of the electronic application when the analytical engine server effects one or more of the following: classifies the electronic application into the first high risk group, determines the smoking binary target, determines the disqualifying medical risk.

20. The system of claim 18 , wherein the predictive machine learning model of the second module is a regression random forest model, and the smoking propensity model of the third module is a random forest classification model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 9, 2023
From: MAIER, MARC; LI, SHANSHAN; CARLOTTO, HAYLEY
To: MASSACHUSETTS MUTUAL LIFE INSURANCE COMPANY
Reel/Frame 062936/0670 →
Continuity (4)
Continuation 15931777 · May 14, 2020
Provisional Application 62993584 · Mar 23, 2020
Provisional Application 62899543 · Sep 12, 2019
Provisional Application 62848397 · May 15, 2019
References Cited (64)
US 8521568B1 · Easley · 2013 [cited by applicant]
US 11257574B1 · Boussios et al. · 2022 [cited by applicant]
US 20030037063A1 · Schwartz · 2003 [cited by applicant]
US 20030177032A1 · Bonissone et al. · 2003 [cited by applicant]
US 20070021987A1 · Binns et al. · 2007 [cited by applicant]
US 20090204446A1 · Simon et al. · 2009 [cited by applicant]
US 20090265190A1 · Ashley et al. · 2009 [cited by applicant]
US 20130173283A1 · Morse et al. · 2013 [cited by applicant]
US 20140172466A1 · Kemp et al. · 2014 [cited by applicant]
US 20150039351A1 · Bell et al. · 2015 [cited by applicant]
US 20150287143A1 · Gabriel et al. · 2015 [cited by applicant]
US 20150294420A1 · Hu · 2015 [cited by applicant]
US 20160048766A1 · McMahon · 2016 [cited by examiner]
US 20160078195A1 · Sarkar et al. · 2016 [cited by applicant]
US 20160196394A1 · Chanthasiriphan et al. · 2016 [cited by applicant]
US 20170124662A1 · Crabtree et al. · 2017 [cited by applicant]
US 20190180379A1 · Nayak · 2019 [cited by examiner]
US 20190180852A1 · Jiao · 2019 [cited by examiner]
US 20190220793A1 · Saarenvirta · 2019 [cited by applicant]
US 20190378210A1 · Merrill et al. · 2019 [cited by applicant]
US 20200020040A1 · Gokhale et al. · 2020 [cited by applicant]
US 20200098048A1 · Kuruvilla et al. · 2020 [cited by applicant]
US 20200104876A1 · Chintakindi et al. · 2020 [cited by applicant]
US 20200160998A1 · Ward et al. · 2020 [cited by applicant]
US 20220391670A1 · Dalli et al. · 2022 [cited by applicant]
WO WO2015084548A1 · 2015 [cited by applicant]
WO WO2017220140A1 · 2017 [cited by applicant]
Argys et al., Killer Debt: The Impact of Debt on Mortality, FRB Atlanta Working Paper No. 2016-14, (Nov. 1, 2016) (Year: 2016). [cited by examiner]
Aggour, K. S.; Bonissone, P. P.; Cheetham, W. E.; and Messmer, R. P. 2006. Automating the underwriting of insurance applications. AI magazine 27(3):36; Sep. 2006; 15 pages. [cited by applicant]
B. Letham, C. Rudin, T. H. McCormick, and D. Madigan; Interpretable classifers using rules and bayesian analysis: Building a better stroke prediction model. Annals of Applied Statistics; Apr. 2015; 22 pages. [cited by applicant]
Breiman, L. 2001. Random forests. Machine learning 45(1):5-32; Apr. 11, 2001; 11 pages. [cited by applicant]
Case, A., and Deaton, A. 2015. Rising morbidity and mortality in midlife among white non-hispanic Americans in the 21st century. Pmc. of the National Academy of Sciences 112(49): 15078-15083; Sep. 17, 2015; 6 pages. [cited by applicant]
Chen, T., and Guestrin, C. 2016. Xgboost: A scalable tree boosting system. In Proceedings of the Twenty-Second ACM SIGKDD In?ternational Conference on Knowledge Discovery and Data Mining, 785-794. ACM; Aug. 13, 2016; 10… [cited by applicant]
Chokshi, D. A.; El-Sayed, A. M.; and Stine, N. W. 2015. J-shaped curves and public health. JAMA 314(13):1339-1340; Oct. 6, 2015; 2 pages. [cited by applicant]
Consumer Financial Protection Bureau. Consumer credit reports: A study of medical and non-medical collec?tions, https://files.consumerfinance.gov/f/201412 cfpb reports consumer-credit-medical-and-non-medical- collection… [cited by applicant]
Cox, D. R. 1972. Regression models and life-tables regression. Journal of the Royal Statistical Society, Series B 34:187-220; Mar. 8, 1972; 34 pages. [cited by applicant]
Cox, H. J.; Bhandari, S.; Rigby, A. S.; and Kilpatrick, E. S. 2008. Mortality at low and high estimated glomerular filtration rate values: A u-shaped curve. Nephron Clinical Practice 110(2):c67-c72; Feb. 19, 2008; 6 pag… [cited by applicant]
D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-Francois Crespo, and Dan Dennison. Hidden technical debt in machine learning systems. In Advances… [cited by applicant]
David J Garrow. Toward a definitive history of [cited by applicant]
Goldwasser, P., and Feldman, J. 1997. Association of serum albumin and mortality risk. Journal of Clinical Epidemiology 50(6):693-703; Feb. 3, 1997; 11 pages. [cited by applicant]
Guizhou Hu, Mortality Assessment Technology: A New Tool for Life Insurance Underwriting, On The Risk, vol. 18, No. 3, https://pdfs.semanticscholar.org/bac0/3b8a85bf89c7a7b65076c082632d2d325519.pdf, 2002, 9 pages. [cited by applicant]
Hemant Ishwaran, Udaya B Kogalur, Eiran Z Gorodeski, Andy J Minn, and Michael S Lauer. High-dimensional variable selection for survival data. Journal of the American Statistical Association, 105(489):205-217, Nov. 2008,… [cited by applicant]
Hemant Ishwaran, Udaya B Kogalur, Eugene H Blackstone, and Michael S Lauer. Random Survival Forests. The Annals of Applied Statistics, 2(3):841-860; Mar. 2008, 22 pages. [cited by applicant]
John Karlen. climbeR: “Calculate Average Minimal Depth of a Maximal Subtree for ‘ranger’ Package Forests”; https://CRAN.R-project.org/package=climbeR. R package version 0.0.1, Nov. 19, 2016; 8 pages. [cited by applicant]
Kalben, B. B. 2000. Why men die younger: Causes of mortality differences by sex. N. Am. Actuarial Journal 4(4):83-111; Jan. 4, 2013; 30 pages. [cited by applicant]
Kaplan, E. L., and Meier, P. 1958. Nonparametric estimation from incomplete observations. Journal of the American Statistical Association 53(282):457-481; Jun. 1958; 25 pages. [cited by applicant]
Katzman, J.; Shaham, U.; Bates, J.; Cloninger, A.; Jiang, T.; and Kluger, Y. 2016. Deep survival: A deep cox proportional hazards network. arXiv preprint arXiv: 1606.00931; Jun. 2016; 11 pages. [cited by applicant]
Kronmal, R. A.; Cain, K. C.; Ye, Z.; and Omenn, G. S. 1993. Total serum cholesterol levels and mortality risk as a function of age: A report based on the Framingham data. Archives of Internal Medicine 153 (9):1065-1073;… [cited by applicant]
Lipton, Z. C. 2016. The mythos of model interpretability. In Proceedings of the ICML Workshop on Human Interpretability in Machine Learning, 96-100; Jun. 10, 2016; 5 pages. [cited by applicant]
Lundberg, Scott M, Gabriel G Erion, and Su-In Lee.;“Consistent Individualized Feature Attribution for Tree Ensembles.” arXiv Preprint arXiv:1802.03888.; Feb. 12, 2018; 9 pages. [cited by applicant]
Marc Maier, Hayley Carlotto, Freddie Sanchez, Sherriff Balogun, Sears Merritt; “Transforming Underwriting in the Life Insurance Industry”; Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence, vol.… [cited by applicant]
Mike Batty et al., Predictive Modeling for Life Insurance: Ways Life Insurance can participate in the Business Analytics Revolution, Deloitte Consulting LLP, Apr. 2010, 22 pages. [cited by applicant]
National Association of Insurance Commissioners. Credit-based insurance scores, 2018. https://www.naic.org/cipr topics/topic credit based insurance score.htm , last updated Dec. 7, 2018, 3 pages. [cited by applicant]
Peter WF Wilson, Ralph B D'Agostino, Daniel Levy, Albert M Belanger, Halit Silbershatz, and William B Kannel. Prediction of coronary heart disease using risk factor categories. Circulation, 97(18):1837-1847, May 1998; 1… [cited by applicant]
Phillip Janz et al., “The Impact On Relative Mortality and Prevalence from Triage in an Accelerated Underwriting Program”, Reinsurance News, Jul. 2019, 5 pages. [cited by applicant]
Ranganath, R.; Perotte, A.; Elhadad, N.; and Blei, D.; Deep survival analysis. arXiv preprint arXiv: 1608.02158; Aug. 6, 2016; 13 pages. [cited by applicant]
Ribeiro, M. T.; Singh, S.; and Guestrin, C.; Why should I trust you?: Explaining the predictions of any classifier. In Proceedings of the Tventy-Second ACM SIGKDD International Conference on Knowledge Discovery and Data… [cited by applicant]
Richard Wright, Mark Ellis, Steven R Holloway, and Sandy Wong. Patterns of racial diversity and segregation in the united states: 1990-2010. The Professional Geographer, 66(2):173-182, https://europepmc.org/articles/pmc… [cited by applicant]
Robert Chen et al., “Patient Stratification Using Electronic Health Records from a Chronic Disease Management Program”, IEEE J Biomed Health Inform., Author manuscript; available in PMC Jul. 4, 2017. [cited by applicant]
Ronen Avraham, Kyle D Logue, and Daniel Schwarcz. Understanding insurance antidiscrimination law. S. Cal. L. Rev., 87:195, Jan. 2014; 81 pages. [cited by applicant]
Rosinger, A.; Carroll, M. D.; Lacher, D.; and Ogden, C. 2017. Trends in total Cholesterol, Triglycerides, and Low-density Lipoprotein in US Adults, 1999-2014. JAMA Cardiology 2(3):339-341; Mar. 2017; 3 pages. [cited by applicant]
Scism, L. 2017. New York regulator seeks details from life insurers using algorithms to issue policies. The Wall Street Journal; Jun. 29, 2017; 2 pages. [cited by applicant]
Scott M Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Informatio… [cited by applicant]
Wright, M. N., and Ziegler, A. 2017. ranger: A fast implementation of random forests for high dimensional data in C++ and R. Journal of Statistical Software 77(1): Mar. 1-17, 2017; 17 pages. [cited by applicant]