IP Library › Granted Patent US 12,694,950
Granted Patent B2
US 12,694,950 · App. 17/804,416 · Granted Jul 28, 2026

Comparatively-refined polygenic risk score generation machine learning frameworks

Inventors: Michael Bridges (Dublin, IE); Paul J. Godden (London, GB)
Assignee: Optum Services (Ireland) Limited
G16B40/00G16B10/00G16B20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,694,950
App. No.
17/804,416
Granted
Jul 28, 2026
Kind
B2
Abstract

Various embodiments of the present invention describe techniques for generating a polygenic risk score generation machine learning framework that integrates an optimal genetic variant refinement model without requiring brute-force traversal of potential parameter spaces defined by various distinct genetic variant sets. In response, various embodiments of the present invention use holistic Bayesian sampling routines to efficiently generate Bayesian evidence numerical estimates for various genetic variant refinement models and select an optimal genetic variant refinement model accordingly. This enables enhancing the accuracy of polygenic risk score generation machine learning frameworks without resorting to computationally resource-intensive traversals of potential parameter spaces defined by various distinct genetic variant sets. In doing so, various embodiments of the present invention enhance the computational efficiency of generating a polygenic risk score generation machine learning framework that integrates an optimal genetic variant refinement model in contrast to computationally-inefficient techniques that require brute-force traversal of potential parameter spaces.

Claims (63)

1 . A computer-implemented method for generating a polygenic risk score for a target phenotype using a comparatively-refined polygenic risk score generation machine learning framework, the computer-implemented method comprising:

identifying, by one or more processors, the comparatively-refined polygenic risk score generation machine learning framework, wherein:

the comparatively-refined polygenic risk score generation machine learning framework comprises an optimal genetic variant refinement model that is selected from a plurality of defined genetic variant refinement models,

each defined genetic variant refinement model: (i) is associated with: (a) a distinct per-model genetic variant set of a group of genetic variants, and (b) a per-model parameter set comprising a per-model effect weight parameter set for the distinct per-model genetic variant set of the group of genetic variants that is associated with the corresponding defined genetic variant refinement model, and (ii) is configured to generate a per-model polygenic risk score based at least in part on a per-model input feature vector corresponding to the distinct per-model genetic variant set of the group of genetic variants for the corresponding defined genetic variant refinement model and the per-model parameter set for the corresponding defined genetic variant refinement model, and

generating the optimal genetic variant refinement model comprises: (i) for each defined genetic variant refinement model, sampling from a per-model posterior probability distribution for the corresponding defined genetic variant refinement model given target genome-wide association data for the target phenotype and by using a holistic Bayesian sampling routine that is configured to generate: (a) a per-model parameter numerical estimate set for the per-model parameter set that is associated with the corresponding defined genetic variant refinement model, and (b) a Bayesian evidence numerical estimate for the corresponding defined genetic variant refinement model, and (ii) selecting the optimal genetic variant refinement model as the corresponding defined

genetic variant refinement model with an optimal Bayesian evidence numerical estimate as generated by the holistic Bayesian sampling routine,

generating, by the one or more processors, the polygenic risk score based at least in part on the per-model polygenic risk score for the optimal genetic variant refinement model; and

performing, by the one or more processors, one or more prediction-based actions based at least in part on the polygenic risk score.

2 . The computer-implemented method of claim 1 , wherein the holistic Bayesian sampling routine comprises a nested sampling sub-routine.

3 . The computer-implemented method of claim 1 , wherein the holistic Bayesian sampling routine comprises a dynamic nested sampling sub-routine.

4 . The computer-implemented method of claim 1 , wherein:

the holistic Bayesian sampling routine comprises a nested sampling sub-routine and a dynamic nested sampling sub-routine, and

the Bayesian evidence numerical estimate for a particular defined genetic variant refinement model is generated based at least in part on a first Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the nested sampling sub-routine and a second Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the dynamic nested sampling sub-routine.

5 . The computer-implemented method of claim 4 , wherein:

the Bayesian evidence numerical estimate for the particular defined genetic variant refinement model is generated based at least in part on a cross-estimate weighted combination of the first Bayesian evidence numerical estimate and the second Bayesian evidence numerical estimate, and

the cross-estimate weighted combination is generated based at least in part on a first historical model performance quality weight for the nested sampling sub-routine and a second historical model performance quality weight for the dynamic nested sampling sub-routine.

6 . The computer-implemented method of claim 1 , wherein:

the comparatively-refined polygenic risk score generation machine learning framework further comprises a cross-model refinement model that is configured to generate a cross-model weighted combination of each per-model polygenic risk score for the plurality of defined genetic variant refinement models,

the cross-model weighted combination is generated based at least in part on a plurality of probabilistic model quality weights for the plurality of defined genetic variant refinement models, and

each probabilistic model quality weight for a respective defined genetic variant refinement model is generated based at least in part on the Bayesian evidence numerical estimate for the respective defined genetic variant refinement model as generated by the holistic Bayesian sampling routine.

7 . The computer-implemented method of claim 6 , wherein generating the polygenic risk score comprises:

adopting the cross-model weighted combination as the polygenic risk score.

8 . A system for generating a polygenic risk score for a target phenotype using a comparatively-refined polygenic risk score generation machine learning framework, the system comprising one or more processors and one or more non-transitory computer readable media storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

identifying the comparatively-refined polygenic risk score generation machine learning framework, wherein:

the comparatively-refined polygenic risk score generation machine learning framework comprises an optimal genetic variant refinement model that is selected from a plurality of defined genetic variant refinement models,

each defined genetic variant refinement model: (i) is associated with: (a) a distinct per-model genetic variant set of a group of genetic variants, and (b) a per-model parameter set comprising a per-model effect weight parameter set for the distinct per-model genetic variant set of the group of genetic variants that is associated with the corresponding defined genetic variant refinement model, and (ii) is configured to generate a per-model polygenic risk score based at least in part on a per-model input feature vector corresponding to the distinct per-model genetic variant set of the group of genetic variants for the corresponding defined genetic variant refinement model and the per-model parameter set for the corresponding defined genetic variant refinement model, and

generating the optimal genetic variant refinement model comprises: (i) for each defined genetic variant refinement model, sampling from a per-model posterior probability distribution for the corresponding defined genetic variant refinement model given target genome-wide association data for the target phenotype and by using a holistic Bayesian sampling routine that is configured to generate: (a) a per-model parameter numerical estimate set for the per-model parameter set that is associated with the corresponding defined genetic variant refinement model, and (b) a Bayesian evidence numerical estimate for the corresponding defined genetic variant refinement model, and (ii) selecting the optimal genetic variant refinement model as the corresponding defined

genetic variant refinement model with an optimal Bayesian evidence numerical estimate as generated by the holistic Bayesian sampling routine,

generating the polygenic risk score based at least in part on the per-model polygenic risk score for the optimal genetic variant refinement model; and

performing one or more prediction-based actions based at least in part on the polygenic risk score.

9 . The system of claim 8 , wherein the holistic Bayesian sampling routine comprises a nested sampling sub-routine.

10 . The system of claim 8 , wherein the holistic Bayesian sampling routine comprises a dynamic nested sampling sub-routine.

11 . The system of claim 8 , wherein:

the holistic Bayesian sampling routine comprises a nested sampling sub-routine and a dynamic nested sampling sub-routine, and

the Bayesian evidence numerical estimate for a particular defined genetic variant refinement model is generated based at least in part on a first Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the nested sampling sub-routine and a second Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the dynamic nested sampling sub-routine.

12 . The system of claim 11 , wherein:

the Bayesian evidence numerical estimate for the particular defined genetic variant refinement model is generated based at least in part on a cross-estimate weighted combination of the first Bayesian evidence numerical estimate and the second Bayesian evidence numerical estimate, and

the cross-estimate weighted combination is generated based at least in part on a first historical model performance quality weight for the nested sampling sub-routine and a second historical model performance quality weight for the dynamic nested sampling sub-routine.

13 . The system of claim 8 , wherein:

the comparatively-refined polygenic risk score generation machine learning framework further comprises a cross-model refinement model that is configured to generate a cross-model weighted combination of each per-model polygenic risk score for the plurality of defined genetic variant refinement models,

the cross-model weighted combination is generated based at least in part on a plurality of probabilistic model quality weights for the plurality of defined genetic variant refinement models, and

each probabilistic model quality weight for a respective defined genetic variant refinement model is generated based at least in part on the Bayesian evidence numerical estimate for the respective defined genetic variant refinement model as generated by the holistic Bayesian sampling routine.

14 . The system of claim 13 , wherein generating the polygenic risk score comprises:

adopting the cross-model weighted combination as the polygenic risk score.

15 . One or more non-transitory computer-readable storage media for generating a polygenic risk score for a target phenotype using a comparatively-refined polygenic risk score generation machine learning framework storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

identifying the comparatively-refined polygenic risk score generation machine learning framework, wherein:

the comparatively-refined polygenic risk score generation machine learning framework comprises an optimal genetic variant refinement model that is selected from a plurality of defined genetic variant refinement models,

each defined genetic variant refinement model: (i) is associated with: (a) a distinct per-model genetic variant set of a group of genetic variants, and (b) a per-model parameter set comprising a per-model effect weight parameter set for the distinct per-model genetic variant set of the group of genetic variants that is associated with the corresponding defined genetic variant refinement model, and (ii) is configured to generate a per-model polygenic risk score based at least in part on a per-model input feature vector corresponding to the distinct per-model genetic variant set of the group of genetic variants for the corresponding defined genetic variant refinement model and the per-model parameter set for the corresponding defined genetic variant refinement model, and

generating the optimal genetic variant refinement model comprises: (i) for each defined genetic variant refinement model, sampling from a per-model posterior probability distribution for the corresponding defined genetic variant refinement model given target genome-wide association data for the target phenotype and by using a holistic Bayesian sampling routine that is configured to generate: (a) a per-model parameter numerical estimate set for the per-model parameter set that is associated with the corresponding defined genetic variant refinement model, and (b) a Bayesian evidence numerical estimate for the corresponding defined genetic variant refinement model, and (ii) selecting the optimal genetic variant refinement model as the corresponding defined genetic variant refinement model with an optimal Bayesian evidence numerical estimate as generated by the holistic Bayesian sampling routine,

generating the polygenic risk score based at least in part on the per-model polygenic risk score for the optimal genetic variant refinement model; and

performing one or more prediction-based actions based at least in part on the polygenic risk score.

16 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the holistic Bayesian sampling routine comprises a nested sampling sub-routine.

17 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the holistic Bayesian sampling routine comprises a dynamic nested sampling sub-routine.

18 . The one or more non-transitory computer-readable storage media of claim 15 , wherein:

the holistic Bayesian sampling routine comprises a nested sampling sub-routine and a dynamic nested sampling sub-routine, and

the Bayesian evidence numerical estimate for a particular defined genetic variant refinement model is generated based at least in part on a first Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the nested sampling sub-routine and a second Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the dynamic nested sampling sub-routine.

19 . The one or more non-transitory computer-readable storage media of claim 18 , wherein:

the Bayesian evidence numerical estimate for the particular defined genetic variant refinement model is generated based at least in part on a cross-estimate weighted combination of the first Bayesian evidence numerical estimate and the second Bayesian evidence numerical estimate, and

the cross-estimate weighted combination is generated based at least in part on a first historical model performance quality weight for the nested sampling routine and a second historical model performance quality weight for the dynamic nested sampling sub-routine.

20 . The one or more non-transitory computer-readable storage media of claim 15 , wherein:

the comparatively-refined polygenic risk score generation machine learning framework further comprises a cross-model refinement model that is configured to generate a cross-model weighted combination of each per-model polygenic risk score for the plurality of defined genetic variant refinement models,

the cross-model weighted combination is generated based at least in part on a plurality of probabilistic model quality weights for the plurality of defined genetic variant refinement models, and

each probabilistic model quality weight for a respective defined genetic variant refinement model is generated based at least in part on the Bayesian evidence numerical estimate for the respective defined genetic variant refinement model as generated by the holistic Bayesian sampling routine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2022
From: BRIDGES, MICHAEL; GODDEN, PAUL J.
To: OPTUM SERVICES (IRELAND) LIMITED
Reel/Frame 060040/0145 →
Continuity (2)
Provisional Application 63202148 · May 28, 2021
Related Publication 20220383982A1 · Dec 1, 2022
References Cited (36)
US 9598733B2 · Gharvari et al. · 2017 [cited by applicant]
US 11985930B2 · Cooper · 2024 [cited by examiner]
US 20140032122A1 · Bader · 2014 [cited by examiner]
US 20190345566A1 · Khera · 2019 [cited by examiner]
US 20200118647A1 · Zhang · 2020 [cited by examiner]
US 20210113536A1 · Natarajan et al. · 2021 [cited by applicant]
US 20210118571A1 · Hsu et al. · 2021 [cited by applicant]
US 20230154618A1 · Enderling · 2023 [cited by examiner]
US 20230391875A1 · Chandler · 2023 [cited by examiner]
US 20240105280A1 · Moore · 2024 [cited by examiner]
US 20250266129A1 · Polcari · 2025 [cited by examiner]
WO 2006099142A2 · 2006 [cited by applicant]
WO 2020172432A1 · 2020 [cited by applicant]
WO 2020264466A1 · 2020 [cited by applicant]
WO WO2021011990A1 · 2021 [cited by examiner]
WO 2021038234A1 · 2021 [cited by applicant]
WO WO2022084554A1 · 2022 [cited by examiner]
Skilling, John. “Nested sampling for general Bayesian computation.” (2006): 833-859. (Year: 2006). [cited by examiner]
Johnson, Rob, Paul Kirk, and Michael PH Stumpf. “SYSBIONS: nested sampling for systems biology.” Bioinformatics 31.4 (2014): 604-605. (Year: 2014). [cited by examiner]
Thijssen, Bram, et al. “BCM: toolkit for Bayesian analysis of computational models using samplers.” BMC Systems Biology 10.1 (2016): 100. (Year: 2016). [cited by examiner]
Kochen, Michael Allen. Mechanistic Hypothesis Exploration of Signaling Network Processes via Bayesian Inference Methods. Diss. Vanderbilt University, 2020. (Year: 2020). [cited by examiner]
Mikelson, Jan, and Mustafa Khammash. “Likelihood-free nested sampling for parameter inference of biochemical reaction networks.” PLoS computational biology 16.10 (2020): e1008264. (Year: 2020). [cited by examiner]
Speagle, Joshua S. “dynesty: a dynamic nested sampling package for estimating Bayesian posteriors and evidences.” Monthly Notices of the Royal Astronomical Society 493.3 (2020): 3132-3158. (Year: 2020). [cited by examiner]
Pasetto, Stefano, Robert A. Gatenby, and Heiko Enderling. “Bayesian framework to augment tumor board decision making.” JCO Clinical Cancer Informatics 5 (2021): 508-517. (Year: 2021). [cited by examiner]
Song, Shuang, Lin Hou, and Jun S. Liu. “A data-adaptive Bayesian regression approach for polygenic risk prediction.” Bioinformatics 38.7 (2022): 1938-1946. (Year: 2022). [cited by examiner]
Gabbutt, Calum, et al. “Reconstruction of Contemporary Human Stem Cell Dynamics with Oscillatory Molecular Clocks.” bioRxiv (2021): Mar. 2021. (Year: 2021). [cited by examiner]
Choi, Shing Wan et al. “A Guide to Performing Polygenic Risk Score Analyses,” Nature Protocols, vol. 15, No. 9, pp. 2759-2772, Sep. 2020 (ePub: Jul. 24, 2020), DOI: 10.1038/s41596-020-0353-1. [cited by applicant]
Feroz, F. et al. “Multinest: An Efficient and Robust Bayesian Inference Tool for Cosmology and Particle Physics,” Monthly Notices of the Royal Astronomical Society, vol. 398, Issue 4, pp. 1601-1614, Oct. 2009 (ePub: Sep… [cited by applicant]
Handley, W.J. et al. “Polychord: Next-Generation Nested Sampling,” Monthly Notices of the Royal Astronomical Society, vol. 453, Issue 4, pp. 4384-4398, Nov. 11, 2015, (ePub: Sep. 16, 2015), DOI: 10.1093/mnras/stv1911. [cited by applicant]
Higson, Edward et al. “Dynamic Nested Sampling: An Improved Algorithm for Parameter Estimation and Evidence Calculation,” Statistics and Computing, vol. 29, Issue 5, pp. 891-913, Sep. 2019, (ePub: Dec. 3, 2018), DOI: 10… [cited by applicant]
Mavaddat, Nasim et al. “Polygenic Risk Scores for Prediction of Breast Cancer and Breast Cancer Subtypes,” The American Journal of Human Genetics, vol. 104, Issue 1, pp. 21-34, Jan. 3, 2019, DOI: 10.1016/j.ajhg.2018.11.… [cited by applicant]
Skilling, John. “Nested Sampling for General Bayesian Computation,” Bayesian Analysis vol. 1, No. 4, pp. 833-860, Dec. 2006. [cited by applicant]
So, Hon-Cheong et al. “Improving Polygenic Risk Prediction From Summary Statistics by an Empirical Bayes Approach, ” Scientific Reports, vol. 7, No. 41262, pp. 1-11, Feb. 1, 2017, DOI: 10.1038/srep41262. [cited by applicant]
Thibodeau, Eric L. “Child Maltreatment, Adaptive Functioning, and Polygenic Risk: A Structural Equation Mixture Model,” Developyment and Psychopathology, vol. 31, No. 2, pp. 433-456, May 2019, DOI: 10.1017/S095479419000… [cited by applicant]
Vilhjalmsson, Bjarni J. et al. “Modeling Linkage Disequilibrium Increases Accuracy of Polygenic Risk Scores,” The American Journal of Human Genetics, vol. 97, No. 4, Oct. 1, 2015, pp. 576-592, DOI: 10.1016/j.ajhg.2015.0… [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/031352; dated Aug. 11, 2022, (12 pages), European Patent Office, Rijswijk, Netherlands. [cited by applicant]