Comparatively-refined polygenic risk score generation machine learning frameworks
Various embodiments of the present invention describe techniques for generating a polygenic risk score generation machine learning framework that integrates an optimal genetic variant refinement model without requiring brute-force traversal of potential parameter spaces defined by various distinct genetic variant sets. In response, various embodiments of the present invention use holistic Bayesian sampling routines to efficiently generate Bayesian evidence numerical estimates for various genetic variant refinement models and select an optimal genetic variant refinement model accordingly. This enables enhancing the accuracy of polygenic risk score generation machine learning frameworks without resorting to computationally resource-intensive traversals of potential parameter spaces defined by various distinct genetic variant sets. In doing so, various embodiments of the present invention enhance the computational efficiency of generating a polygenic risk score generation machine learning framework that integrates an optimal genetic variant refinement model in contrast to computationally-inefficient techniques that require brute-force traversal of potential parameter spaces.
1 . A computer-implemented method for generating a polygenic risk score for a target phenotype using a comparatively-refined polygenic risk score generation machine learning framework, the computer-implemented method comprising:
identifying, by one or more processors, the comparatively-refined polygenic risk score generation machine learning framework, wherein:
the comparatively-refined polygenic risk score generation machine learning framework comprises an optimal genetic variant refinement model that is selected from a plurality of defined genetic variant refinement models,
each defined genetic variant refinement model: (i) is associated with: (a) a distinct per-model genetic variant set of a group of genetic variants, and (b) a per-model parameter set comprising a per-model effect weight parameter set for the distinct per-model genetic variant set of the group of genetic variants that is associated with the corresponding defined genetic variant refinement model, and (ii) is configured to generate a per-model polygenic risk score based at least in part on a per-model input feature vector corresponding to the distinct per-model genetic variant set of the group of genetic variants for the corresponding defined genetic variant refinement model and the per-model parameter set for the corresponding defined genetic variant refinement model, and
generating the optimal genetic variant refinement model comprises: (i) for each defined genetic variant refinement model, sampling from a per-model posterior probability distribution for the corresponding defined genetic variant refinement model given target genome-wide association data for the target phenotype and by using a holistic Bayesian sampling routine that is configured to generate: (a) a per-model parameter numerical estimate set for the per-model parameter set that is associated with the corresponding defined genetic variant refinement model, and (b) a Bayesian evidence numerical estimate for the corresponding defined genetic variant refinement model, and (ii) selecting the optimal genetic variant refinement model as the corresponding defined
genetic variant refinement model with an optimal Bayesian evidence numerical estimate as generated by the holistic Bayesian sampling routine,
generating, by the one or more processors, the polygenic risk score based at least in part on the per-model polygenic risk score for the optimal genetic variant refinement model; and
performing, by the one or more processors, one or more prediction-based actions based at least in part on the polygenic risk score.
2 . The computer-implemented method of claim 1 , wherein the holistic Bayesian sampling routine comprises a nested sampling sub-routine.
3 . The computer-implemented method of claim 1 , wherein the holistic Bayesian sampling routine comprises a dynamic nested sampling sub-routine.
4 . The computer-implemented method of claim 1 , wherein:
the holistic Bayesian sampling routine comprises a nested sampling sub-routine and a dynamic nested sampling sub-routine, and
the Bayesian evidence numerical estimate for a particular defined genetic variant refinement model is generated based at least in part on a first Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the nested sampling sub-routine and a second Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the dynamic nested sampling sub-routine.
5 . The computer-implemented method of claim 4 , wherein:
the Bayesian evidence numerical estimate for the particular defined genetic variant refinement model is generated based at least in part on a cross-estimate weighted combination of the first Bayesian evidence numerical estimate and the second Bayesian evidence numerical estimate, and
the cross-estimate weighted combination is generated based at least in part on a first historical model performance quality weight for the nested sampling sub-routine and a second historical model performance quality weight for the dynamic nested sampling sub-routine.
6 . The computer-implemented method of claim 1 , wherein:
the comparatively-refined polygenic risk score generation machine learning framework further comprises a cross-model refinement model that is configured to generate a cross-model weighted combination of each per-model polygenic risk score for the plurality of defined genetic variant refinement models,
the cross-model weighted combination is generated based at least in part on a plurality of probabilistic model quality weights for the plurality of defined genetic variant refinement models, and
each probabilistic model quality weight for a respective defined genetic variant refinement model is generated based at least in part on the Bayesian evidence numerical estimate for the respective defined genetic variant refinement model as generated by the holistic Bayesian sampling routine.
7 . The computer-implemented method of claim 6 , wherein generating the polygenic risk score comprises:
adopting the cross-model weighted combination as the polygenic risk score.
8 . A system for generating a polygenic risk score for a target phenotype using a comparatively-refined polygenic risk score generation machine learning framework, the system comprising one or more processors and one or more non-transitory computer readable media storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:
identifying the comparatively-refined polygenic risk score generation machine learning framework, wherein:
the comparatively-refined polygenic risk score generation machine learning framework comprises an optimal genetic variant refinement model that is selected from a plurality of defined genetic variant refinement models,
each defined genetic variant refinement model: (i) is associated with: (a) a distinct per-model genetic variant set of a group of genetic variants, and (b) a per-model parameter set comprising a per-model effect weight parameter set for the distinct per-model genetic variant set of the group of genetic variants that is associated with the corresponding defined genetic variant refinement model, and (ii) is configured to generate a per-model polygenic risk score based at least in part on a per-model input feature vector corresponding to the distinct per-model genetic variant set of the group of genetic variants for the corresponding defined genetic variant refinement model and the per-model parameter set for the corresponding defined genetic variant refinement model, and
generating the optimal genetic variant refinement model comprises: (i) for each defined genetic variant refinement model, sampling from a per-model posterior probability distribution for the corresponding defined genetic variant refinement model given target genome-wide association data for the target phenotype and by using a holistic Bayesian sampling routine that is configured to generate: (a) a per-model parameter numerical estimate set for the per-model parameter set that is associated with the corresponding defined genetic variant refinement model, and (b) a Bayesian evidence numerical estimate for the corresponding defined genetic variant refinement model, and (ii) selecting the optimal genetic variant refinement model as the corresponding defined
genetic variant refinement model with an optimal Bayesian evidence numerical estimate as generated by the holistic Bayesian sampling routine,
generating the polygenic risk score based at least in part on the per-model polygenic risk score for the optimal genetic variant refinement model; and
performing one or more prediction-based actions based at least in part on the polygenic risk score.
9 . The system of claim 8 , wherein the holistic Bayesian sampling routine comprises a nested sampling sub-routine.
10 . The system of claim 8 , wherein the holistic Bayesian sampling routine comprises a dynamic nested sampling sub-routine.
11 . The system of claim 8 , wherein:
the holistic Bayesian sampling routine comprises a nested sampling sub-routine and a dynamic nested sampling sub-routine, and
the Bayesian evidence numerical estimate for a particular defined genetic variant refinement model is generated based at least in part on a first Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the nested sampling sub-routine and a second Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the dynamic nested sampling sub-routine.
12 . The system of claim 11 , wherein:
the Bayesian evidence numerical estimate for the particular defined genetic variant refinement model is generated based at least in part on a cross-estimate weighted combination of the first Bayesian evidence numerical estimate and the second Bayesian evidence numerical estimate, and
the cross-estimate weighted combination is generated based at least in part on a first historical model performance quality weight for the nested sampling sub-routine and a second historical model performance quality weight for the dynamic nested sampling sub-routine.
13 . The system of claim 8 , wherein:
the comparatively-refined polygenic risk score generation machine learning framework further comprises a cross-model refinement model that is configured to generate a cross-model weighted combination of each per-model polygenic risk score for the plurality of defined genetic variant refinement models,
the cross-model weighted combination is generated based at least in part on a plurality of probabilistic model quality weights for the plurality of defined genetic variant refinement models, and
each probabilistic model quality weight for a respective defined genetic variant refinement model is generated based at least in part on the Bayesian evidence numerical estimate for the respective defined genetic variant refinement model as generated by the holistic Bayesian sampling routine.
14 . The system of claim 13 , wherein generating the polygenic risk score comprises:
adopting the cross-model weighted combination as the polygenic risk score.
15 . One or more non-transitory computer-readable storage media for generating a polygenic risk score for a target phenotype using a comparatively-refined polygenic risk score generation machine learning framework storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:
identifying the comparatively-refined polygenic risk score generation machine learning framework, wherein:
the comparatively-refined polygenic risk score generation machine learning framework comprises an optimal genetic variant refinement model that is selected from a plurality of defined genetic variant refinement models,
each defined genetic variant refinement model: (i) is associated with: (a) a distinct per-model genetic variant set of a group of genetic variants, and (b) a per-model parameter set comprising a per-model effect weight parameter set for the distinct per-model genetic variant set of the group of genetic variants that is associated with the corresponding defined genetic variant refinement model, and (ii) is configured to generate a per-model polygenic risk score based at least in part on a per-model input feature vector corresponding to the distinct per-model genetic variant set of the group of genetic variants for the corresponding defined genetic variant refinement model and the per-model parameter set for the corresponding defined genetic variant refinement model, and
generating the optimal genetic variant refinement model comprises: (i) for each defined genetic variant refinement model, sampling from a per-model posterior probability distribution for the corresponding defined genetic variant refinement model given target genome-wide association data for the target phenotype and by using a holistic Bayesian sampling routine that is configured to generate: (a) a per-model parameter numerical estimate set for the per-model parameter set that is associated with the corresponding defined genetic variant refinement model, and (b) a Bayesian evidence numerical estimate for the corresponding defined genetic variant refinement model, and (ii) selecting the optimal genetic variant refinement model as the corresponding defined genetic variant refinement model with an optimal Bayesian evidence numerical estimate as generated by the holistic Bayesian sampling routine,
generating the polygenic risk score based at least in part on the per-model polygenic risk score for the optimal genetic variant refinement model; and
performing one or more prediction-based actions based at least in part on the polygenic risk score.
16 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the holistic Bayesian sampling routine comprises a nested sampling sub-routine.
17 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the holistic Bayesian sampling routine comprises a dynamic nested sampling sub-routine.
18 . The one or more non-transitory computer-readable storage media of claim 15 , wherein:
the holistic Bayesian sampling routine comprises a nested sampling sub-routine and a dynamic nested sampling sub-routine, and
the Bayesian evidence numerical estimate for a particular defined genetic variant refinement model is generated based at least in part on a first Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the nested sampling sub-routine and a second Bayesian evidence numerical estimate for the particular defined genetic variant refinement model as generated by the dynamic nested sampling sub-routine.
19 . The one or more non-transitory computer-readable storage media of claim 18 , wherein:
the Bayesian evidence numerical estimate for the particular defined genetic variant refinement model is generated based at least in part on a cross-estimate weighted combination of the first Bayesian evidence numerical estimate and the second Bayesian evidence numerical estimate, and
the cross-estimate weighted combination is generated based at least in part on a first historical model performance quality weight for the nested sampling routine and a second historical model performance quality weight for the dynamic nested sampling sub-routine.
20 . The one or more non-transitory computer-readable storage media of claim 15 , wherein:
the comparatively-refined polygenic risk score generation machine learning framework further comprises a cross-model refinement model that is configured to generate a cross-model weighted combination of each per-model polygenic risk score for the plurality of defined genetic variant refinement models,
the cross-model weighted combination is generated based at least in part on a plurality of probabilistic model quality weights for the plurality of defined genetic variant refinement models, and
each probabilistic model quality weight for a respective defined genetic variant refinement model is generated based at least in part on the Bayesian evidence numerical estimate for the respective defined genetic variant refinement model as generated by the holistic Bayesian sampling routine.