IP Library › Granted Patent US 11,967,430
Granted Patent B2
US 11,967,430 · App. 16/863,403 · Granted Apr 23, 2024

Cross-variant polygenic predictive data analysis

Inventors: Kenneth Bryan (Dublin, IE); Megan O'Brien (Kildare, IE); David S. Monaghan (Dublin, IE); Chirag Chadha (Dublin, IE)
Assignee: Optum Services (Ireland) Limited
G16H50/30G06F17/18G06F18/211G06F18/22G06N20/00G16B40/00G16B20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,967,430
App. No.
16/863,403
Granted
Apr 23, 2024
Kind
B2
Abstract

There is a need for more effective and efficient predictive data analysis solutions for processing genetic sequencing data. This need can be addressed by, for example, techniques for performing predictive data analysis based on genetic sequences that utilize at least one of cross-variant polygenic risk modeling using genetic risk profiles, cross-variant polygenic risk modeling using functional genetic risk profiles, per-condition polygenic clustering operations, cross-condition polygenic predictive inferences, and cross-condition polygenic diagnoses.

Claims (69)

1. A computer-implemented method comprising:

generating, by one or more processors, a plurality of cross-variant polygenic risk data objects based at least in part on using a trained polygenic machine learning model trained based at least in part on a plurality of training genetic sequences comprising genetic sequences associated with a plurality of individuals;

identifying, by the one or more processors and from the plurality of cross-variant polygenic risk data objects, one or more primary cross-variant polygenic risk data objects for a primary medical condition;

identifying, by the one or more processors and from the plurality of cross-variant polygenic risk data objects, one or more secondary cross-variant polygenic risk data objects for a secondary medical condition;

generating, by the one or more processors, a cross-condition polygenic similarity measure between the primary medical condition and the secondary medical condition based at least in part on comparing the one or more primary cross-variant polygenic risk data objects and the one or more secondary cross-variant polygenic risk data objects; and

initiating, by the one or more processors, the performance of one or more prediction-based actions based at least in part on the cross-condition polygenic similarity measure.

2. The computer-implemented method of claim 1 , wherein:

the one or more primary cross-variant polygenic risk data objects comprise a primary genetic risk profile associated with a primary individual; and

the one or more secondary cross-variant polygenic risk data objects comprise a secondary genetic risk profile associated with a secondary individual.

3. The computer-implemented method of claim 2 , wherein:

the primary genetic risk profile comprises one or more primary per-variant genetic risk scores for a primary set of correlated genetic variants for the primary individual with respect to the primary medical condition in accordance with a primary chromosome-based grouping of the primary set of correlated genetic variants; and

the secondary genetic risk profile comprises one or more secondary per-variant genetic risk scores for a secondary set of correlated genetic variants for the secondary individual with respect to the secondary medical condition in accordance with a secondary chromosome-based grouping of the secondary set of correlated genetic variants.

4. The computer-implemented method of claim 1 , wherein:

the one or more primary cross-variant polygenic risk data objects comprise a primary functional genetic risk profile associated with a primary individual; and

the one or more secondary cross-variant polygenic risk data objects comprise a secondary functional genetic risk profile associated with a secondary individual.

5. The computer-implemented method of claim 4 , wherein:

the primary functional genetic risk profile comprises one or more primary per-variant genetic risk scores for a primary set of correlated genetic variants for the primary individual with respect to the primary medical condition in accordance with a primary functional-grouping-based grouping of the primary set of correlated genetic variants; and

the secondary functional genetic risk profile comprises one or more secondary per-variant genetic risk scores for a secondary set of correlated genetic variants for the secondary individual with respect to the secondary medical condition in accordance with a secondary functional-grouping-based grouping of the secondary set of correlated genetic variants.

6. The computer-implemented method of claim 1 , wherein generating cross-condition polygenic similarity measure comprises:

for each object pair of a plurality of object pairs that comprises a primary cross-variant polygenic risk data object of the one or more primary cross-variant polygenic risk data objects and a secondary cross-variant polygenic risk data object of the one or more secondary cross-variant polygenic risk data objects, determining a pairwise polygenic similarity measure of a plurality of pairwise polygenic similarity measures; and

generating the cross-condition polygenic similarity measure based at least in part on each pairwise polygenic similarity measure of the plurality of pairwise polygenic similarity measures.

7. The computer-implemented method of claim 6 , wherein determining the pairwise polygenic similarity measure for a particular object pair comprises:

determining an intersectional variant count of the primary cross-variant polygenic risk data object in the particular object pair and the secondary cross-variant polygenic risk data object in the particular object pair;

determining a maximal variant count of the primary cross-variant polygenic risk data object in the particular object pair and the secondary cross-variant polygenic risk data object in the particular object pair; and

determining the pairwise polygenic similarity measure based at least in part on the intersectional variant count and the maximal variant count.

8. The computer-implemented method of claim 6 , wherein determining the pairwise polygenic similarity measure for a particular object pair comprises:

determining an intersectional variant count of the primary cross-variant polygenic risk data object in the particular object pair and the secondary cross-variant polygenic risk data object in the particular object pair;

determining a minimal variant count of the primary cross-variant polygenic risk data object in the particular object pair and the secondary cross-variant polygenic risk data object in the particular object pair; and

determining the pairwise polygenic similarity measure based at least in part on the intersectional variant count and the minimal variant count.

9. The computer-implemented method of claim 6 , wherein determining the pairwise polygenic similarity measure for a particular object pair comprises:

determining an intersectional variant count of the primary cross-variant polygenic risk data object in the particular object pair and the secondary cross-variant polygenic risk data object in the particular object pair;

determining a union variant count of the primary cross-variant polygenic risk data object in the particular object pair and the secondary cross-variant polygenic risk data object in the particular object pair; and

determining the pairwise polygenic similarity measure based at least in part on the intersectional variant count and the union variant count.

10. The computer-implemented method of claim 6 , wherein determining the pairwise polygenic similarity measure for a particular object pair comprises:

identifying a plurality of genetic variants each associated with at least one of the primary cross-variant polygenic risk data object in the particular object pair and the secondary cross-variant polygenic risk data object in the particular object pair;

selecting a comparative variant subset of the plurality of genetic variants;

for each genetic variant in the comparative variant subset, determining a per-variant pairwise polygenic similarity measure based at least in part on comparing a primary per-variant genetic risk score for the genetic variant inferred from the primary cross-variant polygenic risk data object and a secondary per-variant genetic risk score for the genetic variant inferred from the secondary cross-variant polygenic risk data object; and

determining the pairwise polygenic similarity measure based at least in part on each per-variant pairwise polygenic similarity measure for a genetic variant in the comparative variant subset.

11. The computer-implemented method of claim 10 , wherein selecting the comparative variant subset comprises adopting the plurality of genetic variants as the comparative variant subset.

12. The computer-implemented method of claim 10 , wherein selecting the comparative variant subset comprises adopting an intersectional variant set of the primary cross-variant polygenic risk data object in the particular object pair and the secondary cross-variant polygenic risk data object in the particular object pair as the comparative variant subset.

13. The computer-implemented method of claim 10 , wherein selecting the comparative variant subset comprises adopting a symmetric difference variant set of the primary cross-variant polygenic risk data object in the particular object pair and the secondary cross-variant polygenic risk data object in the particular object pair as the comparative variant sub set.

14. The computer-implemented method of claim 1 , wherein initiating the performance of the one or more prediction-based actions comprises:

determining whether the cross-condition polygenic similarity measure exceeds a cross-condition polygenic similarity threshold, and

in response to determining that the cross-condition polygenic similarity measure exceeds the cross-condition polygenic similarity threshold, generating an inferred drug prescription profile of the secondary medical condition based at least in part on an existing drug prescription profile of the primary medical condition.

15. An apparatus comprising one or more processors and at least one memory including program code, the at least one memory and the program code configured to, with the one or more processors, cause the apparatus to at least:

generate a plurality of cross-variant polygenic risk data objects based at least in part on using a trained polygenic machine learning model trained based at least in part on a plurality of training genetic sequences comprising genetic sequences associated with a plurality of individuals;

identify, from the plurality of cross-variant polygenic risk data objects, one or more primary cross-variant polygenic risk data objects for a primary medical condition;

identify, from the plurality of cross-variant polygenic risk data objects, one or more secondary cross-variant polygenic risk data objects for a secondary medical condition;

generate a cross-condition polygenic similarity measure between the primary medical condition and the secondary medical condition based at least in part on comparing the one or more primary cross-variant polygenic risk data objects and the one or more secondary cross-variant polygenic risk data objects; and

initiate the performance of one or more prediction-based actions based at least in part on the cross-condition polygenic similarity measure.

16. The apparatus of claim 15 , wherein:

the one or more primary cross-variant polygenic risk data objects comprise a primary genetic risk profile associated with a primary individual; and

the one or more secondary cross-variant polygenic risk data objects comprise a secondary genetic risk profile associated with a secondary individual.

17. The apparatus of claim 16 , wherein:

the primary genetic risk profile comprises one or more primary per-variant genetic risk scores for a primary set of correlated genetic variants for the primary individual with respect to the primary medical condition in accordance with a primary chromosome-based grouping of the primary set of correlated genetic variants; and

the secondary genetic risk profile comprises one or more secondary per-variant genetic risk scores for a secondary set of correlated genetic variants for the secondary individual with respect to the secondary medical condition in accordance with a secondary chromosome-based grouping of the secondary set of correlated genetic variants.

18. The apparatus of claim 15 , wherein:

the one or more primary cross-variant polygenic risk data objects comprise a primary functional genetic risk profile associated with a primary individual; and

the one or more secondary cross-variant polygenic risk data objects comprise a secondary functional genetic risk profile associated with a secondary individual.

19. The apparatus of claim 18 , wherein:

the primary functional genetic risk profile comprises one or more primary per-variant genetic risk scores for a primary set of correlated genetic variants for the primary individual with respect to the primary medical condition in accordance with a primary functional-grouping-based grouping of the primary set of correlated genetic variants; and

the secondary functional genetic risk profile describes comprises one or more secondary per-variant genetic risk scores for a secondary set of correlated genetic variants for the secondary individual with respect to the secondary medical condition in accordance with a secondary functional-grouping-based grouping of the secondary set of correlated genetic variants.

20. At least one non-transitory computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions configured to:

generate a plurality of cross-variant polygenic risk data objects based at least in part on using a trained polygenic machine learning model trained based at least in part on a plurality of training genetic sequences comprising genetic sequences associated with a plurality of individuals;

identify, from the plurality of cross-variant polygenic risk data objects, one or more primary cross-variant polygenic risk data objects for a primary medical condition;

identify, from the plurality of cross-variant polygenic risk data objects, one or more secondary cross-variant polygenic risk data objects for a secondary medical condition;

generate a cross-condition polygenic similarity measure between the primary medical condition and the secondary medical condition based at least in part on comparing the one or more primary cross-variant polygenic risk data objects and the one or more secondary cross-variant polygenic risk data objects;

and

initiate the performance of one or more a prediction-based actions based at least in part on the cross-condition polygenic similarity measure.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2020
From: BRYAN, KENNETH; O'BRIEN, MEGAN; MONAGHAN, DAVID S.; CHADHA, CHIRAG
To: OPTUM SERVICES (IRELAND) LIMITED
Reel/Frame 052540/0847 →
Continuity (1)
Related Publication 20210343362A1 · Nov 4, 2021