IP Library Granted Patent US 12,626,783
Granted Patent B2
US 12,626,783 · App. 17/753,271 · Granted May 12, 2026

Computer-implemented method and apparatus for analysing genetic data

Inventors: Vincent Yann Marie Plagnol (Oxford, GB); Rachel Moore (Oxford, GB); Eva Maria Laura Krapohl (Oxford, GB); Christopher Charles Alan Spencer (Oxford, GB)
Assignee: GENOMICS LIMITED
G16B40/00C12Q1/6883G16B50/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,783
App. No.
17/753,271
Granted
May 12, 2026
Kind
B2
Abstract

The disclosure relates to analysing genetic data. In one arrangement, a method operates on input data comprising strengths of association between one or more phenotypes including a target phenotype and a plurality of genetic variants. A fine-mapping algorithm is applied to all or a subset of the input data to identify one or more independent phenotype-variant associations. A set of one or more fine-mapped variants is identified for each association. A fine-mapping predictive model is calculated on the basis of the input data and the set of fine-mapped variants. The effect on the target phenotype of the set of fine-mapped variants is subtracted from the input data to obtain residual association data. A machine learning algorithm is applied to the residual association data to identify further predictive correlations between the target phenotype and the plurality of genetic variants.

Claims (122)

1 . A computer-implemented method of analysing genetic data about an organism to obtain information about the organism, the method comprising:

receiving input data comprising strengths of association between one or more phenotypes including a target phenotype and a plurality of genetic variants in a region of interest of the genome of the organism;

applying a fine-mapping algorithm to all or a subset of the input data to identify one or more independent phenotype-variant associations within the region of interest, comprising identifying for each association a set of one or more fine-mapped variants from the plurality of genetic variants, and determining for each fine-mapped variant an estimated probability of being causal for the phenotype-variant association, the sum of the probabilities for the fine-mapped variants within the set adding to one;

generating, on the basis of the input data and the set of one or more fine-mapped variants, a fine-mapping predictive model quantifying an effect on the target phenotype of the set of one or more fine-mapped variants;

subtracting from the input data, using the fine-mapping predictive model, the effect on the target phenotype of the set of one or more fine-mapped variants to obtain residual association data, wherein the subtracting comprises subtracting a weighted sum of effect sizes from an estimated effect size of each of the plurality of genetic variants on the target phenotype to obtain a residual effect size for each of the plurality of genetic variants, and wherein the residual association data comprises the residual effect size for each of the plurality of genetic variants;

inputting, into a machine learning algorithm, at least the residual association data;

outputting, by the machine learning algorithm, predicted weight values for non-fine mapped variants, wherein the predicted weight values indicate a significance assigned to the non-fine mapped variants based on residual signals, while accounting for correlation between the non-fine mapped variants, wherein the non-fine mapped variants are variants included in the plurality of genetic variants but are not identified by the fine-mapping algorithm as the one or more fine-mapped variants, and wherein the outputting comprises iterating through multiple selections of variants from the plurality of genetic variants and, as the variants are selected, estimating the residual signal for each of the variants based on the residual association data;

generating a polygenic risk score model based on the fine-mapping predictive model and the predicted weight values for the non-fine mapped variants; and

applying the polygenic risk score model to genetic data from an individual to determine a polygenic risk score for the individual for the target phenotype.

2 . The method of claim 1 , wherein the strengths of association comprise an estimated effect size of each of the plurality of genetic variants on the target phenotype, and a standard error of each of the estimated effect sizes.

3 . The method of claim 1 , wherein the step of receiving input data comprises:

receiving individual level data comprising genotypes and corresponding phenotypes for each of a plurality of individuals; and

determining using the individual level data an estimated effect size of each of the plurality of genetic variants on the target phenotype and a standard error of each of the estimated effect sizes.

4 . The method of claim 1 , wherein the identifying of the set of one or more fine-mapped variants is performed using an iterative method, wherein each iteration comprises:

identifying, on the basis of the input data, a fine-mapped variant within the region of the genome different from any previously identified fine-mapped variant;

updating the input data to account for the effect on the target phenotype of the fine-mapped variants already identified, using a matrix of correlations between the genetic variants within the region of the genome; and

determining whether to perform a further iteration on the basis of the updated input data.

5 . The method of claim 1 , wherein the identifying of the set of one or more fine-mapped variants comprises using a plurality of instrument traits known to affect the target phenotype, the use of the instrument traits comprising:

determining an initial set of fine-mapped variants for the instrument traits; and

determining whether to include each fine-mapped variant of the initial set of fine-mapped variants for the instrument traits in the set of one or more fine-mapped variants for the target phenotype on the basis of a relationship between the plurality of instrument traits and the target phenotype.

6 . The method of claim 5 , wherein the generating of the fine-mapping predictive model comprises:

determining effect sizes on the one or more instrument traits of the initial set of fine-mapped variants for the one or more directly causal instrument traits, and

determining an effect size for the target phenotype of each fine-mapped variant of the initial set of fine-mapped variants for the instrument traits included in the set of one or more fine-mapped variants for the target phenotype on the basis of a predetermined relationship between effect sizes for the instrument traits and effect sizes for the target phenotype.

7 . The method of claim 1 , wherein the identifying of the set of one or more fine-mapped variants comprises identifying an initial set of fine-mapped variants for one or more directly causal instrument traits known to affect the target phenotype.

8 . The method of claim 1 , wherein:

the strengths of association comprise an estimated effect size of each of the plurality of genetic variants on the target phenotype, and a standard error of each of the estimated effect sizes; and

the fine-mapping predictive model comprises a fine-mapped effect size on the target phenotype for each of the fine-mapped variants, the fine-mapped effect size being calculated from the estimated effect size of the fine-mapped variants taking account of the estimated probability of the fine-mapped variants being causal for the phenotype-variant association.

9 . The method of claim 1 , wherein the effect on the target phenotype of the set of one or more fine-mapped variants is inferred using a machine learning algorithm.

10 . The method of claim 9 , wherein the set of one or more fine-mapped variants further comprises one or more variants known to have a high likelihood of being causal for the target phenotype.

11 . The method of claim 1 , wherein:

the strengths of association comprise an estimated effect size of each of the plurality of genetic variants on the target phenotype, and a standard error of each of the estimated effect sizes; and

the step of subtracting from the input data the effect on the target phenotype of the set of one or more fine-mapped variants comprises obtaining the residual effect size for each of a plurality of the genetic variants in the input data, the residual association data comprising the residual effect sizes,

wherein, after appropriate renormalisation of the effect sizes to ensure equal variance, the residual effect size {circumflex over (β)} i for genetic variant i is given by:

β

^

i

=

β

i

-

j

=

1

N

p

j

r

ij

β

~

j

where β i is the estimated marginal effect size of genetic variant i, N is the number of fine-mapped variants, p j is the probability that variant j is causal, {tilde over (β)}j is the fine-mapped effect size of the j th fine-mapped variant on the target phenotype, and r ij is a correlation between the j th fine-mapped variant and genetic variant i.

12 . The method of claim 1 , wherein the input data are derived from a plurality of different genetic studies, and the step of inputting into the machine learning algorithm comprises using a prior probability for each of the plurality of genetic variants of being causal for the target phenotype that is dependent on the consistency of the strength of association between each genetic variant and the target phenotype between the different genetic studies.

13 . The method of claim 1 , wherein the step of inputting into the machine learning algorithm comprises using a prior probability for each of the plurality of genetic variants of being causal for the target phenotype that is dependent on genomic annotations of the plurality of genetic variants in the region of interest.

14 . The method of claim 1 , wherein the step of applying the polygenic risk score model to genetic data from the individual further comprises applying the fine-mapping predictive model and the non-fine mapped variants identified by the machine learning algorithm.

15 . The method of claim 14 , wherein the polygenic risk score is given by the followed weighted sum:

PRS

=

l

=

1

L

α

l

x

l

where L is the number of variants that contribute to the PRS, each variant being included either in the fine-mapping predictive model or in the non-fine mapped variants from the machine learning algorithm, α l quantifies a strength of association of variant l on the target phenotype, the strength of association being specified by the fine-mapping predictive model or by the non-fine mapped variants from the machine learning algorithm, and x l is the genotype for variant l.

16 . The method of claim 14 , wherein the polygenic risk score for the individual is derived from a combination of a first partial polygenic risk score provided by applying the fine-mapping predictive model to genetic data from the individual and a second partial polygenic risk score provided by applying the non-fine mapped variants of the machine learning algorithm to the genetic data from the individual.

17 . The method of claim 1 , wherein the input data are derived from a plurality of different populations of the organism, and either or both of the following is satisfied:

the generating of the fine-mapping predictive model is performed separately for portions of the input data corresponding to different populations to obtain multiple respective population-matched fine-mapping predictive models; and

the inputting into the machine learning algorithm at least the residual association data is performed separately for portions of the input data corresponding to different populations to obtain multiple respective sets of population-matched further predictive correlations.

18 . The method of claim 17 , further comprising:

receiving input data from an individual having genes from a mixture of the different populations; and

generating a polygenic risk score for the individual by performing either or both of:

matching each of multiple population-matched fine-mapping predictive models to a corresponding portion of the input data that matches the population of the population-matched fine-mapping predictive model and applying each matched fine-mapping predictive model to the corresponding portion of the input data; and

matching each of multiple sets of population-matched further predictive correlations to a corresponding portion of the input data that matches the population of the set of population-matched further predictive correlations and applying each population-matched set of further predictive correlations to the corresponding portion of the input data.

19 . The method of claim 18 , wherein the matching of each of multiple sets of population-matched further predictive correlations is performed and the matching of each of multiple population-matched fine-mapping predictive models is not performed, the generation of the polygenic risk score further comprising applying a shared population-consistent fine-mapping predictive model to the input data from the individual.

20 . The method of claim 17 , further comprising:

receiving input data from an individual having genes predominantly from one of the different populations; and

generating a polygenic risk score for the individual by performing either or both of:

applying a population-matched fine-mapping predictive model to all of the input data from the individual, the population-matched fine-mapping predictive model being matched to the population of the individual; and

applying a set of population-matched further predictive correlations to all of the input data from the individual, the set of population-matched further predictive correlations being matched to the population of the individual.

21 . The method of claim 20 , wherein the applying of the set of population-matched further predictive correlations is performed and the applying of the population-matched fine-mapping predictive model is not performed, the calculation of the polygenic risk score comprising applying a shared population-consistent fine-mapping predictive model to the input data from the individual.

22 . The method of claim 1 , wherein the identifying of the set of one or more fine-mapped variants by the fine-mapping algorithm takes account of associations between the plurality of genetic variants and phenotypes other than the target phenotype.

23 . The method of claim 1 , wherein the organism is a human.

24 . An apparatus comprising:

one or more processors; and

one or more computer-readable media storing instructions which, when executed by the one or more processors, cause units of the apparatus to perform operations, the units comprising:

a receiving unit configured to receive input data comprising strengths of association between one or more phenotypes including a target phenotype and a plurality of genetic variants in a region of interest of the genome of the organism; and

a data processing unit configured to:

apply a fine-mapping algorithm to all or a subset of the input data to identify one or more independent phenotype-variant associations within the region of interest, by identifying for each association a set of one or more fine-mapped variants from the plurality of genetic variants, and determining for each fine-mapped variant an estimated probability of being causal for the phenotype-variant association, the sum of the probabilities for the fine-mapped variants within the set adding to one;

generate, on the basis of the input data and the set of one or more fine-mapped variants, a fine-mapping predictive model quantifying an effect on the target phenotype of the set of one or more fine-mapped variants;

subtract, using the fine-mapping predictive model, the effect on the target phenotype of the set of one or more fine-mapped variants to obtain residual association data, wherein the subtracting comprises subtracting a weighted sum of effect sizes from an estimated effect size of each of the plurality of genetic variants on the target phenotype to obtain a residual effect size for each of the plurality of genetic variants, and wherein the residual association data comprises the residual effect size for each of the plurality of genetic variants;

input, into a machine learning algorithm, at least the residual association data;

output, by the machine learning algorithm, predicted weight values for non-fine mapped variants, wherein the predicted weight values indicate a significance assigned to the non-fine mapped variants based on residual signals, while accounting for correlation between the non-fine mapped variants, wherein the non-fine mapped variants are variants included in the plurality of genetic variants but are not identified by the fine-mapping algorithm as the one or more fine-mapped variants, and wherein the outputting comprises iterating through multiple selections of variants from the plurality of genetic variants and, as the variants are selected, estimating the residual signal for each of the variants based on the residual association data;

generate a polygenic risk score model based on the fine-mapping predictive model and the predicted weight values for the non-fine mapped variants; and

apply the polygenic risk score model to genetic data from an individual to determine a polygenic risk score for the individual for the target phenotype.

25 . A system comprising:

one or more processors; and

one or more computer-readable media storing instructions which, when executed by the one or more processors, cause the system to perform operations comprising:

receiving input data comprising strengths of association between one or more phenotypes including a target phenotype and a plurality of genetic variants in a region of interest of the genome of the organism;

applying a fine-mapping algorithm to all or a subset of the input data to identify one or more independent phenotype-variant associations within the region of interest, comprising identifying for each association a set of one or more fine-mapped variants from the plurality of genetic variants, and determining for each fine-mapped variant an estimated probability of being causal for the phenotype-variant association, the sum of the probabilities for the fine-mapped variants within the set adding to one;

generating, on the basis of the input data and the set of one or more fine-mapped variants, a fine-mapping predictive model quantifying an effect on the target phenotype of the set of one or more fine-mapped variants;

subtracting from the input data, using the fine-mapping predictive model, the effect on the target phenotype of the set of one or more fine-mapped variants to obtain residual association data, wherein the subtracting comprises subtracting a weighted sum of effect sizes from an estimated effect size of each of the plurality of genetic variants on the target phenotype to obtain a residual effect size for each of the plurality of genetic variants, and wherein the residual association data comprises the residual effect size for each of the plurality of genetic variants;

inputting, into a machine learning algorithm, at least the residual association data;

outputting, by the machine learning algorithm, predicted weight values for non-fine mapped variants, wherein the predicted weight values indicate a significance assigned to the non-fine mapped variants based on residual signals, while accounting for correlation between the non-fine mapped variants, wherein the non-fine mapped variants are variants included in the plurality of genetic variants but are not identified by the fine-mapping algorithm as the one or more fine-mapped variants, and wherein the outputting comprises iterating through multiple selections of variants from the plurality of genetic variants and, as the variants are selected, estimating the residual signal for each of the variants based on the residual association data;

generating a polygenic risk score model based on the fine-mapping predictive model and the predicted weight values for the non-fine mapped variants; and

applying the polygenic risk score model to genetic data from an individual to determine a polygenic risk score for the individual for the target phenotype.

26 . One or more non-transitory computer-readable media storing instructions which, when executed by one or more processors, cause a system to perform operations comprising:

receiving input data comprising strengths of association between one or more phenotypes including a target phenotype and a plurality of genetic variants in a region of interest of the genome of the organism;

applying a fine-mapping algorithm to all or a subset of the input data to identify one or more independent phenotype-variant associations within the region of interest, comprising identifying for each association a set of one or more fine-mapped variants from the plurality of genetic variants, and determining for each fine-mapped variant an estimated probability of being causal for the phenotype-variant association, the sum of the probabilities for the fine-mapped variants within the set adding to one;

generating, on the basis of the input data and the set of one or more fine-mapped variants, a fine-mapping predictive model quantifying an effect on the target phenotype of the set of one or more fine-mapped variants;

subtracting from the input data, using the fine-mapping predictive model, the effect on the target phenotype of the set of one or more fine-mapped variants to obtain residual association data, wherein the subtracting comprises subtracting a weighted sum of effect sizes from an estimated effect size of each of the plurality of genetic variants on the target phenotype to obtain a residual effect size for each of the plurality of genetic variants, and wherein the residual association data comprises the residual effect size for each of the plurality of genetic variants;

inputting, into a machine learning algorithm, at least the residual association data;

outputting, by the machine learning algorithm, predicted weight values for non-fine mapped variants, wherein the predicted weight values indicate a significance assigned to the non-fine mapped variants based on residual signals, while accounting for correlation between the non-fine mapped variants, wherein the non-fine mapped variants are variants included in the plurality of genetic variants but are not identified by the fine-mapping algorithm as the one or more fine-mapped variants, and wherein the outputting comprises iterating through multiple selections of variants from the plurality of genetic variants and, as the variants are selected, estimating the residual signal for each of the variants based on the residual association data;

generating a polygenic risk score model based on the fine-mapping predictive model and the predicted weight values for the non-fine mapped variants; and

applying the polygenic risk score model to genetic data from an individual to determine a polygenic risk score for the individual for the target phenotype.

Assignments (3)
CHANGE OF NAME Recorded Dec 16, 2024
From: GENOMICS PLC
To: GENOMICS LIMITED
Reel/Frame 069712/0519 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2023
From: SPENCER, CHRISTOPHER CHARLES ALAN
To: GENOMICS PLC
Reel/Frame 065090/0664 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2022
From: PLAGNOL, VINCENT YANN MARIE; MOORE, RACHEL; KRAPOHL, EVA MARIA LAURA
To: GENOMICS PLC
Reel/Frame 059514/0774 →
Priority Claims (1)
GB 1912331 · Aug 28, 2019 · national
Continuity (1)
Related Publication 20220367009A1 · Nov 17, 2022
References Cited (35)
US 20090307180A1 · Colby · 2009 [cited by examiner]
US 20150066378A1 · Robison · 2015 [cited by examiner]
US 20160085909A1 · Reese · 2016 [cited by examiner]
US 20160092631A1 · Yandell · 2016 [cited by examiner]
US 20170286594A1 · Reid · 2017 [cited by examiner]
US 20190005192A1 · Kermani · 2019 [cited by examiner]
US 20190087534A1 · Zhang · 2019 [cited by examiner]
US 20190311785A1 · Torkamani · 2019 [cited by examiner]
CN 1809644A · 2006 [cited by applicant]
CN 109196590A · 2019 [cited by applicant]
CN 109637582A · 2019 [cited by applicant]
WO WO2015042496A1 · 2015 [cited by examiner]
WO WO2018051072A1 · 2018 [cited by examiner]
WO 2019166792 · 2019 [cited by applicant]
WO WO2021011990A1 · 2021 [cited by examiner]
Kichaev G, Yang WY, Lindstrom S, Hormozdiari F, Eskin E, Price AL, Kraft P, Pasaniuc B. Integrating functional data to prioritize causal variants in statistical fine-mapping studies. PLoS Genet. Oct. 30, 2014;10(10):e10… [cited by examiner]
Spain SL, Barrett JC. Strategies for fine-mapping complex traits. Hum Mol Genet. Oct. 15, 2015;24(R1):R111-9. doi: 10.1093/hmg/ddv260. Epub Jul. 8, 2015. PMID: 26157023; PMCID: PMC4572002. (Year: 2015). [cited by examiner]
Pepke S, Ver Steeg G. Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer. BMC Med Genomics. Mar. 15, 2017;10(1):12. doi: 10.1186/s12920-017-024… [cited by examiner]
Emmert-Streib F, Tripathi S, de Matos Simoes R. Harnessing the complexity of gene expression data from cancer: from single gene to structural pathway methods. Biol Direct. Dec. 10, 2012;7:44. doi: 10.1186/1745-6150-7-44… [cited by examiner]
Shepard SS, Meno S, Bahl J, Wilson MM, Barnes J, Neuhaus E. Viral deep sequencing needs an adaptive approach: IRMA, the iterative refinement meta-assembler. BMC Genomics. Sep. 5, 2016;17(1):708. doi: 10.1186/s12864-016-… [cited by examiner]
Spain SL, Barrett JC. Strategies for fine-mapping complex traits. Hum Mol Genet. Oct. 15, 2015;24(R1):R111-9. doi: 10.1093/hmg/ddv260. Epub Jul. 8, 2015. PMID: 26157023; PMCID: PMC4572002. (Year: 2015) (Year: 2015). [cited by examiner]
Chinese Application No. 202080061338.1, Office Action mailed on Apr. 21, 2025, 11 pages (6 pages of original document and 5 pages of English Translation). [cited by applicant]
Pan, Two-stage Design and Analysis for Genome-wide Association Studies, Medical and Health Technology, Oct. 15, 2012, 106 pages. [cited by applicant]
Benner et al., FineMap: Efficient Variable Selection Using Summary Data from Genome-Wide Association Studies, Bioinformatics, vol. 32, No. 10, Jan. 14, 2016, pp. 1493-1501. [cited by applicant]
Chatterjee et al., Developing and Evaluating Polygenic Risk Prediction Models for Stratified Disease Prevention, Nature Reviews Genetics, vol. 17, No. 7, Jul. 2016, pp. 392-406. [cited by applicant]
Choi et al., A Guide to Performing Polygenic Risk Score Analyses, Nature Protocols, Sep. 14, 2018, 22 pages. [cited by applicant]
Lambert et al., Towards Clinical Utility of Polygenic Risk Scores, Human Molecular Genetics, vol. 28, No. R2, Nov. 21, 2019, pp. R133-R142. [cited by applicant]
International Application No. PCT/GB2019/050525, International Preliminary Report on Patentability mailed on Sep. 3, 2020, 15 pages. [cited by applicant]
International Application No. PCT/GB2019/050525, International Search Report and Written Opinion mailed on Jun. 6, 2019, 18 pages. [cited by applicant]
International Application No. PCT/GB2020/052060, International Preliminary Report on Patentability mailed on Mar. 10, 2022, 9 pages. [cited by applicant]
International Application No. PCT/GB2020/052060, International Search Report and Written Opinion mailed on Nov. 3, 2020, 9 pages. [cited by applicant]
Spain et al., Strategies for Fine-Mapping Complex Traits, Human Molecular Genetics, vol. 24, No. R1, Jul. 8, 2015, pp. R111-R119. [cited by applicant]
Vilhjalmsson, Modeling Linkage Disequilibrium Increases Accuracy of Polygenic Risk Scores, The American Journal of Human Genetics, vol. 97, Oct. 1, 2015, pp. 576-592. [cited by applicant]
Yang et al., Conditional and Joint Multiple-SNP Analysis of GWAS Summary Statistics Identifies Additional Variants Influencing Complex Traits, Nature Genetics, vol. 44, No. 4, Mar. 18, 2012, pp. 369-375. [cited by applicant]
Zhu et al., Bayesian Large-Scale Multiple Regression with Summary Statistics from Genome-wide Association Studies, The Annals of Applied Statistics, vol. 11, No. 3, Available Online at: https://projecteuclid.org/euclid.… [cited by applicant]