IP Library Granted Patent US 12,367,948
Granted Patent B2
US 12,367,948 · App. 15/733,547 · Granted Jul 22, 2025

Computer-implemented method of analysing genetic data about an organism

Inventors: Christopher Charles Alan Spencer (Oxford, GB); Gerard Anton Lunter (Oxford, GB); Peter James Donnelly (Oxford, GB); Vincent Yann Marie Plagnol (Oxford, GB)
Assignee: GENOMICS LIMITED
G16B20/40G06F18/295G06N3/002G16B5/20G16H20/13G16H70/40C12Q2600/112
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,948
App. No.
15/733,547
Granted
Jul 22, 2025
Kind
B2
Abstract

Methods are disclosed for analysing genetic data about an organism. In one arrangement, input units are derived from studies that provide information about the association between genetic variants and phenotypes. Input units are assigned to one of a plurality of clusters, based on an assessment of the extent to which input units share genetic variants that affect any aspect of the phenotype corresponding to each input unit or any of the underlying biological mechanisms of the phenotype, thereby identifying phenotypes that share underlying biological mechanisms.

Claims (90)

1. A computer-implemented method of analysing genetic data about an organism, comprising:

accessing a plurality of genome-wide association studies;

deriving, using the plurality of genome-wide association studies, at least 50 input units, wherein:

each input unit is a data structure derived from a genome-wide association study of the plurality of genome-wide association studies that provides summary statistic data including an inferred effect size of each of a plurality of genetic variants along a genome of the organism on a phenotype corresponding to the input unit and a standard error of the inferred effect size, and

the deriving comprises:

completing each genome-wide association study having missing data about an association between one or more genetic variants and the phenotype corresponding to the input unit by modeling the association between the one or more genetic variants and the phenotype corresponding to the input unit, and

determining a set of characteristics for each input unit based on the summary statistic data associated with the genome-wide association study the set of characteristics comprising probability metrics which quantify evidence for each of the plurality of genetic variants being causal for the phenotype corresponding to the input unit;

selecting a region or regions of the genome of the organism;

for each of the selected region or regions, assigning each of the input units to one or more of a plurality of clusters, the assigning being an iterative process for exploring a space of possible assignment using a Markov Chain Monte Carlo (MCMC) algorithm designed to effectively explore an exponentially large space of study assignments, the iterative process comprising:

determining a degree of similarity between i) the set of characteristics of the input unit, and ii) a set of characteristics of each of the plurality of clusters, wherein the set of characteristics of each cluster is either pre-determined or calculated by combining the sets of characteristics of input units already assigned to the cluster, and

either a) assigning the input unit to one or more existing clusters of the plurality of clusters with a probability dependent on the corresponding degree of similarity, or

b) creating a new cluster in the plurality of clusters and assigning the input unit to the new cluster with a probability dependent on the set of characteristics of the input unit and the sets of characteristics of the existing clusters of the plurality of clusters;

identifying that the phenotypes corresponding to the input units assigned to the same cluster share underlying biological mechanisms; and

displaying, for each of the selected region or regions, an analysis across all input units using a computer-based interface that summarizes a membership of each of the clusters and the assignment of each of the genetic variants to each cluster based on the assigning and the identifying.

2. The method of claim 1 , wherein the assignment of each of the input units to one or more of the plurality of clusters is based on assessing said assignment using a Bayesian probabilistic model.

3. The method of claim 1 , wherein the assignment of each of the input units to one or more of the plurality of clusters is based on an analysis of a path taken by the Markov chain.

4. The method of claim 1 , wherein one step of the Markov chain Monte Carlo algorithm comprises the assignment of one of the input units to one or more of the plurality of clusters with a probability dependent on the degree of similarity between the set of characteristics of the input unit and the set of characteristics of the cluster.

5. The method of claim 1 , wherein the plurality of clusters comprises a null cluster.

6. The method of claim 1 , wherein the assignment of each of the input units to one or more of the plurality of clusters further comprises using a prior distribution on a number and/or size of clusters.

7. The method of claim 6 , wherein the prior distribution on the number and/or size of clusters follows a Chinese Restaurant Process.

8. The method of claim 1 , wherein:

the sets of characteristics of the input units are determined by calculating Bayes factors of each of the plurality of genetic variants in the input unit;

the set of characteristics of each of the plurality of clusters is calculated by combining the sets of characteristics of input units already assigned to the cluster; and

combining the sets of characteristics of input units already assigned to the cluster comprises calculating a product of the Bayes factors of the input units assigned to the cluster.

9. The method of claim 1 , wherein calculating the set of characteristics of a cluster comprises using a prior distribution of the probabilities of each of the plurality of genetic variants being causal.

10. The method of claim 9 , wherein the prior distribution of the probabilities incorporates pre-existing information about variation in functionality relevant to causality over the genetic variants.

11. The method of claim 1 , wherein calculating the set of characteristics of a cluster further comprises taking account of correlations or other known relationships between the genome-wide association studies used to derive the input units.

12. The method of claim 1 , wherein the assignment of each of the input units to one or more of the plurality of clusters further comprises using a prior distribution on pairs or larger collections of input units being assigned to the same cluster.

13. The method of claim 1 , further comprising outputting a probability distribution over aspects of cluster membership, and characterising certainty around cluster assignment.

14. The method of claim 1 , further comprising iteratively repeating the assignment of the input units to one or more of the plurality of clusters until a predetermined convergence threshold is reached or a predetermined number of iterations has been performed.

15. The method of claim 1 , further comprising, for each of one or more of the plurality of clusters, identifying one or more lead genetic variants from the plurality of genetic variants that are causal for one or more phenotypes corresponding to the cluster.

16. The method of claim 15 , further comprising:

calculating a size of an effect of the one or more lead genetic variants on a target phenotype;

determining a genotype of an individual organism with respect to each of the one or more lead genetic variants; and

calculating a contribution of the genotype of the individual organism to the target phenotype of the individual organism on the basis of the sizes of the effect of the one or more lead genetic variants on the target phenotype.

17. The method of claim 16 , further comprising combining the contribution of the genotype of the individual organism to the target phenotype with other clinical and/or genetic data about the individual organism to improve a determination of the target phenotype of the individual organism.

18. The method of claim 15 , wherein two or more lead genetic variants are identified for each of one or more clusters, and identifying the two or more lead genetic variants comprises calculating a matrix of correlations between genetic variants.

19. The method of claim 15 , wherein two or more lead genetic variants are identified for each of one or more clusters, and the identifying of two or more lead genetic variants comprises:

calculating a set of characteristics of the cluster quantifying the probability of each of the genetic variants being causal for the cluster;

determining a first lead genetic variant for the cluster on the basis of the set of characteristics of the cluster;

updating the set of characteristics of the cluster to account for the effect of the first lead genetic variant; and

determining a second lead genetic variant for the cluster on the basis of the updated set of characteristics.

20. The method of claim 1 , wherein the plurality of genetic variants are chosen from genetic variants in the selected region of the genome, and the selected region comprises a predetermined number of base pairs and includes a gene of interest.

21. The method of claim 20 , wherein the selected region is chosen by minimising the correlation between genetic variants in the selected region, and genetic variants in one or more other regions of the genome.

22. The method of claim 20 , wherein the method is performed for each of a plurality of selected regions of the genome, thereby providing a plurality of sets of clusters containing assigned input units, each set of clusters corresponding to a different one of the selected regions, the method further comprising associating a first subset of one or more clusters from a first set with a second subset of one or more clusters from each of one or more other sets by assessing a similarity between phenotypes corresponding to input units in the first subset of clusters with phenotypes corresponding to input units in the one or more second subsets of clusters.

23. The method of claim 22 , further comprising identifying a group of phenotypes corresponding to input units which are assigned to the same cluster across the plurality of sets with a frequency above a predetermined frequency threshold.

24. The method of claim 1 , wherein one or more of the input units further comprise additional information about the association between one or more genetic variants in the selected region or regions of the genome and the phenotype corresponding to the input unit, and the additional information is obtained by modelling the association between the one or more genetic variants and the phenotype corresponding to the input unit.

25. The method of claim 1 , wherein the phenotypes corresponding to the input units include one or more of: a level of expression of a gene; regulation of expression of a gene; epigenetic characteristics; a level of abundance of a protein or peptide; the function and/or molecular structure of a protein or peptide; a quantity of a molecule in the organism; characteristics of biochemical and metabolic processes; cellular properties, including morphology and function; tissue properties, including morphology and function; organ and organ system properties, including morphology and function; any response to an external stimulus or stimuli; any response to exposure to a substance or pathogen; behavioural and lifestyle characteristics; reproductive and life course characteristics and function; the onset, trajectory, and prognosis of a disease or condition; a measurable anatomical characteristic; a measurable physiological or functional characteristic; and measurable psychological or cognitive characteristics.

26. The method of claim 1 , wherein the identification of phenotypes that share underlying biological mechanisms comprises identification of one or more of the following combinations of phenotypes:

i) an occurrence of a disease or condition and one or more of a level of expression of a gene, a level of expression of a protein or peptide, a quantity of a biological molecule in the organism;

ii) a measurable characteristic of the organism and one or more of a level of expression of a gene, a level of expression of a protein or peptide, a quantity of a biological molecule in the organism;

iii) an occurrence of a disease or condition and a measurable characteristic of the organism;

iv) a response to an input of a substance to the organism and one or more of an occurrence of a disease or condition, a measurable characteristic of the organism, a quantity of a biological molecule in the organism; and

v) a level of expression of a gene and one or more of a level of expression of a different gene, a level of expression of a protein or peptide, a quantity of a biological molecule in the organism.

27. The method of claim 1 , wherein the method comprises, on the basis of the identification of phenotypes that share underlying biological mechanisms, one or more of:

(i) determining clinical support strategies for the organism;

(ii) determining a probability of the organism developing a disease or condition;

(iii) stratifying the organism in a clinical or pre-clinical trial;

(iv) identifying a biological pathway in the organism; or

(v) where the phenotypes corresponding to the input units include a response of the organism to input of one or more new or existing drugs, one or more of:

(a) determining the mechanism of action of the one or more drugs;

(b) determining the clinical and/or adverse response of the organism to the one or more drugs;

(c) determining biomarkers of the one or more drugs; or

(d) determining to which subsets of patients the drug should or should not be given to achieve particular clinical or safety outcomes.

28. The method of claim 1 , wherein the phenotypes corresponding to the input units include the response of an organism to input of two or more new or existing drugs, and the method comprises, on the basis of the identification of phenotypes directly or indirectly causally associated with each other, one or more of:

(i) determining interactions between the two or more drugs; or

(ii) determining the effect on the organism of interactions between the two or more drugs.

29. The method of claim 1 , wherein further analysis of the sets of clusters obtained when the method is applied to different regions of the genome is used to identify biological features, properties, or mechanisms, which are shared between more than one of the sets of clusters in order to identify any of:

(i) intermediate molecular, cellular, or other phenotypes which occur on the path to a particular disease outcome or set of outcomes;

(ii) the use of phenotypes identified in (i) above as therapeutic targets;

(iii) the use of phenotypes identified in (i) or (ii) above as readouts in assays in drug development or as biomarkers in drug development;

(iv) the use of phenotypes identified in any of (i)-(iii) above to stratify patients in clinical trials; or

(v) the use of phenotypes identified in any of (i)-(iv) above to determine a plurality of subsets of patients to whom particular therapeutics or treatments or combinations of therapeutics or treatments should or should not be given in order to aim to achieve particular outcomes.

30. A computer-implemented method of analysing genetic data about an organism to provide treatment or prevention of a disease or condition, the method comprising:

accessing a plurality of genome-wide association studies;

receiving input data comprising a plurality of deriving, using the plurality of genome-wide association studies, at least 50 input units, wherein:

each input unit is a data structure derived from a genome-wide association study of the plurality of genome-wide association studies that provides summary statistic data including an inferred effect size of each of a plurality of genetic variants along a genome of the organism on a phenotype corresponding to the input unit and a standard error of the inferred effect size, and

the deriving comprises:

completing each genome-wide association study having missing data about an association between one or more genetic variants and the phenotype corresponding to the input unit by modeling the association between the one or more genetic variants and the phenotype corresponding to the input unit, and

determining a set of characteristics for each input unit based on the summary statistic data associated with the genome-wide association study, the set of characteristics comprising probability metrics which quantify evidence for each of the plurality of genetic variants being causal for the phenotype corresponding to the input unit;

selecting a region or regions of the genome of the organism;

for each of the selected region or regions, assigning each of the input units to one or more of a plurality of clusters, the assigning being an iterative process for exploring a space of possible assignment using a Markov Chain Monte Carlo (MCMC) algorithm designed to effectively explore an exponentially large space of study assignments, the iterative process comprising:

determining a degree of similarity between i) the set of characteristics of the input unit, and ii) a set of characteristics of each of the plurality of clusters, wherein the set of characteristics of each cluster is either pre-determined or calculated by combining the sets of characteristics of input units already assigned to the cluster; and

either a) assigning the input unit to one or more existing clusters of the plurality of clusters with a probability dependent on the corresponding degree of similarity; or

b) creating a new cluster in the plurality of clusters and assigning the input unit to the new cluster with a probability dependent on the set of characteristics of the input unit and the sets of characteristics of the existing clusters of the plurality of clusters;

identifying that the phenotypes corresponding to input units assigned to the same cluster share underlying biological mechanisms; and

on the basis of the identification of phenotypes that share underlying biological mechanisms, one or more of:

(i) outputting clinical support strategies for the organism;

(ii) determining outputting a probability of the organism developing a disease or condition; and

(iii) where the phenotypes corresponding to the input units include the response of the organism to input of one or more new or existing drugs, outputting the clinical and/or adverse response of the organism to the one or more drugs.

Assignments (2)
CHANGE OF NAME Recorded Dec 16, 2024
From: GENOMICS PLC
To: GENOMICS LIMITED
Reel/Frame 069712/0519 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2021
From: SPENCER, CHRISTOPHER CHARLES ALAN; LUNTER, GERARD ANTON; DONNELLY, PETER JAMES; PLAGNOL, VINCENT YANN MARIE
To: GENOMICS PLC
Reel/Frame 057215/0901 →
Priority Claims (1)
GB 1803202.9 · Feb 27, 2018 · national
Continuity (1)
Related Publication 20200402614A1 · Dec 24, 2020
References Cited (33)
US 20140032122A1 · Bader · 2014 [cited by examiner]
CN 107301330A · 2017 [cited by applicant]
Aldous, Exchangeability and Related Topics, vol. 1117, Nov. 15, 2006, 98 pages. [cited by applicant]
Band et al., Imputation-Based Meta-Analysis of Severe Malaria in Three African Populations, PLOS Genetics, vol. 9, No. 5, May 2013, 13 pages. [cited by applicant]
Bellenguez et al., Genome-wide Association Study Identifies a Variant in HDAC9 Associated with Large Vessel Ischemic Stroke, Nature Genetics, vol. 44, No. 3, Mar. 2012, pp. 328-335. [cited by applicant]
Benner et al., Finemap: Efficient Variable Selection Using Summary Data from Genome-wide Association Studies, Bioinformatics, vol. 32, No. 10, 2016, pp. 1493-1501. [cited by applicant]
Chun et al., Limited Statistical Evidence for Shared Genetic Effects of eQTLs and Autoimmune-Disease-Associated Loci in Three Major Immune-Cell Types, Nature Genetics, vol. 49, No. 4, Apr. 2017, pp. 600-608. [cited by applicant]
Chung et al., GPA: A Statistical Approach to Prioritizing GWAS Results by Integrating Pleiotropy and Annotation, PLOS Genetics, vol. 10, No. 11, Nov. 2014, 14 pages. [cited by applicant]
Fearnhead, Exact and Efficient Bayesian Inference for Multiple Changepoint Problems, Statistics and Computing, vol. 16, 2006, pp. 203-213. [cited by applicant]
Giambartolomei et al., Bayesian Test for Colocalisation between Pairs of Genetic Association Studies Using Summary Statistics, PLOS Genetics, vol. 10. No. 5, May 2014, 15 pages. [cited by applicant]
Hackinger et al., Statistical Methods to Detect Pleiotropy in Human Complex Traits, Open Biology, vol. 7, No. 11, Nov. 1, 2017, 13 pages. [cited by applicant]
Han et al., A Method to Decipher Pleiotropy by Detecting Underlying Heterogeneity Driven by Hidden Subgroups Applied to Autoimmune and Neuropsychiatric Diseases, Nature Genetics, vol. 48, No. 7, Jul. 2016, pp. 803-812. [cited by applicant]
Hormozdiari et al., Colocalization of GWAS and eQTL Signals Detects Target Genes, The American Journal of Human Genetics, vol. 99, Dec. 1, 2016, pp. 1245-1260. [cited by applicant]
Kichaev et al., Integrating Functional Data to Prioritize Causal Variants in Statistical Fine-Mapping Studies, PLOS Genetics, vol. 10, No. 10, Oct. 2014, 16 pages. [cited by applicant]
Kwak et al., Gene- and Pathway-based Association Tests for Multiple Traits with Gwas Summary Statistics, Bioinformatics, vol. 33, No. 1, Jan. 2017, pp. 64-71. [cited by applicant]
Li et al., An Empirical Bayes Approach for Multiple Tissue eQTL Analysis, Biostatistics, vol. 19, No. 3, 2018, pp. 391-406. [cited by applicant]
Liu et al., Eps: An Empirical Bayes Approach to Integrating Pleiotropy and Tissue-specific Information for Prioritizing Risk Genes, Bioinformatics, vol. 32, No. 12, Feb. 15, 2016, pp. 1856-1864. [cited by applicant]
Mahmoud et al., TWIST1 Integrates Endothelial Responses to Flow in Vascular Dysfunction and Atherosclerosis, Circulation Research, Available online at: http://circres.ahajournals.org, Jul. 22, 2016, pp. 450-462. [cited by applicant]
Maller et al., Bayesian Refinement of Association Signals for 14 Loci in 3 Common Diseases, Nature Genetics, vol. 44, No. 12, Dec. 2012, pp. 1294-1302. [cited by applicant]
Solovieff et al., Pleiotropy in Complex Traits: Challenges and Strategies, Nature Reviews Genetics, vol. 14, Jul. 2013, pp. 483-495. [cited by applicant]
Turley et al., Multi-trait Analysis of Genome-Wide Association Summary Statistics Using MTAG, Nature Genetics, vol. 50, Feb. 2018, pp. 229-237. [cited by applicant]
Wakefield, Bayes Factors for Genome-Wide Association Studies: Comparison with P-values, Genetic Epidemiology, vol. 33, 2009, pp. 79-86. [cited by applicant]
Wen et al., Integrating Molecular QTL Data into Genomewide Genetic Association Analysis: Probabilistic Assessment of Enrichment and Colocalization, PLOS Genetics, Available Online At: https://doi.org/10.1371/journal.pge… [cited by applicant]
Zhu et al., Bayesian Large-scale Multiple Regression with Summary Statistics from Genome-wide Association Studies, The Annals of Applied Statistics, vol. 11, No. 3, 2017, pp. 1561-1592. [cited by applicant]
Zhu et al., Integration of Summary Data from Gwas and eQTL Studies Predicts Complex Trait Gene Targets, Nature Genetics, vol. 48, Mar. 28, 2016, 9 pages. [cited by applicant]
Allen et al., Hundreds of Variants Clustered in Genomic Loci and Biological Pathways Affect Human Height, Nature, vol. 467, No. 7317, Sep. 29, 2010, pp. 832-838. [cited by applicant]
Farh et al., Genetic and Epigenetic Fine Mapping of Causal Autoimmune Disease Variants, Nature, vol. 518, No. 7539, XP055591899, Oct. 29, 2014, 21 pages. [cited by applicant]
Giambartolomei et al., A Bayesian Framework for Multiple Trait Colocalization From Summary Association Statistics, bioRxiv, Available Online At: URL: https://www.biorxiv.org/content/biorxiv/early/2018/02/13/155481.full.… [cited by applicant]
Li et al., A Probabilistic Framework to Dissect Functional Cell-type-specific Regulatory Elements and Risk Loci Underlying the Genetics of Complex Traits, bioRxiv, Available Online At: https://www.biorxiv.org/content/bi… [cited by applicant]
Nieuwboer et al., Gwis: Genome-wide Inferred Statistics for Functions of Multiple Phenotypes, American Journal of Human Genetics, vol. 99, No. 4, Oct. 6, 2016, pp. 917-927. [cited by applicant]
International Application No. PCT/GB2019/050525, International Search Report and Written Opinion, mailed on Jun. 6, 2019, 4 pages. [cited by applicant]
Pickrell et al., Detection and Interpretation of Shared Genetic Influences on 42 Human Traits, Nature Genetics, vol. 48, No. 7, XP055591430, Jul. 2016, 21 pages. [cited by applicant]
Saeed, Novel Linkage Disequilibrium Clustering Algorithm Identifies New Lupus Genes on Meta-analysis of Gwas Datasets, Immunogenetics, vol. 69, No. 5, XP036217074, Feb. 28, 2017, pp. 295-302. [cited by applicant]