IP Library Granted Patent US 9,443,056
Granted Patent B2
US 9,443,056 · App. 13/445,925 · Granted Sep 13, 2016

Phased whole genome genetic risk in a family quartet

Inventors: Frederick Dewey (Redwood City, CA); Euan A. Ashley (Menlo Park, CA); Jake Byrnes (San Francisco, CA); Carlos Daniel Bustamante (Emerald Hills, CA); Atul J. Butte (Menlo Park, CA); Rong Chen (Fremont, CA)
Assignee: The Board of Trustees of the Leland Stanford Junior University
G06F19/12G06F19/18G06F19/24G06N5/02G06F19/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,443,056
App. No.
13/445,925
Granted
Sep 13, 2016
Kind
B2
Abstract

An embodiment of the present invention is a methodology for prioritizing variants relevant to inherited Mendelian (“single gene”) disease syndromes according to disease phenotype, gene, and variant level information.

Claims (167)

1. A method for resolving haplotype phase, comprising:

receiving allele data describing allele information regarding genotypes for a family comprising at least a mother, a father, and at least two children of the mother and the father, where the genotypes for the family contain single nucleotide variants and storing the allele data on a computer system comprising a processor and a memory;

receiving pedigree data for the family describing information regarding a pedigree for the family and storing the pedigree data on a computer system comprising a processor and a memory;

determining an inheritance state for the allele information described in the allele data based on identity between single nucleotide variants contained in the genotypes for the family using a Hidden Markov Model having hidden states implemented on a computer system comprising a processor and a memory,

wherein the hidden states comprise inheritance states, a compression fixed error state, and an MIE-rich fixed error state,

wherein the inheritance states are maternal identical, paternal identical, identical, and non-identical;

receiving transition probability data describing transition probabilities for inheritance states and storing the transition probability data on a computer system comprising a processor and a memory;

receiving population linkage disequilibrium data and storing the population disequilibrium data on a computer system comprising a processor and a memory;

determining a haplotype phase for at least one member of the family based on the pedigree data for the family, the inheritance state for the information described in the allele data, the transition probability data, and the population linkage disequilibrium data using a computer system comprising a processor and a memory;

storing the haplotype phase for at least one member of the family using a computer system comprising a processor and a memory; and

providing the stored haplotype phase for at least one member of the family in response to a request using a computer system comprising a processor and a memory.

2. The method of claim 1 , further comprising determining a most likely state path in view of the received allele data using a computer system comprising a processor and a memory.

3. The method of claim 1 , wherein the long-range haplotype phase is determined for each of the mother and father in contigs according to passage of allele contigs to one, both, or neither of the children using a computer system comprising a processor and a memory.

4. The method of claim 1 , further comprising determining a risk prediction for passage of at least one gene from at least one of the mother or the father to a selected child using a computer system comprising a processor and a memory.

5. The method of claim 1 , further comprising:

determining whether at least one genetic variant associated with disease is within the stored haplotype phase by utilizing the haplotype phase to query a disease associated-single nucleotide polymorphism database using a computer system comprising a processor and a memory;

determining a contribution to inherited disease syndromes from multigenic causes in view of the determination whether the at least one genetic variant associated with disease is within the stored haplotype phase and the pedigree data using a computer system comprising a processor and a memory,

storing the determined contribution using a computer system comprising a processor and a memory; and

providing the determined contribution in response to a request using a computer system comprising a processor and a memory.

6. The method of claim 1 , further comprising:

determining whether at least one genetic variant associated with disease is within the stored haplotype phase by utilizing the haplotype phase to query a disease associated-single nucleotide polymorphism database using a computer system comprising a processor and a memory;

determining a diagnosis for at least one member of the family based on whether the at least one disease-associated genetic variant is within the stored haplotype phase using a computer system comprising a processor and a memory;

storing the determined diagnosis using a computer system comprising a processor and a memory; and

providing the determined diagnosis in response to a request using a computer system comprising a processor and a memory.

7. The method of claim 1 , further comprising:

determining whether at least one genetic variant associated with disease is within the stored haplotype phase by utilizing the haplotype phase to query a disease associated-single nucleotide polymorphism database using a computer system comprising a processor and a memory;

determining a drug for treatment of at least one member of the family based on information regarding drug-variant-phenotype associations from a pharmacogenomics database and the determination whether the at least one genetic variant associated with disease is within the stored haplotype phase using a computer system comprising a processor and a memory;

storing the determined drug using a computer system comprising a processor and a memory; and

providing the determined drug in response to a request using a computer system comprising a processor and a memory.

8. The method of claim 1 , further comprising:

determining whether at least one genetic variant associated with disease is within the stored haplotype phase by utilizing the haplotype phase to query a disease associated-single nucleotide polymorphism database using a computer system comprising a processor and a memory;

determining genetic risk for single nucleotide genetic variants and genetic variants associated with disease using a computer system comprising a processor and a memory;

determining a prognosis for at least one member of the family responsive to the determined haplotype phase using a computer system comprising a processor and a memory

storing the prognosis for at least one member of the family using a computer system comprising a processor and a memory; and

providing the prognosis for at least one member of the family in response to a user request using a computer system comprising a processor and a memory.

9. A non-transitory computer-readable medium including instructions that, when executed by a processing unit, cause the processing unit to resolve haplotype phase, by performing the steps comprising:

receiving allele data describing allele information regarding genotypes for a family comprising at least a mother, a father, and at least two children of the mother and the father, where the genotypes for the family contain single nucleotide variants and storing the allele data in a memory;

receiving pedigree data for the family describing information regarding a pedigree for the family and storing the pedigree data in a memory;

determining an inheritance state for the allele information described in the allele data based on identity between single nucleotide variants contained in the genotypes for the family using a Hidden Markov Model having hidden states,

wherein the hidden states comprise inheritance states, a compression fixed error state, and an MIE-rich fixed error state,

wherein the inheritance states are maternal identical, paternal identical, identical, and non-identical;

receiving transition probability data describing transition probabilities for inheritance states and storing the transition probability data in a memory;

receiving population linkage disequilibrium data and storing the population disequilibrium data on a memory;

determining a haplotype phase for at least one member of the family based on the pedigree data for the family, the inheritance state for the information described in the allele data, the transition probability data, and the population linkage disequilibrium data

storing the haplotype phase for at least one member of the family in a memory; and

providing the stored haplotype phase in response to a request.

10. The non-transitory computer-readable medium of claim 9 , further comprising determining a most likely state path in view of the received allele data.

11. The non-transitory computer-readable medium of claim 9 , wherein the haplotype phase is determined for each of the mother and father in contigs according to passage of allele contigs to one, both, or neither of the children.

12. The non-transitory computer-readable medium of claim 9 , further comprising determining a risk prediction for passage of at least one gene from at least one of the mother or the father to a selected.

13. The non-transitory computer-readable medium of claim 9 , further comprising:

determining whether at least one genetic variant associated with disease is within the stored haplotype phase by utilizing the stored haplotype phase to query a disease associated-single nucleotide polymorphism database;

determining a contribution to inherited disease syndromes from multigenic causes in view of the determination whether the at least one genetic variant associated with disease is within the stored haplotype phase and the pedigree data;

storing the determined contribution in a memory; and

providing the determined contribution in response to a user request.

14. The non-transitory computer-readable medium of claim 9 , further comprising:

determining whether at least one genetic variant associated with disease is within the stored haplotype phase by utilizing the haplotype phase to query a disease associated-single nucleotide polymorphism database;

determining a diagnosis for at least one member of the family based on whether the at least one disease-associated genetic variant is within the stored haplotype phase;

storing the determined diagnosis for at least one member of the family in a memory; and

providing the determined diagnosis for at least one member of the family in response to a user request.

15. The non-transitory computer-readable medium of claim 9 , further comprising:

determining whether at least one genetic variant associated with disease is within the stored haplotype phase by utilizing the stored haplotype phase to query a disease associated-single nucleotide polymorphism database;

determining a drug for treatment of at least one member of the family based on the information regarding drug-variant-phenotype associations from a pharmacogenomics database and the determination whether the at least one genetic variant associated with disease is within the stored haplotype phase;

storing the drug for treatment of at least one member of the family in a memory; and

providing the drug for treatment of at least one member of the family in response to a user request.

16. The non-transitory computer-readable medium of claim 9 , further comprising

determining whether at least one genetic variant associated with disease is within the stored haplotype phase by utilizing the haplotype phase to query the disease-single nucleotide polymorphism database;

determining genetic risk for single nucleotide genetic variants and at least one genetic variant associated with disease;

determining a prognosis for at least one member of the family responsive to the stored haplotype phase;

storing the prognosis for at least one member of the family in a memory; and

providing the prognosis for at least one member of the family in response to a user request.

17. A computing device comprising: a data bus; a memory unit coupled to the data bus; a processing unit coupled to the data bus and configured to:

receive allele data describing allele information regarding genotypes for a family comprising at least a mother, a father, and at least two children of the mother and the father, where the genotypes for the family contain single nucleotide variants and store the allele data in a memory;

receive pedigree data for the family describing information regarding a pedigree for the family and store the pedigree data in a memory;

determine an inheritance state for the allele information described in the allele data based on identity between single nucleotide variants contained in the genotypes for the family using a Hidden Markov Model having hidden states,

wherein the hidden states comprise inheritance states, a compression fixed error state, and an MIE-rich fixed error state,

wherein the inheritance states are maternal identical, paternal identical, identical, and non-identical;

receive transition probability data describing transition probabilities for inheritance states and store the transition probability data in a memory;

receive population linkage disequilibrium data and store the population disequilibrium data on a memory;

determine a haplotype phase for at least one member of the family based on the pedigree data for the family, the inheritance state for the information described in the allele data, the transition probability data, and the population linkage disequilibrium data;

store the haplotype phase for at least one member of the family in a memory; and

provide the stored haplotype phase in response to a request.

18. The method of claim 1 , wherein determining the haplotype phase for at least one member of the family based on the population linkage disequilibrium data comprises assigning a minor allele to either a paternal haplotype scaffold or a maternal haplotype scaffold based on maximizing a correlation.

19. The method of claim 18 , wherein the correlation is determined by calculating an r 2 value.

20. The method of claim 18 , wherein the single nucleotide variants have r 2 values greater than 0.3 and are within 250 kb of a point at which the haplotype phase for at least one member of the family is being determined.

21. The method of claim 18 , wherein the minor allele is placed at a locus l on the maternal haplotype scaffold h m or the paternal haplotype scaffold h p according to L max for a single nucleotide variant, where L max =max(L m , L p ), and where

L

(

l

;

h

)

=

{

1

,

r

i

2

=

1

,

1

n

n

r

i

2

,

r

i

2

1

}

,

where

h is a haplotype scaffold,

i is a heterozygous loci position,

n is a number of heterozygous loci on h,

r i 2 is a value for r 2 between the minor allele at l and the minor allele at heterozygous loci i on h.

22. The non-transitory computer-readable medium of claim 9 , wherein determining the haplotype phase for at least one member of the family based on the population linkage disequilibrium data comprises assigning a minor allele to either a paternal haplotype scaffold or a maternal haplotype scaffold based on maximizing a correlation.

23. The non-transitory computer readable medium of claim 22 , wherein the correlation is determined by calculating an r 2 value.

24. The non-transitory computer readable medium of claim 22 , wherein the single nucleotide variants have r 2 values greater than 0.3 and are within 250 kb of a point at which the haplotype phase for at least one member of the family is being determined.

25. The non-transitory computer readable medium of claim 22 , wherein the minor allele is placed at a locus l on the maternal haplotype scaffold h m or the paternal haplotype scaffold h p according to L max for a single nucleotide variant, where L max =max(L m , L p ), and where

L

(

l

;

h

)

=

{

1

,

r

i

2

=

1

,

1

n

n

r

i

2

,

r

i

2

1

}

,

where

h is a haplotype scaffold,

i is a heterozygous loci position,

n is a number of heterozygous loci on h,

r i 2 is a value for r 2 between the minor allele at l and the minor allele at heterozygous loci i on h.

Assignments (2)
CONFIRMATORY LICENSE Recorded Jul 2, 2015
From: STANFORD UNIVERSITY
To: NATIONAL INSTITUTES OF HEALTH (NIH), U.S. DEPT. OF HEALTH AND HUMAN SERVICES (DHHS), U.S. GOVERNMENT
Reel/Frame 036054/0126 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2015
From: DEWEY, FREDERICK; ASHLEY, EUAN; BUSTAMANTE, CARLOS; BUTTE, ATUL; BYRNES, JAKE; CHEN, RONG
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
Reel/Frame 035674/0378 →
Continuity (4)
Provisional Application 61474749 · Apr 13, 2011
Provisional Application 61502280 · Jun 28, 2011
Related Publication 20130080068A1 · Mar 28, 2013
Related Publication 20150370959A9 · Dec 24, 2015