SYSTEM AND METHOD FOR CLEANING NOISY GENETIC DATA FROM TARGET INDIVIDUALS USING GENETIC DATA FROM GENETICALLY RELATED INDIVIDUALS
A system and method for determining the genetic data for one or a small set of cells, or from fragmentary DNA, where a limited quantity of genetic data is available, are disclosed. Genetic data for the target individual is acquired and amplified using known methods, and poorly measured base pairs, missing alleles and missing regions are reconstructed using expected similarities between the target genome and the genome of genetically related subjects. In accordance with one embodiment of the invention, incomplete genetic data is acquired from embryonic cells, fetal cells, or cell-free fetal DNA isolated from the mother's blood, and the incomplete genetic data is reconstructed using the more complete genetic data from a larger sample diploid cells from one or both parents, with or without genetic data from haploid cells from one or both parents, and/or genetic data taken from other related individuals.
1 . A method for determining genetic data for DNA from a first individual, the method comprising:
isolating cell-free DNA from a biological sample of a second individual and amplifying at least 70 loci from the isolated cell-free DNA in a single reaction volume to obtain amplification products;
using high-throughput DNA sequencing to detect genetic material from the amplification products and produce genetic data for the at least 70 loci, wherein the genetic data is noisy due to a small amount of genetic material from a first individual; and
determining the most likely genetic data for DNA from the first individual.
2 . The method of claim 1 , wherein the small amount of genetic material is from 0.3 ng or less of DNA.
3 . The method of claim 1 , wherein the biological sample is a blood sample, and the small amount of genetic material is from cell-free DNA in the blood sample.
4 . The method of claim 1 , wherein the noisy data comprises allele drop out errors.
5 . The method of claim 1 , wherein the noisy data comprises measurement bias.
6 . The method of claim 1 , wherein the noisy data comprises incorrect measurements.
7 . The method of claim 1 , wherein the loci are SNP loci.
8 . The method of claim 7 , wherein the confidence that each SNP is correctly called is at least 95%.
9 . The method of claim 7 , wherein the confidence that each SNP is correctly called is at least 99%.
10 . The method of claim 1 , wherein the genetic data has been obtained by amplifying and/or measuring the genetic material using tools and/or techniques selected from the group consisting of Polymerase Chain Reaction (PCR), Ligation-mediated PCR, degenerative oligonucleotide primer PCR, Multiple Displacement Amplification, allele-specific amplification techniques, and combinations thereof, and wherein one or more of the individual's genetic data has been measured using tools and or techniques selected from the group consisting of MOLECULAR INVERSION PROBES (MIPs) circularizing probes, other circularizing probes, Genotyping Microarrays, the TAQMAN SNP Genotyping Assay, other hydrolysis probes, the ILLUMINA Genotyping System, other genotyping assays, Sanger DNA sequencing, pyrosequencing, other methods of DNA sequencing, other high through-put genotyping platforms, fluorescent in-situ hybridization (FISH), and combinations thereof.
11 . The method of claim 1 , further comprising normalizing the genetic data for differences in amplification and/or measurement efficiency between the loci.
12 . The method of claim 1 , wherein the probability of each of the hypotheses is determined without use of a reference sample.