IP Library Granted Patent US 12,229,141
Granted Patent B2
US 12,229,141 · App. 17/868,775 · Granted Feb 18, 2025

Linking individual datasets to a database

Inventors: Shiya Song (San Mateo, CA); Jingwen Pei (San Mateo, CA); Brett Frederick Jorgensen (Draper, UT); Aaron James Stern (Berkeley, CA); Ross E. Curtis (Cedar Hills, UT)
Assignee: Ancestry.com DNA, LLC
G06F16/24558G06F16/2246G06F16/24578G16B10/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,229,141
App. No.
17/868,775
Granted
Feb 18, 2025
Kind
B2
Abstract

The disclosed system links an individual dataset to a database. The system receives a target individual dataset associated with a target individual and identifies candidate individual datasets that are potentially related to the target individual dataset. The system identifies a related individual dataset that has data bits that match some data bits in the target individual dataset. The system then identifies a parent node that is a common parent node to both the target individual dataset and the related individual dataset. The system retrieves a data tree that the parent node belongs to with the data tree containing information describing inter-relationships among datasets in the data tree. A node in the data tree is identified to assign the target individual dataset based on strings of matched data bits and number of the matched strings between the target individual dataset and the datasets in the data tree.

Claims (95)

1. A computer-implemented method comprising:

receiving a target individual dataset associated with a target individual;

identifying a plurality of candidate individual datasets that are potentially related to the target individual dataset;

identifying a related individual dataset from the plurality of candidate individual datasets, wherein the related individual dataset has data bits that match at least a portion of data bits in the target individual dataset, wherein identifying the related individual dataset comprises:

phasing a genotype corresponding to the target individual and a genotype corresponding to the related individual;

identifying identity by descent (IBD) segments shared between the phased genotype of the target individual dataset and the phased genotype of the related individual dataset;

determining a total length of the IBD segments shared between the phased genotype of the target individual dataset and the phased genotype of the related individual dataset; and

determining that the total length of shared IBD segments exceeds a threshold;

identifying a parent node that is a common parent node for both the related individual dataset and the target individual dataset;

retrieving a data tree to which the parent node belongs, the data tree describing inter-relationships among datasets in the data tree;

identifying, based on strings of matched data bits and number of the strings of matched data bits between the target individual dataset and the datasets in the data tree, a descendant position to the common parent node in the data tree to which the target individual dataset is assigned, wherein identifying the descendant position to the common parent node in the data tree comprises: generating a plurality of candidate data trees that have the individual dataset assigned to different candidate descendant positions to the common parent node, wherein generating the plurality of candidate data trees comprises:

generating a first candidate data tree that is based on the data tree to which the parent node belongs, the first candidate data tree being the data tree with a first new descendant position, the first candidate data tree adding the individual dataset to the first new descendant position, and

generating a second candidate data tree that is based on the data tree to which the parent node belongs, the second candidate data tree being the data tree with a second new descendant position, the second candidate data tree adding the individual dataset to the second new descendant position,

calculating, for each of the candidate data trees with new descendant positions added for the individual dataset to the common parent node, a plurality of pairwise relationship likelihoods, each pairwise relationship likelihood measuring a likelihood between a candidate descendant position and another dataset that also represents a descendant of the common parent node, and

selecting a candidate data tree as the data tree based on the plurality of pairwise relationship likelihoods; and

outputting the data tree with the target individual dataset located in the descendant position.

2. The method of claim 1 , wherein identifying a parent node further comprises:

identifying a plurality of candidate parent nodes, wherein a candidate parent node represents a candidate common ancestor for both the target individual dataset and one of the candidate individual datasets;

calculating confidence scores for the candidate parent nodes; and

selecting one of the candidate parent nodes as the parent node based on a ranking of the candidate parent nodes by the confidence scores.

3. The method of claim 2 , wherein identifying the parent node further comprises a pruning process, the pruning process comprising:

retrieving a meiosis separation between the related individual represented by the related individual dataset and the target individual;

determining a generation value between the related individual and one of the candidate parent nodes;

determining a range for the generation value based on the meiosis; and

removing the one of the candidate parent nodes as a candidate in response to the generation value out of the range.

4. The method of claim 1 , wherein generating the plurality of candidate data trees that have the individual dataset assigned to different candidate descendant positions to the common parent node comprises one or more of the following:

(i) assigning the target individual dataset at an existing node in the data tree as the one of the new descendant positions, the one of the new descendant positions replacing the existing node;

(ii) adding a child node that descends from a leaf node in the data tree as the one of the new descendant positions of the target individual dataset; and/or

(iii) adding a child node that descends from an inner node in the data tree wherein the child node is in a new branch descending from the inner node, the child node being the one of the new descendant positions of the target individual dataset.

5. The method of claim 1 , wherein each candidate data is associated with a likelihood score determined from the pairwise relationship likelihoods, the likelihood score for each candidate data tree is a composite likelihood calculated based on individual datasets in each candidate data tree, wherein the individual datasets contain DNA information.

6. The method of claim 5 , wherein the composite likelihood for each candidate data tree is determined based on steps comprising:

determining a likelihood for each pairwise individual datasets between the target individual and other individuals in the candidate data tree, the pairwise individual datasets containing DNA information in the candidate data tree, the likelihood calculated based on matched DNA information and positions of the pair of individual datasets in the candidate data tree; and

generating the composite likelihood based on a product of the likelihood of each pair of individual datasets.

7. The method of claim 1 , wherein identifying the descendant position to the common parent node in the data tree to which the target individual dataset is assigned is further based on a relationship between the target individual dataset and the related individual dataset determined based on matched DNA information.

8. The method of claim 1 , wherein:

the data bits contain information associated with DNA;

the strings of matched data bits contain information associated with matched DNA segments; and

the number of the strings contain information associated with number of matched DNA segments.

9. The computer-implemented method of claim 1 , wherein identifying a parent node that is a common parent node for both the related individual dataset and the target individual dataset comprises:

identifying a data tree corresponding to the target individual dataset and a data tree corresponding to the related individual dataset in a large-scale network comprising concatenated data trees corresponding to a plurality of individuals, the concatenated data trees having been linked in the large-scale network by identifying common individuals in different data trees, and

identifying a parent node by defining a path connecting the target individual dataset and the related individual dataset in the large-scale network.

10. A system comprising:

a computing server comprising one or more processors and memory for storing computer code comprising instructions, wherein the instructions, when executed by one or more computer processors, cause the one or more computer processors to perform steps comprising:

receiving a target individual dataset associated with a target individual;

identifying a plurality of candidate individual datasets that are potentially related to the target individual dataset;

identifying a related individual dataset from the plurality of candidate individual datasets, wherein the related individual dataset has data bits that match at least a portion of data bits in the target individual dataset, wherein identifying the related individual dataset comprises:

phasing a genotype corresponding to the target individual and a genotype corresponding to the related individual;

identifying identity by descent (IBD) segments shared between the phased genotype of the target individual dataset and the phased genotype of the related individual dataset;

determining a total length of the IBD segments shared between the phased genotype of the target individual dataset and the phased genotype of the related individual dataset; and

determining that the total length of shared IBD segments exceeds a threshold;

identifying a parent node that is a common parent node for both the related individual dataset and the target individual dataset;

retrieving a data tree to which the parent node belongs, the data tree describing inter-relationships among datasets in the data tree;

identifying, based on strings of matched data bits and number of the strings of matched data bits between the target individual dataset and the datasets in the data tree, a descendant position to the common parent node in the data tree to which the target individual dataset is assigned, wherein identifying the descendant position to the common parent node in the data tree comprises:

generating a plurality of candidate data trees that have the individual dataset assigned to different candidate descendant positions to the common parent node, wherein generating the plurality of candidate data trees comprises:

generating a first candidate data tree that is based on the data tree to which the parent node belongs, the first candidate data tree being the data tree with a first new descendant position, the first candidate data tree adding the individual dataset to the first new descendant position, and generating a second candidate data tree that is based on the data tree to which the parent node belongs, the second candidate data tree being the data tree with a second new descendant position, the second candidate data tree adding the individual dataset to the second new descendant position,

calculating, for each of the candidate data trees with new descendant positions added for the individual dataset to the common parent node, a plurality of pairwise relationship likelihoods, each pairwise relationship likelihood measuring a likelihood between a candidate descendant position and another dataset that also represents a descendant of the common parent node, and

selecting a candidate data tree as the data tree based on the plurality of pairwise relationship likelihoods; and

outputting the data tree with the target individual dataset located in the descendant position; and

a graphical user interface in communication with the computing server, the graphical user interface configured to display the data tree with the target individual dataset located in the descendant position.

11. The system of claim 10 , wherein identifying a parent node further comprises:

identifying a plurality of candidate parent nodes, wherein a candidate parent node represents a candidate common ancestor for both the target individual dataset and one of the candidate individual datasets;

calculating confidence scores for the candidate parent nodes; and

selecting one of the candidate parent nodes as the parent node based on a ranking of the candidate parent nodes by the confidence scores.

12. The system of claim 11 , wherein identifying the parent node further comprises a pruning process, the pruning process comprising:

retrieving a meiosis separation between the related individual represented by the related individual dataset and the target individual;

determining a generation value between the related individual and one of the candidate parent nodes;

determining a range for the generation value based on the meiosis; and

removing the one of the candidate parent nodes as a candidate in response to the generation value out of the range.

13. The system of claim 10 , wherein generating one or more candidate data trees that have the individual dataset assigned to different candidate descendant positions to the common parent node comprises one or more of the following:

(i) assigning the target individual dataset at an existing node in the data tree as the one of the new descendant positions, the new descendant positions replacing the existing node;

(ii) adding a child node that descends from a leaf node in the data tree as the one of the new descendant positions of the target individual dataset; and/or

(iii) adding a child node that descends from an inner node in the data tree wherein the child node is in a new branch descending from the inner node, the child node being the one of the new descendant positions of the target individual dataset.

14. The system of claim 10 , wherein each candidate data is associated with a likelihood score determined from the pairwise relationship likelihoods, the likelihood score for each candidate data tree is a composite likelihood calculated based on individual datasets in each candidate data tree, wherein the individual datasets contain DNA information.

15. The system of claim 14 , wherein the composite likelihood for each candidate data tree is determined based on steps comprising:

determining a likelihood for each pairwise individual datasets between the target individual and other individuals in the candidate data tree, the pairwise individual datasets containing DNA information in the candidate data tree, the likelihood calculated based on matched DNA information and positions of the pair of individual datasets in the candidate data tree; and

generating the composite likelihood based on a product of the likelihood of each pair of individual datasets.

16. The system of claim 10 , wherein identifying the descendant position to the common parent node in the data tree to which the target individual dataset is assigned is further based on metadata associated with the datasets.

17. The system of claim 10 , wherein identifying the descendant position to the common parent node in the data tree to which the target individual dataset is assigned is further based on a relationship between the target individual dataset and the related individual dataset determined based on matched DNA information.

18. A non-transitory computer readable medium for storing computer code comprising instructions for linking an individual dataset to a database, the instructions, when executed by one or more computer processors, cause the one or more computer processors to perform steps comprising:

receiving a target individual dataset associated with a target individual;

identifying a plurality of candidate individual datasets that are potentially related to the target individual dataset;

identifying a related individual dataset from the plurality of candidate individual datasets, wherein the related individual dataset has data bits that match at least a portion of data bits in the target individual dataset, wherein identifying the related individual dataset comprises:

phasing a genotype corresponding to the target individual and a genotype corresponding to the related individual;

identifying identity by descent (IBD) segments shared between the phased genotype of the target individual dataset and the phased genotype of the related individual dataset;

determining a total length of the IBD segments shared between the phased genotype of the target individual dataset and the phased genotype of the related individual dataset; and

determining that the total length of shared IBD segments exceeds a threshold;

identifying a parent node that is a common parent node for both the related individual dataset and the target individual dataset;

retrieving a data tree to which the parent node belongs, the data tree describing inter-relationships among datasets in the data tree;

identifying, based on strings of matched data bits and number of the strings of matched data bits between the target individual dataset and the datasets in the data tree, a descendant position to the common parent node in the data tree to which the target individual dataset is assigned, wherein identifying the descendant position to the common parent node in the data tree comprises:

generating a plurality of candidate data trees that have the individual dataset assigned to different candidate descendant positions to the common parent node, wherein generating the plurality of candidate data trees comprises:

generating a first candidate data tree that is based on the data tree to which the parent node belongs, the first candidate data tree being the data tree with a first new descendant position, the first candidate data tree adding the individual dataset to the first new descendant position, and

generating a second candidate data tree that is based on the data tree to which the parent node belongs, the second candidate data tree being the data tree with a second new descendant position, the second candidate data tree adding the individual dataset to the second new descendant position,

calculating, for each of the candidate data trees with new descendant positions added for the individual dataset to the common parent node, a plurality of pairwise relationship likelihoods, each pairwise relationship likelihood measuring a likelihood between a candidate descendant position and another dataset that also represents a descendant of the common parent node, and

selecting a candidate data tree as the data tree based on the plurality of pairwise relationship likelihoods; and

outputting the data tree with the target individual dataset located in the descendant position.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2022
From: SONG, SHIYA; PEI, JINGWEN; JORGENSEN, BRETT FREDERICK; STERN, AARON JAMES; CURTIS, ROSS E.
To: ANCESTRY.COM DNA, LLC
Reel/Frame 061748/0881 →
Continuity (3)
Continuation 17128009 · Dec 19, 2020
Provisional Application 62951646 · Dec 20, 2019
Related Publication 20220365934A1 · Nov 17, 2022
References Cited (145)
US 6570567B1 · Eaton · 2003 [cited by applicant]
US 7062752B2 · Simpson et al. · 2006 [cited by applicant]
US 7249129B2 · Cookson et al. · 2007 [cited by applicant]
US 7818281B2 · Kennedy et al. · 2010 [cited by applicant]
US 8510057B1 · Avey et al. · 2013 [cited by applicant]
US 8769438B2 · Mangum et al. · 2014 [cited by applicant]
US 9116882B1 · Macpherson et al. · 2015 [cited by applicant]
US 9213947B1 · Do et al. · 2015 [cited by applicant]
US 9239835B1 · Tiwari et al. · 2016 [cited by applicant]
US 9336177B2 · Hawthorne et al. · 2016 [cited by applicant]
US 9367800B1 · Do et al. · 2016 [cited by applicant]
US 9390225B2 · Barber et al. · 2016 [cited by applicant]
US 9836576B1 · Do et al. · 2017 [cited by applicant]
US 9864835B2 · Avey et al. · 2018 [cited by applicant]
US 11113609B2 · Roy et al. · 2021 [cited by applicant]
US 20020019746A1 · Rienhoff et al. · 2002 [cited by applicant]
US 20020143578A1 · Cole et al. · 2002 [cited by applicant]
US 20030101000A1 · Bader et al. · 2003 [cited by applicant]
US 20030113727A1 · Girn et al. · 2003 [cited by applicant]
US 20030172065A1 · Sorenson et al. · 2003 [cited by applicant]
US 20040083226A1 · Eaton · 2004 [cited by applicant]
US 20040093334A1 · Scherer · 2004 [cited by applicant]
US 20040126840A1 · Cheng et al. · 2004 [cited by applicant]
US 20040267458A1 · Judson et al. · 2004 [cited by applicant]
US 20050089852A1 · Lee et al. · 2005 [cited by applicant]
US 20050147947A1 · Cookson, Jr. · 2005 [cited by examiner]
US 20050164704A1 · Winsor · 2005 [cited by applicant]
US 20050164705A1 · Rajkotia et al. · 2005 [cited by applicant]
US 20050192008A1 · Desai et al. · 2005 [cited by applicant]
US 20070050354A1 · Rosenberg · 2007 [cited by applicant]
US 20070260599A1 · McGuire et al. · 2007 [cited by applicant]
US 20080040046A1 · Chakraborty · 2008 [cited by examiner]
US 20080081331A1 · Myres et al. · 2008 [cited by applicant]
US 20080082955A1 · Andreessen et al. · 2008 [cited by applicant]
US 20080111716A1 · Artan et al. · 2008 [cited by applicant]
US 20080113727A1 · Vallejo et al. · 2008 [cited by applicant]
US 20080154566A1 · Myres et al. · 2008 [cited by applicant]
US 20080162510A1 · Baio et al. · 2008 [cited by applicant]
US 20080243398A1 · Rabinowitz et al. · 2008 [cited by applicant]
US 20080255768A1 · Martin et al. · 2008 [cited by applicant]
US 20090030985A1 · Yuan · 2009 [cited by applicant]
US 20090287660A1 · Shinjo et al. · 2009 [cited by applicant]
US 20100199066A1 · Artan et al. · 2010 [cited by applicant]
US 20100223281A1 · Hon et al. · 2010 [cited by applicant]
US 20100256917A1 · McVean et al. · 2010 [cited by applicant]
US 20100287213A1 · Rolls et al. · 2010 [cited by applicant]
US 20120054190A1 · Peters · 2012 [cited by applicant]
US 20120191903A1 · Araki et al. · 2012 [cited by applicant]
US 20130085728A1 · Tang et al. · 2013 [cited by applicant]
US 20130149707A1 · Sorenson et al. · 2013 [cited by applicant]
US 20140067355A1 · Noto et al. · 2014 [cited by applicant]
US 20140082568A1 · Hulet et al. · 2014 [cited by applicant]
US 20140108527A1 · Aravanis et al. · 2014 [cited by applicant]
US 20140194300A1 · Song et al. · 2014 [cited by applicant]
US 20140278138A1 · Barber · 2014 [cited by examiner]
US 20150106115A1 · Hu et al. · 2015 [cited by applicant]
US 20160070859A1 · Ignatenko · 2016 [cited by applicant]
US 20160350479A1 · Han et al. · 2016 [cited by applicant]
US 20170017752A1 · Noto et al. · 2017 [cited by applicant]
US 20170213127A1 · Duncan · 2017 [cited by applicant]
US 20170262577A1 · Ball et al. · 2017 [cited by applicant]
US 20190361923A1 · Joseph et al. · 2019 [cited by applicant]
WO WO0217190A1 · 2002 [cited by applicant]
WO WO2016061568A1 · 2016 [cited by applicant]
Alexander, D.H., et al., “Fast model-based estimation of ancestry in unrelated individuals,” Genome research, Sep. 2009, vol. 19, No. 9, pp. 1655-1664. [cited by applicant]
Ball, C. et al., “Ancestry DNA Matching White Paper,” Ancestry.com., 2016, [Online] [Retrieved Sep. 18, 2019], Retrieved from the internet, URL:<<https://www.ancestry.com/corporate/sites/default/files/AncestryDNA-Matchi… [cited by applicant]
Baran, Y. et al., “Fast and accurate inference of local ancestry in Latino populations.” Bioinformatics, May 2012, vol. 28, No. 10, pp. 1359-1367. [cited by applicant]
Bastian, M. et al., “Gephi: an open source software for exploring and manipulating networks,” Third international AAAI conference on weblogs and social media, May 2009, 361-362. [cited by applicant]
Bercovici, S. et al., “Ancestry inference in complex admixtures via variable-length Markov chain linkage models,” Annual International Conference on Research in Computational Molecular Biology, Springer, Berlin, Heidelb… [cited by applicant]
Brisbin, A. et al. “PCAdmix: principal components-based assignment of ancestry along each chromosome in individuals with admixed ancestry from two or more populations,” Human biology, Aug. 2012, vol. 84, No. 4, 343-364. [cited by applicant]
Browning, B.I. et al., “Detecting Identity by Descent and Estimating Genotype Error Rates in Sequence Data,” The American Journal of Human Genetics, Nov. 7, 2013, pp. 840-851. [cited by applicant]
Browning, B.L., “A Fast, Powerful Method for Detecting Identity by Descent,” The American Journal of Human Genetics, Feb. 11, 2011, vol. 88, pp. 173-182. [cited by applicant]
Browning, B.L., “A Unified Approach to Genotype Imputation and Haplotype-Phase Inference for Large Data Sets of Trios and Unrelated Individuals,” The American Journal of Human Genetics, Feb. 13, 2009, vol. 84, pp. 210-2… [cited by applicant]
Browning, B.L et al., “Efficient Multilocus Association Testing for Whole Genome Association Studies Using Localized Haplotype Clustering,” Genetic Epidemiology, vol. 31, Feb. 26, 2007, pp. 365-375. [cited by applicant]
Browning, B.L., “Genotype Imputation with Millions of Reference Samples,” The American Journal of Human Genetics, Jan. 7, 2016, vol. 98, pp. 116-126. [cited by applicant]
Browning, S. R., “Rapid and Accurate Haplotype Phasing and Missing-Data Inference for Whole-Genome Association Studies by Use of Localized Haplotype Clustering,” The American Journal of Human Genetics, Nov. 2007, vol. 8… [cited by applicant]
Browning, S.R. et al., “Haplotype phasing: Existing methods and new developments,” Nat Rev Genet, Apr. 1, 2012, vol. 12, No. 10, pp. 703-714. [cited by applicant]
Browning, S.R., “Multilocus Association Mapping Using Variable-Length Markov Chains,” The American Journal of Human Genetics, Jun. 2006, vol. 78, pp. 903-913. [cited by applicant]
Cann, H.M. et al., “A human genome diversity cell line panel,” Science, Apr. 2002, vol. 296, No. 5566, pp. 261-262. [cited by applicant]
Cavalli-Sforza, L.L. “The human genome diversity project: past, present and future,” Nature Reviews Genetics, Apr. 2005, vol. 6, No. 4. [cited by applicant]
De Roos, A.P.W., “Genomic selection in dairy cattle,” PHD Thesis at Wageningen University, Jan. 2011, 185 pages. [cited by applicant]
Dudoit, et al., “A score test for the linkage analysis of qualitative and quantitative traits based on identity by descent data from sib-pairs,” Biostatistics, vol. 1, Iss. 1, Mar. 2000, pp. 1-26. [cited by applicant]
Falush, D. et al., “Inference of Population Structure Using Multilocus Genotype Data: Linked Loci and Correlated Allele Frequencies,” Genetics, vol. 164, Aug. 2003, pp. 1567-1587. [cited by applicant]
Ghahramani, Z. “An Introduction to Hidden Markov Models and Bayesian Networks,” International Journal of Pattern recognition and Artificial Intelligence, Jun. 2001, vol. 15, No. 1, pp. 9-42. [cited by applicant]
Gravel, S., “Population genetics models of local ancestry,” Genetics, Jun. 2012, vol. 191, No. 2, pp. 607-619. [cited by applicant]
Guan, Y. “Detecting structure of haplotypes and local ancestry,” Genetics, Mar. 2014, vol. 196, No. 3, pp. 625-642. [cited by applicant]
Halperin, E. et al., “Haplotype reconstruction from genotype data using Imperfect Phylogeny,” Bioinformatics, Aug. 2004, vol. 20, No. 12, pp. 1842-1849. [cited by applicant]
Han, E. et al., “Clustering of 770,000 genomes reveals post-colonial population structure of North America,” Nature communications, Feb. 2017, vol. 8, pp. 1-12. [cited by applicant]
Harvard.edu, “Plink . . . Whole genome assocaition analysis toolset,” [Online] [Retrieved Sep. 19, 2019], Last edited Jan. 25, 2017, Retrieved from the internet , URL:<<http://zzz.bwh.harvard.edu/plink/>>, 4 pages. [cited by applicant]
Hellenthal, G. et al., “A genetic atlas of human admixture history,” Science, Feb. 2014, vol. 343, No. 6172, pp. 747-751. [cited by applicant]
Howie, B. N. et al., “A Flexible and Accurate Genotype Imputation Method for the Next Generation of Genome-Wide Association Studies,” PLoS Genetics, Jun. 2009, vol. 5, No. 6, pp. 1-15. [cited by applicant]
International HapMap Consortium, “A haplotype map of the human genome,” Nature, Oct. 2005, vol. 437, No. 27, pp. 1299-1320. [cited by applicant]
International HapMap Consortium, “A second generation human haplotype map of over 3.1 million SNPs,” Nature, Oct. 2007, vol. 449, No. 7164, pp. 1-30. [cited by applicant]
Itan, Y. et al., “The origins of lactase persistence in Europe,” PLoS computational biology, Aug. 2009, vol. 5, No. 8, pp. 1-13. [cited by applicant]
Jarvis, J.P. et al., “Patterns of Ancestry of Natural Selection and Genetic Association with Stature in Western African Pygmies,” PLoS Genetics, vol. 8, Iss. 4, Apr. 26, 2012, pp. 1-15. [cited by applicant]
Ke, X. et al. “Singleton SNPs in the human genome and implications for genome-wide association studies,” European Journal of Human Genetics, Jan. 2008, vol. 16, No. 4, 10 pages. [cited by applicant]
Lawson, D.J. et al., “Inference of population structure using dense haplotype data,” PLoS genetics, Jan. 2012, vol. 8, No. 1, pp. 1-16. [cited by applicant]
Li, N. et al., “Modeling Linkage disequilibrium and Identifying Recombination Hotspots Using Single-Nucleotide Polymorphism Data,” the Genetics Society of America, Dec. 2003, vol. 165, pp. 2213-2233. [cited by applicant]
Li, Y. et al., “MaCH: Using Sequence and Genotype Data to Estimate haplotypes and Unobserved Genotypes,” Genetic Epidemiology, Dec. 2010, vol. 34, pp. 816-834. [cited by applicant]
Loh, P.R. et al., “Inferring admixture histories of human populations using linkage disequilibrium,” Genetics, Apr. 2013, vol. 193, No. 4, pp. 1233-1254. [cited by applicant]
Ma, P. et al., “Comparison of different methods for imputing genome-wide marker genotypes in Swedish and Finnish Red Cattle,” J. Dairy Sci., Jul. 2013, vol. 96, pp. 4666-4677. [cited by applicant]
Ma, Y. et al. “Accurate inference of local phased ancestry of modern admixed populations,” Scientific reports, Jul. 2014, vol. 4, No. 5800 , pp. 1-5. [cited by applicant]
Maples, B.K. et al., “RFMix: a discriminative modeling approach for rapid and robust local-ancestry inference,” The American Journal of Human Genetics, Aug. 2013, vol. 93, No. 2, pp. 278-288. [cited by applicant]
McPeek, M. S. et al., “Assessment of Linkage Disequilibrium by the Decay of Haplotype Sharing with Application to Fine-Scale Genetic Mapping,” American Journal of Human Genetics, Sep. 1999, vol. 65, pp. 858-875. [cited by applicant]
Moreno-Estrada, A. et al., Reconstructing the Population Genetic History of the Caribbean, PLOS Genetics, Nov. 2013, vol. 9, No. 11, pp. 1-19. [cited by applicant]
Morrison, A.C. et al., “Prediction of Coronary Heart Disease Risk using a Genetic Risk Score: The Atherosclerosis Risk in Communities Study,” American Journal of Epidemiology, vol. 166, No. 1, Apr. 18, 2007, pp. 28-35. [cited by applicant]
Noto, K. et al., “A novel approach for estimating local and global admicture proportion based on rich haplotype models,” Invited Talk at the American Society of Human Genetics (ASHG) annual meeting, Baltimore, MD, Oct. … [cited by applicant]
Noto, K. et al., Abstract, “322 Polly: A novel approach for estimating local and global admixture proportion based on rich haplotype models,” ASHG 2015 Abstracts, The American Society of Human Genetics 65th Annual Meeti… [cited by applicant]
Noto, K., et al. “Underdog: a fully-supervised phasing algorithm that learns from hundreds of thousands of samples and phases in minutes. Invited Talk,” 64th Annual Meeting of the American Society of Human Genetics, 201… [cited by applicant]
Palin, K. et al., “Identity-by-Descent-Based Phasing and Imputation in Founder Populations Using Graphical Models,” Genetic Epidemiology, vol. 35, Oct. 17, 2011, pp. 853-860. [cited by applicant]
Paşaniuc, B. et al., “Imputation-based local ancestry inference in admixed populations,” International Symposium on Bioinformatics Research and Applications, Springer, Berlin, Heidelberg, May 2009, pp. 221-233. [cited by applicant]
Paşaniuc, B.et al. “Inference of locus-specific ancestry in closely related populations,” Bioinformatics, May 2009, vol. 25, No. 12, pp. i213-i221. [cited by applicant]
Patterson, N. et al., “Population structure and eigenanalysis,” PLoS genetics, Dec. 2006, vol. 2, No. 12, pp. 2074-2093. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Patent Application No. PCT/1B2019/057667, Jan. 10, 2020, 10 pages. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Patent Application No. PCT/1B2020/062256, Mar. 22, 2021, 14 pages. [cited by applicant]
Platt, J.C., “Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods,” Mar. 26, 1999, pp. 1-11. [cited by applicant]
Price, A.L. et al., “Sensitive Detection of Chromosomal Segments of Distinct Ancestry in Admixed Populations,” PLoS Genetics, vol. 5, Iss. 6, Jun. 2009, pp. 1-18. [cited by applicant]
Pritchard, J.K. et al., “Inference of population structure using multilocus genotype data,” Genetics Society of America, Jun. 2000, vol. 155, No. 2, pp. 945-959. [cited by applicant]
Purcell, S. et al., “PLINK: a tool set for whole-genome association and population-based linkage analyses,” The American journal of human genetics, Sep. 2007, vol. 81, No. 3, pp. 559-575. [cited by applicant]
Qian, Y. et al., “Efficient clustering of identity-by-descent between multiple individuals,” Bioinformatics, vol. 30, No. 7, Dec. 19, 2013, pp. 915-922. [cited by applicant]
Rabiner, L.R., “A Tutorial on Hidden Markov Models and Selected Applications in Speech Recognition,” Proceedings of the IEEE, Feb. 1989, vol. 77, No. 2, pp. 257-286. [cited by applicant]
Ranciaro, A. et al., “Genetic origins of lactase persistence and the spread of pastoralism in Africa,” The American Journal of Human Genetics, Apr. 2014, vol. 94, No. 4, pp. 496-510. [cited by applicant]
Roach, J.C. et al., “Analysis of genetic inheritance in a family quartet by whole-genome sequencing,” Science, Apr. 2010, vol. 328, No. 5978, pp. 636-639. [cited by applicant]
Ron, D., “On the Learnability and Usage of Acyclic Probabilistic Finite Automata,” Journal of Computer and System Sciences, Apr. 1998, vol. 56, pp. 133-152. [cited by applicant]
Sankararaman, S. et al., “Estimating local ancestry in admixed populations,” The American Journal of Human Genetics, Feb. 2008, vol. 82, No. 2, pp. 290-303. [cited by applicant]
Scheet, P. et al., “A Fast and Flexible Statistical model for Large-Scale Population Genotype Data: Applications to Inferring Missing Genotypes and Hplotypic Phase,” The American Journal of Human Genetics, Apr. 2006, vo… [cited by applicant]
Seligsohn, U. et al., “Genetic Susceptibility to Venous Thrombosis,” The New England Journal of Medicine, vol. 344, No. 16, Apr. 19, 2001, pp. 1222-1231. [cited by applicant]
Staples, J. et al., “PRIMUS: Rapid Reconstruction of Pedigrees from Genome-wide Estimates of Identity by Descent,” The American Journal of Human Genetics, vol. 95, Nov. 6, 2014, pp. 553-564. [cited by applicant]
Stephens, M. et al., “Accounting for Decay of Linkage Disequilibrium in Haplotype Inference and Missing-Data Imputation,” American Journal of Human Genetics, Mar. 2005, vol. 76, pp. 449-462. [cited by applicant]
Sturm, R.A. et al., “A single SNP in an evolutionary conserved region within intron 86 of the HERC2 gene determines human blue-brown eye color,” The American Journal of Human Genetics, Feb. 2008, vol. 82, No. 2, pp. 424… [cited by applicant]
Sundquist, A. et al., “Effect of genetic divergence in identifying ancestral origin using HAPAA,” Genome Res., Mar. 18, 2008, vol. 18, pp. 676-682. [cited by applicant]
Tang, H. et al., “Reconstructing Genetic Ancestry Blocks in Admixed Individuals,” The American journal of Human Genetics, Jul. 2006, vol. 79, pp. 1-12. [cited by applicant]
The 1000 Genomes Project Consortium, “A global reference for human genetic variation,” Macmillan Publishers Limited, Nature, Oct. 1, 2015, vol. 526, No. 7571, pp. 68-74. [cited by applicant]
The International HAPMAP 3 Consortium, “Integrating common and rare genetic variation in diverse human populations,” Nature, vol. 467, Sep. 2, 2010, pp. 52-58. [cited by applicant]
Tipping, M.E., “Sparse Bayesian Learning and the Relevance Vector Machine,” Journal of Machine Learning Research, Jun. 2001, pp. 211-244. [cited by applicant]
U.S. Appl. No. 61/724,228, filed Nov. 8, 2012, Inventor Chuong Do. [cited by applicant]
U.S. Appl. No. 61/724,236, filed Nov. 8, 2012, Inventor Chuong Do. [cited by applicant]
Visscher, P.M. et al., “Heritability in the genomics era-concepts and misconceptions,” Nature Reviews Genetics, Mar. 4, 2008, pp. 255-266. [cited by applicant]
Weedon, M.N. et al., “Combining Information from Common Type 2 Diabetes Risk Polymorphisms Improves Disease Prediction,” PLoS Med., vol. 3, Iss. 10, Oct. 2006, pp. 1877-1882. [cited by applicant]
Wikipedia, “Inverse distance weighting,” [Online] [Retrieved Sep. 18, 2019], Last edited Mar. 4, 2019, Retrieved from the internet ,URL:<<https://en.wikipedia.org/wiki/Inverse_distance_weighting>>. [cited by applicant]
Williams, A.L. et al., “Phasing of Many Thousands of Genotyped Samples,” The American Journal of Human Genetics, Aug. 10, 2012, vol. 91, pp. 238-251. [cited by applicant]
Yang, Q. et al., “Improving the Prediction of Complex Diseases by Testing for Multiple Disease-Susceptibility Genes,” American Journal of Human Genetics, vol. 72, Feb. 14, 2003, pp. 636-649. [cited by applicant]
Yoon, B.J., “Hidden Markov Models and their Applications in Biological Sequence Analysis,” Current Genomics, Nov. 2009, vol. 10, pp. 402-415. [cited by applicant]
Zhao, H. et al., “Haplotype analysis in population genetics and association studies,” Pharmacogenomics, Mar. 2003, vol. 4, No. 2, pp. 171-178. [cited by applicant]
U.S. Appl. No. 17/128,009, filed Dec. 24, 2021, 16 pages. [cited by applicant]