IP Library › Granted Patent US 12,332,974
Granted Patent B2
US 12,332,974 · App. 18/759,587 · Granted Jun 17, 2025

Determination of data-source influence on data manifestations

Inventors: Andre Everson Kim (Upland, CA); Alisa Elnaz Sedghifar (San Francisco, CA); Ross Eugene Curtis (Cedar Hills, UT); Caitlyn Elizabeth Bruns (Saratoga Springs, UT)
Assignee: Ancestry.com DNA, LLC
G06F18/2415
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,332,974
App. No.
18/759,587
Granted
Jun 17, 2025
Kind
B2
Abstract

Disclosed are methods and system for predicting data-source influences on one or more data manifestations of a named entity. The method includes receiving an inheritance dataset of the named entity. The method determines first and second portions of the inheritance dataset of the named entity. The method determines an aggregated data-bit association score for the named entity based on the inheritance dataset at an identified subset of the data-bit regions. The method determines aggregated data-bit association scores associated with the first and second data source based on the first and second portions of the inheritance dataset at the identified subset of the data-bit regions. The method selects one of the first and second data sources as having a measure of influence on a data manifestation of the named entity corresponding to the identified subset of the data-bit regions.

Claims (78)

1. A computer-implemented method for predicting data-source influences on a data manifestation of a named entity, the computer-implemented method comprising:

receiving an inheritance dataset of the named entity, the inheritance dataset comprising one or more reads at a plurality of data-bit regions;

determining a first portion and a second portion of the inheritance dataset of the named entity, the first portion inherited from a first data source, and the second portion inherited from a second data source;

identifying a subset of the data-bit regions that are associated with the data manifestation;

determining a first source-specific association score representing a first measurement of association between the first data source and the data manifestation, wherein the first source-specific association score is determined based on the first portion of the inheritance dataset at the identified subset of the data-bit regions;

determining a second source-specific association score representing a second measurement of association between the second data source and the data manifestation, wherein the second source-specific association score is determined based on the second portion of the inheritance dataset at the identified subset of the data-bit regions;

comparing the first source-specific association score to the second source-specific association score; and

identifying at least one of the first data source and the second data source as having a measure of influence on the data manifestation of the named entity.

2. The method of claim 1 , wherein determining the first portion and the second portion of the inheritance dataset of the named entity comprises:

using a phasing algorithm to generate the first portion and the second portion of the inheritance dataset.

3. The method of claim 1 , wherein determining an association score based on the inheritance dataset at the identified subset of the data-bit regions comprises:

calculating a weight for each data-bit region of the identified subset of the data-bit regions based on a p-value score for the data-bit region; and

calculating an aggregated data-bit association score by summing over each product of data bit pattern at the data-bit region at the data-bit region and a corresponding weight for the data-bit region, wherein the aggregated data-bit association score is the association score.

4. The method of claim 1 , wherein identifying at least one of the first data source and the second data source as having the measure of influence on the data manifestation of the named entity comprises:

applying a machine learning model to predict the data manifestation of the named entity, wherein applying the machine learning model to predict the data manifestation of the named entity comprises inputting a feature vector comprising one or more association scores corresponding to the named entity into the machine learning model and receiving, from the machine learning model, an output indicating the data manifestation associated with the named entity; and

responsive to receiving the data manifestation associated with the named entity, determining one of the first and second data sources as having the measure of influence on the data manifestation associated with the named entity.

5. The method of claim 4 , wherein identifying at least one of the first data source and the second data source as having the measure of influence on the data manifestation of the named entity comprises:

selecting a threshold as a metric for discerning the measure of influence of each one of the first and second data sources;

comparing the source-specific association scores associated with the first and second data sources;

responsive to a difference between the source-specific association scores associated with the first and second data sources greater than the threshold, selecting the data source with the higher association score as having a dominant influence on the data manifestation of the named entity; and

responsive to the difference between the source-specific association scores between the first and second data sources falling within the threshold, selecting both data sources as having equal influence on the data manifestation of the named entity.

6. The method of claim 5 , wherein selecting the threshold as the metric for discerning the measure of influence of each one of the first and second data sources comprises:

determining, based on comparing the source-specific association scores of the first and second data sources, a first decile dataset and a second decile dataset, the first decile dataset representing a first decile of score similarity and the second decile dataset representing a second decile of score similarity;

determining data manifestation scores for each data source based on the first and second decile datasets, wherein a distribution of the data manifestation scores provides a third dataset representing the data sources with substantially similar association scores and a fourth dataset representing the data sources with substantially different association scores; and

setting the threshold based on the third dataset or a combination of the third and fourth datasets.

7. The method of claim 6 , wherein setting the threshold based on the third dataset or a combination of the third and fourth datasets comprises:

determining a mean of the third dataset; and

setting the threshold as one standard deviation below the determined mean.

8. The method of claim 4 , wherein the feature vector further comprises an age, sex, or ethnicity of the named entity to improve an accuracy of data manifestation predictions.

9. The method of claim 4 , wherein the machine learning model comprises any one of:

a logistic regression model;

an extreme gradient boosting; or

a gradient boosting machine.

10. The method of claim 4 , further comprising

training the machine learning model with training data comprising association scores of training named entities.

11. A system comprising:

one or more processors; and

memory configured to store instructions for predicting data-source influences on a data manifestation of a named entity, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform steps comprising:

receiving an inheritance dataset of the named entity, the inheritance dataset comprising one or more reads at a plurality of data-bit regions;

determining a first portion and a second portion of the inheritance dataset of the named entity, the first portion inherited from a first data source, and the second portion inherited from a second data source;

identifying a subset of the data-bit regions that are associated with the data manifestation;

determining a first source-specific association score representing a first measurement of association between the first data source and the data manifestation, wherein the first source-specific association score is determined based on the first portion of the inheritance dataset at the identified subset of the data-bit regions;

determining a second source-specific association score representing a second measurement of association between the second data source and the data manifestation, wherein the second source-specific association score is determined based on the second portion of the inheritance dataset at the identified subset of the data-bit regions;

comparing the first source-specific association score to the second source-specific association score; and

identifying at least one of the first data source and the second data source as having a measure of influence on the data manifestation of the named entity.

12. The system of claim 11 , wherein determining the first portion and the second portion of the inheritance dataset of the named entity comprises:

using a phasing algorithm to generate the first portion and the second portion of the inheritance dataset.

13. The system of claim 11 , wherein determining an association score based on the inheritance dataset at the identified subset of the data-bit regions comprises:

calculating a weight for each data-bit region of the identified subset of the data-bit regions based on a p-value score for the data-bit region; and

calculating an aggregated data-bit association score by summing over each product of data bit pattern at the data-bit region at the data-bit region and a corresponding weight for the data-bit region, wherein the aggregated data-bit association score is the association score.

14. The system of claim 11 , wherein identifying at least one of the first data source and the second data source as having the measure of influence on the data manifestation of the named entity comprises:

applying a machine learning model to predict the data manifestation of the named entity, wherein applying the machine learning model to predict the data manifestation of the named entity comprises inputting a feature vector comprising one or more association scores corresponding to the named entity into the machine learning model and receiving, from the machine learning model, an output indicating the data manifestation associated with the named entity; and

responsive to receiving the data manifestation associated with the named entity, determining one of the first and second data sources as having the measure of influence on the data manifestation associated with the named entity.

15. The system of claim 14 , wherein identifying at least one of the first data source and the second data source as having the measure of influence on the data manifestation of the named entity comprises:

selecting a threshold as a metric for discerning the measure of influence of each one of the first and second data sources;

comparing the source-specific association scores associated with the first and second data sources;

responsive to a difference between the source-specific association scores associated with the first and second data sources greater than the threshold, selecting the data source with the higher association score as having a dominant influence on the data manifestation of the named entity; and

responsive to the difference between the source-specific association scores between the first and second data sources falling within the threshold, selecting both data sources as having equal influence on the data manifestation of the named entity.

16. The system of claim 15 , wherein selecting the threshold as the metric for discerning the measure of influence of each one of the first and second data sources comprises:

determining, based on comparing the source-specific association scores of the first and second data sources, a first decile dataset and a second decile dataset, the first decile dataset representing a first decile of score similarity and the second decile dataset representing a second decile of score similarity;

determining data manifestation scores for each data source based on the first and second decile datasets, wherein a distribution of the data manifestation scores provides a third dataset representing the data sources with substantially similar association scores and a fourth dataset representing the data sources with substantially different association scores; and

setting the threshold based on the third dataset or a combination of the third and fourth datasets.

17. The system of claim 16 , wherein setting the threshold based on the third dataset or a combination of the third and fourth datasets comprises:

determining a mean of the third dataset; and

setting the threshold as one standard deviation below the determined mean.

18. The system of claim 14 , wherein the feature vector further comprises an age, sex, or ethnicity of the named entity to improve an accuracy of data manifestation predictions.

19. The system of claim 14 , wherein the machine learning model comprises any one of:

a logistic regression model;

an extreme gradient boosting; or

a gradient boosting machine.

20. A non-transitory computer readable medium for storing computer code comprising instructions, when executed by one or more computer processors, causing one or more computer processors to perform steps comprising:

receiving an inheritance dataset of the named entity, the inheritance dataset comprising one or more reads at a plurality of data-bit regions;

determining a first portion and a second portion of the inheritance dataset of the named entity, the first portion inherited from a first data source, and the second portion inherited from a second data source;

identifying a subset of the data-bit regions that are associated with the data manifestation;

determining a first source-specific association score representing a first measurement of association between the first data source and the data manifestation, wherein the first source-specific association score is determined based on the first portion of the inheritance dataset at the identified subset of the data-bit regions;

determining a second source-specific association score representing a second measurement of association between the second data source and the data manifestation, wherein the second source-specific association score is determined based on the second portion of the inheritance dataset at the identified subset of the data-bit regions;

comparing the first source-specific association score to the second source-specific association score; and

identifying at least one of the first data source and the second data source as having a measure of influence on the data manifestation of the named entity.

Assignments (3)
PATENT SECURITY AGREEMENT Recorded Aug 3, 2026
From: ANCESTRY.COM OPERATIONS INC.; ANCESTRY.COM DNA, LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 076116/0447 →
PATENT SECURITY AGREEMENT Recorded Aug 3, 2026
From: ANCESTRY.COM OPERATIONS INC.; ANCESTRY.COM DNA, LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 076144/0726 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2024
From: KIM, ANDRE EVERSON; SEDGHIFAR, ALISA ELNAZ; CURTIS, ROSS EUGENE; BRUNS, CAITLYN ELIZABETH
To: ANCESTRY.COM DNA, LLC
Reel/Frame 067987/0782 →
Continuity (2)
Provisional Application 63511084 · Jun 29, 2023
Related Publication 20250005108A1 · Jan 2, 2025
References Cited (153)
US 4201386A · Seale et al. · 1980 [cited by applicant]
US 5115504A · Belove et al. · 1992 [cited by applicant]
US 5246374A · Boodram · 1993 [cited by applicant]
US 5413908A · Jeffreys · 1995 [cited by applicant]
US 5467471A · Bader · 1995 [cited by applicant]
US 5978811A · Smiley · 1999 [cited by applicant]
US 6049803A · Szalwinski · 2000 [cited by applicant]
US 6105147A · Molloy · 2000 [cited by applicant]
US 6277567B1 · Graziosi · 2001 [cited by applicant]
US 6528260B1 · Blumenfeld et al. · 2003 [cited by applicant]
US 6570567B1 · Eaton · 2003 [cited by applicant]
US 6886015B2 · Notargiacomo et al. · 2005 [cited by applicant]
US 6950753B1 · Rzhetsky et al. · 2005 [cited by applicant]
US 7957907B2 · Sorenson et al. · 2011 [cited by applicant]
US 8185557B2 · Slinker · 2012 [cited by applicant]
US 8224821B2 · Graham et al. · 2012 [cited by applicant]
US 8510057B1 · Avey et al. · 2013 [cited by applicant]
US 8738297B2 · Sorenson et al. · 2014 [cited by applicant]
US 8855935B2 · Myres et al. · 2014 [cited by applicant]
US 9116882B1 · Macpherson et al. · 2015 [cited by applicant]
US 9213947B1 · Do et al. · 2015 [cited by applicant]
US 9367800B1 · Do et al. · 2016 [cited by applicant]
US 9836576B1 · Do et al. · 2017 [cited by applicant]
US 9940433B2 · Han et al. · 2018 [cited by applicant]
US 10025877B2 · Macpherson · 2018 [cited by applicant]
US 10223498B2 · Han et al. · 2019 [cited by applicant]
US 11238957B2 · Byrnes et al. · 2022 [cited by applicant]
US 20020143578A1 · Cole et al. · 2002 [cited by applicant]
US 20030032015A1 · Toivonen et al. · 2003 [cited by applicant]
US 20030113727A1 · Girn et al. · 2003 [cited by applicant]
US 20030113756A1 · Mertz · 2003 [cited by applicant]
US 20030172065A1 · Sorenson et al. · 2003 [cited by applicant]
US 20030195707A1 · Schork et al. · 2003 [cited by applicant]
US 20030204418A1 · Ledley · 2003 [cited by applicant]
US 20040122705A1 · Sabol et al. · 2004 [cited by applicant]
US 20040229231A1 · Frudakis et al. · 2004 [cited by applicant]
US 20040243531A1 · Dean · 2004 [cited by applicant]
US 20050147947A1 · Cookson et al. · 2005 [cited by applicant]
US 20050149522A1 · Cookson et al. · 2005 [cited by applicant]
US 20060020398A1 · Vernon et al. · 2006 [cited by applicant]
US 20060136143A1 · Avinash et al. · 2006 [cited by applicant]
US 20060161535A1 · Holbrook · 2006 [cited by applicant]
US 20070037182A1 · Gaskin et al. · 2007 [cited by applicant]
US 20080027656A1 · Parida · 2008 [cited by applicant]
US 20080154566A1 · Myres et al. · 2008 [cited by applicant]
US 20080228751A1 · Kenedy et al. · 2008 [cited by applicant]
US 20080255768A1 · Martin et al. · 2008 [cited by applicant]
US 20090100030A1 · Isakson et al. · 2009 [cited by applicant]
US 20090299645A1 · Colby et al. · 2009 [cited by applicant]
US 20100218228A1 · Walter · 2010 [cited by applicant]
US 20110093448A1 · Rafi et al. · 2011 [cited by applicant]
US 20110161168A1 · Dubnicki · 2011 [cited by applicant]
US 20120218289A1 · Rasmussen et al. · 2012 [cited by applicant]
US 20130085728A1 · Tang et al. · 2013 [cited by applicant]
US 20130149707A1 · Sorenson et al. · 2013 [cited by applicant]
US 20140067355A1 · Noto et al. · 2014 [cited by applicant]
US 20140108527A1 · Aravanis et al. · 2014 [cited by applicant]
US 20140278138A1 · Barber et al. · 2014 [cited by applicant]
US 20150100243A1 · Myres et al. · 2015 [cited by applicant]
US 20160350479A1 · Han et al. · 2016 [cited by applicant]
US 20170011042A1 · Kermany et al. · 2017 [cited by applicant]
US 20170277827A1 · Granka et al. · 2017 [cited by applicant]
US 20170329891A1 · Macpherson et al. · 2017 [cited by applicant]
US 20190034587A1 · Anderson et al. · 2019 [cited by applicant]
US 20190147973A1 · Han et al. · 2019 [cited by applicant]
US 20210057041A1 · Byrnes et al. · 2021 [cited by applicant]
CN 109121436A · 2019 [cited by examiner]
WO WO2008042232A2 · 2008 [cited by applicant]
WO WO2016073953A1 · 2016 [cited by applicant]
WO WO2016193891A1 · 2016 [cited by applicant]
Alexander, D.H. et al., “Enhancements to the ADMIXTURE Algorithm for Individual Ancestry Estimation,” BMC Bioinformatics 12(1), 246, Jun. 18, 2011, pp. 1-6. [cited by applicant]
Alexander, D.H., et al., “Fast model-based estimation of ancestry in unrelated individuals,” Genome research, Sep. 2009, vol. 19, No. 9, pp. 1655-1664. [cited by applicant]
Atzmon, G. et al., “Abraham's Children in the Genome Era: Major Jewish Diaspora Populations Comprise Distinct Genetic Clusters with Shared Middle Eastern Ancestry,” American Journal of Human Genetics 86(6), Jun. 11, 201… [cited by applicant]
Belkin, M. et al., “Laplacian Eigenmaps for Dimensionality Reduction and Data Representation,” Neural Computation 15, Jun. 2003, pp. 1373-1396. [cited by applicant]
Bengio, Y. et al., “Out-of-Sample Extensions for LLE, Isomap, MOS, Eigenmaps and Spectral Clustering,” NIPS'03: Proceedings of the 16th International Conference on Neural Information Processing Systems, Dec. 2003, pp. 1… [cited by applicant]
Blondel, V.D. et al., “Fast Unfolding of Community Hierarchies in Large Networks,” arXiv Preprint arXiv:0803.0476v1, Mar. 4, 2008, pp. 1-6. [cited by applicant]
Browning, B.L. et al., “Efficient Multilocus Association Testing for Whole Genome Association Studies Using Localized Haplotype Clustering,” Genetic Epidemiology, 2007, Vo. 31, pp. 365-375. [cited by applicant]
Browning, S.R. et al., “High-resolution detection of Identity by Descent in unrelated individuals,” The American Journal of Human Genetics, vol. 86, Apr. 9, 2010, pp. 526-539. [cited by applicant]
Browning, S.R. et al., “Rapid and Accurate Haplotype Phasing and Missing-Data Inference for Whole-Genome Association Studies by Use of Localized Haplotype Clustering,” The American Journal of Human Genetics, Nov. 2007, … [cited by applicant]
Browning, S.R., “Multilocus Association Mapping Using Variable-Length Markov Chains,” The American Journal of Human Genetics, Jun. 2006, vol. 78, pp. 903-913. [cited by applicant]
Butler, John M., “Commonly Used Short Tandem Repeat Markers,” Forensic DNA Typing, Chapter 5, 2001, pp. 53-54, Academic Press. [cited by applicant]
Cann, H.M. et al., “A human genome diversity cell line panel,” Science, Apr. 2002, vol. 296, No. 5566, pp. 261-262. [cited by applicant]
Capocci, A., et al. “Detecting communities in large networks,” Physica A: Statistical Mechanics and its Applications, vol. 352, Nos. 2-4, Jul. 15, 2005, 2005, pp. 669-676. [cited by applicant]
Carmi, S. et al., “Sequencing an Ashkenazi Reference Panel Supports Population-Targeted Personal Genomics and Illuminates Jewish and European Origins,” Nature Communications 5:4835, Sep. 9, 2014, pp. 1-9. [cited by applicant]
Carmi, S. et al., “The Variance of Identity-by-Descent Sharing in the Wright-Fisher Model,” Genetics 193(3), Mar. 2013, pp. 911-928. [cited by applicant]
Cavalli-Sforza, L.L. “The human genome diversity project: past, present and future,” Nature Reviews Genetics, Apr. 2005, vol. 6, No. 4. pp. 333-340. [cited by applicant]
Coifman, R.R. et al., “Diffusion Maps,” Applied and Computational Harmonic Analysis, vol. 21, No. 1, Jul. 2006, pp. 5-30. [cited by applicant]
Corach et al., “Mass disasters: Rapid molecular screening of human remains by means of short tandem repeats typing,” Electrophoresis (1995) vol. 16, pp. 1617-1623. [cited by applicant]
Curtis, R.E. et al., “Estimation of recent ancestral origins of individuals on a large scale,” KDD '17, Aug. 2017, pp. 1417-1425. [cited by applicant]
Durand, E.Y. et al., “Reducing Pervasive False-Positive Identical-by-Descent Segments Detected by Large-Scale Pedigree Analysis,” Molecular Biology and Evolution 31(8), Apr. 30, 2014, pp. 2212-2222. [cited by applicant]
European Patent Office, Extended European Search Report and Opinion, EP Patent Application No. 19782160.6, dated Nov. 12, 2021, nine pages. [cited by applicant]
Falush, D. et al., “Inference of Population Structure Using Multilocus Genotype Data: Linked Loci and Correlated Allele Frequencies,” Genetics Society of America, 2003, vol. 164, pp. 1567-1587. [cited by applicant]
Family Tree DNA. “Family Tree DNA.” Family Tree DNA: Genealogy by Genetics, Ltd., Feb. 5, 2001, 2 pages, [Online] [Retrieved Aug. 1, 2023], Retrieved from the Internet Archive <URL:https://web.archive.org/web/2001020500… [cited by applicant]
Fortunato, S. et al., “Resolution Limit in Community Detection,” Proceedings of the National Academy of Sciences 104(1), Jan. 2, 2007, pp. 36-41. [cited by applicant]
Fortunato, S., “Community Detection in Graphs,” Physics Reports 486, No. 3-5, Feb. 2010, pp. 75-174. [cited by applicant]
Francioli, L. et al., “Whole-Genome Sequence Variation, Population Structure and Demographic History of the Dutch Population,” Nature Genetics 46(8), Aug. 2014, pp. 818-825. [cited by applicant]
Gauvin, H. et al., “Genome-Wide Patterns of Identity-by-Descent Sharing in the French Canadian Founder Population,” European Journal of Human Genetics 22, Oct. 16, 2013, pp. 814-821. [cited by applicant]
Genealogy Blog With Attitude, “Surnames, 23andMe, and AncestryDNA: Making the Most of Match Counts and “Enrichment”—Genealogy and Genomics,” Apr. 4, 2015, 24 pages, [Online] Retrieved Jan. 23, 2019, Retrieved from the I… [cited by applicant]
Genealogy definition, Merriam-Webster Online Dictionery, 2004, http://www.mw.com/cqibin/dictionary?book=Dictionary&va=genealogy (1 page). [cited by applicant]
Girvan, M. et al., “Community Structure in Social and Biological Networks,” PNAS 99(12), Jun. 11, 2002, pp. 7821-7826. [cited by applicant]
Good, B.H. et al., “Performance of Modularity Maximization in Practical Contexts,” Physical Review E 81(4), Apr. 15, 2010, pp. 1-19. [cited by applicant]
Gusev, A. et al., “Whole Population, Genome-Wide Mapping of Hidden Relatedness,” Genome Research, 2009, pp. 318-326, vol. 19. [cited by applicant]
Han, L. et al., “Identity by descent estimation with dense genome-wide genotype data,” Genet Epidemiol., vol. 35, No. 6, Sep. 2011, pp. 557-567. [cited by applicant]
Hao, W. et al., “Probabilistic Models of Genetic Variation in Structured Populations Applied to Global Human Studies,” arXiv:1312.2041, Dec. 7, 2013, pp. 1-35. [cited by applicant]
Jarvis, J.P. et al., “Patterns of Ancestry, Signatures of Natural Selection, and Genetic Association with Stature in Western African Pygmies,” PLoS Genetics, 2012, vol. 8, No. 4, pp. 1-15. [cited by applicant]
King, T et al., “What's in a Name? Y Chromosomes, Surnames and the Genetic Genealogy Revolution,” Trends in Genetics, 2009, pp. 351-360, vol. 25, No. 8. [cited by applicant]
Lee, A.B. et al., “A Spectral Graph Approach to Discovering Genetic Ancestry,” Annals of Applied Statistics 4(1), Mar. 2010, pp. 179-202. [cited by applicant]
Lee, A.B. et al., “Discovering Genetic Ancestry Using Spectral Graph Theory,” Genetic Epidemiology 34, May 19, 2009, pp. 51-59. [cited by applicant]
McGraw, P.N. et al., “Laplacian Spectra as a Diagnostic Tool for Network Structure and Dynamics,” arXiv Preprint arXiv:0708.4206v1, Aug. 30, 2007, pp. 1-13. [cited by applicant]
McVean, G., “A Genealogical Interpretation of Principal Components Analysis,” PLoS Genetics 5(10), Oct. 16, 2009, pp. 1-10. [cited by applicant]
Meirmans, P.G., “The Trouble with Isolation by Distance,” Molecular Ecology 21(12), May 11, 2012, pp. 2839-2846. [cited by applicant]
Morrison, AC et al., “Prediction of Coronary Heart Disease Risk using a Genetic Risk Score: The Atherosclerosis Risk in Communities Study,” American Journal of Epidemiology, 2007, vol. 166, No. 1, pp. 28-35. [cited by applicant]
Newman, M.E.J., “Communities, Modules and Large-Scale Structure in Networks,” Nature Physics 8(1), Jan. 2012, pp. 25-31. [cited by applicant]
Noto, K. et al., “Underdog: A Fully-Supervised Phasing Algorithm That Learns from Hundreds of Thousands of Samples and Phases in Minutes,” Oct. 20, 2014, 1 page. [cited by applicant]
Novembre, J. et al., “Genes Mirror Geography within Europe,” Nature 456(7218), Nov. 2008, pp. 98-101. [cited by applicant]
Oxford Ancestors. “Oxford Ancestors: We Put the Genes in Genealogy.” Oxfordancestors.com, Feb. 24, 2001, 3 pages, [Online] [Retrieved Aug. 1, 2023], Retrieved from the Internet Archive <URL:https://web.archive.org/web/2… [cited by applicant]
Palamara, P.F. et al., “Inference of Historical Migration Rates Via Haplotype Sharing,” Bioinformatics 29(13), Jul. 2013, pp. 180-188. [cited by applicant]
Palamara, P.F. et al., “Length Distributions of Identity by Descent Reveal Fine-Scale Demographic History,” American Journal of Human Genetics 91(5), Nov. 2, 2012, pp. 809-822. [cited by applicant]
Palin, K. et al., “Identity-by-Descent-Based Phasing and Imputation in Founder Populations Using Graphical Models,” Genetic Epidemiology, vol. 35, Oct. 17, 2011, pp. 853-860. [cited by applicant]
Patterson, N. et al., “Population structure and eigenanalysis,” PLoS genetics, Dec. 2006, vol. 2, No. 12, pp. 2074-2093. [cited by applicant]
PCT International Search Report and Written Opinion, International Application No. PCT/IB2019/052788, dated Aug. 9, 2019, eight pages. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/IB2016/053166, dated Sep. 6, 2016, 11 pages. [cited by applicant]
Platt, J.C. et al., “Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods,” Microsoft Research, Mar. 26, 1999, pp. 1-11. [cited by applicant]
Price, Al. et al., “Sensitive Detection of Chromosomal Segments of Distinct Ancestry in Admixed Populations,” PLoS Genetics, 2009, vol. 5, No. 6, pp. 1-18. [cited by applicant]
Pritchard, J.K. et al., “Inference of population structure using multilocus genotype data,” Genetics Society of America, Jun. 2000, vol. 155, No. 2, pp. 945-959. [cited by applicant]
Pugh, M. B. et al. “Stedman's Medical Dictionary.” 27th Edition, 2000, p. 703. [cited by applicant]
Purcell, S. et al., “PLINK: A tool set for whole-genome association and population-based linkage analyses,” The American Journal of Human Genetics, vol. 81, Sep. 2007, pp. 559-575. [cited by applicant]
Purcell, S., “Plink (1.07) Documentation,” Harvard.edu, May 10, 2010, 293 pages, [Online] [Retrieved Oct. 14, 2024], Retrieved from the Internet <URL:http://zzz.bwh.harvard.edu/plink/dist/plink-doc-1.07.pdf.>. [cited by applicant]
Qian, Y. et al., “Efficient clustering of identity-by-descent between multiple individuals,” Bioinformatics, vol. 30, No. 7, Dec. 19, 2013, pp. 915-922. [cited by applicant]
Rabiner, L.R. et al., “A Tutorial on hidden Markov Models and Selected Application in Speech Recognition,” Proceedings of the IEEE, Feb. 1989, vol. 77, No. 2, pp. 257-286. [cited by applicant]
Raj, A. et al., “FastSTRUCTURE: Variational Inference of Population Structure in Large SNP Data Sets,” Genetics 197(2), Jun. 2014, pp. 573-589. [cited by applicant]
Ron, D. et al., “On the Learnability and Usage of Acyclic Probabilistic Finite Automata,” Journal of Computer and System Sciences, vol. 56, 1998, pp. 133-152. [cited by applicant]
Scheet, P. et al., “A Fast and Flexible Statistical Model for Large-Scale Population Genotype Data: Applications to Inferring Missing Genotypes and Haplotypic Phase,” The American journal of Human Genetics, Apr. 2006, v… [cited by applicant]
Seligsohn, U. et al., “Genetic Susceptibility to Venous Thrombosis,” The New England Journal of Medicine, Apr. 19, 2001, vol. 344, No. 16, pp. 1222-1231. [cited by applicant]
Staples, J. et al., “PRIMUS: Rapid Reconstruction of Pedigrees from Genome-wide Estimates of Identity by Descent,” The American Journal of Human Genetics, vol. 95, Nov. 6, 2014, pp. 553-564. [cited by applicant]
Stevens, E. L., et al., “Inference of Relationships in Population Data Using Identity-by-Descent,” PLOS Genetics, vol. 7, Iss. 9, Sep. 22, 2011, pp. 1-15. [cited by applicant]
Sundquist, A et al., “Effect of genetic divergence in identifying ancestral origin using HAPPA,” Genome Research, 2008, Vo. 18, 8 pages. [cited by applicant]
The 1000 Genomes Project Consortium, “An Integrated Map of Genetic Variation from 1,092 Human Genomes,” Nature 491, Nov. 2012, pp. 56-65. [cited by applicant]
The International Hapmap 3 Consortium. “Integrating common and rare genetic variation in diverse human populations,” Nature, Sep. 2, 2010, vol. 467, pp. 52-58. [cited by applicant]
The International Hapmap Consortium, “A haplotype map of the human genome,” Nature, Oct. 2005, vol. 437, No. 27, pp. 1299-1320. [cited by applicant]
The International Hapmap Consortium, “A second generation human haplotype map of over 3.1 million SNPs,” Nature, Oct. 2007, vol. 449, No. 7164, pp. 1-30. [cited by applicant]
Tipping, M.E., “Sparse Bayesian Learning and the Relevance Vector Machine,” Journal of Machine Learning Research, 2001, vol. 1, pp. 211-244. [cited by applicant]
Von Luxburg, U., “A Tutorial on Spectral Clustering,” Statistics and Computing 17(4), Aug. 22, 2007, pp. 1-32. [cited by applicant]
Weedon, M.N. et al., “Combining Information from Common Type 2 Diabetes Risk Polymorphisms Improves Disease Prediction,” PLoS Medicine, Oct. 2006, vol. 3, No. 10, pp. 1877-1882. [cited by applicant]
Welch, B.L., “The Generalization of “Student's” Problem when Several Different Population Variances are Involved,” Biometrika 34(1-2), 1947, pp. 28-35. [cited by applicant]
Wikipedia, “Identity by descent,” Wikipedia: The Free Encyclopedia, Feb. 13, 2015, 5 pages, [Online] [Retrieved Oct. 14, 2024], Retrieved from the Internet <URL:https://en.wikipedia.org/w/index.php?title=Identity_by_des… [cited by applicant]
Williams, A.L. et al., “Phasing of Many Thousands of Genotyped Samples,” The American Journal of Human Genetics, Aug. 10, 2012, vol. 91, pp. 238-251. [cited by applicant]
Wilson et al., “Genealogical Inference from Microsatellite Data,” Genetics (1998) 150:499-510. [cited by applicant]
Yang, Q. et al., “Improving the Prediction of Complex Diseases by Testing for Multiple Disease—Susceptibility Genes,” American Journal of Human Genetics, 2003, vol. 72, pp. 636-649. [cited by applicant]
Yoon, B.J., “Hidden Markov Models and their Applications in Biological Sequence Analysis,” Current Genomics, 2009, vol. 10, pp. 402-415. [cited by applicant]
Zelnik-Manor, L. et al., “Self-Tuning Spectral Clustering,” In Advances in Neural Information Processing Systems 17, Jan. 2004, pp. 1601-1608. [cited by applicant]
Zhang, J., “Ancestral Informative Marker Selection and Population Structure Visualization Using Sparse Laplacian Eigenfunctions,” PLoS One 5(11), Nov. 4, 2010, pp. 1-12. [cited by applicant]
Zhao, F. et al., “Spectral Clustering with Eigenvector Selection Based on Entropy Ranking,” Neurocomputing 73(10-12), Mar. 12, 2010, pp. 1704-1717. [cited by applicant]
Cited By (1)
US 12,737,988