IP Library Granted Patent US 12,332,902
Granted Patent B2
US 12,332,902 · App. 18/137,232 · Granted Jun 17, 2025

Filtering individual datasets in a database

Inventors: Milos Pavlovic (Sandy, UT); Ross Eugene Curtis (Cedar Hills, UT)
Assignee: Ancestry.com DNA, LLC
G06F16/2457G06N3/0464G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,332,902
App. No.
18/137,232
Filed
Apr 20, 2023
Granted
Jun 17, 2025
Kind
B2
Art Unit
2153
USPC
707/754
Abstract

A user of a genetic database may create and build upon their family tree in the database. For new users, creating a family tree can be difficult and time consuming. Even for users with established family trees, extending their family tree is a challenge requiring extensive research. Disclosed herein are embodiments for assisting users of a genetic database with building their family trees. In some embodiments, a method for assisting with constructing family trees includes receiving a target individual's genetic dataset. The method identifies a plurality of matched individuals who genetically match the target individual. The method identifies potential ancestors who are potential common ancestors between the target individual and one of the matched individuals. The method inputs a set of features related to the target individual to a machine learning model and filters the potential common ancestors to determine a subset of likely common ancestors for the target individual.

Claims (72)

1. A computer-implemented method, comprising:

receiving a target individual genetic dataset associated with a target individual;

identifying a plurality of matched individuals who genetically match with the target individual;

identifying a plurality of potential ancestors who are potential common ancestors between the target individual and one of the matched individuals;

inputting a set of features related to the target individual to a machine learning model, wherein training of the machine learning model comprises:

generating a plurality of training samples, wherein generating one of the plurality of training samples comprises:

identifying a training target individual, the training target individual having a genetic dataset,

identifying ancestors of the training target individual from one or more family trees of the training target individual,

determining, based on existing family trees that includes the training target individuals, whether each of the ancestors is a direct-line ancestor of the training target individual, and

assigning a positive label to a particular ancestor responsive to the particular ancestor being a direct-line ancestor of the training target individual; and

training the machine learning model using the plurality of training samples;

filtering the plurality of potential common ancestors using the machine learning model to identify a subset of the potential common ancestors of the target individual.

2. The computer-implemented method of claim 1 , wherein the target individual and the plurality of matched individuals are related by identity by descent (IBD).

3. The computer-implemented method of claim 1 , wherein a potential common ancestor of the plurality of potential common ancestors is documented in a family tree of one of the matched individuals.

4. The computer-implemented method of claim 1 , wherein identifying a matched individual of the plurality of matched individuals comprises:

receiving a candidate individual genetic dataset of a candidate individual;

identifying matched genetic segments between the candidate individual genetic dataset and the target individual;

measuring a total length of the matched genetic segments in centimorgans; and

classifying the candidate individual as the matched individual responsive to the total length exceeding a threshold.

5. The computer-implemented method of claim 4 , wherein the set of features includes the total length of the matched genetic segments in centimorgans.

6. The computer-implemented method of claim 1 , wherein the set of features includes a generation difference between a potential common ancestor and a matched individual.

7. The computer-implemented method of claim 1 , wherein the set of features includes an age difference between the target individual and a potential common ancestor.

8. The computer-implemented method of claim 1 , wherein the set of features includes a percentage of descendants in a family tree of a potential common ancestor who are matched individuals of the target individual.

9. The computer-implemented method of claim 1 , wherein the machine learning model is a supervised learning model.

10. The computer-implemented method of claim 1 , wherein the potential common ancestors in the subset of the potential common ancestors are more likely to be direct-line ancestors of the target individual than other potential common ancestors that are filtered out by the machine learning model.

11. The computer-implemented method of claim 1 , wherein the machine learning model is configured to provide at least a label that a potential common ancestor is a direct-line ancestor or not.

12. The computer-implemented method of claim 1 , wherein training of the machine learning model comprises:

generating a plurality of training samples, the plurality of training samples comprising a set of positive training samples and a set of negative training samples;

extracting training features of the plurality of training samples to generate a plurality of feature vectors, each feature vector corresponding to one of the training samples;

inputting the plurality of feature vectors to the machine learning model;

using the machine learning model to predict labels of one or more ancestors in the training samples;

determining an objective function that compares the predicted labels to actual labels of the training samples; and

adjusting one or more weights of the machine learning model based on the objective function.

13. A non-transitory computer readable medium configured to store computer code comprising instructions, the instructions, when executed by one or more processors, cause the one or more processors to perform steps comprising:

receiving a target individual genetic dataset associated with a target individual;

identifying a plurality of matched individuals who genetically match with the target individual;

identifying a plurality of potential ancestors who are potential common ancestors between the target individual and one of the matched individuals;

inputting a set of features related to the target individual to a machine learning model, wherein training of the machine learning model comprises:

generating a plurality of training samples, wherein generating one of the plurality of training samples comprises:

identifying a training target individual, the training target individual having a genetic dataset,

identifying ancestors of the training target individual from one or more family trees of the training target individual,

determining, based on existing family trees that includes the training target individuals, whether each of the ancestors is a direct-line ancestor of the training target individual, and

assigning a positive label to a particular ancestor responsive to the particular ancestor being a direct-line ancestor of the training target individual; and

training the machine learning model using the plurality of training samples; and

filtering the plurality of potential common ancestors using the machine learning model to identify a subset of the potential common ancestors of the target individual.

14. The non-transitory computer readable medium of claim 13 , wherein a potential common ancestor of the plurality of potential common ancestors is documented in a family tree of one of the matched individuals.

15. The non-transitory computer readable medium of claim 13 , wherein identifying a matched individual of the plurality of matched individuals comprises:

receiving a candidate individual genetic dataset of a candidate individual;

identifying matched genetic segments between the candidate individual genetic dataset and the target individual;

measuring a total length of the matched genetic segments in centimorgans; and

classifying the candidate individual as the matched individual responsive to the total length exceeding a threshold.

16. The non-transitory computer readable medium of claim 13 , wherein the machine learning model is a supervised learning model.

17. The non-transitory computer readable medium of claim 13 , wherein training of the machine learning model comprises:

generating a plurality of training samples, the plurality of training samples comprising a set of positive training samples and a set of negative training samples;

extracting training features of the plurality of training samples to generate a plurality of feature vectors, each feature vector corresponding to one of the training samples;

inputting the plurality of feature vectors to the machine learning model;

using the machine learning model to predict labels of one or more ancestors in the training samples;

determining an objective function that compares the predicted labels to actual labels of the training samples; and

adjusting one or more weights of the machine learning model based on the objective function.

18. A system comprising:

one or more processors; and

memory configured to store computer code comprising instructions, the instructions, when executed by one or more processors, cause the one or more processors to perform steps comprising:

receiving a target individual genetic dataset associated with a target individual;

identifying a plurality of matched individuals who genetically match with the target individual;

identifying a plurality of potential ancestors who are potential common ancestors between the target individual and one of the matched individuals;

inputting a set of features related to the target individual to a machine learning model, wherein training of the machine learning model comprises:

generating a plurality of training samples, wherein generating one of the plurality of training samples comprises:

identifying a training target individual, the training target individual having a genetic dataset,

identifying ancestors of the training target individual from one or more family trees of the training target individual,

determining, based on existing family trees that includes the training target individuals, whether each of the ancestors is a direct-line ancestor of the training target individual, and

assigning a positive label to a particular ancestor responsive to the particular ancestor being a direct-line ancestor of the training target individual, and training the machine learning model using the plurality of training samples; and

filtering the plurality of potential common ancestors using the machine learning model to identify a subset of the potential common ancestors of the target individual.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2023
From: PAVLOVIC, MILOS; CURTIS, ROSS EUGENE
To: ANCESTRY.COM DNA, LLC
Reel/Frame 064363/0883 →
Continuity (2)
Provisional Application 63332811 · Apr 20, 2022
Related Publication 20230342364A1 · Oct 26, 2023
References Cited (176)
US D169994S · Soffer et al. · 1953 [cited by applicant]
US D175257S · Hopkins · 1955 [cited by applicant]
US 2793776A · Lipari · 1957 [cited by applicant]
US D196112S · Esser · 1963 [cited by applicant]
US 3831742A · Gardella et al. · 1974 [cited by applicant]
US 4131016A · Layton · 1978 [cited by applicant]
US 4184483A · Greenspan · 1980 [cited by applicant]
US 4217798A · McCarthy et al. · 1980 [cited by applicant]
US 4301812A · Layton et al. · 1981 [cited by applicant]
US 4312950A · Snyder et al. · 1982 [cited by applicant]
US D277736S · Long · 1985 [cited by applicant]
US D286546S · Funahashi · 1986 [cited by applicant]
US 4935342A · Seligson et al. · 1990 [cited by applicant]
US 4982553A · Itoh · 1991 [cited by applicant]
US D330011S · Miller et al. · 1992 [cited by applicant]
US 5283038A · Seymour · 1994 [cited by applicant]
US 5393496A · Seymour · 1995 [cited by applicant]
US 5396986A · Fountain et al. · 1995 [cited by applicant]
US D362623S · Ma · 1995 [cited by applicant]
US 5714341A · Thieme et al. · 1998 [cited by applicant]
US D392187S · King · 1998 [cited by applicant]
US 5736322A · Goldstein · 1998 [cited by applicant]
US 5736355A · Dyke et al. · 1998 [cited by applicant]
US 5830154A · Goldstein et al. · 1998 [cited by applicant]
US 5830410A · Thieme et al. · 1998 [cited by applicant]
US D412107S · Bosshardt · 1999 [cited by applicant]
US 5927549A · Wood · 1999 [cited by applicant]
US 5933498A · Schneck et al. · 1999 [cited by applicant]
US 6003728A · Elliott · 1999 [cited by applicant]
US 6048091A · McIntyre et al. · 2000 [cited by applicant]
US 6152296A · Shih · 2000 [cited by applicant]
US D437786S · van Swieten et al. · 2001 [cited by applicant]
US 6228323B1 · Asgharian et al. · 2001 [cited by applicant]
US 6362473B1 · Germanus · 2002 [cited by applicant]
US 6428962B1 · Naegele · 2002 [cited by applicant]
US 6458546B1 · Baker · 2002 [cited by applicant]
US D470240S · Niedbala et al. · 2003 [cited by applicant]
US D471234S · Okutani · 2003 [cited by applicant]
US 6543612B2 · Lee et al. · 2003 [cited by applicant]
US 6548256B2 · Lienau et al. · 2003 [cited by applicant]
US D474280S · Niedbala et al. · 2003 [cited by applicant]
US 6627152B1 · Wong · 2003 [cited by applicant]
US 6760731B2 · Huff · 2004 [cited by applicant]
US 6786330B2 · Mollstam et al. · 2004 [cited by applicant]
US D507351S · Birnboim · 2005 [cited by applicant]
US 6939672B2 · Lentrichia et al. · 2005 [cited by applicant]
US 6992182B1 · Müller et al. · 2006 [cited by applicant]
US D515435S · Muehlhausen · 2006 [cited by applicant]
US 7055685B1 · Patterson et al. · 2006 [cited by applicant]
US D537416S · Fortin et al. · 2007 [cited by applicant]
US 7178683B2 · Birkmayer et al. · 2007 [cited by applicant]
US 7214484B2 · Weber et al. · 2007 [cited by applicant]
US 7297485B2 · Bornarth et al. · 2007 [cited by applicant]
US 7303876B2 · Greenfield et al. · 2007 [cited by applicant]
US D573465S · Kogure et al. · 2008 [cited by applicant]
US D574507S · Muir et al. · 2008 [cited by applicant]
US D584357S · Oka · 2009 [cited by applicant]
US 7482116B2 · Birnboim · 2009 [cited by applicant]
US D586856S · Yagyu · 2009 [cited by applicant]
US 7537132B2 · Marple et al. · 2009 [cited by applicant]
US 7544468B2 · Goldstein et al. · 2009 [cited by applicant]
US 7589184B2 · Hogan et al. · 2009 [cited by applicant]
US 7645424B2 · O'Donovan · 2010 [cited by applicant]
US D612730S · Rushe · 2010 [cited by applicant]
US 7748550B2 · Cho · 2010 [cited by applicant]
US 7854104B2 · Cronin et al. · 2010 [cited by applicant]
US 7858396B2 · Corstjens et al. · 2010 [cited by applicant]
US D631350S · Beach et al. · 2011 [cited by applicant]
US D631553S · Niedbala et al. · 2011 [cited by applicant]
US D640794S · Sunstrum et al. · 2011 [cited by applicant]
US D640795S · Jackson et al. · 2011 [cited by applicant]
US 7998757B2 · Darrigrand et al. · 2011 [cited by applicant]
US 8038668B2 · Scott et al. · 2011 [cited by applicant]
US 8062908B2 · Mink et al. · 2011 [cited by applicant]
US 8158357B2 · Birnboim et al. · 2012 [cited by applicant]
US 8221381B2 · Muir et al. · 2012 [cited by applicant]
US D673265S · Nonnemacher et al. · 2012 [cited by applicant]
US 8425864B2 · Haywood et al. · 2013 [cited by applicant]
US 8431384B2 · Hogan et al. · 2013 [cited by applicant]
US 8463554B2 · Hon et al. · 2013 [cited by applicant]
US 8470536B2 · Birnboim et al. · 2013 [cited by applicant]
US D693682S · Bahri et al. · 2013 [cited by applicant]
US 8673239B2 · Niedbala et al. · 2014 [cited by applicant]
US 8728414B2 · Beach et al. · 2014 [cited by applicant]
US 8855935B2 · Myres et al. · 2014 [cited by applicant]
US D718127S · Moriyama et al. · 2014 [cited by applicant]
US 9040675B2 · Bales et al. · 2015 [cited by applicant]
US 9072499B2 · Birnboim et al. · 2015 [cited by applicant]
US 9079181B2 · Curry et al. · 2015 [cited by applicant]
US D743044S · Jackson et al. · 2015 [cited by applicant]
US D743571S · Jackson et al. · 2015 [cited by applicant]
US 9207164B2 · Muir et al. · 2015 [cited by applicant]
US D757546S · Seifer · 2016 [cited by applicant]
US 9390225B2 · Barber et al. · 2016 [cited by applicant]
US 9410147B2 · Gundling · 2016 [cited by applicant]
US 9416356B2 · Gundling · 2016 [cited by applicant]
US 9523115B2 · Birnboim · 2016 [cited by applicant]
US D775953S · Ruthe-Steinsiek · 2017 [cited by applicant]
US D777111S · Zantout et al. · 2017 [cited by applicant]
US 9732376B2 · Oyler et al. · 2017 [cited by applicant]
US 9757179B2 · Formica · 2017 [cited by applicant]
US D811882S · Gundersen · 2018 [cited by applicant]
US 10000795B2 · Birnboim et al. · 2018 [cited by applicant]
US D843834S · Gundersen · 2019 [cited by applicant]
US D850647S · Jackson et al. · 2019 [cited by applicant]
US 10435735B2 · Birnboim et al. · 2019 [cited by applicant]
US 12217393B2 · Girshick · 2025 [cited by examiner]
US 20010041327A1 · Gross · 2001 [cited by applicant]
US 20020032687A1 · Huff · 2002 [cited by applicant]
US 20030089627A1 · Chelles et al. · 2003 [cited by applicant]
US 20030172065A1 · Sorenson et al. · 2003 [cited by applicant]
US 20040132091A1 · Ramsey et al. · 2004 [cited by applicant]
US 20050147947A1 · Cookson et al. · 2005 [cited by applicant]
US 20050267903A1 · Golze · 2005 [cited by applicant]
US 20060025929A1 · Eglington · 2006 [cited by applicant]
US 20060201948A1 · Ellson et al. · 2006 [cited by applicant]
US 20070168368A1 · Stone · 2007 [cited by applicant]
US 20070170142A1 · Cho · 2007 [cited by applicant]
US 20070178500A1 · Martin et al. · 2007 [cited by applicant]
US 20070218429A1 · Kolo et al. · 2007 [cited by applicant]
US 20070239802A1 · Razdow et al. · 2007 [cited by applicant]
US 20080027656A1 · Parida · 2008 [cited by applicant]
US 20080068401A1 · Albrecht et al. · 2008 [cited by applicant]
US 20080081331A1 · Myres et al. · 2008 [cited by applicant]
US 20080108027A1 · Sallin · 2008 [cited by applicant]
US 20080189047A1 · Wong et al. · 2008 [cited by applicant]
US 20080270431A1 · Garbero · 2008 [cited by applicant]
US 20090024060A1 · Darrigrand et al. · 2009 [cited by applicant]
US 20090216213A1 · Muir et al. · 2009 [cited by applicant]
US 20100049736A1 · Rolls et al. · 2010 [cited by applicant]
US 20100099149A1 · Birnboim et al. · 2010 [cited by applicant]
US 20100223281A1 · Hon et al. · 2010 [cited by applicant]
US 20100258457A1 · Seelhofer · 2010 [cited by applicant]
US 20110020195A1 · Luotola · 2011 [cited by applicant]
US 20110207621A1 · Montagu et al. · 2011 [cited by applicant]
US 20110212002A1 · Curry et al. · 2011 [cited by applicant]
US 20120024861A1 · Otsuka et al. · 2012 [cited by applicant]
US 20120024862A1 · Otsuka et al. · 2012 [cited by applicant]
US 20120046574A1 · Skakoon · 2012 [cited by applicant]
US 20120061392A1 · Beach et al. · 2012 [cited by applicant]
US 20130092690A1 · Skakoon · 2013 [cited by applicant]
US 20130149707A1 · Sorenson et al. · 2013 [cited by applicant]
US 20130164738A1 · Becker et al. · 2013 [cited by applicant]
US 20140006433A1 · Hon et al. · 2014 [cited by applicant]
US 20140025308A1 · Jorde et al. · 2014 [cited by applicant]
US 20140278138A1 · Barber et al. · 2014 [cited by applicant]
US 20140316302A1 · Nonnemacher et al. · 2014 [cited by applicant]
US 20150056716A1 · Oyler et al. · 2015 [cited by applicant]
US 20160262679A1 · Ivosevic et al. · 2016 [cited by applicant]
US 20160350479A1 · Han et al. · 2016 [cited by applicant]
US 20170001191A1 · Biadillah et al. · 2017 [cited by applicant]
US 20170072393A1 · Jackson et al. · 2017 [cited by applicant]
US 20170130219A1 · Birnboim et al. · 2017 [cited by applicant]
US 20170166955A1 · Birnboim et al. · 2017 [cited by applicant]
US 20170213127A1 · Duncan · 2017 [cited by applicant]
US 20170226469A1 · Birnboim et al. · 2017 [cited by applicant]
US 20170228498A1 · Hon et al. · 2017 [cited by applicant]
US 20170277827A1 · Granka et al. · 2017 [cited by applicant]
US 20170329891A1 · Macpherson et al. · 2017 [cited by applicant]
US 20190151842A1 · Williams et al. · 2019 [cited by applicant]
US 20190194742A1 · Wu · 2019 [cited by examiner]
US 20190210778A1 · Muir et al. · 2019 [cited by applicant]
US 20190358628A1 · Curry et al. · 2019 [cited by applicant]
US 20200135296A1 · Girshick · 2020 [cited by examiner]
US 20200273542A1 · Song et al. · 2020 [cited by applicant]
US 20200380015A1 · Gray · 2020 [cited by applicant]
US 20210082167A1 · Jewett · 2021 [cited by examiner]
EP 2370929A1 · 2011 [cited by applicant]
EP 3276526A1 · 2018 [cited by applicant]
KR 1020090022519A · 2009 [cited by applicant]
WO WO2010077336A1 · 2010 [cited by applicant]
Glodzik, D. et al. “Inference of Identity by Descent in Population Isolates and Optimal Sequencing Studies.” European Journal of Human Genetics, vol. 21, Jan. 30, 2013, pp. 1140-1145. [cited by applicant]
Gusev, A. et al. “Whole Population, Genome-Wide Mapping of Hidden Relatedness.” Genome Research, vol. 19, No. 2, Oct. 29, 2008, pp. 318-326. [cited by applicant]
Huff, C. D. et al. “Maximum-Likelihood Estimation of Recent Shared Ancestry (ERSA).” Genome Research, vol. 21, Feb. 8, 2011, pp. 768-774. [cited by applicant]
Li, X. et al. “Efficient Identification of Identical-by-Descent Status in Pedigrees with Many Untyped Individuals.” vol. 26, No. 12, Jun. 2010, pp. i191-i198. [cited by applicant]
Meulenbelt I. et al. “High-Yield Noninvasive Human Genomic DNA Isolation Method for Genetic Studies in Geographically Dispersed Families and Populations.” American Journal of Human Genetics, vol. 57, No. 5, Nov. 1995, p… [cited by applicant]