IP Library › Granted Patent US 12,326,894
Granted Patent B2
US 12,326,894 · App. 18/736,429 · Granted Jun 10, 2025

Systems and methods for determining ethnicity subregions

Inventors: Alisa Elnaz Sedghifar (San Francisco, CA); Andre Everson Kim (Upland, CA); Ju Zhang (San Jose, CA); Ross Eugene Curtis (Cedar Hills, UT); Natalie Anne Swinford (Saratoga Springs, UT); Jeffrey Adrion (Salt Lake City, UT); Yong Wang (San Mateo, CA)
Assignee: Ancestry.com DNA, LLC
G06F16/35
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,326,894
App. No.
18/736,429
Filed
Jun 6, 2024
Granted
Jun 10, 2025
Kind
B2
Art Unit
2166
USPC
707/737
Abstract

A computing device may receive an inheritance dataset of a target named entity. The device may access a plurality of clusters associated with a region, each cluster comprising inheritance data for a plurality of reference panel named entities. The device may determine that the inheritance dataset of the target named entity has at least a threshold amount of inheritance sequences that are classified to the region. The device may compare, for each cluster, the inheritance dataset of the target named entity to the reference panel named entities in the cluster to identify similarities and shared inheritance segments between the target named entity and the reference panel named entities. The device may determine, for each cluster, a metric based on the inheritance segments shared. The device may assign the target named entity to one or more ethnicities based on the comparison between the metric and the threshold specific to the cluster.

Claims (80)

1. A computer-implemented method, comprising:

receiving an inheritance dataset of a target named entity;

accessing a plurality of clusters that are associated with a region, each cluster comprising inheritance data for a plurality of reference panel named entities;

determining that the inheritance dataset of the target named entity has at least a threshold amount of data strings that are classified to the region;

comparing, for each cluster, the inheritance dataset of the target named entity to the reference panel named entities in the cluster to identify shared data string segments between the target named entity and the reference panel named entities;

determining, for each cluster, a metric based on the data string segments shared between the target named entity and the reference panel named entities included in the cluster;

comparing, for each cluster, the metric to a threshold specific to the cluster; and

assigning the target named entity to one or more data origins based on the comparison between the metric and the threshold specific to each cluster.

2. The computer-implemented method of claim 1 , wherein accessing the plurality of clusters that are associated with the region comprises:

filtering samples to exclude samples from pre-determined regions;

organizing a genetic network formed of the filtered samples to identify the clusters;

excluding reference panel named entities in one or more clusters based on a distribution of matches; and

accessing the one or more clusters upon excluding the reference panel named entities.

3. The computer-implemented method of claim 2 , wherein filtering the samples to exclude the samples from the pre-determined regions comprises:

excluding samples with admixture samples that are associated with a plurality of data origins or samples with less than 95% of a pre-determined data origin.

4. The computer-implemented method of claim 2 , wherein organizing the genetic network formed of the filtered samples to identify the clusters comprises:

determining a modality of a plurality of candidate clusters, wherein the modality is based on genetic relatedness among samples within each candidate cluster; and

adjusting boundaries of the candidate clusters to increase the modality of the genetic network.

5. The computer-implemented method of claim 2 , wherein excluding the reference panel named entities in the one or more clusters based on the distribution of matches comprises:

removing named entities from the reference panel named entities based on tribe and/or language information to refine the one or more clusters.

6. The computer-implemented method of claim 1 , wherein comparing, for each cluster, the inheritance dataset of the target named entity to the reference panel named entities in the cluster to identify shared data string segments between the target named entity and the reference panel named entities comprises:

identifying matched segments shared between the target named entity and the reference panel named entities, wherein the matched segments comprise identity-by-descent (IBD) segments shared between the target named entity and one of the reference panel named entities.

7. The computer-implemented method of claim 1 , wherein determining, for each cluster, the metric based on the data string segments shared between the target named entity and the reference panel named entities included in the cluster comprises:

generating a metric based on a total length or an average length of genetic matches between the target named entity and the reference panel named entities.

8. The computer-implemented method of claim 1 , wherein comparing, for each cluster, the metric to the threshold specific to the cluster comprises:

conducting admixture simulations to generate inheritance data for named entities with known cluster components;

determining matches between the simulated inheritance data of the named entities and the cluster reference samples;

assessing a performance for each cluster based on the metric and different potential thresholds of the determined matches; and

selecting a suitable threshold for each cluster based on the assessment.

9. The computer-implemented method of claim 8 , wherein conducting the admixture simulations to generate the inheritance data for the named entities with the known cluster components comprises:

conducting the admixture simulations to generate single-origin inheritance data for named entities associated with one cluster; or

conducting the admixture simulations to generate inheritance data for named entities associated with two or more clusters.

10. The computer-implemented method of claim 1 , further comprising:

maintaining a hierarchy of data origin regions, the hierarchy including a data origin region and a plurality of data origin subregions associated with the data origin region;

determining that the inheritance dataset of the target named entity has at least a threshold amount of data strings that are classified to the data origin region;

identifying, based on the hierarchy, the plurality of data origin subregions associated with the data origin region; comparing, for each data origin subregion, the inheritance dataset of the target named entity to reference panel named entities of the data origin subregion to identify identity-by-descent (IBD) segments shared between the target named entity and the reference panel named entities of the data origin subregion;

generating, for each data origin subregion, a score based on the IBD segments shared between the target named entity and the reference panel named entities of the data origin subregion; and

assigning the target named entity to one of the data origin subregions based on generated scores.

11. A system comprising:

one or more processors; and

memory configured to store instructions, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform steps comprising:

receiving an inheritance dataset of a target named entity;

accessing a plurality of clusters that are associated with a region, each cluster comprising inheritance data for a plurality of reference panel named entities;

determining that the inheritance dataset of the target named entity has at least a threshold amount of data strings that are classified to the region;

comparing, for each cluster, the inheritance dataset of the target named entity to the reference panel named entities in the cluster to identify shared data string segments between the target named entity and the reference panel named entities;

determining, for each cluster, a metric based on the data string segments shared between the target named entity and the reference panel named entities included in the cluster;

comparing, for each cluster, the metric to a threshold specific to the cluster; and

assigning the target named entity to one or more data origins based on the comparison between the metric and the threshold specific to each cluster.

12. The system of claim 11 , wherein accessing the plurality of clusters that are associated with the region comprises:

filtering samples to exclude samples from pre-determined regions;

organizing a genetic network formed of the filtered samples to identify the clusters;

excluding reference panel named entities in one or more clusters based on a distribution of matches; and

accessing the one or more clusters upon excluding the reference panel named entities.

13. The system of claim 12 , wherein filtering the samples to exclude the samples from the pre-determined regions comprises:

excluding samples with admixture samples that are associated with a plurality of data origins or samples with less than 95% of a pre-determined data origin.

14. The system of claim 12 , wherein organizing the genetic network formed of the filtered samples to identify the clusters comprises:

determining a modality of a plurality of candidate clusters, wherein the modality is based on genetic relatedness among samples within each candidate cluster; and

adjusting boundaries of the candidate clusters to increase the modality of the genetic network.

15. The system of claim 12 , wherein excluding the reference panel named entities in the one or more clusters based on the distribution of matches comprises:

removing named entities from the reference panel named entities based on tribe and/or language information to refine the one or more clusters.

16. The system of claim 11 , wherein comparing, for each cluster, the inheritance dataset of the target named entity to the reference panel named entities in the cluster to identify shared data string segments between the target named entity and the reference panel named entities comprises:

identifying matched segments shared between the target named entity and the reference panel named entities, wherein the matched segments comprise identity-by-descent (IBD) segments shared between the target named entity and one of the reference panel named entities.

17. The system of claim 11 , wherein determining, for each cluster, the metric based on the data string segments shared between the target named entity and the reference panel named entities included in the cluster comprises:

generating a metric based on a total length or an average length of genetic matches between the target named entity and the reference panel named entities.

18. The system of claim 11 , wherein comparing, for each cluster, the metric to the threshold specific to the cluster comprises:

conducting admixture simulations to generate inheritance data for named entities with known cluster components;

determining matches between the simulated inheritance data of the named entities and the cluster reference samples;

assessing a performance for each cluster based on the metric and different potential thresholds of the determined matches; and

selecting a suitable threshold for each cluster based on the assessment.

19. The system of claim 18 , wherein conducting the admixture simulations to generate the inheritance data for the named entities with the known cluster components comprises:

conducting the admixture simulations to generate single-origin inheritance data for named entities associated with one cluster; or

conducting the admixture simulations to generate inheritance data for named entities associated with two or more clusters.

20. A non-transitory computer readable medium for storing computer code comprising instructions, when executed by one or more computer processors, causing one or more computer processors to perform steps comprising:

receiving an inheritance dataset of a target named entity;

accessing a plurality of clusters that are associated with a region, each cluster comprising inheritance data for a plurality of reference panel named entities;

determining that the inheritance dataset of the target named entity has at least a threshold amount of data strings that are classified to the region;

comparing, for each cluster, the inheritance dataset of the target named entity to the reference panel named entities in the cluster to identify shared data string segments between the target named entity and the reference panel named entities;

determining, for each cluster, a metric based on the data string segments shared between the target named entity and the reference panel named entities included in the cluster;

comparing, for each cluster, the metric to a threshold specific to the cluster; and

assigning the target named entity to one or more data origins based on the comparison between the metric and the threshold specific to each cluster.

Assignments (4)
PATENT SECURITY AGREEMENT Recorded Aug 3, 2026
From: ANCESTRY.COM OPERATIONS INC.; ANCESTRY.COM DNA, LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 076116/0447 →
PATENT SECURITY AGREEMENT Recorded Aug 3, 2026
From: ANCESTRY.COM OPERATIONS INC.; ANCESTRY.COM DNA, LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 076144/0726 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2025
From: WANG, YONG
To: ANCESTRY.COM DNA, LLC
Reel/Frame 071058/0878 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2024
From: SEDGHIFAR, ALISA ELNAZ; KIM, ANDRE EVERSON; ZHANG, JU; CURTIS, ROSS EUGENE; SWINFORD, NATALIE ANNE; ADRION, JEFFREY
To: ANCESTRY.COM DNA, LLC
Reel/Frame 069334/0690 →
Continuity (2)
Provisional Application 63506722 · Jun 7, 2023
Related Publication 20240411793A1 · Dec 12, 2024
References Cited (4)
US 8713434B2 · Ford · 2014 [cited by examiner]
US 9152705B2 · Lamba · 2015 [cited by examiner]
US 20070250487A1 · Reuther · 2007 [cited by examiner]
US 20220245154A1 · Gylfason · 2022 [cited by examiner]