IP Library › Granted Patent US 12,461,970
Granted Patent B2
US 12,461,970 · App. 18/235,465 · Granted Nov 4, 2025

Catalog-based data inheritance determination

Inventor: Keith D. Noto (San Francisco, CA)
G06F16/90344
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,461,970
App. No.
18/235,465
Granted
Nov 4, 2025
Kind
B2
Abstract

A computing server may generate a catalog of overrepresented data strings from a database that stores a plurality of data instances. An overrepresented data string is a data string that matches to a number of data instances and the number exceeds a number threshold. The computing server may receive a target data instance that is to be compared to a related data instance. The computing server may determine one or more matched data strings that match between the target data instance and the related data instance. The computing server may compare the matched data strings to the catalog to exclude a subset of matched data strings that are matched to the overrepresented data strings. The computing server may determine a total length of the matched data strings excluding the subset of matched data strings that are matched to the overrepresented data strings.

Claims (55)

1 . A computer-implemented method for determining a normalized data inheritance between two data instances, the computer-implemented method comprising:

generating a catalog of overrepresented haplotypes from a database that stores a plurality of data instances in a population, wherein an overrepresented haplotype is a haplotype that matches to a number of data instances in the population and the number exceeds a number threshold, wherein the catalog of overrepresented haplotypes comprises a version of haplotype string sequences and corresponding tallies of the haplotype string sequences in the population, and wherein the number threshold is based on an overall distribution of haplotype sequence repetitiveness;

receiving a target data instance that is to be compared to a related data instance;

determining one or more matched haplotypes that match between the target data instance and the related data instance;

comparing the one or more matched haplotypes to the catalog of overrepresented haplotypes to exclude a subset of the one or more matched haplotypes that are matched to the overrepresented haplotypes;

determining a normalized data inheritance between the target data instance and the related data instance, the normalized data inheritance corresponding to a total length of the one or more matched haplotypes excluding the subset of the one or more matched haplotypes that are matched to the overrepresented haplotypes; and

storing, responsive to the total length being longer than a length threshold, information regarding the target data instance and the related data instance being a pair of identity-by-descent matches.

2 . The computer-implemented method of claim 1 , wherein the overrepresented haplotypes in the catalog are arranged by windows of data locality within which the overrepresented haplotypes are located.

3 . The computer-implemented method of claim 2 , wherein the catalog comprises a plurality of overrepresented haplotypes in one of the windows of data locality.

4 . The computer-implemented method of claim 1 , wherein generating the catalog comprises:

receiving the plurality of data instances that correspond to named entities, wherein each of the plurality of data instances is phased;

dividing each data instance into a plurality of windows of data locality;

tallying, for each window of data locality, the data instance that has a particular data bit sequence; and

determining whether a tally of the particular data bit sequence exceeds the number threshold.

5 . The computer-implemented method of claim 1 , wherein information regarding the target data instance and the related data instance being a pair of matches includes the total length and an indication that the target data instance and the related data instance are related by inheritance of a real-life event.

6 . The computer-implemented method of claim 1 , wherein the catalog is generated before receiving the target data instance and is stored in a second database.

7 . The computer-implemented method of claim 1 , wherein comparing the matched haplotypes to the catalog to exclude the subset of matched haplotypes that are matched to the overrepresented haplotypes comprises:

dividing the target data instance into a plurality of windows of data locality;

receiving, from the catalog, a subset of the windows of data locality that include the overrepresented haplotypes; and

comparing the target data instance and the related data instance in the windows that are not under the subset.

8 . The computer-implemented method of claim 1 , wherein determining one or more matched haplotypes that match between the target data instance and the related data instance is based on one or more weak matches of haplotypes.

9 . The computer-implemented method of claim 1 , wherein determining one or more matched haplotypes that match between the target data instance and the related data instance comprises:

phasing the target data instance into a pair of data sequences;

comparing the pair of data sequences with data sequences in the related data instance.

10 . The computer-implemented method of claim 1 , wherein the database that stores the plurality of data instances comprises over 10,000 data instances and each data instance includes over 10,000 data bits.

11 . A system for determining a normalized data inheritance between two data instances, the system comprising:

a data store configured to store a catalog of overrepresented haplotypes among a plurality of data instances in a population, wherein an overrepresented haplotype is a haplotype that matches to a number of data instances in the population and the number exceeds a number threshold, wherein the catalog of overrepresented haplotypes comprises a version of haplotype string sequences and corresponding tallies of the haplotype string sequences in the population, and wherein the number threshold is based on an overall distribution of haplotype sequence repetitiveness;

a computing server comprising memory and one or more processors, the memory configured to store code comprising instructions, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform steps comprising:

receiving a target data instance that is to be compared to a related data instance;

determining one or more matched haplotypes that match between the target data instance and the related data instance;

comparing the one or more matched haplotypes to the catalog of overrepresented haplotypes to exclude a subset of the one or more matched haplotypes that are matched to the overrepresented haplotypes;

determining a normalized data inheritance between the target data instance and the related data instance, the normalized data inheritance corresponding to a total length of the one or more matched haplotypes excluding the subset of the one or more matched haplotypes that are matched to the overrepresented haplotypes; and

storing, responsive to the total length being longer than a length threshold, information regarding the target data instance and the related data instance being a pair of identity-by-descent matches.

12 . The system of claim 11 , wherein the overrepresented haplotypes in the catalog are arranged by windows of data locality within which the overrepresented haplotypes are located.

13 . The system of claim 12 , wherein the catalog comprises a plurality of overrepresented haplotypes in one of the windows of data locality.

14 . The system of claim 11 , wherein generating the catalog comprises:

receiving the plurality of data instances that correspond to named entities, wherein each of the plurality of data instances is phased;

dividing each data instance into a plurality of windows of data locality;

tallying, for each window, the phased data instance that has a particular data bit sequence; and

determining whether a tally of the particular data bit sequence exceeds the number threshold.

15 . The system of claim 11 , wherein information regarding the target data instance and the related data instance being a pair of match includes the total length and an indication that the target data instance and the related data instance are related by inheritance of a real-life event.

16 . The system of claim 11 , wherein the catalog is generated before receiving the target data instance and is stored in a second database.

17 . The system of claim 11 , wherein comparing the matched haplotypes to the catalog to exclude the subset of matched haplotypes that are matched to the overrepresented haplotypes comprises:

dividing the target data instance into a plurality of windows of data locality;

receiving, from the catalog, a subset of the windows that include the overrepresented haplotypes; and

comparing the target data instance and the related data instance in the windows that are not under the subset.

18 . The system of claim 11 , wherein determining one or more matched haplotypes that match between the target data instance and the related data instance is based on one or more weak matches of haplotypes.

19 . A non-transitory computer readable medium configured to store code for determining a normalized data inheritance between two data instances, the code comprising instructions, wherein the instructions, when executed by one or more processors, cause the one or more processors to perform steps comprising:

generating a catalog of overrepresented haplotypes from a database that stores a plurality of data instances in a population, wherein an overrepresented haplotype is a haplotype that matches to a number of data instances in the population and the number exceeds a number threshold, wherein the catalog of overrepresented haplotypes comprises a version of haplotype string sequences and corresponding tallies of the haplotype string sequences in the population, and wherein the number threshold is based on an overall distribution of haplotype sequence repetitiveness;

receiving a target data instance that is to be compared to a related data instance;

determining one or more matched haplotypes that match between the target data instance and the related data instance;

comparing the one or more matched haplotypes to the catalog of overrepresented haplotypes to exclude a subset of the one or more matched haplotypes that are matched to the overrepresented haplotypes;

determining a normalized data inheritance between the target data instance and the related data instance, the normalized data inheritance corresponding to a total length of the one or more matched haplotypes excluding the subset of the one or more matched haplotypes that are matched to the overrepresented haplotypes; and

storing, responsive to the total length being longer than a length threshold, information regarding the target data instance and the related data instance being a pair of identity-by-descent matches.

20 . The non-transitory computer readable medium of claim 19 , wherein the overrepresented haplotypes in the catalog are arranged by windows of data locality within which the overrepresented haplotypes are located.

Assignments (3)
PATENT SECURITY AGREEMENT Recorded Aug 3, 2026
From: ANCESTRY.COM OPERATIONS INC.; ANCESTRY.COM DNA, LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 076116/0447 →
PATENT SECURITY AGREEMENT Recorded Aug 3, 2026
From: ANCESTRY.COM OPERATIONS INC.; ANCESTRY.COM DNA, LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 076144/0726 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2023
From: NOTO, KEITH D.
To: ANCESTRY.COM DNA, LLC
Reel/Frame 064876/0702 →
Continuity (2)
Provisional Application 63371875 · Aug 19, 2022
Related Publication 20240061886A1 · Feb 22, 2024
References Cited (29)
US 10114922B2 · Byrnes et al. · 2018 [cited by applicant]
US 10720229B2 · Barber et al. · 2020 [cited by applicant]
US 20020172948A1 · Perlin · 2002 [cited by examiner]
US 20100169338A1 · Kenedy et al. · 2010 [cited by applicant]
US 20100190264A1 · Pericak-Vance · 2010 [cited by examiner]
US 20100223281A1 · Hon et al. · 2010 [cited by applicant]
US 20120053845A1 · Bruestle et al. · 2012 [cited by applicant]
US 20140025308A1 · Jorde et al. · 2014 [cited by applicant]
US 20140067355A1 · Noto et al. · 2014 [cited by applicant]
US 20140278138A1 · Barber et al. · 2014 [cited by applicant]
US 20140378138A1 · Chang et al. · 2014 [cited by applicant]
US 20150363481A1 · Haynes · 2015 [cited by applicant]
US 20170213127A1 · Duncan · 2017 [cited by applicant]
US 20170220738A1 · Barber et al. · 2017 [cited by applicant]
US 20180044730A1 · Pickrell et al. · 2018 [cited by applicant]
US 20190205502A1 · Staples et al. · 2019 [cited by applicant]
US 20200286591A1 · Barber et al. · 2020 [cited by applicant]
US 20210034647A1 · Nguyen · 2021 [cited by examiner]
Browning, S.R. et al., “Identity by Descent Between Distant Relatives: Detection and Applications,” Annu. Rev. Genet., 2012, vol. 46, pp. 617-633. [cited by applicant]
Browning, S.R. et al., “Rapid and Accurate Haplotype Phasing and Missing-Data Inference for Whole-Genome Association Studies by Use of Localized Haplotype Clustering,” The American Journal of Human Genetics, Nov. 2007, … [cited by applicant]
Gusev, A. et al., “The Architecture of Long-Range Haplotypes Shared Within and Across Populations,” Mol. Biol. Evol., 2012, vol. 29, No. 2, pp. 473-486. [cited by applicant]
Gusev, A. et al., “Whole Population, Genome-Wide Mapping of Hidden Relatedness,” Genome Research, Feb. 2009, vol. 19, No. 2, pp. 318-326. [cited by applicant]
Henn, B.M., et al., “Cryptic Distant Relatives Are Common in Both Isolated and Cosmopolitan Genetic Samples,” PLoS One, 2D12, vol. 7, No. e34267 (14 pages). [cited by applicant]
Nelder, J.A et al., “A Simplex Algorithm for Function Minimization,” The Computer Journal, Apr. 1964-Jan. 1965, vol. 7, pp. 308-313. [cited by applicant]
Padhukasahasram, B., “Inferring Ancestry from Population Genomic Data and Its Applications,” Frontiers in Genetics, Jul. 3, 2014, vol. 5, pp. 1-5. [cited by applicant]
Pfeil, M. “What is Data Persistence & Why Does it Matter.” Datastax, Oct. 22, 2010, 6 pages, [Online] [Retrieved Feb. 7, 2024], Retrieved from the Internet <URL:https://www.datastax.com/blog/what-persistence-and-why-doe… [cited by applicant]
Purcell, S. et al., “PLINK: a tool set for whole-genome association and population-based linkage analyses,” The American Journal of Human Genetics, 2007, vol. 81, No. 3, pp. 559-575. [cited by applicant]
Sabeti, P.C. et al., “Genome-wide detection and characterization of positive selection in human populations,” Nature, 2007, vol. 449, No. 7164, pp. 1-16. [cited by applicant]
The International Hapmap Consorium, “A second generation human haplotype map of over 3.1 million SNPs,” Nature, 2007, vol. 449, No. 164, pp. 1-30. [cited by applicant]