IP Library Granted Patent US 12,159,690
Granted Patent B2
US 12,159,690 · App. 18/377,219 · Granted Dec 3, 2024

Ancestry composition determination

Inventors: Peter Richard Wilton (Los Gatos, CA); Gabriel David Poznik (Menlo Park, CA); Kimberly Faith McManus (San Francisco, CA); Ethan Macneil Jewett (San Jose, CA); William Allen Freyman (Menlo Park, CA); Adam Auton (Menlo Park, CA)
Assignee: 23andMe, Inc.
G16B10/00G06F16/285G06N7/01G16B5/20G16B40/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,159,690
App. No.
18/377,219
Granted
Dec 3, 2024
Kind
B2
Abstract

Presenting ancestral origin information, comprising: receiving a request to display ancestry data of an individual; obtaining ancestry composition information of the individual, the ancestry composition information including information pertaining to a proportion of the individual's genotype data that is deemed to correspond to a specific ancestry; and presenting the ancestry composition information to be displayed.

Claims (47)

1. A method comprising:

obtaining, by way of one or more processors and from a classifier, initial ancestry classifications based on chromosome data of an individual;

identifying, by way of the one or more processors, an initial Hidden Markov Model (HMM);

refining, by way of the one or more processors, parameters of the initial HMM with the chromosome data of the individual;

performing, by way of the one or more processors, error correction on the initial ancestry classifications, wherein the error correction involves applying an error correction model based on the initial HMM as refined to the initial ancestry classifications to form corrected ancestry classifications; and

providing, by way of the one or more processors, the corrected ancestry classifications for storage or display.

2. The method of claim 1 , wherein the chromosome data of the individual is from a plurality of chromosomes of the individual.

3. The method of claim 1 , wherein obtaining the initial ancestry classifications based on the chromosome data of the individual comprises:

obtaining unphased chromosome data of the individual; and

determining, using dynamic programming and based on a predetermined reference haplotype graph, phased chromosome data of the individual from the unphased chromosome data of the individual, wherein the chromosome data of the individual is based on the phased chromosome data of the individual.

4. The method of claim 3 , wherein obtaining the initial ancestry classifications based on the chromosome data of the individual further comprises:

performing trio-based phasing of the phased chromosome data of the individual, wherein the trio-based phasing incorporates genotyping data of one or more biological parents of the individual.

5. The method of claim 3 , wherein the predetermined reference haplotype graph is a probabilistic representation of two or more possible haplotype substrings as a directed acyclic graph.

6. The method of claim 5 , wherein the directed acyclic graph is modified to include additional edges that account for recombination and genotyping errors.

7. The method of claim 5 , wherein the directed acyclic graph is modified to exclude paths with an overall probability less than a given threshold.

8. The method of claim 1 wherein the error correction model is also based on a templated positional Burrows-Wheeler transform.

9. The method of claim 1 , wherein obtaining the initial ancestry classifications based on the chromosome data of the individual comprises:

obtaining a plurality of ancestries, wherein each of the ancestries is respectively associated with reference chromosome data of unadmixed individuals;

training the classifier based on the reference chromosome data and the respectively associated ancestries;

dividing the chromosome data of the individual into segments; and

determining the initial ancestry classifications from applying the classifier to the segments.

10. The method of claim 9 , wherein each of the unadmixed individuals has self-identified as being of a specific ancestry or has four grandparents of the specific ancestry.

11. The method of claim 9 , wherein the classifier is configured to provide, for a given segment of the chromosome data, two or more probabilities associated with the given segment being of corresponding ancestries, and wherein determining the initial ancestry classifications from applying the classifier to the segments comprises selecting a corresponding ancestry with a highest probability for the given segment.

12. The method of claim 9 , wherein training the classifier based on the reference chromosome data and the respectively associated ancestries comprises training the classifier based on the reference chromosome data from at least 14,400 unrelated individuals each with unadmixed ancestry.

13. The method of claim 9 , wherein the segments are each of a size between 100 markers and 500 markers.

14. The method of claim 1 , wherein the initial HMM is a Pair Hidden Markov Model (PHMM) in which an observed state corresponds to the initial ancestry classifications associated with a portion of one of two haplotypes of the individual, and a hidden state corresponds to ancestries associated with a corresponding portion of the two haplotypes of the individual.

15. The method of claim 1 , wherein the initial HMM is an Autoregressive Pair Hidden Markov Model (PHMM) in which an observed state depends on its corresponding hidden state and its previous observed state.

16. The method of claim 1 , wherein applying the error correction model based on the initial HMM as refined to the initial ancestry classifications to form the corrected ancestry classifications comprises:

determining transition parameters of the error correction model; and

based on error correction results of the error correction model, repeatedly updating the transition parameters until the error correction results converge.

17. The method of claim 1 , wherein the initial HMM is selected from a pool of pre-trained HMMs, and wherein each of the pre-trained HMMs has respective transition parameters based on an unsupervised training procedure performed on samples of approximately 1000 individuals of respective regional ancestries.

18. The method of claim 1 , wherein the corrected ancestry classifications are a most probable sequence of ancestry assignments for the chromosome data of the individual and posterior probabilities associated with each of the ancestry assignments, and wherein the posterior probabilities have been recalibrated based on reference chromosome data of unadmixed individuals or simulated admixed individuals until a given threshold confidence level of the posterior probabilities is reached.

19. A non-transitory computer-readable medium storing program instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations comprising:

obtaining, from a classifier, initial ancestry classifications based on chromosome data of an individual;

identifying an initial Hidden Markov Model (HMM);

refining parameters of the initial HMM with the chromosome data of the individual;

performing error correction on the initial ancestry classifications, wherein the error correction involves applying an error correction model based on the initial HMM as refined to the initial ancestry classifications to form corrected ancestry classifications; and

providing the corrected ancestry classifications for storage or display.

20. A computing system comprising:

one or more processors;

memory; and

program instructions, stored in the memory, that upon execution by the one or more processors cause the computing system to perform operations comprising:

obtaining, from a classifier, initial ancestry classifications based on chromosome data of an individual;

identifying an initial Hidden Markov Model (HMM);

refining parameters of the initial HMM with the chromosome data of the individual;

performing error correction on the initial ancestry classifications, wherein the error correction involves applying an error correction model based on the initial HMM as refined to the initial ancestry classifications to form corrected ancestry classifications; and

providing the corrected ancestry classifications for storage or display.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE APP. NO. 63806415 TO 63806145 AND APPL NO. 17721779 TO 17731779 PREVIOUSLY RECORDED ON REEL 73168 FRAME 531. ASSIGNOR(S) HEREBY CONFIRMS THE CHANGE OF NAME. Recorded Jan 6, 2026
From: 23ANDME PGS LLC
To: 23ANDME GENOMICS LLC
Reel/Frame 074434/0334 →
CHANGE OF NAME Recorded Oct 22, 2025
From: 23ANDME PGS LLC
To: 23ANDME GENOMICS LLC
Reel/Frame 073168/0531 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2025
From: 23ANDME, INC.
To: 23ANDME PGS LLC
Reel/Frame 072562/0795 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: WILTON, PETER RICHARD; POZNIK, GABRIEL DAVID; MCMANUS, KIMBERLY FAITH; JEWETT, ETHAN MACNEIL; FREYMAN, WILLIAM ALLEN; AUTON, ADAM
To: 23ANDME, INC.
Reel/Frame 065146/0966 →
Continuity (4)
Continuation 17444989 · Aug 12, 2021
Provisional Application 63093039 · Oct 16, 2020
Provisional Application 62706396 · Aug 13, 2020
Related Publication 20240062845A1 · Feb 22, 2024