IP Library Granted Patent US 10,658,071
Granted Patent B2
US 10,658,071 · App. 14/938,111 · Granted May 19, 2020

Scalable pipeline for local ancestry inference

Inventors: Chuong Do (Mountain View, CA); Eric Yves Jean-Marc Durand (San Francisco, CA); John Michael Macpherson (Mountain View, CA)
Assignee: 23andMe, Inc.
G16B40/00G06N5/04G06N7/005G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,658,071
App. No.
14/938,111
Granted
May 19, 2020
Kind
B2
Abstract

Ancestry deconvolution includes obtaining unphased genotype data of an individual; phasing, using one or more processors, the unphased genotype data to generate phased haplotype data; using a learning machine to classify portions of the phased haplotype data as corresponding to specific ancestries respectively and generate initial classification results; and correcting errors in the initial classification results to generate modified classification results.

Claims (82)

1. A system, comprising:

one or more processors configured to:

train a learning machine using a training set comprising genetic information of a plurality of individuals with known ancestries;

obtain unphased genotype data of an individual whose ancestry composition is to be determined;

phase the unphased genotype data of the individual to generate phased haplotype data of the individual;

use the trained learning machine to classify portions of the phased haplotype data of the individual as corresponding to specific ancestries and generate initial ancestry classification results;

correct one or more errors in the initial ancestry classification results to generate modified ancestry classification results, wherein the modified ancestry classification results include ancestry assignments and posterior probabilities association with the ancestry assignments;

recalibrate the modified ancesry classification results to establish confidence levels associated with the ancestry assignments;

determine if the confidence levels associated with the ancestry assignments meet a threshold level; and

one or more memories coupled with the one or more processors, configured to provide the one or more processors with instructions.

2. The system of claim 1 , wherein to phase the unphased genotype data of the individual to generate the phased haplotype data of the individual includes to perform out-of-sample phasing where the unphased genotype data of the individual is not included in a reference population.

3. The system of claim 1 , wherein the learning machine includes a neural network.

4. The system of claim 1 , wherein the learning machine includes a support vector machine (SVM).

5. The system of claim 1 , wherein the one or more processors are further configured to recalibrate the modified ancestry classification results to establish confidence levels associated with ancestry assignments associated with the modified ancestry classification results.

6. The system of claim 1 , wherein the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments.

7. The system of claim 1 , wherein:

the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments; and

the one or more processors are further configured to recalibrate the modified ancestry classification results to establish confidence levels associated with the ancestry assignments.

8. A system, comprising:

one or more memories coupled with one or more processors, configured to provide the one or more processors with instructions; and

the one or more processors configured to:

train a learning machine using a training set comprising genetic information of a plurality of individuals with known ancestries;

obtain unphased genotype data of an individual whose ancestry composition is to be determined;

phase the unphased genotype data of the individual to generate phased haplotype data of the individual;

use the trained learning machine to classify portions of the phased haplotype data of the individual as corresponding to specific ancestries and generate initial ancestry classification results; and

correct one or more errors in the initial ancestry classification results to generate modified ancestry classification results, wherein the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments;

recalibrate the modified ancestry classification results to establish confidence levels associated with the ancestry assignments; and

store the recalibrated modified ancestry classification results and the confidence levels associated with the ancestry assignments to a database, output the recalibrated modified ancestry classification results and the confidence levels associated with the ancestry assignments to another application, or both.

9. The system of claim 1 , wherein:

the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments; and

the one or more processors are further configured to:

recalibrate the modified ancestry classification results to establish confidence levels associated with the ancestry assignments; and

determine if the confidence levels associated with the ancestry assignments meet a threshold level.

10. A system, comprising:

one or more memories coupled with one or more processors, configured to provide the one or more processors with instructions; and

the one or more processors configured to:

train a learning machine using a training set comprising genetic information of a plurality of individuals with known ancestries;

obtain unphased genotype data of an individual whose ancestry composition is to be determined;

phase the unphased genotype data of the individual to generate phased haplotype data of the individual;

use the trained learning machine to classify portions of the phased haplotype data of the individual as corresponding to specific ancestries and generate initial ancestry classification results;

correct one or more errors in the initial ancestry classification results to generate modified ancestry classification results, wherein the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments;

recalibrate the modified ancestry classification results to establish confidence levels associated with the ancestry assignments;

determine if the confidence levels associated with the ancestry assignments meet a threshold level; and

in response to the confidence levels associated with the ancestry assignments not meeting the threshold level, cluster at least some of the ancestry assignments associated with corresponding confidence levels to form one or more new probabilities associated with broader geographical regions.

11. A method, comprising:

training a learning machine using a training set comprising genetic information of a plurality of individuals with known ancestries;

obtaining unphased genotype data of an individual whose ancestry composition is to be determined;

phasing the unphased genotype data of the individual to generate phased haplotype data of the individual;

using the trained learning machine to classify portions of the phased haplotype data of the individual as corresponding to specific ancestries and to generate initial ancestry classification results;

correcting one or more errors in the initial ancestry classification results to generate modified ancestry classification results, wherein the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments;

recalibrating the modified ancestry classification results to establish confidence levels associated with the ancestry assignments; and storing the recalibrated modified ancestry classification results and the confidence levels associated with the ancestry assignments to a database, outputting the recalibrated modified ancestry classification results and the confidence levels associated with the ancestry assignments to another application, or both.

12. The method of claim 11 , wherein phasing the unphased genotype data of the individual to generate the phased haplotype data of the individual includes performing out-of-sample phasing where the unphased genotype data of the individual is not included in a reference population.

13. The method of claim 11 , wherein the learning machine includes a neural network.

14. The method of claim 11 , wherein the learning machine includes a support vector machine (SVM).

15. The method of claim 11 , further comprising recalibrating the modified ancestry classification results to establish confidence levels associated with ancestry assignments associated with the modified ancestry classification results.

16. The method of claim 11 , wherein the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments.

17. The method of claim 11 , wherein:

the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments; and

the method further comprises recalibrating the modified ancestry classification results to establish confidence levels associated with the ancestry assignments.

18. The method of claim 11 , wherein:

the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments; and

the method further comprises:

recalibrating the modified ancestry classification results to establish confidence levels associated with the ancestry assignments; and

storing the recalibrated modified ancestry classification results and the confidence levels associated with the ancestry assignments to a database, outputting the recalibrated modified ancestry classification results and the confidence levels associated with the ancestry assignments to another application, or both.

19. The method of claim 11 , wherein:

the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments; and

the method further comprises:

recalibrating the modified ancestry classification results to establish confidence levels associated with the ancestry assignments; and

determining if the confidence levels associated with the ancestry assignments meet a threshold level.

20. The method of claim 11 , wherein:

the modified ancestry classification results include ancestry assignments and posterior probabilities associated with the ancestry assignments; and

the method further comprises:

recalibrating the modified ancestry classification results to establish confidence levels associated with the ancestry assignments; and

determining if the confidence levels associated with the ancestry assignments meet a threshold level; and

in response to the confidence levels associated with the ancestry assignments not meeting the threshold level, clustering at least some of the ancestry assignments associated with corresponding confidence levels to form one or more new probabilities associated with broader geographical regions.

21. The system of claim 1 , wherein the initial ancestry classification results or the modified ancestry classification results comprise local ancestry estimates.

22. The system of claim 21 , wherein the local ancestry estimates comprise ancestry estimates for a defined segment size or window size.

23. The system of claim 10 , wherein the ancestry assignments and the broader geographical regions comprise regions in a geographical hierarchy.

24. The system of claim 23 , the geographical hierarchy comprises the world, continents, subcontinents, and individual countries or specific geographical regions.

25. The system of claim 10 , the one or more processors are further configured to:

in response to the confidence levels associated with the ancestry assignments meeting the threshold level, display the ancestry classification results of the ancestry assignments in an interactive GUI; or

in response the one or more new probabilities associated with the broader geographical regions meeting the threshold level, display the ancestry classification results of the broader geographical regions in the interactive GUI.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE APP. NO. 63806415 TO 63806145 AND APPL NO. 17721779 TO 17731779 PREVIOUSLY RECORDED ON REEL 73168 FRAME 531. ASSIGNOR(S) HEREBY CONFIRMS THE CHANGE OF NAME. Recorded Jan 6, 2026
From: 23ANDME PGS LLC
To: 23ANDME GENOMICS LLC
Reel/Frame 074434/0334 →
CHANGE OF NAME Recorded Oct 22, 2025
From: 23ANDME PGS LLC
To: 23ANDME GENOMICS LLC
Reel/Frame 073168/0531 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2025
From: 23ANDME, INC.
To: 23ANDME PGS LLC
Reel/Frame 072562/0795 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2020
From: DO, CHUONG; DURAND, ERIC; MACPHERSON, JOHN MICHAEL
To: 23ANDME, INC.
Reel/Frame 051528/0240 →
Continuity (4)
Continuation 13801056 · Mar 13, 2013
Provisional Application 61724236 · Nov 8, 2012
Provisional Application 61724228 · Nov 8, 2012
Related Publication 20160171155A1 · Jun 16, 2016
Cited By (7)
US 12,243,654 US 12,260,936 US 12,293,268 US 12,327,615 US 12,354,710 US 12,431,221 US 12,580,048