IP Library Granted Patent US 10,385,391
Granted Patent B2
US 10,385,391 · App. 13/497,069 · Granted Aug 20, 2019

Entangled mate sequencing

Inventors: Frederick P. Roth (Newton, MA); Joseph C. Mellor (Brookline, MA); Yong Lu (College Park, MD); Mark Chee (Encinitas, CA)
Assignee: President and Fellows of Harvard College
C12Q1/6874
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,385,391
App. No.
13/497,069
Granted
Aug 20, 2019
Kind
B2
Abstract

Methods and compositions are provided for performing a set of N DNA sequencing reaction cycles whereby sequence information is obtained for approximately 2*N nucleotide bases.

Claims (46)

1. A method of determining nucleotide sequence identities of heterogeneous, tethered nucleic acid template sequences comprising the steps of:

providing a mosaic target nucleic acid sequence including a first template nucleic acid sequence and a second template nucleic acid sequence, wherein the first and second template nucleic acid sequences are tethered and heterogeneous and of relatively homogeneous length;

simultaneously reading the first template nucleic acid sequence and the second template nucleic acid sequence to obtain a mixed entangled sequencing signal representative of each of the first template nucleic acid sequence and the second template nucleic acid sequence; and

disentangling the mixed sequencing signal into its constituent individual template sequences by matching the mixed entangled sequencing signal to its closest match in a reference collection of known sequence signatures wherein the mixed entangled sequencing signal is obtained by annealing a first primer sequence to a portion of the first template nucleic acid sequence and a second primer sequence to a portion of the second template nucleic acid sequence;

extending the annealed primers simultaneously; and

sequentially determining the identity of each set of two bases extended from the 3′ ends of the first and second primers to obtain mixed base signatures and using said mixed base signatures to obtain a mixed nucleic acid sequence signature of the tethered first and second template nucleic acid sequences, wherein the identity of each set of two bases extended from the 3′ ends of the first and second primers is sequentially determined by nucleotide addition sequencing.

2. The method of claim 1 , wherein the heterogeneous, tethered nucleic acid template sequences represent one or more of a genome, a proteome, a transcriptome or a cellular pathway.

3. A method of determining nucleotide sequence identities of tethered, heterogeneous nucleic acid template sequences comprising the steps of:

providing a heterogeneous library of mosaic target nucleic acid sequences, in which a plurality include a first template nucleic acid sequence and a second template nucleic acid sequence, wherein the first and second template nucleic acid sequences are tethered and heterogeneous and of relatively homogeneous length;

simultaneously reading the first template nucleic acid sequence and the second template nucleic acid sequence to obtain a mixed entangled sequencing signal representative of each of the first template nucleic acid sequence and the second template nucleic acid sequence; and

disentangling the mixed sequencing signal into its constituent individual template sequences by matching the mixed entangled sequencing signal to its closest match in a reference collection of known sequence signatures wherein the mixed entangled sequencing signal is obtained by annealing a first primer sequence to a portion of the first template nucleic acid sequence and a second primer sequence to a portion of the second template nucleic acid sequence;

extending the annealed primers simultaneously; and

sequentially determining the identity of each set of two bases extended from the 3′ ends of the first and second primers to obtain mixed base signatures and using said mixed base signatures to obtain a mixed nucleic acid sequence signature of the tethered first and second template nucleic acid sequences, wherein the identity of each set of two bases extended from the 3′ ends of the first and second primers is sequentially determined by nucleotide addition sequencing.

4. The method of claim 3 , wherein the first template nucleic acid sequence is a barcode sequence and the second template nucleic acid sequence is a genomic sequence.

5. The method of claim 3 , wherein the reference collection is in a database.

6. The method of claim 3 , wherein the plurality includes the first template nucleic acid sequence, the second template nucleic acid sequence, and a third template nucleic acid sequence wherein the first, second and third template nucleic acid sequences are tethered and heterogeneous and of relatively homogeneous length;

simultaneously reading the first template nucleic acid sequence, the second template nucleic acid sequence and the third template nucleic acid sequence to obtain a mixed entangled sequencing signal representative of each of the first template nucleic acid sequence the second template nucleic acid sequence and the third template nucleic acid sequence; and

disentangling the mixed sequencing signal into its constituent individual template sequences by matching the mixed entangled sequencing signal to its closest match in a reference collection of known sequence signatures.

7. The method of claim 3 , wherein the mosaic target nucleic acid sequence includes 3 or more tethered, nucleic acid sequences.

8. The method of claim 3 , wherein the mosaic target nucleic acid sequence is present on an array.

9. A method of determining nucleotide sequence identities of two tethered nucleic acid sequences comprising the steps of:

providing a heterogeneous library of mosaic target nucleic acid sequences, in which a plurality include a first template nucleic acid sequence and a second template nucleic acid sequence, wherein the first and second template nucleic acid sequences are tethered and heterogeneous and of relatively homogeneous length;

simultaneously reading the first template nucleic acid sequence, the second template nucleic acid sequence and the third template nucleic acid sequence to obtain a mixed entangled sequencing signal representative of each of the first template nucleic acid sequence the second template nucleic acid sequence and the third template nucleic acid sequence; and

disentangling the mixed sequencing signal into its constituent individual template sequences by matching the mixed entangled sequencing signal to its closest match in a reference collection of known sequence signatures

wherein the mixed entangled sequencing signal is obtained by annealing a first primer sequence to a portion of the first template nucleic acid sequence and a second primer sequence to a portion of the second template nucleic acid sequence;

extending the annealed primers simultaneously; and

sequentially determining the identity of each set of two bases extended from the 3′ ends of the first and second primers to obtain mixed base signatures and using said mixed base signatures to obtain a mixed nucleic acid sequence signature of the tethered first and second template nucleic acid sequences, wherein the identity of each set of two bases extended from the 3′ ends of the first and second primers is sequentially determined by nucleotide addition sequencing.

10. The method of claim 9 , wherein the reference collection is a collection of sequence signatures and the second nucleic acid sequence is not present in the reference collection of sequence signatures.

11. The method of claim 9 , wherein each mixed base signature sequence includes a set of two bases selected from the group consisting of AA, CC, GG, TT, AC, AG, AT, CG, CT and GT.

12. The method of claim 9 , wherein the first and second template nucleic acid sequences are barcode sequences.

13. The method of claim 9 , wherein the first and second template nucleic acid sequences are genomic sequences.

14. The method of claim 9 , wherein the first and second primer sequences have the same sequence identity.

15. The method of claim 9 , wherein the first and second primer sequences have a different sequence identity.

16. A method of identifying a cell having heterogeneous nucleic acid sequences comprising the steps of:

providing a cell having a mosaic target nucleic acid sequence including a first barcode sequence and a second barcode sequence, wherein the first and second barcode sequences are tethered and heterogeneous and of relatively homogeneous length;

simultaneously reading the first barcode sequence and the second barcode sequence to obtain a mixed entangled sequencing signal representative of each of the first barcode sequence and second barcode sequence;

disentangling the mixed sequencing signal into its constituent individual barcode sequences by matching the mixed entangled sequencing signal to its closest match in a reference collection of known sequence signatures;

wherein the mixed entangled sequencing signal is obtained by annealing a first primer sequence to a portion of the first template nucleic acid sequence and a second primer sequence to a portion of the second template nucleic acid sequence;

extending the annealed primers simultaneously;

sequentially determining the identity of each set of two bases extended from the 3′ ends of the first and second primers to obtain mixed base signatures and using said mixed base signatures to obtain a mixed nucleic acid sequence signature of the tethered first and second template nucleic acid sequences, wherein the identity of each set of two bases extended from the 3′ ends of the first and second primers is sequentially determined by nucleotide addition sequencing; and

determining the nucleotide sequence identity of one or both of the tethered first and second barcode sequences to identify the cell from the mixed nucleic acid barcode sequence signature.

17. The method of claim 16 , wherein the cell is a yeast cell.

18. The method of claim 17 , wherein the yeast cell contains two or more alterations.

19. The method of claim 16 , wherein the heterogeneous nucleic acid sequences represent one or more of a genome, a proteome, a transcriptome or a cellular pathway.

20. The method of claim 16 , wherein the heterogeneous nucleic acid sequences represent Homo sapiens nucleic acid sequences.

21. The method of claim 1 , wherein the step of determining the sequence is performed by comparing the nucleic acid sequence signature to a reference collection of seed sequences.

Assignments (1)
CONFIRMATORY LICENSE Recorded Apr 9, 2012
From: HARVARD UNIVERSITY
To: NATIONAL INSTITUTES OF HEALTH (NIH), U.S. DEPT. OF HEALTH AND HUMAN SERVICES (DHHS), U.S. GOVERNMENT
Reel/Frame 028010/0133 →
Continuity (2)
Provisional Application 61244503 · Sep 22, 2009
Related Publication 20120245039A1 · Sep 27, 2012