IP Library › Granted Patent US 10,870,882
Granted Patent B2
US 10,870,882 · App. 15/978,208 · Granted Dec 22, 2020

Method for accurate sequencing of DNA

Inventors: Zbyszek Otwinowski (Dallas, TX); Dominika Borek (Dallas, TX)
Assignee: Board of Regents, The University of Texas System
C12Q1/6869C12Q1/6844
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,870,882
App. No.
15/978,208
Granted
Dec 22, 2020
Kind
B2
Abstract

DNA is sequenced by (a) independently sequencing first and second strands of a dsDNA to obtain corresponding first and second sequences; and (b) combining the first and second sequences to generate a consensus sequence of the dsDNA. By independently sequencing first and second strands the error probability of the consensus sequence approximates a multiplication of those of the first and second sequences.

Claims (33)

1. A method of sequencing DNA using a sequencer, comprising:

(a) generating a sequencing library comprising the steps—

attaching sequencing adapters to a plurality of DNA fragments to form a plurality of adapter-DNA molecules; and

amplifying the adapter-DNA molecules to generate clonal strands of each adapter-DNA molecule;

(b) flowing a plurality of clonal strands over a sequencer flow cell such that multiple clonal strands of adapter-DNA molecules bind to the flow cell;

(c) bridge amplifying the bound strands of adapter-DNA molecules to form one or more polonies;

(d) sequencing the polonies to obtain a preliminary sequence read for each polony, wherein for each polony, sequencing includes—

obtaining fluorescence intensity measurements for each nucleotide position of the polony; and

performing a first round of base calling by converting the fluorescence intensity measurements into a probable base call for each nucleotide position of the polony to provide a preliminary sequence read, the preliminary sequence read having a preliminary error probability;

(e) grouping preliminary sequence reads generated from the clonal strands derived from a particular adapter-DNA molecule;

(f) calculating a fluorescence intensity measurement average for each nucleotide position from the grouped preliminary sequence reads; and

(g) performing a second round of base calling for each nucleotide position of the grouped preliminary sequence reads -using the fluorescence intensity measurement averages to provide a sequence read;

(h) wherein an error probability of the sequence read corresponding to the particular adapter-DNA molecule is lower than any one of the preliminary error probabilities of the individual preliminary sequence reads in the group.

2. The method of claim 1 wherein each of the preliminary sequence reads is a composite sequence generated from a plurality of sequencer readouts of each polony.

3. The method of claim 2 wherein a plurality of sequencer readouts of each polony is obtained by paired-end sequencing.

4. The method of claim 1 further comprising the step of identifying one or more correspondences between nucleotide positions along a plurality of sequence reads generated from the clonal strands derived from a particular adapter-DNA molecule by comparing fluorescence intensity measurements at said nucleotide positions.

5. The method of claim 1 , further comprising mapping the preliminary sequencing reads to a reference genome, wherein the second round of base calling is on regions corresponding to the reference genome identified in the preliminary sequence reads as suspicious of having a phasing error.

6. The method of claim 5 , wherein the regions of the reference genome identified in the preliminary sequence reads correspond to one or more polonies on the flow cell comprising inconsistent fluorescence intensities.

7. The method of claim 5 , wherein the regions of the reference genome comprises an inverted repeat.

8. The method of claim 1 , further comprising:

identifying a phasing error in a first group of preliminary sequence reads derived from a first adapter-DNA molecule having a first nucleic acid sequence; and

providing a phasing error estimate for a second group of preliminary sequence reads derived from a second adapter-DNA molecule, the second adapter DNA molecule having a second nucleic acid sequence that at least partially overlaps the first nucleic acid sequence.

9. The method of claim 1 , further comprising analyzing a first group of preliminary sequence reads corresponding to a first adapter-DNA molecule to identify inconsistencies between individual preliminary sequence reads belonging to the first group.

10. The method of claim 9 , wherein the analysis identified no inconsistencies between preliminary sequence reads of the first group and the preliminary sequence reads are determined to be an accurate sequence of the bound strands of the adapter-DNA molecule.

11. The method of claim 1 , further comprising comparing one or more sequence reads to a reference sequence to identify a presence of a variant.

12. The method of claim 1 , wherein, following the second round of base-calling, the method further comprises comparing a first sequence read derived from a first strand of a particular adapter-DNA molecule and a second sequence read derived from a second strand of the same adapter-DNA molecule to identify one or more inconsistencies.

13. The method of claim 12 , wherein the inconsistencies are identified as amplification errors.

14. The method of claim 1 , wherein, following the second round of base-calling, the method further comprises detecting an error when a sequence read derived from a particular adapter-DNA molecule has one or more inconsistent base calls, and wherein the error is identified as an amplification error.

15. The method of claim 1 , wherein, following the second round of base-calling, the method further comprises detecting an error when a sequence reads derived from an original DNA fragment has one or more inconsistent base calls when compared to a second sequence reads derived from the same original DNA fragment, and wherein the error is identified as a chemical base alteration occurring prior to the bridge amplification step.

16. The method of claim 1 , wherein one or more probable DNA base calls is identified as an incorrect base call, and wherein the method further comprises correcting the probable base call in the second round of base calling.

17. The method of claim 1 , wherein for a plurality of sequence reads, the method further comprises identifying a sequence-specific desynchronization event.

18. The method of claim 17 , wherein the sequence-specific-desynchronization event is identified by a steep cumulative Q-value decrease along the plurality of sequence reads.

19. The method of claim 18 , wherein the plurality of sequence reads having the sequence-specific desynchronization event comprises one or more of high GC content, two GGC triplets, and a secondary structure formation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2018
From: OTWINOWSKI, ZBYSZEK; BOREK, DOMINIKA
To: BOARD OF REGENTS, THE UNIVERSITY OF TEXAS SYSTEM
Reel/Frame 046330/0016 →
Continuity (4)
Continuation 14550517 · Nov 21, 2014
Continuation PCTUS2013042949 · May 28, 2013
Provisional Application 61654069 · May 31, 2012
Related Publication 20180282802A1 · Oct 4, 2018
Cited By (2)
US 12,529,101 US 12,680,131