COMPOSITIONS AND METHODS FOR ACCURATELY IDENTIFYING MUTATIONS
The present disclosure provides compositions and methods for accurately detecting mutations by uniquely tagging double stranded nucleic acid molecules with dual cyphers such that sequence data obtained from a sense strand can be linked to sequence data obtained from an anti-sense strand when sequenced, for example, by massively parallel sequencing methods.
1 .- 38 . (canceled)
39 . A method for reducing an error rate in sequence reads of a double-stranded target nucleic acid molecule, comprising:
a) ligating the double-stranded target nucleic acid molecule to a double-stranded cypher to form a cypher-target nucleic acid complex, wherein the double-stranded cypher comprises a random or partially random identifier sequence that alone or in combination with an end of the target nucleic acid molecule uniquely labels the double-stranded target nucleic acid molecule;
b) amplifying each strand of the cypher-target nucleic acid complex to produce a plurality of cypher-target amplification products from each of a first strand and a complementary second strand of the cypher-target nucleic acid complex;
(c) sequencing the cypher-target amplification products to produce a plurality of first-strand sequencing reads and a plurality of second-strand sequencing reads; and
(d) comparing the first-strand sequencing reads with the second-strand sequencing reads, and generating an error-corrected sequence of the double-stranded target nucleic acid molecule by distinguishing erroneous nucleotides in one strand that lack a matched base change in the complementary strand.
40 . The method of claim 39 , wherein the double-stranded target nucleic acid molecule comprises (i) a DNA molecule, (ii) an RNA molecule, or (iii) a cDNA molecule derived from an RNA molecule.
41 . The method of claim 39 , wherein the cypher-target nucleic acid complex comprises at least two nucleic acid molecule priming sites.
42 . The method of claim 39 , wherein the cypher-target nucleic acid complex comprises an identifier sequence on both strands.
43 . The method of claim 39 , wherein the cypher-target nucleic acid complex comprises an identifier sequence at each end.
44 . The method of claim 42 , wherein the random or partially-random identifier sequence is double-stranded.
45 . The method of claim 44 , wherein the random or partially-random identifier sequence comprises about 5 to about 20 nucleotides.
46 . The method of claim 39 , wherein the random or partially-random identifier sequence uniquely labels the double-stranded target nucleic acid molecule.
47 . The method of claim 39 , wherein the double-stranded cypher uniquely links each strand of the double-stranded target nucleic acid molecule relative to its original complementary strand.
48 . The method of claim 39 , wherein the target nucleic acid molecule is ligated to a distinct cypher on each end, thereby providing a unique pair of identifiers for the target nucleic acid molecule.
49 . The method of claim 39 , wherein a unique pair of identifiers is provided for each target nucleic acid molecule in a plurality of target nucleic acid molecules in the ligating step.