COMPOSITIONS AND METHODS FOR ACCURATELY IDENTIFYING MUTATIONS
The present disclosure provides compositions and methods for accurately detecting mutations by uniquely tagging double stranded nucleic acid molecules with dual cyphers such that sequence data obtained from a sense strand can be linked to sequence data obtained from an anti-sense strand when sequenced, for example, by massively parallel sequencing methods.
1 .- 38 . (canceled)
39 . A method of determining an error-corrected sequence of a double-stranded target nucleic acid molecule, comprising:
(a) ligating the double-stranded target nucleic acid molecule to at least one cypher polynucleotide, to form a cypher-target nucleic acid complex, wherein the at least one cypher polynucleotide comprises:
(i) a random or partially-random identifier sequence that alone or in combination with an end of the target nucleic acid molecule uniquely labels the double-stranded target nucleic acid molecule; and
(ii) a nucleotide sequence that tags each strand of the cypher-target nucleic acid complex such that each strand of the cypher-target nucleic acid complex has a distinct nucleotide sequence relative to its complementary strand;
(b) amplifying each strand of the cypher-target nucleic acid complex to produce a plurality of cypher-target amplification products from each of a first strand and a complementary second strand of the cypher-target nucleic acid complex;
(c) sequencing the cypher-target amplification products to produce a plurality of first-strand sequencing reads and a plurality of second-strand sequencing reads; and
(d) comparing the first-strand sequencing reads with the second-strand sequencing reads, and generating an error-corrected sequence of the double-stranded target nucleic acid molecule by distinguishing erroneous nucleotides in one strand that lack a matched base change in the complementary strand.
40 . The method of claim 39 , wherein the double-stranded target nucleic acid molecule comprises (i) a DNA molecule, or (ii) an RNA molecule.
41 . The method of claim 39 , wherein the cypher-target nucleic acid complex comprises at least two nucleic acid molecule priming sites.
42 . The method of claim 39 , wherein the cypher-target nucleic acid complex comprises an identifier sequence on both strands.
43 . The method of claim 39 , wherein the cypher-target nucleic acid complex comprises an identifier sequence at each end.
44 . The method of claim 42 , wherein the random or partially-random identifier sequence is double-stranded.
45 . The method of claim 44 , wherein the random or partially-random identifier sequence comprises about 5 to about 20 nucleotides.