COMPOSITIONS AND METHODS FOR ACCURATELY IDENTIFYING MUTATIONS
The present disclosure provides compositions and methods for accurately detecting mutations by uniquely tagging double stranded nucleic acid molecules with dual cyphers such that sequence data obtained from a sense strand can be linked to sequence data obtained from an anti-sense strand when sequenced, for example, by massively parallel sequencing methods.
1 .- 38 . (canceled)
39 . A method of sequencing DNA, the method comprising:
(a) attaching cypher polynucleotides to double-stranded DNA fragments to generate double-stranded cypher-target nucleic acid complexes, wherein the cypher polynucleotides comprise bar codes selected from a plurality of distinct bar code sequences;
(b) amplifying original strands of the cypher-target nucleic acid complexes to produce a plurality of cypher-target amplification products from first strands and complementary second strands of the cypher-target nucleic acid complexes;
(c) sequencing the cypher-target amplification products to produce a plurality of first-strand sequencing reads and a plurality of second-strand sequencing reads; and
(d) for each of a plurality of the cypher-target nucleic acid complexes:
(i) comparing the first-strand sequencing reads and second-strand sequencing reads to a reference sequence to identify one or more sequence correspondences to the reference sequence;
(ii) analyzing the one or more sequence correspondences to identify a sequence variation; and
(iii) identifying the sequence variation as a true mutation or an artifact mutation, wherein a true mutation is a sequence variation relative to the reference sequence that is consistent between the first strand sequencing reads and second strand sequencing reads, and wherein an artifact mutation is a sequence variation relative to the reference sequence that is not consistent between the first strand sequencing reads and second strand sequencing reads.
40 . The method of claim 39 , wherein:
(a) the artifact mutation is a processing error or a site of DNA damage;
(b) comparing the first-strand sequencing reads and second-strand sequencing reads to the reference sequence comprises comparing the first-strand sequencing reads to the second-strand sequencing reads; and
(c) the method further comprises:
(i) identifying nucleotide bases that are not consistent between the first strand sequencing reads and second strand sequencing reads; and
(ii) identifying nucleotide bases that are consistent between the first strand sequencing reads and second strand sequencing reads.
41 . The method of claim 40 , wherein prior to comparing the first-strand sequencing reads and second-strand sequencing reads to the reference sequence, the method comprises grouping the first-strand sequencing reads and second-strand sequencing reads based on at least the bar code sequences.
42 . The method of claim 39 , further comprising generating an error-corrected sequence for a plurality of the cypher-target nucleic acid molecules, wherein each error-corrected sequence comprises nucleotide bases at which the first-strand sequencing reads and second-strand sequencing reads are in agreement.
43 . The method of claim 42 , further comprising comparing the error-corrected sequence to the reference sequence, and identifying a mutation occurring at a particular position in the error-corrected sequence as a true mutation.
44 . The method of claim 43 , further comprising comparing the error corrected sequence to the reference sequence and identifying a mutation type of the true mutation.
45 . The method of claim 44 , wherein the mutation type is a transition, a substitution, an insertion, or a mutation of a single nucleotide.
46 . The method of claim 42 , further comprising identifying a nucleotide sequence at a particular position in the error-corrected sequence as a true nucleotide sequence.
47 . The method of claim 46 , wherein the true nucleotide sequence comprises a true mutation relative to the reference sequence.
48 . The method of claim 39 , wherein:
(a) comparing the first-strand sequencing reads and second-strand sequencing reads to the reference sequence comprises comparing the first-strand sequencing reads to the second-strand sequencing reads; and
(b) the method further comprises identifying non-complementary bases between the first-strand sequencing reads and the second-strand sequencing reads as experimental errors or sites of DNA damage.
49 . The method of claim 39 , wherein amplifying original strands comprises amplifying original strands via bridge amplification, emulsion amplification, nano-ball amplification, or PCR amplification.
50 . The method of claim 39 , wherein the double-stranded DNA fragments comprise a deaminated cytosine.
51 . The method of claim 50 , wherein the method further comprises enzymatically treating the double-stranded DNA molecules to repair damaged ends thereof prior to the attaching.
52 . The method of claim 39 , wherein the cypher-target nucleic acid complexes subjected to the amplifying step comprise double-stranded DNA fragments that range in size from 100 to 1,000 nucleotides.
53 . The method of claim 52 , wherein the cypher-target nucleic acid complexes subjected to the amplifying step comprise double-stranded DNA fragments that range in size from 150 to 500 nucleotides.
54 . The method of claim 39 , further comprising providing a sample comprising the double-stranded DNA fragments from a patient tissue.
55 . The method of claim 39 , wherein prior to comparing the first-strand sequencing reads and second-strand sequencing reads to a reference sequence, the method comprises grouping the first-strand sequencing reads and second-strand sequencing reads based on at least the bar code sequences.
56 . The method of claim 39 , wherein the double-stranded DNA fragments were generated by nuclease cleavage.
57 . The method of claim 56 , wherein the nuclease is a restriction endonuclease.
58 . The method of claim 39 , further comprising grouping the first-strand sequencing reads and second-strand sequencing reads for a particular cypher-target nucleic acid complex based on at least a bar code sequence.
59 . The method of claim 39 , wherein prior to sequencing, the method further comprises purifying a plurality of cypher-target nucleic acid complexes, wherein the purified cypher-target nucleic acid complexes comprise nucleic acid molecules that map to specific genomic regions.
60 . The method of claim 39 , wherein the double-stranded DNA fragments range in size from 100 to 1,000 nucleotides.
61 . The method of claim 60 , wherein the double-stranded DNA fragments range in size from 150 to 500 nucleotides.
62 . A method of sequencing DNA, the method comprising:
(a) attaching partially single-stranded cypher polynucleotides comprising bar codes selected from a plurality of distinct bar code sequences to double-stranded DNA fragments obtained from a patient sample, wherein attachment of the adapters to the double-stranded DNA fragments generates a library of double-stranded cypher-target nucleic acid complexes;
(b) amplifying the cypher-target nucleic acid complexes in the library to produce a plurality of cypher-target amplification products from first strands and complementary second strands of the cypher-target nucleic acid complexes;
(c) sequencing the cypher-target amplification products to produce a plurality of sequencing reads comprising a bar code sequence and DNA fragment-specific sequence; and
(d) for at least some of the cypher-target nucleic acid complexes:
(i) grouping the sequencing reads based on the bar code sequence and the DNA fragment-specific sequence;
(ii) comparing sequencing reads within the groups to generate an error-corrected sequence for each of a plurality of the double-stranded DNA fragments;
(iii) comparing the error-corrected sequences to a reference sequence; and
(iv) analyzing one or more sequence correspondences between the error-corrected sequence and the reference sequence to identify a true mutation, wherein the true mutation is a mutation present in both the first strand and complementary second strand of the cypher-target nucleic acid complex.
63 . The method of claim 62 , wherein the patient sample comprises tissue obtained from the patient.
64 . The method of claim 62 , wherein at least some of the double-stranded DNA fragments are derived from a tumor or circulating tumor cells.
65 . The method of claim 62 , wherein the patient sample is derived from a patient having tumor cells, wherein the true mutation is a mutation that confers resistance to therapy, and wherein the true mutation is present in an error-corrected sequence derived from one of the double-stranded DNA fragments in the patient sample.
66 . The method of claim 65 , wherein the double-stranded DNA fragments in the patient sample comprise double-stranded DNA fragments obtained from the tumor cells.
67 . The method of claim 62 , wherein
(i) at least two of the bar codes are identical in sequence and are attached to different double-stranded DNA fragments, thereby non-uniquely tagging the different double-stranded DNA fragments; and
(ii) the different double-stranded DNA fragments that are non-uniquely tagged comprise distinguishable end sequences.
68 . The method of claim 62 , further comprising purifying a plurality of cypher-target nucleic acid complexes, wherein the purified cypher-target nucleic acid complexes comprise nucleic acid molecules that map to specific genomic regions.
69 . The method of claim 68 , wherein (a) prior to sequencing, the cypher-target nucleic acid complexes or amplification products thereof are selectively enriched by hybridization to substrate bound oligonucleotides; and (b) the sequencing produces sequencing reads for the molecules that map to the specific genomic regions.
70 . The method of claim 62 , wherein the bar code sequences are 6 nucleotides in length.