COMPOSITIONS AND METHODS FOR ACCURATELY IDENTIFYING MUTATIONS
The present disclosure provides compositions and methods for accurately detecting mutations by uniquely tagging double stranded nucleic acid molecules with dual cyphers such that sequence data obtained from a sense strand can be linked to sequence data obtained from an anti-sense strand when sequenced, for example, by massively parallel sequencing methods.
1 .- 38 . (canceled)
39 . A method comprising:
(a) providing a sample comprising a set of double-stranded polynucleotide molecules, each double-stranded polynucleotide molecule including first and second complementary strands;
(b) tagging said double-stranded polynucleotide molecules with a set of double-stranded cypher polynucleotides to form double-stranded cypher-target nucleic acid complexes, wherein the cypher polynucleotides comprise double-stranded bar codes;
(c) sequencing the cypher-target nucleic acid complexes to produce a plurality of first-strand sequencing reads and a plurality of second-strand sequencing reads;
(d) grouping sequencing reads based on the bar codes, wherein a group comprises sequencing reads for a first tagged strand and a second differently-tagged complementary strand derived from an original double-stranded polynucleotide molecule in said set; and
(e) quantifying said groups of sequencing reads and the read depth of said groups of sequencing reads.
40 . The method of claim 39 , further comprising quantifying one or more mutations in the sample based on said quantification of said groups of said sequencing reads that map to one or more genetic loci.
41 . The method of claim 39 , wherein said cypher polynucleotides are not sequencing adapters.
42 . The method of claim 39 , wherein for a plurality of cypher-target nucleic acid complexes, the method further comprises comparing first-strand sequencing reads with second-strand sequencing reads produced from amplified products of one of the cypher-target nucleic acid complexes to form an error-corrected sequence of the original double-stranded polynucleotide molecule.
43 . The method of claim 42 , further comprising identifying double-stranded polynucleotide molecules comprising a sequence variant at one or more genetic loci.
44 . The method of claim 42 , further comprising quantifying a mutation by mapping error-corrected sequences to a reference sequence, and quantifying the error-corrected sequences corresponding to one or more genetic loci of the reference sequence.
45 . The method of claim 39 , further comprising quantifying said groups of sequencing reads that map to a genetic locus, wherein the reads for the first and second strand of a group comprise a sequence variant.
46 . The method of claim 39 , wherein said set of double-stranded polynucleotide molecules comprises double-stranded circulating nucleic acid molecules obtained from a patient sample.
47 . The method of claim 46 , wherein said double-stranded circulating nucleic acid molecules comprise a mutation present at a frequency of 2.1×10 −6 or lower.
48 . The method of claim 47 , wherein said mutation: (i) is a single nucleotide mutation, (ii) is a cancer biomarker, and (iii) maps to a cancer-associated genetic locus in a reference genome.
49 . The method of claim 48 , further comprising quantifying the single nucleotide mutation cancer biomarker.
50 . The method of claim 49 , wherein quantifying the single nucleotide mutation cancer biomarker comprises quantifying groups of sequencing reads having the single nucleotide mutation cancer biomarker that maps to the cancer-associated genetic locus.
51 . The method of claim 46 , wherein the circulating nucleic acid molecules comprise genomic DNA originating from one or more of a healthy cell, a tumor cell, and a cancer cell.
52 . The method of claim 51 , wherein the circulating nucleic acid molecules comprise plasma DNA biomarkers.
53 . The method of claim 51 , wherein the patient sample comprises a blood sample.
54 . The method of claim 51 , wherein the circulating nucleic acid molecules are obtained from plasma.
55 . The method of claim 39 , wherein the set of double-stranded polynucleotide molecules were generated by nuclease cleavage.
56 . The method of claim 55 , wherein the nuclease is a restriction endonuclease.
57 . The method of claim 55 , wherein the double-stranded polynucleotide molecules comprise overhangs or blunt ends.
58 . The method of claim 39 , wherein
(i) the bar codes are selected from a plurality of distinct bar code sequences;
(ii) at least two of the bar codes are identical in sequence and are ligated to different double-stranded polynucleotide molecules, thereby non-uniquely tagging the different double-stranded polynucleotide molecules; and
(iii) the different double-stranded polynucleotide molecules that are non-uniquely tagged comprise distinguishable end sequences.
59 . The method of claim 58 , wherein the bar code sequences comprise known oligonucleotide sequences.
60 . The method of claim 58 , wherein the bar code sequences comprise random or partially random sequences.
61 . The method of claim 39 , wherein the tagging step comprises attaching double-stranded cypher polynucleotides to both ends of each of the double-stranded polynucleotide molecules, and wherein individual cypher-target nucleic acid complexes can be distinguished by different pairs of bar codes.
62 . The method of claim 39 , wherein:
(a) the tagging step comprises attaching double-stranded cypher polynucleotides to both ends of each of the double-stranded polynucleotide molecules;
(b) at least two of the bar codes are identical in sequence and are ligated to different double-stranded polynucleotide molecules, thereby non-uniquely tagging the different double-stranded polynucleotide molecules; and
(c) said cypher-target nucleic acid complexes can be differentiated from other cypher-target nucleic acid complexes using:
(i) a combination of a bar code and sequence information derived from the original double-stranded polynucleotide molecule,
(ii) a combination of a first bar code at a first end of the double-stranded polynucleotide molecule and a second bar code at a second end of the double-stranded polynucleotide molecule, or
(iii) a combination of (i) and (ii).
63 . The method of claim 62 , wherein in (c) (i) the bar code is a non-unique bar code, and in (c) (ii) the first bar code is a first non-unique bar code and the second bar code is a second non-unique bar code.
64 . The method of claim 39 , wherein prior to sequencing, the method further comprises purifying a plurality of cypher-target nucleic acid complexes comprising double-stranded polynucleotide molecules from specific genomic regions.
65 . The method of claim 39 , further comprising determining a total number of original double-stranded polynucleotide molecules in the sample based on the quantification of said groups of sequencing reads.