Methods and compositions for phased sequencing
The present disclosure provides methods and compositions for molecular tagging of complex populations of nucleic acid molecules. The disclosure provides methods and compositions to obtain phase information of tagged nucleic acid molecules from high-throughput nucleic acid sequencing data.
1 . A method comprising:
a) providing a sample comprising a plurality of nucleic acids, wherein said plurality of nucleic acids comprises a nucleic acid strand, wherein said nucleic acid strand comprises an adaptor comprising a molecular barcode;
b) appending an elongation sequence to said nucleic acid strand, wherein said elongation sequence is complementary to at least a portion of a nucleic acid sequence in said nucleic acid strand;
c) annealing said elongation sequence to said portion of said nucleic acid sequence in said nucleic acid strand, thereby generating a partially-duplexed nucleic acid strand, wherein said partially-duplexed nucleic acid strand comprises a 5′ portion comprising a single-stranded region and a 3′ portion comprising said elongation sequence in an intramolecular duplex with said portion of said nucleic acid sequence; and
d) extending said elongation sequence with a polymerase using said 5′ portion of said partially-duplexed nucleic acid strand as a template, thereby generating an extended nucleic acid,
wherein said extended nucleic acid comprises a stem-loop structure comprising a hybridized region and an unhybridized region,
wherein said hybridized region comprises a first strand and a second strand, wherein said first strand comprises a 5′ end of said extended nucleic acid and said second strand comprises a 3′ end of said extended nucleic acid, and
wherein a 3′ end of said unhybridized region comprises said molecular barcode.
2 . The method of claim 1 , wherein a 3′ end of said nucleic acid strand comprises said adaptor.
3 . The method of claim 1 , wherein a 3′ end of said adaptor comprises said elongation sequence.
4 . The method of claim 1 , wherein a 5′ end of said second strand comprises said elongation sequence.
5 . The method of claim 1 , wherein said unhybridized region is 3′ to said first strand.
6 . The method of claim 1 , wherein said nucleic acid strand comprises DNA.
7 . The method of claim 1 , wherein said nucleic acid strand is generated from RNA, and wherein the method further comprises reverse transcribing said RNA before step a).
8 . The method of claim 1 , further comprising appending said adaptor to a nucleic acid molecule to generate said nucleic acid strand comprising said adaptor.
9 . The method of claim 8 , wherein said appending is performed by ligation or by polymerase chain reaction (PCR).
10 . The method of claim 8 , further comprising purifying said nucleic acid strand comprising said adaptor after said appending.
11 . The method of claim 10 , wherein said purifying comprises use of solid phase reversible immobilization, column-based solid phase extraction, or gel filtration to remove one or more unappended adaptors.
12 . The method of claim 1 , further comprising amplifying said nucleic acid strand comprising said adaptor prior to step a).
13 . The method of claim 1 , further comprising denaturing a double-stranded DNA molecule comprising said adaptor prior to step a), thereby generating a single-stranded nucleic acid comprising said nucleic acid strand comprising said adaptor.
14 . The method of claim 13 , wherein said denaturing comprises:
a) biotinylating a strand of said double-stranded DNA molecule to generate a biotinylated double-stranded DNA molecule;
b) binding said biotinylated double-stranded DNA molecule to a streptavidin-coated surface; and
c) washing said streptavidin-coated surface to release a non-biotinylated DNA strand, thereby denaturing said double-stranded DNA molecule.
15 . The method of claim 13 , wherein said denaturing comprises heating said double-stranded DNA molecule or alkaline denaturation.
16 . The method of claim 1 , wherein said elongation sequence comprises a random sequence.
17 . The method of claim 1 , wherein said elongation sequence is substantially or completely complementary to said portion of said nucleic acid sequence.
18 . The method of claim 1 , wherein said molecular barcode comprises a random or semi-random sequence.
19 . The method of claim 1 , wherein said plurality of nucleic acids in said sample comprises a first adaptor comprising a unique molecular barcode.
20 . The method of claim 19 , wherein said first adaptor in said plurality of nucleic acids further comprises a second barcode common to individual single-stranded nucleic acids within said plurality of single-stranded nucleic acids.
21 . The method of claim 1 , further comprising appending an additional adaptor to said extended nucleic acid.
22 . The method of claim 21 , wherein said appending is performed by ligating or by PCR.
23 . The method of claim 21 , wherein said additional adaptor is appended at a 3′ end of said extended nucleic acid.
24 . The method of claim 21 , further comprising amplifying said extended nucleic acid appended to said additional adaptor.
25 . The method of claim 1 , further comprising annealing a padlock probe to said extended nucleic acid, wherein said padlock probe comprises a 5′ end complementary to a region of interest primer sequence within said adaptor and a 3′ end complementary to a region of said extended nucleic acid, wherein the 5′ end and the 3′ end are connected by a linker sequence.
26 . The method of claim 25 , further comprising extending said 3′ end of said padlock probe to generate an extended nucleic acid comprising said padlock probe and sequence complementary to said portion of said nucleic acid sequence.
27 . The method of claim 26 , further comprising ligating a 5′ end and a 3′ end of said extended nucleic acid comprising said padlock probe and said sequence complementary to said portion of said nucleic acid sequence, thereby generating a circularized nucleic acid comprising said padlock probe and said sequence complementary to said portion of said nucleic acid sequence.
28 . The method of claim 27 , further comprising amplifying said circularized nucleic acid using a universal primer complementary to a universal sequencing adaptor sequence within said linker sequence of said padlock probe, thereby generating linearized nucleic acids comprising said molecular barcode of said nucleic acid and a sequence complementary to said universal primer.
29 . A method comprising:
a) appending a first adaptor to a nucleic acid in a plurality of nucleic acids, thereby generating a barcoded nucleic acid comprising said first adaptor, wherein said first adaptor comprises a molecular barcode, wherein said nucleic acid comprises a first target region and a second target region;
b) amplifying said barcoded nucleic acid, thereby generating amplified barcoded nucleic acids;
c) appending an elongation sequence to a barcoded nucleic acid in said amplified barcoded nucleic acids, thereby generating a barcoded nucleic acid comprising said elongation sequence, wherein said elongation sequence is complementary to at least a portion of a nucleic acid sequence in a strand of said barcoded nucleic acid, wherein said strand comprises said elongation sequence and said first adaptor;
d) annealing said elongation sequence to said portion of said nucleic acid sequence in said strand of said barcoded nucleic acid, thereby generating a partially-duplex nucleic acid, wherein said partially-duplex nucleic acid comprises a 5′ portion comprising a single-stranded region and a 3′ portion comprising said elongation sequence in an intramolecular duplex with said portion of said nucleic acid sequence;
e) extending said elongation sequence with a polymerase using said 5′ portion of said partially-duplex nucleic acid as a template, thereby generating an extended nucleic acid;
f) appending a second adaptor to said extended nucleic acid, thereby generating an extended nucleic acid comprising said first adaptor and said second adaptor, wherein said second adaptor comprises a sequence complementary to a sequencing primer; and
g) amplifying said extended nucleic acid comprising said first adaptor and said second adaptor with a first primer and a second primer, wherein said first primer anneals to said first adaptor or a complement thereof, and wherein said second primer anneals to said second adaptor or a complement thereof.