Methods and compositions for enrichment of target polynucleotides
High-fidelity, high-throughput nucleic acid sequencing enables healthcare practitioners and patients to gain insight into genetic variants and potential health risks. However, previous methods of nucleic acid sequencing often introduce sequencing errors (for example, mutations that arise during the preparation of a nucleic acid library, during amplification, or sequencing). Provided herein are methods and compositions for sequencing nucleic acids. Further provided are methods of identifying an error in a nucleic acid sequence.
1 . A method for sequencing a target nucleic acid duplex molecule, comprising:
(a) ligating an adaptor to each end of a target nucleic acid duplex,
wherein the nucleic acid duplex comprises first and second nucleic acid strands that are complementary to one another,
wherein each of said adaptors comprises: (i) a double stranded region that comprises a molecular barcode sequence; and (ii) first and second single stranded regions, wherein the molecular barcode sequence does not include a homopolymer sequence and is not self-complementary,
wherein the first single stranded region and a portion of the double stranded region of each of said adaptors comprises a sequence S1 that is 5′ of the molecular barcode sequence and the second single stranded region and a portion of the double stranded region of each adaptor comprises sequence S2′ that is 3′ of the molecular barcode sequence, wherein sequences S1 and S2′ are different;
(b) amplifying the ligated nucleic acids produced in (a) using primers with sequence S1 and the complement of S2′, thereby producing (i) amplified copies of the first strand that comprise sequence S1 at the 5′ end and a first molecular barcode sequence A between S1 and the target nucleic acid sequence of the first strand, and sequence S2′ at the 3′ end and a second molecular barcode sequence B between S2′ and the target nucleic acid sequence of the first strand; (ii) amplified copies of the second strand that comprise sequence S1 at the 5′ end and the complement B′ of the second molecular barcode sequence between S1 and the target nucleic acid sequence of the second strand, and sequence S2′ at the 3′ end and the complement A′ of the first molecular barcode sequence between S2′ and the target nucleic acid sequence of the second strand; and amplified complements of (i) and (ii);
(c) hybridizing and extending a primer that comprises: (i) a probe sequence that is complementary to a portion of the target nucleic acid sequence of the first and/or second strand, and (ii) a sequence S3, thereby producing primer extension products complementary to the second strand that comprise S3 at the 5′ end and either S1′ or S2′ at the 3′ end and that comprise molecular barcode sequence B between the target nucleic acid sequence and S1′ or S2′, and/or primer extension products complementary to the first strand that comprise S3 at the 5′ end and either S1′ or S2′ at the 3′ end and that comprise molecular barcode sequence A′ between the target nucleic acid sequence and S1′ or S2′;
(d) differentially amplifying the primer extension products,
wherein a first reaction comprises amplification using a first primer that comprises a sequence complementary to S3 and one or more sample index sequence(s), and a second primer that comprises S2 and one or more sample index sequence(s), and
wherein a second reaction comprises amplification using a first primer that comprises a sequence complementary to S3 and one or more sample index sequence(s), and a second primer that comprises S1 and one or more sample index sequence(s); and
(e) sequencing the amplified primer extension products.
2 . The method according to claim 1 , wherein the adaptors are selected from the group consisting of Y-shaped adaptors having first and second single stranded regions on separate polynucleotides, and U-shaped adaptors having first and second single stranded regions on the same polynucleotide.
3 . The method according to claim 1 , wherein step (c) comprises inclusion of blocking oligonucleotides that comprise sequences S1 and S2, and that each comprise a modification at the 3′ end to prevent extension by a polymerase.
4 . The method according to claim 1 , wherein the molecular barcode sequences are 4-15 nucleotides in length.
5 . The method according to claim 1 , further comprising combining the primer extension products produced in separate amplification reactions in (d), prior to sequencing.
6 . The method according to claim 1 , wherein barcode sequences A and B are different.
7 . The method according to claim 1 , wherein barcode sequences A and B are the same.
8 . The method according to claim 1 , wherein the sample index sequence(s), if any, on the first primer are different from the sample index sequence(s) on the second primer in step (d).
9 . The method according to claim 1 , wherein the sample index sequence(s), if any, on the first primer are the same as the sample index sequence(s) on the second primer in step (d).
10 . The method according to claim 1 , wherein said amplifying in step (b) comprises polymerase chain reaction (PCR) or a linear amplification method.
11 . The method according to claim 1 , wherein said differentially amplifying in step (d) comprises temporal or spatial separation of said first and second reactions.
12 . The method according to claim 1 , wherein said amplifying in step (d) comprises PCR or a linear amplification method.
13 . The method according to claim 1 , comprising performing step (c) with a plurality of different probes, in the same or different reaction mixtures, to produce a plurality of primer extension products that will provide different start points for sequencing of the target nucleic acid sequence.
14 . The method according to claim 1 , wherein the target nucleic acid duplex comprises cell-free DNA selected from cell-free tumor DNA or cell-free fetal DNA.
15 . The method according to claim 1 , wherein the target nucleic acid duplex is enriched from a nucleic acid library using a set of capture probes for a region of interest.
16 . The method according to claim 1 , comprising performing a first read of a first strand of the target sequence, comprising sequencing with first primers that comprise sequence S1 and second primers that comprise sequence S2, in the same or different reaction mixtures.
17 . The method according to claim 16 , wherein the first read with one of the primers begins 5′ of the molecular barcode sequence and the first read with the other primer begins at the molecular barcode sequence, or wherein the first read with both of the primers begins 5′ of the molecular barcode sequence.
18 . The method according to claim 16 , wherein the first read begins at the terminus or within a sample index sequence.
19 . The method according to claim 16 , comprising performing second reads to read sample index sequence(s).
20 . The method according to claim 16 , comprising compiling a set of first reads to construct a consensus sequence of the first strand of the target nucleic acid duplex.
21 . The method according to claim 20 , wherein the set of first strand reads is compiled based on sequence distance or alignment to a reference sequence.
22 . The method according to claim 20 , wherein constructing the first strand consensus sequence comprises:
comparing the first strand reads in the set of first strand reads;
identifying and removing errors in the set of first strand reads; and
constructing an error-corrected first strand consensus sequence.
23 . The method according to claim 22 , comprising identifying a mutation by comparison of the error-corrected consensus sequence to a reference sequence.
24 . The method according to claim 20 , further comprising sequencing the second strand of the target nucleic acid duplex and constructing a consensus sequence of the second strand of the target nucleic acid duplex.
25 . The method according to claim 24 , further comprising:
comparing the first strand consensus sequence and the second strand consensus sequence;
identifying and removing errors in the set of first strand reads and the set of second strand reads; and
constructing an error-corrected duplex consensus sequence.
26 . The method according to claim 25 , comprising identifying a chemical lesion by comparison of the sequences of the two strands in the error-corrected duplex consensus sequence.
27 . The method according to claim 25 , comprising distinguishing between (i) a chemical lesion or introduced sequence error, and (ii) a mutation, by comparison of the sequences of the two strands in the error-corrected duplex consensus sequence, wherein an error present in one strand indicates a chemical lesion or introduced sequence error, and an error present on both strands indicates a mutation.