Nucleotide cleavable linkers with rigid spacers and uses thereof
Disclosed herein, inter alia, are compounds, compositions, and methods of use thereof for sequencing a nucleic acid.
1 . A compound having the formula:
wherein
B is a divalent nucleobase;
L 100 is a polymerase-compatible cleavable linker;
R 1 is a polyphosphate moiety, monophosphate moiety, 5′-O-nucleoside protecting group, nucleic acid moiety, hydrogen, or —OH;
R 2 is hydrogen, a polymerase-compatible cleavable moiety, or —OH;
R 3 is an —O-polymerase-compatible cleavable moiety, a polymerase-compatible cleavable moiety, hydrogen, —OH, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl;
R 4 is an anchor moiety or a detectable moiety;
L 200 has the formula:
R 201 is independently hydrogen, —CCl 3 , —CBr 3 , —CF 3 , —Cl 3 , —CHCl 2 , —CHBr 2 , —CHF 2 , —CHI 2 , —CH 2 Cl, —CH 2 Br, —CH 2 F, —CH 2 I, —CN, —OH, —COOH, —CONH 2 , —OCCl 3 , —OCF 3 , —OCBr 3 , —OCl 3 , —OCHCl 2 , —OCHBr 2 , —OCHI 2 , —OCHF 2 , —OCH 2 Cl, —OCH 2 Br, —OCH 2 I, —OCH 2 F, substituted or unsubstituted alkyl, or substituted or unsubstituted heteroalkyl;
R 202 is independently —SO 3 H, —SO 4 H, —SO 2 NH 2 , —PO 3 H, —PO 4 H, halogen, —CCl 3 , —CBr 3 , —CF 3 , —Cl 3 , —CHCl 2 , —CHBr 2 , —CHF 2 , —CHI 2 , —CH 2 Cl, —CH 2 Br, —CH 2 F, —CH 2 I, —CN, —OH, —NH 2 , —COOH, —CONH 2 , —NO 2 , —SH, —NHNH 2 , —ONH 2 , —NHC(O)NHNH 2 , —NHC(O)NH 2 , —NHSO 2 H, —NHC(O)H, —NHC(O)OH, —NHOH, —OCCl 3 , —OCF 3 , —OCBr 3 , —OC 3 , —OCHCl 2 , —OCHBr 2 , —OCHI 2 , —OCHF 2 , —OCH 2 Cl, —OCH 2 Br, —OCH 2 I, —OCH 2 F, —N 3 , —SF 5 , —SO 2 Cl, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, or L 212 -R 212A ;
L 202 is independently a covalent linker;
R 202A is independently a detectable moiety, anchor moiety, triplet state quencher moiety, or protein moiety;
z202 is independently an integer from 0 to 2;
R 201 and R 202 may optionally be joined to form a substituted or unsubstituted heterocycloalkyl;
W 203 and W 204 are independently CH, N, or C(R 202);
L 205 is independently a bond or —CH 2 NH—; and
z206 is an integer from 1 to 100;
wherein, when R 4 is a detectable moiety, then R 202 is L 202 -R 202A , and R 4 and R 202A are a FRET pair of detectable moieties.
2 . A compound having the formula:
wherein
B is a divalent nucleobase;
L 100 is a polymerase-compatible cleavable linker;
R 1 is a polyphosphate moiety, monophosphate moiety, 5′-O-nucleoside protecting group, nucleic acid moiety, hydrogen, or —OH;
R 2 is hydrogen, a polymerase-compatible cleavable moiety, or —OH;
R 3 is an —O-polymerase-compatible cleavable moiety, a polymerase-compatible cleavable moiety, hydrogen, —OH, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl;
R 4 is an anchor moiety or a detectable moiety; and
L 200 is a divalent polymer, divalent double-stranded nucleic acid, or divalent polypeptide.
3 . The compound of claim 2 , wherein the divalent polypeptide comprises the amino acid sequences (EAAAK) n1 (SEQ ID NO:1), (EP) n2 (SEQ ID NO:5), (KP) n3 (SEQ ID NO:6), (AP) n4 (SEQ ID NO:7), or (TPR) n5 (SEQ ID NO:8), wherein n1, n2, n3, n4, and n5 are each independently an integer from 2 to 20.
4 . The compound of claim 1 , wherein L 200 has the formula:
5 . The compound of claim 1 , wherein the triplet state quencher moiety is a monovalent ascorbic acid, monovalent cyclooctatetraene (COT), monovalent nitrobenzyl alcohol, monovalent methyl viologen, monovalent Trolox, or monovalent Trolox-quinone.
6 . The compound of claim 1 , wherein the polymerase-compatible cleavable moiety is independently -(substituted or unsubstituted alkylene)-SS-(unsubstituted alkyl).
7 . The compound of claim 1 , wherein the polymerase-compatible cleavable moiety is independently:
8 . The compound of claim 1 , wherein B is a divalent cytosine or a derivative thereof, divalent guanine or a derivative thereof, divalent adenine or a derivative thereof, divalent thymine or a derivative thereof, divalent uracil or a derivative thereof, divalent hypoxanthine or a derivative thereof, divalent xanthine or a derivative thereof, divalent 7-methylguanine or a derivative thereof, divalent 5,6-dihydrouracil or a derivative thereof, divalent 5-methylcytosine or a derivative thereof, or divalent 5-hydroxymethylcytosine or a derivative thereof.
9 . The compound of claim 1 , wherein B is
10 . The compound of claim 1 , wherein L 100 is a polymerase-compatible cleavable linker comprising:
wherein R 9 is substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl.
11 . The compound of claim 1 , wherein L 100 is a polymerase-compatible cleavable linker comprising
wherein R 102 is unsubstituted C 1 -C 4 alkyl.
12 . The compound of claim 1 , wherein L 100 is -L 101 -L 102 -L 103 -L 104 -L 105 -; and
L 101 , L 102 , L 103 , L 104 , and L 105 are independently a bond, —NH—, —O—, —C(O)—, —C(O)NH—, —NHC(O)—, —NHC(O)NH—, —C(O)O—, —OC(O)—, —SS—, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, or substituted or unsubstituted heteroarylene;
wherein at least one of L 101 , L 102 , L 103 , L 104 , and L 105 is not a bond.
13 . The compound of claim 12 , wherein L 100 is -L 101 -O—CH(—SR 100 )-L 103 -L 104 -L 105 -, -L 101 -O—C(CH 3 )(—SR 100 )-L 103 -L 104 -L 105 -, -L 101 -O—CH(N 3 )-L 103 -L 104 -L 105 -, or -L 101 -SS-L 103 -L 104 -L 105 -;
L 101 , L 103 , L 104 , and L 105 are independently a bond, —NH—, —O—, —C(O)—, —C(O)NH—, —NHC(O)—, —NHC(O)NH—, —C(O)O—, —OC(O)—, substituted or unsubstituted alkylene, substituted or unsubstituted heteroalkylene, substituted or unsubstituted cycloalkylene, substituted or unsubstituted heterocycloalkylene, substituted or unsubstituted arylene, or substituted or unsubstituted heteroarylene;
R 100 is —SR 102 or —CN; and
R 102 is unsubstituted C 1 -C 4 alkyl.
14 . The compound of claim 12 , wherein L 100 is -L 101 -O—CH(—SR 100 )-L 103 -L 104 -L 105 -, -L 101 -O—C(CH 3 )(—SR 100 )-L 103 -L 104 -L 105 -, -L 101 -O—CH(N 3 )-L 103 -L 104 -L 105 -, or -L 101 -SS-L 103 -L 104 -L 105 -;
L 101 is a substituted or unsubstituted C 1 -C 4 alkylene or substituted or unsubstituted 8 to 20 membered heteroalkylene;
L 103 is a bond or substituted or unsubstituted 2 to 10 membered heteroalkylene;
L 104 is a bond, substituted or unsubstituted 4 to 18 membered heteroalkylene, or substituted or unsubstituted phenylene;
L 105 is a bond or substituted or unsubstituted 4 to 18 membered heteroalkylene;
R 100 is —SR 102 or —CN; and
R 102 is unsubstituted C 1 -C 4 alkyl.
15 . The compound of claim 12 , wherein L 100 is
wherein
R 9 is substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl; and
R 102 is unsubstituted C 1 -C 4 alkyl.
16 . The compound of claim 12 , wherein L 100 is
wherein
R 9 is substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, or substituted or unsubstituted heteroaryl; and
R 102 is unsubstituted C 1 -C 4 alkyl.
17 . The compound of claim 12 , wherein L 100 is
18 . The compound of claim 1 , wherein R 4 is an anchor moiety.
19 . The compound of claim 18 , wherein the anchor moiety is biotin, azide, transcyclooctene (TCO), or phenyl boric acid (PBA).
20 . A method for sequencing a nucleic acid, comprising:
(i) incorporating in series with a nucleic acid polymerase, within a reaction vessel, one of four different compounds into a primer to create an extension strand, wherein said primer is hybridized to said nucleic acid and wherein each of the four different compounds comprises a unique detectable moiety or a unique anchor moiety;
(ii) if the compound of step (i) above comprises a unique anchor moiety, further adding to said reaction vessel a complementary anchor compound comprising a complementary anchor moiety to said unique anchor moiety bonded to a unique detectable moiety; and
(iii) detecting the unique detectable moiety of each incorporated compound or incorporated compound-complementary anchor compound complex, so as to thereby identify each incorporated compound in said extension strand, thereby sequencing the nucleic acid;
wherein each of said four different compounds is independently a compound of claim 1 .
21 . The method of claim 20 , further comprising adding to said reaction vessel a photodamage mitigating agent.