Conjugates of guide RNA-Cas protein complex
Provided herein are compositions of conjugates of a guide RNA(s)-CRISPR Cas protein (RNP) complex. The conjugate comprises a guide RNA(s)-CRISPR Cas protein (RNP) complex and one or more molecules selected from PEG, non-PEG polymers, ligands for cellular receptors, lipids, oligonucleotides, polysaccharides and peptides and chemically linked to the Cas protein and/or guide RNA(s). The conjugates are delivered to targeted cells as RNP complexes, or formed in targeted cells from guide RNA conjugates and an mRNA or a viral vector encoding a Cas protein, or formed in targeted cells from a crRNA conjugates and a viral vector encoding both a Cas protein and a tracrRNA. Also provided are preparation methods and uses of these conjugates.
1 . A conjugate comprising:
i. a clustered regularly interspaced short palindromic repeats (CRISPR)-Cas protein-chemically ligated guide RNA (lgRNA) complex (RNP); and
ii. one or more molecules selected from the group consisting of antibodies, aptamers, and a single strand nucleic acid template,
wherein said lgRNA comprises:
(i) a spacer of an oligonucleotide that targets a DNA sequence, and
(ii) at least one non-nucleotide linker located after its first nucleotide and before its last nucleotide,
and wherein said one or more molecules are covalently linked to said lgRNA by one or more non-nucleotide linkers at one or more positions other than the 5′-end of said lgRNA.
2 . Said conjugate of claim 1 , wherein each of said non-nucleotide linkers is independently selected from a group consisting of chemical structures comprising:
i. an optionally substituted M core structure of Formula M-1 to M-18:
wherein X=O, S, NH, or CH 2 ; X 1 =N or CH; X 2 =N or CH; R M =H, CH 3 , alkyl, aryl, or heteroaryl, m=0 to 3 and n=0 to 3, and
ii. two L linkers, and each said L linker comprises one or more structures selected from the group consisting of L-1 to L-24:
wherein m=0 to 16 and n=0 to 16,
and wherein
said L linkers and said M core structure are joined as L-M-L, wherein the two L linkers are the same or different, and attached to two terminal nucleotides of Formula Nuc-1 to Nuc-19:
wherein the attached positions are
to L-M-L and
to upstream and downstream oligonucleotides, respectively, and wherein R is H, OH,
CH 2 OH,
F, NH 2 , OMe, CH 2 OMe, OCH 2 CH 2 OMe, an alkyl, a cycloalkyl, an aryl, or a heteroaryl; R′ is H, OH,
CH 2 OH
F, NH 2 , OMe, CH 2 OMe, OCH 2 CH 2 OMe, an alkyl, a cycloalkyl, an aryl, or a heteroaryl, and Q is a natural or a non-natural nucleic acid base.
3 . Said conjugate of claim 1 , wherein said Cas protein is covalently linked with one or more molecules selected from the group consisting of polyethylene glycol (PEG), non-PEG polymers, ligands of cellular receptors, lipids, oligonucleotides, antibodies, polysaccharides, aptamers, glycans and peptides to form a Cas protein conjugate, and the said more molecules can be the same or different.
4 . Said conjugate of claim 3 , wherein said one or more molecules covalently linked to said Cas protein comprise molecules for targeted cellular delivery.
5 . Said conjugate of claim 1 , comprising one or more nucleotides selected from the group consisting of:
wherein R is H, OH, F, OMe, or OCH 2 CH 2 OCH 3 and Q is a nucleobase or modified nucleobase selected from the group consisting of B-1 through 38 and B-40 through 45:
wherein
(i) Z is N or CR 10 ; and
(ii) R 9 , R 10 , R 11 and R 12 are independently H, F, Cl, Br, I, OH, OR′, SH, SR′, SeH, SeR′, NH 2 , NHR′, NHOH, NHOR′, NR′OR′, NR′ 2 , NHNH 2 , NR′NH 2 , NR′NHR′, NHNR′ 2 , NR′NR′ 2 , alkyl of C 1 -C 6 , halogenated alkyl of C 1 -C 6 , alkenyl of C 2 -C 6 , halogenated alkenyl of C 2 -C 6 , CN, alkynyl of C 2 -C 6 , halogenated alkynyl of C 2 -C 6 , alkoxy of C 1 -C 6 , halogenated alkoxy of C 1 -C 6 , CN, CO 2 H, CO 2 R′, CONH 2 , CONHR′, CONR′ 2 , CH═CHCO 2 H, or CH═CHCO 2 R′, wherein R′ is H, an optionally substituted alkyl, an optionally substituted cycloalkyl, an optionally substituted alkynyl of C 2 -C 6 , an optionally substituted alkenyl of C 2 -C 6 , an optionally substituted aryl, or an optionally substituted heteroaryl, or when R′ is part of NR′ 2 and NR′ 2 is a heterocyclic group linked at its N.
6 . Said conjugate of claim 1 , wherein said spacer is selected from sequences of 12-20 nt in human immunodeficiency virus (HIV) genomes, of which each thymine is replaced with uracil.
7 . Said conjugate of claim 1 , wherein said spacer is selected from sequences of 12-20 nt in hepatitis B virus (HBV) genomes, of which each thymine is replaced with uracil.
8 . Said conjugate of claim 1 , wherein said spacer is selected from sequences of 12-20 nt in herpes simplex virus (HSV) genomes, of which each thymine is replaced with uracil.
9 . Said conjugate of claim 1 , wherein said spacer is selected from sequences of 12-20 nt in Epstein-Barr virus (EBV) genomes, of which each thymine is replaced with uracil.
10 . Said conjugate of claim 1 , wherein said spacer is selected from sequences of 12-20 nt of a host genome to be edited, of which each thymine is replaced with uracil.
11 . Said conjugate of claim 1 , wherein said nucleic acid template comprises a gene editing sequence flanked with two homology arms, said homology arms each overlapping with the target strand or the non-target strand of a target DNA duplex to be edited, and wherein said homology arms are optionally modified.
12 . Said conjugate of claim 11 , wherein said gene editing sequence comprises an oligonucleotide sequence to introduce one or more stop codons selected from the group consisting of 5′-(tga)-3′, 5′-(taa)-3′, 5′-(tag)-3′, 5′-(tga-ntga-ntga)-3′ (Sequence No. 84), 5′-(tga-ntga-ntaa)-3′ (Sequence No. 85), 5′-(tga-ntga-ntag)-3′ (Sequence No. 86), 5′-(tga-ntaa-ntga)-3′ (Sequence No. 87), 5′(tga-ntaa-ntaa)-3′ (Sequence No. 88), 5′-(tga-ntaa-ntag)-3′ (Sequence No. 89), 5′-(tga-ntga-ntga)-3′ (Sequence No. 90), 5′-(tga-ntga-ntaa)-3′ (Sequence No. 91), 5′-(tga-ntga-ntag)-3′ (Sequence No. 92), 5′-(taa-ntga-ntga)-3′ (Sequence No. 93), 5′-(taa-ntga-ntaa)-3′ (Sequence No. 94), 5′-(taa-ntga-ntag)-3′ (Sequence No. 95), 5′-(taa-ntaa-ntga)-3′ (Sequence No. 96), 5′-(taa-ntaa-ntaa)-3′ (Sequence No. 97), 5′-(taa-ntaa-ntag)-3′ (Sequence No. 98), 5′-(taa-ntga-ntga)-3′ (Sequence 99), 5′-(taa-ntga-ntaa)-3′ (Sequence 100), 5′-(taa-ntga-ntag)-3′ (Sequence No. 101), 5′-(tag-ntga-ntga)-3′ (Sequence No. 102), 5′-(tag-ntga-ntaa)-3′ (Sequence No. 103), 5′-(tag-ntga-ntag)-3′ (Sequence No. 104), 5′-(tag-ntaa-ntga)-3′ (Sequence No. 105), 5′-(tag-ntaa-ntaa)-3′ (Sequence No. 106), 5′-(tag-ntaa-ntag)-3′ (Sequence No. 107), 5′-(tag-ntga-ntga)-3′ (Sequence 108), 5′-(tag-ntga-ntaa)-3′ (Sequence No. 109) and 5′-(tag-ntga-ntag)-3′ (Sequence No. 110), wherein n is any nucleotide.
13 . Said conjugate of claim 11 , wherein said gene editing sequence comprises an oligonucleotide sequence to introduce one or more transcription cis-regulatory elements.
14 . Said conjugate of claim 1 comprising a mixture of lgRNA conjugates of various spacers targeting at different loci of target genomes, and/or sequences of variants or viral quasispecies of a single locus of target genomes.
15 . Said conjugate of claim 1 , wherein said Cas protein is a recombinant engineered class 2 endonuclease, catalytically inactive class 2 Cas endonuclease, nickase, or Cas-effector fusion protein.
16 . Said conjugate of claim 15 , wherein said Cas protein is a recombinant engineered endonuclease comprising at least two cysteines, and at least one of said cysteines are introduced by site directed mutations of internal amino acids, and said cysteines are conjugated with molecules for epitope masking and/or targeted delivery, and wherein the wild type cysteines are optionally mutated to avoid deactivation of said Cas protein due to covalent conjugations.
17 . Said conjugate of claim 15 , wherein said Cas-effector fusion protein is a Cas protein fused with a DNA polymerase, said nucleic acid template is a DNA, and said DNA polymerase uses said DNA template for gene editing.
18 . Said conjugate of claim 1 , wherein said complex is PEGylated.
19 . A method of editing a target gene using the conjugate of claim 1 , comprising the following steps:
i. cleaving a target gene DNA to be edited by said Cas protein of said conjugate, thereby generating a double-strand break or a nick; and
ii. hybridizing a resulting single-stranded DNA of the cleavage product with a 3′-homology arm of said covalently linked nucleic acid template of said conjugate and extending the 3′-end of said single DNA strand using said template as a template, thereby editing the target gene.
20 . Said conjugate of claim 1 , wherein said conjugate is self-assembled in a cell from (i) a Cas protein expressed from a viral plasmid or from mRNA encoding said Cas protein, and (ii) the lgRNA to which said one or more molecules are covalently linked.
21 . Said conjugate of claim 1 , wherein said CRISPR-Cas protein is a fusion protein of nickase or catalytically inactive Cas further comprising one or more functional protein domains selected from the group consisting of FokI, DNA polymerase, reverse transcriptase, DNA methyltransferase, nucleic acid deaminases, transcription activator(s), transcription repressor(s), and histone acetyltransferase and deacetylase.
22 . Said conjugate of claim 1 , wherein said covalently-linked nucleic acid template is used by a host protein for gene editing.
23 . Said conjugate of claim 22 , wherein said host protein is endogenous.
24 . Said conjugate of claim 22 , wherein both said host protein and said Cas protein are parts of a fusion protein.
25 . Said conjugate of claim 1 , wherein said covalently-linked nucleic acid template is an ssDNA.
26 . Pharmaceutical agents comprising:
i. an lgRNA; and
ii. one or more molecules selected from the group consisting of antibodies, aptamers, and a single strand nucleic acid template,
wherein said lgRNA is a guide RNA comprising:
(i) a spacer of an oligonucleotide that targets a DNA sequence; and
(ii) at least one non-nucleotide linker located after its first nucleotide and before its last nucleotide,
wherein said one or more molecules are covalently linked to said lgRNA by one or more non-nucleotide linkers at one or more positions other than the 5′-end of said LgRNA.
27 . Pharmaceutical agents of claim 26 , further comprising:
i. a Cas protein or a conjugate thereof;
ii. an mRNA encoding a Cas protein;
iii. a plasmid encoding a Cas protein; or
iv. a viral vector encoding a Cas protein.
28 . Pharmaceutical agents of claim 26 further comprising:
i. a fusion protein of nickase or catalytically inactive Cas and one or more functional protein domains selected from the group consisting of FokI, DNA polymerase, reverse transcriptase, DNA methyltransferase, nucleic acid deaminases, transcription activator(s), transcription repressor(s), histone acetyltransferase and histone deacetylase, or a conjugate thereof;
ii. an mRNA encoding said fusion protein; or
iii. a plasmid or a viral vector encoding said fusion protein.
29 . Pharmaceutical agents of claim 26 further comprising:
i. a Cas-DNA polymerase fusion protein or a conjugate thereof;
ii. an mRNA encoding a Cas-DNA polymerase fusion protein; or
iii. a plasmid or a viral vector encoding a Cas-DNA polymerase fusion protein, wherein the Cas-DNA polymerase fusion protein forms a complex with a guide RNA in targeted cells, and wherein the polymerase uses said nucleic acid template for gene editing.
30 . Pharmaceutical agents of claim 26 further comprising lipids.
31 . Pharmaceutical agents of claim 26 further comprising cell penetration peptides (CPP).