SINGLE DOMAIN ANTIBODY LIBRARIES WITH MAXIMIZED ANTIBODY DEVELOPABILITY CHARACTERISTICS
Described here are VHH antibody libraries with heavy chain variable domains framework scaffolds having complementary determining regions (CDRs) found in naturally-occurring human antibodies, and methods of making such antibody libraries. The antibody libraries are free of members that comprise one or more liabilities affecting one or more features of such members.
1 . A VHH antibody library, comprising:
nucleic acids encoding a framework region 1, a framework region 2, a framework region 3 and a framework region 4; and
a plurality of nucleic acids encoding a population of VHH genes comprising a population of CDR1s, a population of CDR2s, and a population of CDR3s located at the CDR1 region, the CDR2 region, and the CDR3 region of the VHH gene, wherein amino acid sequences of the CDR1s, the CDR2s, and the CDR3s are naturally replicated from naturally-occurring antibodies, wherein at least 90% of the population of CDR1s and at least 90% of the population of CDR2s are completely free of members comprising one or more of:
(i) a glycosylation site comprising the motif NXS, NXT, or NXC, in which X represents any naturally-occurring amino acid residue except for proline;
(ii) a deamidation site comprising the motif of NG, NS, NT, NN, NA, NH, ND, NQ, NF, NW or NY;
(iii) an isomerization site comprising the motif of DT, DH, DS, DO, DN, DR, DY or DD;
(iv) any cysteines,
(v) net charge greater than 1,
(vi) a tripeptide motif containing at least two residues with aromatic side chains comprising F, H, W or Y;
(vii) a polyspecificity site comprising the motif GO, GGG, RR, VG, W, WV, WW, WWW, YY, or WXW, in which X represents any amino acid residue;
(viii) a protease sensitive or hydrolysis prone site comprising the motif of DX, in which X is P, G, S, V, Y, F, Q, K, L, or D;
(ix) an integrin binding site comprising RGD, RYD, LDV, or KGD;
(x) a lysine glycation site comprising KE, EK, or ED;
(xi) a metal catalyzed fragmentation site comprising the motif of HS, SH, KT, HXS, or SXH, in which X represents any amino acid residue;
(xii) a polyspecificity aggregation site comprising a motif of X 1 X 2 × 1 , wherein each of X 1 , X 2 , and X 3 is independently selected from the group consisting of F, I, L, V, W and Y;
(xiii) a streptavidin binding motif comprises the motif HPQ, EPDW (SEQ ID NO: 49), PWXWL (SEQ ID NO: 50), in which X represents any amino acid residue, GDWVFI (SEQ ID NO: 51), or PWPWLG (SEQ ID NO: 52);
(xiv) one or more arginines;
(xv) a hydrophobic CDR sequence summed using reference numbers from Parker J M et al. Biochemistry. 1986 Sep. 23; 25(19):5425-32 and results in a value smaller than zero;
(xvi) a CDR mutation that reduces binding to protein A said CDR mutation 5 comprising any mutation in the last amino acid of the CDR2, according to the IMGT definition, to A, G, C, D, E, F, G, H, I, L, M, N, P, Q, S, V, W or Y;
wherein at least 95% of the population of CDR1s, at least 95% of the population of CDR2s, and at least 95% of the population of CDR3s are completely free of non-functional members;
wherein framework region 1, framework region 2, framework region 3 and framework region 4 are from a single therapeutic antibody VHH; and
wherein each framework region can contain up to 5 amino acid substitutions.
2 . The VHH antibody library of claim 1 , wherein the single VHH domain is selected from the therapeutic antibody selected from the group consisting of Caplacizumab, Envafolimab, Gontivimab, Isecarosmab, Ozoralizumab, Sonelokimab and Vobarilizumab.
3 . The VHH antibody library of claim 1 or claim 2 , wherein the naturally replicated nucleic acid sequences of the CDR1s, the CDR2s, and the CDR3s are synthesized.
4 . The VHH antibody library of claim 1 or claim 2 , wherein the nucleic acid sequences of the CDR3s are naturally replicated by obtaining the sequences from heavy chain CDR3s from donor lymphocytes.
5 . A method for producing an antibody library, comprising:
(a) providing a first plurality of nucleic acids encoding a population of naturally-replicated CDR1 fragments;
(b) providing a second plurality of nucleic acids encoding a population of naturally-replicated CDR2 fragments;
(c) providing a third plurality of nucleic acids encoding a population of naturally-replicated CDR3 fragments;
(d) providing a nucleic acid gene encoding a common VHH domain comprising heavy chain framework region 1, heavy chain framework region 2, heavy chain framework region 3 and heavy chain framework region 4, and inserting the first plurality of nucleic acids, the second plurality of nucleic acids, and the third plurality of nucleic acids into the CDR1 region, the CDR2 region, and the CDR3 region, respectively, of the common VHH domain, thereby producing a population of nucleic acids encoding a VHH domain library; wherein at least 90% of the population of heavy chain CDR1s, and at least 90% of the population of heavy chain CDR2s are completely free of members comprising one or more of:
(i) a glycosylation site comprising the motif NXS, NXT, or NXC, in which X represents any naturally-occurring amino acid residue except for proline;
(ii) a deamidation site comprising the motif of NG, NS, NT, NN, NA, NH, ND, NQ, NF, NW or NY;
(iii) an isomerization site comprising the motif of DT, DH, DS, DO, DN, DR, DY or DD;
(iv) any cysteines,
(v) net charge greater than 1,
(vi) a tripeptide motif containing at least two residues with aromatic side chains comprising F, H, W or Y;
(vii) a polyspecificity site comprising the motif GO, GGG, RR, VG, W, WV, WW, WWW, YY, or WXW, in which X represents any amino acid residue;
(viii) a protease sensitive or hydrolysis prone site comprising the motif of DX, in which X is P, G, S, V, Y, F, Q, K, L, or D;
(ix) an integrin binding site comprising RGD, RYD, LDV, or KGD;
(x) a lysine glycation site comprising KE, EK, or ED;
(xi) a metal catalyzed fragmentation site comprising the motif of HS, SH, KT, HXS, or SXH, in which X represents any amino acid residue;
(xii) a polyspecificity aggregation site comprising a motif of X 1 X 2 X 3 , wherein each of X 1 , X 2 , and X 3 is independently selected from the group consisting of F, I, L, V, W and Y;
(xiii) a streptavidin binding motif comprises the motif HPQ, EPDW (SEQ ID NO: 49), PWXWL (SEQ ID NO: 50), in which X represents any amino acid residue, GDWVFI (SEQ ID NO: 51), or PWPWLG (SEQ ID NO: 52);
(xiv) one or more arginines;
(xv) a hydrophobic CDR sequence summed using reference numbers from Parker J M et al. Biochemistry. 1986 Sep. 23; 25(19):5425-32 and results in a value smaller than zero;
(xvi) a CDR mutation that reduces binding to protein A said CDR mutation comprising any mutation in the last amino acid of the CDR2, according to the IMGT definition, to A, G, C, D, E, F, G, H, I, L, M, N, P, Q, S, V, W or Y;
wherein at least 95% of the population of heavy chain CDR1s, the population of 5 heavy chain CDR2s, and the population of heavy chain CDR3s are completely free of non-functional members;
wherein framework 1, framework 2, framework 3 and framework 4 are from a single VHH domain from a therapeutic comprised of one or more VHH domains; and
wherein each framework region can contain up to 5 amino acid substitutions.
6 . The method of claim 5 wherein the therapeutic comprised of one or more VHH domains is selected from the group consisting of Caplacizumab, Envafolimah, Gontivimab, Isecarosmab, Ozoralizumab, Sonelokimab and Vobarilizumab.
7 . The method of claims 5 or 6 , wherein the first plurality of nucleic acids, the second plurality of nucleic acids and the third plurality of nucleic acids is produced by a process comprising:
(a) obtaining amino acid sequences of the heavy chain CDR1 regions, heavy chain CDR2 regions and heavy chain CDR3 regions of a population of naturally-occurring antibodies;
(b) excluding from (a) the heavy chain CDR1 amino acid sequences, the heavy chain CDR2 amino acid sequences, and the heavy chain CDR3 amino acid sequences amino acid sequences that comprise any or all of (i) to (xvi) to obtain liability-free heavy chain CDR1 sequences, liability-free heavy chain CDR2 sequences and liability-free heavy chain CDR3 sequences; and
(c) synthesizing the first plurality of nucleic acids that encode the liability-free heavy chain CDR1 regions, the second plurality of nucleic acids that encode the liability-free heavy chain CDR2 regions, and the third plurality of nucleic acids that encode the liability-free heavy chain CDR2 regions.
8 . The method of claim 7 , wherein the third plurality of nucleic acids is produced by a process comprising:
(a) amplifying the heavy chain CDR3 regions from a population of B cells; and
(b) combining the third plurality of nucleic acids that encode the heavy chain CDR3 regions obtained in (a) with the remaining CDRs.
9 . The method of claim 7 , wherein the processes for producing the first plurality of nucleic acids, the second plurality of nucleic acids, and the third plurality of nucleic acids further comprise isolating functional members from the liability-free heavy chain CDR1 and CDR2 regions, and/or from the CDR3 regions, wherein:
(i) the functional members of the liability-free heavy chain CDR1 and CDR2 regions or the functional members of the CDR3 regions are isolated by expressing antibodies comprising the liability-free heavy chain CDR1 and CDR2 regions, and/or the CDR3 regions in host cells in a manner that the antibodies are displayed on surface of the host cells, isolating the antibodies that display on the host cells, and identifying the CDR1, CDR2, and/or CDR3 regions in the displayed antibodies, which are functional members of the CDR1, CDR2, and/or CDR3 regions; or
(ii) the functional members of the liability-free heavy chain CDR1 and CDR2 regions, and/or the CDR3 regions are isolated by expressing antibodies comprising the liability-free heavy chain CDR1 and CDR2 regions, and/or the CDR3 regions in fusion with a folding reporter, which optionally is β-lactamase or green fluorescent protein, or fragments thereof, to obtain members with improved folding.
10 . The antibody library of any of claims 1-4 wherein the CDR1 comprises human homolog sequences.
11 . The antibody library of claim 10 , wherein the CDR2 comprises human homolog sequences.
12 . The antibody library of claim 11 , wherein the framework region is derived from a single VHH domain from a therapeutic comprised of one or more VHH domains.
13 . The method of any one of claims 5-9 , wherein the heavy chain CDR1, CDR2, and CDR3 fragment and, the heavy chain variable domain gene are derived from naturally-occurring antibodies of a mammalian species.
14 . The method of claim 13 , wherein the mammalian species is human or camelid.