IP Library › Granted Patent US 12,367,329
Granted Patent B1
US 12,367,329 · App. 18/736,232 · Granted Jul 22, 2025

Protein binder search

Inventors: Salvatore J. Candido (San Francisco, CA); Alexander W. Rives (Brooklyn, NY); Thomas F. Hayes (San Francisco, CA); Jun Gong (Mountain View, CA)
Assignee: EvolutionaryScale, PBC
G06F30/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,329
App. No.
18/736,232
Granted
Jul 22, 2025
Kind
B1
Abstract

A specification of a binding target protein is received. A machine learning model is used to predict a plurality of candidates for a property of a selected amino acid position of a binder protein to bind to the binding target protein. For each selected property candidate of the plurality of property candidates, the selected property candidate is used as a model input to predict properties for one or more other amino acid positions into a corresponding candidate set of properties. The corresponding candidate sets are evaluated using a binding quality evaluation. Based on the evaluation, one of the plurality of property candidates is selected as a determined result property of the selected amino acid position. The determined result property is used as a model input to predict a plurality of candidates for a property of a different selected amino acid position included in the binder protein.

Claims (55)

1. A method, comprising:

receiving a specification of a binding target protein;

receiving an identification of a known binding location of the binding target protein where a binder protein is to bind;

based at least in part on a property of the binding target protein at the known binding location, using a machine learning model to predict a plurality of candidates for a property of a selected amino acid position of the binder protein to bind to the binding target protein;

for each selected property candidate of the plurality of property candidates, using the selected property candidate as an input to the machine learning model to predict properties for one or more other amino acid positions of the binder protein into a corresponding candidate set of properties;

evaluating the corresponding candidate sets of properties for the plurality of property candidates using a binding quality evaluation;

based on the evaluation of the corresponding candidate set of properties, selecting one of the plurality of property candidates as a determined result property of the selected amino acid position of the binder protein in defining the binding target protein; and

iteratively refining at least a portion of the predicted properties for the one or more other amino acid positions of the binder protein including by discarding at least a portion of the predicted properties for the one or more other amino acid positions of the binder protein and feeding back the determined result property of the selected amino acid position as a subsequent feedback input to the machine learning model to predict a plurality of candidates for a property of a different selected amino acid position included in the binder protein.

2. The method of claim 1 , wherein the property of the selected amino acid position of the binder protein to bind to the binding target protein corresponds to an amino acid sequence token.

3. The method of claim 1 , wherein the property of the selected amino acid position of the binder protein to bind to the binding target protein corresponds to an amino acid structure token.

4. The method of claim 3 , wherein the structure token references neighboring amino acids of the binder protein and the binding target protein.

5. The method of claim 1 , further comprising selecting the selected amino acid position based on a search criterion.

6. The method of claim 5 , wherein the search criterion includes randomly selecting the selected amino acid position of the binder protein to bind to the binding target protein.

7. The method of claim 1 , further comprising decoding one or more structure tokens associated with the corresponding candidate set of properties.

8. The method of claim 7 , wherein using the binding quality evaluation includes analyzing the decoded one or more structure tokens.

9. The method of claim 1 , wherein using the binding quality evaluation includes determining distances between one or more atoms of the binder protein and one or more atoms of the binding target protein.

10. The method of claim 1 , wherein the known binding location corresponds to a specified epitope of the binding target protein.

11. The method of claim 1 , wherein using the binding quality evaluation includes accessing a computational oracle, a binding evaluation service, or a binding prediction model.

12. The method of claim 1 , further comprising:

creating an amino acid sequence token input sequence using at least one or more amino acid sequence tokens associated with the corresponding candidate set of properties;

providing the amino acid sequence token input sequence to the machine learning model to generate a proposed binder protein result;

predicting a structure result associated with the proposed binder protein result and the binding target protein; and

evaluating the structure result for a binding quality evaluation score.

13. The method of claim 1 , further comprising:

creating a structure token input sequence based on at least one or more structure tokens associated with the corresponding candidate set of properties;

providing the structure token input sequence to the machine learning model to generate a proposed binder protein result;

determining a structure result associated with the proposed binder protein result and the binding target protein; and

evaluating the structure result for a binding quality evaluation score.

14. The method of claim 1 , further comprising:

determining a first set of structure tokens representing the binding target protein, wherein the first set of structure tokens corresponds to a non-binding state of the binding target protein;

using the first set of structure tokens to determine a second set of structure tokens representing the binding target protein, wherein the second set of structure tokens correspond to a binding state of the binding target protein;

identifying one or more epitope structure tokens associated with an epitope of the binding target protein from the second set of structure tokens; and

including the identified one or more epitope structure tokens as an additional input to the machine learning model.

15. A system, comprising:

one or more processors configured to:

receive a specification of a binding target protein;

receive an identification of a known binding location of the binding target protein where a binder protein is to bind;

based at least in part on a property of the binding target protein at the known binding location, predict using a machine learning model a plurality of candidates for a property of a selected amino acid position of the binder protein to bind to the binding target protein;

for each selected property candidate of the plurality of property candidates, predict properties for one or more other amino acid positions of the binder protein into a corresponding candidate set of properties using the selected property candidate as an input to the machine learning model;

evaluate the corresponding candidate sets of properties for the plurality of property candidates using a binding quality evaluation;

based on the evaluation of the corresponding candidate set of properties, select one of the plurality of property candidates as a determined result property of the selected amino acid position of the binder protein in defining the binding target protein; and

iteratively refine at least a portion of the predicted properties for the one or more other amino acid positions of the binder protein including by being configured to discard at least a portion of the predicted properties for the one or more other amino acid positions of the binder protein and feed back the determined result property of the selected amino acid position as a subsequent feedback input to the machine learning model to predict a plurality of candidates for a property of a different selected amino acid position included in the binder protein; and

a memory coupled to at least one of the one or more processors and configured to provide instructions.

16. The system of claim 15 , wherein the property of the selected amino acid position of the binder protein to bind to the binding target protein corresponds to an amino acid sequence token.

17. The system of claim 15 , wherein the property of the selected amino acid position of the binder protein to bind to the binding target protein corresponds to an amino acid structure token.

18. The system of claim 17 , wherein the structure token references neighboring amino acids of the binder protein and the binding target protein.

19. The system of claim 15 , wherein using the binding quality evaluation includes determining distances between one or more atoms of the binder protein and one or more atoms of the binding target protein.

20. A computer program product, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

receiving a specification of a binding target protein;

receiving an identification of a known binding location of the binding target protein where a binder protein is to bind;

based at least in part on a property of the binding target protein at the known binding location, using a machine learning model to predict a plurality of candidates for a property of a selected amino acid position of the binder protein to bind to the binding target protein;

for each selected property candidate of the plurality of property candidates, using the selected property candidate as an input to the machine learning model to predict properties for one or more other amino acid positions of the binder protein into a corresponding candidate set of properties;

evaluating the corresponding candidate sets of properties for the plurality of property candidates using a binding quality evaluation;

based on the evaluation of the corresponding candidate set of properties, selecting one of the plurality of property candidates as a determined result property of the selected amino acid position of the binder protein in defining the binding target protein; and

iteratively refining at least a portion of the predicted properties for the one or more other amino acid positions of the binder protein including by discarding at least a portion of the predicted properties for the one or more other amino acid positions of the binder protein and feeding back the determined result property of the selected amino acid position as a subsequent feedback input to the machine learning model to predict a plurality of candidates for a property of a different selected amino acid position included in the binder protein.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2026
From: CHAN ZUCKERBERG INITIATIVE, LLC
To: CHAN ZUCKERBERG BIOHUB, INC.
Reel/Frame 075160/0394 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2025
From: EVOLUTIONARYSCALE, PBC
To: CHAN ZUCKERBERG INITIATIVE, LLC
Reel/Frame 072807/0171 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2024
From: CANDIDO, SALVATORE J.; RIVES, ALEXANDER W.; HAYES, THOMAS F.; GONG, JUN
To: EVOLUTIONARYSCALE, PBC
Reel/Frame 069363/0801 →
References Cited (102)
US 5577249A · Califano · 1996 [cited by examiner]
US 11017882B2 · Cameron · 2021 [cited by examiner]
US 11450407B1 · Laniado · 2022 [cited by examiner]
US 11773143B2 · Dubois · 2023 [cited by examiner]
US 12211588B1 · Aithani · 2025 [cited by examiner]
US 20040083091A1 · Ie · 2004 [cited by examiner]
US 20050130224A1 · Saito · 2005 [cited by examiner]
US 20060217891A1 · Tanuma · 2006 [cited by examiner]
US 20070231329A1 · Lazar · 2007 [cited by examiner]
US 20080014646A1 · Kuroda · 2008 [cited by examiner]
US 20080215301A1 · Eyal · 2008 [cited by examiner]
US 20090006059A1 · Arora · 2009 [cited by examiner]
US 20090144209A1 · Miyakawa · 2009 [cited by examiner]
US 20200168293A1 · Baker · 2020 [cited by examiner]
US 20200342953A1 · Morrone · 2020 [cited by examiner]
US 20210090690A1 · Plumbley · 2021 [cited by examiner]
US 20210107936A1 · Varanasi · 2021 [cited by examiner]
US 20210166779A1 · Jumper · 2021 [cited by examiner]
US 20210249105A1 · Madani · 2021 [cited by examiner]
US 20210304847A1 · Senior · 2021 [cited by examiner]
US 20220036966A1 · Meyers · 2022 [cited by examiner]
US 20220104515A1 · Hume · 2022 [cited by examiner]
US 20220108766A1 · Mackinnon · 2022 [cited by examiner]
US 20220246233A1 · Morrone · 2022 [cited by examiner]
US 20220246239A1 · Cho · 2022 [cited by examiner]
US 20220359045A1 · Manica · 2022 [cited by examiner]
US 20220375539A1 · Alvarez · 2022 [cited by examiner]
US 20230034425A1 · Laniado · 2023 [cited by examiner]
US 20230093507A1 · Pei · 2023 [cited by examiner]
US 20230110719A1 · Krause · 2023 [cited by examiner]
US 20230114905A1 · Egertson · 2023 [cited by examiner]
US 20230154573A1 · Roy · 2023 [cited by examiner]
US 20230161996A1 · Bepler · 2023 [cited by examiner]
US 20230297812A1 · Shin · 2023 [cited by examiner]
US 20230298687A1 · Figurnov · 2023 [cited by examiner]
US 20230326545A1 · Narayanan · 2023 [cited by examiner]
US 20230335228A1 · Van Hoorn · 2023 [cited by examiner]
US 20230360734A1 · Evans · 2023 [cited by examiner]
US 20230368915A1 · Abraham · 2023 [cited by examiner]
US 20230386610A1 · Singh · 2023 [cited by examiner]
US 20230395186A1 · Kohl · 2023 [cited by examiner]
US 20230410938A1 · Pritzel · 2023 [cited by examiner]
US 20230420084A1 · Langevin · 2023 [cited by examiner]
US 20240013854A1 · Kyriacou · 2024 [cited by examiner]
US 20240029819A1 · Merbl · 2024 [cited by examiner]
US 20240029820A1 · Song · 2024 [cited by examiner]
US 20240038337A1 · Laniado · 2024 [cited by examiner]
US 20240047012A1 · Kozintsev · 2024 [cited by examiner]
US 20240087674A1 · Gligorijevic · 2024 [cited by examiner]
US 20240105277A1 · Kumar · 2024 [cited by examiner]
US 20240120022A1 · Senior · 2024 [cited by examiner]
US 20240145026A1 · Zhang · 2024 [cited by examiner]
US 20240153590A1 · Essaghir · 2024 [cited by examiner]
US 20240177798A1 · Min · 2024 [cited by examiner]
US 20240203523A1 · Leem · 2024 [cited by examiner]
US 20240203532A1 · Madani · 2024 [cited by examiner]
US 20240212785A1 · El Hibouri · 2024 [cited by examiner]
US 20240220685A1 · Xu · 2024 [cited by examiner]
US 20240257902A1 · Zhao · 2024 [cited by examiner]
US 20240257907A1 · Yang · 2024 [cited by examiner]
US 20240282408A1 · Lee · 2024 [cited by examiner]
US 20240321386A1 · Evans · 2024 [cited by examiner]
US 20240355413A1 · Duplay · 2024 [cited by examiner]
US 20240379248A1 · Hemmatian · 2024 [cited by examiner]
US 20250029680A1 · Vural · 2025 [cited by examiner]
Huang et al., Protein Structure Prediction: Challenges, Advances, and the Shift of Research Paradigms, Genomics Proteomics Bioinformatics, vol. 21, 2023, pp. 913-925. [cited by applicant]
Abramson et al., Accurate structure prediction of biomolecular interactions with AlphaFold 3, Nature, vol. 630, Jun. 13, 2024, 24 pages, https://doi.org/10.1038/s41586-024-07487-w. [cited by applicant]
Alamdari et al., Protein generation with evolutionary diffusion: sequence is all you need, bioRxiv preprint, Sep. 12, 2023, pp. 1-62, https://doi.org/10.1101/2023.09.11.556673. [cited by applicant]
Alley et al., Unified rational protein engineering with sequence-based deep representation learning, Nature Methods, vol. 16, Dec. 2019, 14 pages, https://doi.org/10.1038/s41592-019-0598-1. [cited by applicant]
Chang et al., MaskGIT: Masked Generative Image Transformer, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Feb. 8, 2022, 23 pages. [cited by applicant]
Chen et al., xTrimoPGLM: Unified 100B-Scale Pre-trained Transformer for Deciphering the Language of Protein, arXiv preprint arXiv:2401.06199v1, Jan. 15, 2024, 86 pages. [cited by applicant]
Elnaggar et al., ProtTrans: Towards Cracking the Language of Life's Code Through Self-Supervised Learning, IEEE Trans Pattern Analysis & Machine Intelligence, vol. 14, No. 8, Aug. 2021, pp. 1-29. [cited by applicant]
Ferruz et al., ProtGPT2 is a deep unsupervised language model for protein design, Nature Communications, Jul. 27, 2022, pp. 1-10, https://doi.org/10.1038/s41467-022-32007-7. [cited by applicant]
Gaujac et al., Learning the Language of Protein Structure, arXiv preprint arXiv:2405.15840v1, May 24, 2024, pp. 1-20. [cited by applicant]
Heinzinger et al., Bilingual Language Model for Protein Sequence and Structure, bioRxiv preprint, Mar. 24, 2024, pp. 1-23, https://doi.org/10.1101/2023.07.23.550085. [cited by applicant]
Heinzinger et al., Modeling aspects of the language of life through transfer-learning protein sequences, BMC Bioinformatics, 2019, pp. 1-17, https://doi.org/10.1186/s12859-019-3220-8. [cited by applicant]
Hesslow et al., RITA: a Study on Scaling Up Generative Protein Sequence Models, arXiv preprint arXiv:2205.05789v2, Jul. 14, 2022, 10 pages. [cited by applicant]
Ingraham et al., Generative models for graph-based protein design, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, pp. 1-12. [cited by applicant]
Ingraham et al., Illuminating protein space with a programmable generative model, Nature, vol. 623, Nov. 30, 2023, 15 pages, https://doi.org/10.1038/s41586-023-06728-8. [cited by applicant]
Jumper et al., Highly accurate protein structure prediction with AlphaFold, Nature, vol. 596, Aug. 26, 2021, 12 pages, https://doi.org/10.1038/s41586-021-03819-2. [cited by applicant]
Li et al., DeProt: Protein language modeling with quantizied structure and disentangled attention, bioRxiv preprint, Apr. 17, 2024, 15 pages, https://doi.org/10.1101/2024.04.15.589672. [cited by applicant]
Lin et al., Evolutionary-scale prediction of atomic level protein structure with a language model, bioRxiv preprint, Oct. 31, 2022, 28 pages, https://doi.org/10.1101/2022.07.20.500902. [cited by applicant]
Lin et al., Out of Many, One: Designing and Scaffolding Proteins at the Scale of the Structural Universe with Genie 2, arXiv preprint arXiv:2405.15489v1, May 24, 2024, 27 pages. [cited by applicant]
Miyato et al., GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers, arXiv preprint arXiv:2310.10375v3, Jun. 7, 2024, 37 pages. [cited by applicant]
Nijkamp et al., ProGen2: Exploring the Boundaries of Protein Language Models, arXiv preprint arXiv:2206.13517v1, Jun. 27, 2022, 20 pages. [cited by applicant]
Rao et al., Transformer Protein Language Models Are Unsupervised Structure Learners, bioRxiv preprint, Dec. 15, 2020, 24 pages, https://doi.org/10.1101/2020.12.15.422761. [cited by applicant]
Razavi et al., Generating Diverse High-Fidelity Images with VQ-VAE-2, arXiv preprint arXiv:1906.00446v1, Jun. 2, 2019, 15 pages. [cited by applicant]
Rives et al., Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences, Proceedings of the National Academy of Sciences, vol. 118, No. 15, 2021, pp. 1-12, https://doi.… [cited by applicant]
Su et al., Saprot: Protein Language Modeling with Structure-Aware Vocabulary, bioRxiv preprint, Oct. 2, 2023, 22 pages, https://doi.org/10.1101/2023.10.01.560349. [cited by applicant]
Van Den Oord et al., Neural Discrete Representation Learning, 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017, 11 pages. [cited by applicant]
Van Kempen et al., Fast and accurate protein structure search with Foldseek, Nature Biotechnology, vol. 42, Feb. 2024, 17 pages, https://doi.org/10.1038/s41587-023-01773-0. [cited by applicant]
Watson et al., De novo design of protein structure and function with RFdiffusion, Nature, vol. 620, Aug. 31, 2023, 38 pages, https://doi.org/10.1038/s41586-023-06415-8. [cited by applicant]
Yu et al., Scaling Autoregressive Models for Content-Rich Text-to-Image Generation, arXiv preprint arXiv:2206.10789v1, Jun. 22, 2022, 49 pages. [cited by applicant]
Costa et al., Ophiuchus: Scalable Modeling of Protein Structures through Hierarchical Coarse-Graining SO(3)-Equivariant Autoencoders, arXiv preprint arXiv:2310.02508v2, Dec. 27, 2023, 29 pages. [cited by applicant]
Lee et al., ShapeProt: Top-down Protein Design with 3D Protein Shape Generative Model, bioRxiv preprint, Dec. 3, 2023, 19 pages, https://doi.org/10.1101/203.12.03.567710. [cited by applicant]
Liu et al., Euclidean transformers for macromolecular structures: Lessons learned, 2022 ICML Workshop on Computational Biology, 2022, 8 pages. [cited by applicant]
Luo et al., Spherical Rotation Dimension Reduction with Geometric Loss Functions, arXiv preprint arXiv:2204.10975v2, Apr. 27, 2023, 60 pages. [cited by applicant]
Pereira et al., Step-by-Step design of proteins for small molecule interaction: A review on recent milestones, Protein Science, vol. 30, Apr. 23, 2021, pp. 1502-1520, wileyonlinelibrary.com/journal/pro. [cited by applicant]
Sekmen et al., Mathematical and Machine Learning Approaches for Classification of Protein Secondary Structure Elements from Cα Coordinates, Biomolecules 13, No. 6, May 3, 2023, 19 pages, https://doi.org/10.3390.biom1306… [cited by applicant]
Wu et al., Surface-VQMAE: Vector-quantized Masked Auto-encoders on Molecular Surfaces, International Conference on Machine Learning, 2024, pp. 1-16, https://openreview.net/forum?id=szxtVHOhOC. [cited by applicant]
Zhang et al., A survey on masked autoencoder for self-supervised learning in vision and beyond, Journal of Latex Class Files, vol. 14, No. 8, Aug. 2015, pp. 1-13. [cited by applicant]
Laine et al., Protein sequence-to-structure learning: Is this the end(-to-end revolution)?, HAL open science, Sep. 13, 2021, 21 pages, https://hal.science/hal-03342140v1. [cited by applicant]
Cited By (2)
US 12,633,373 US 12,712,051