IP Library › Granted Patent US 12,655,176
Granted Patent B2
US 12,655,176 · App. 17/781,805 · Granted Jun 16, 2026

System and method for generating a protein sequence

Inventors: Philip M. Kim (Toronto, CA); Alexey V. Strokach (Toronto, CA); David B. Romero (Etobicoke, CA); Carlos Corbi Verge (Toronto, CA); Alberto Perez Riba (Toronto, CA)
Assignee: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
C07K1/00G16B15/20G16B40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,655,176
App. No.
17/781,805
Granted
Jun 16, 2026
Kind
B2
Abstract

A method and system for generating a protein sequence is implemented using a computer-implemented neural network. An empty or partially filed sequence of node elements, representing amino acid positions of the protein sequence, and an edge index, having edge elements defining physical interaction between amino acid positions, are received. The computer-implemented neural network operates to determine enhanced edge attribute values for edge elements of the edge index and enhanced amino acid values for node elements of the sequence. Amino acid values are generated for elements of the partially filed sequence having missing values.

Claims (62)

1 . A method for generating a protein sequence, the method comprising:

receiving an empty or partially filled sequence of node elements, each node element of the sequence representing an amino acid position of the protein sequence and having an amino acid value chosen from one of a predetermined amino acid value and a missing value;

receiving an edge index of edge elements, each edge element of the edge index being associated to a respective pair of the node elements and having at least a presence value, a positive presence value indicating the presence of a physical interaction between the amino acid positions corresponding to the pair of node elements, the edge index thereby defining a desired fold structure of the protein sequence;

processing, by a computer-implemented neural network, a working copy of the partially filled sequence and the edge index to:

determine, for each given edge element of the edge index having a positive presence value, an enhanced edge attribute value set based on the amino acid values of the pair of node elements of the working copy sequence associated to the given edge element;

determine, for each given node element of the working copy sequence, an enhanced amino acid value based on the enhanced edge attribute value sets for every edge element having a positive presence value associated to the given node element; and

for each given node element of the partially filled sequence having the missing value, determining a generated amino acid value based on the enhanced amino acid value corresponding to the given node element.

2 . The method of claim 1 , further comprising:

receiving an edge attribute dataset defining, for each edge element having a positive presence value, an edge attribute value set representing at least one characteristic of the physical interaction between the amino acid positions corresponding to the pair of node elements associated to the given edge element; and

wherein the enhanced edge attribute value set for each given edge element of the edge index is initially determined based on the edge attribute value set for the given edge element in combination with the amino acid values of the pair of node elements of the working copy sequence associated to the given edge element.

3 . The method of claim 2 , wherein:

the edge attribute value set for each given edge element defines a Euclidian distance between the amino acid positions corresponding to the pair of node elements associated to the given edge element.

4 . The method of claim 1 , wherein:

the enhanced edge attribute value set for each given edge element is determined using one or more trained multi-layer perceptrons.

5 . The method of claim 4 , wherein:

the enhanced amino acid value for each given node element of the working copy sequence is determined in a pooling module from summing the enhanced edge attribute value sets for the edge elements having the positive presence value associated to the given node element.

6 . The method of claim 5 , wherein:

the neural network comprises a plurality of repeated residual blocks each encompassing a respective one of the one or more trained multi-layer perceptrons and a respective summing block; and

wherein the enhanced edge attribute value sets and the enhanced amino acid values determined in a preceding one of the residual blocks are provided as inputs to a subsequent one of the residual blocks.

7 . The method of claim 6 , wherein:

in the subsequent one of the residual blocks, the enhanced edge attribute value set for each given edge element of the edge index is determined based on the enhanced edge attribute value set for the given edge element determined in the preceding one of the residual blocks in combination with the enhanced amino acid values of the pair of node elements associated to the given edge element determined in the preceding one of the residual blocks.

8 . The method of claim 4 , wherein:

the enhanced amino acid value for each given node element of the working copy sequence is determined in a pooling module from applying a weighted sum of the enhanced edge attribute value sets for the edge elements having the positive presence value associated to the given node element.

9 . The method of claim 1 , wherein:

the working copy sequence is mapped into an embedding space; and

the processing by the neural network is carried out on the working copy sequence as mapped into the embedding space.

10 . The method of claim 1 , wherein:

the enhanced edge attribute value sets for the edge elements of the edge index are mapped into an embedding space; and

wherein the processing by the neural network is carried out on the edge attribute value sets as mapped into the embedding space.

11 . The method of claim 1 , further comprising:

for each given node element of the partially filled sequence having the missing value, replacing the missing value by the generated amino acid value corresponding to the given node element; and

whereby the partially filled sequence formed of the node elements having predetermined amino acid values and node elements having missing values replaced by the generated amino acid values represents a completed protein sequence.

12 . The method of claim 1 , wherein:

determining the generated amino acid value for each given node element of the partially filled sequence having the missing value comprises:

determining, for each of a plurality of possible amino acid values, a probability metric based on the enhanced amino acid value of the corresponding node element; and

selecting the possible amino acid value having the highest probability metric as the generated amino acid value for the given node element of the partially filed sequence.

13 . The method of claim 1 , wherein:

an annotated training dataset for training the neural network comprises a plurality of protein sequence entries, each entry comprising:

an amino acid sequence of position elements each having a respective amino acid value;

an edge index indicating pairs of position elements of the amino acid sequence having a physical interaction; and

the annotated training dataset further comprises an input subset and a target subset, wherein:

in the input subset, the amino acid values of a portion of the position elements of the amino acid sequences are masked; and

in the target subset, the amino acid values of the position elements of amino acid sequences corresponding to the sequences of the input subset are non-masked.

14 . The method of claim 13 , wherein:

each entry of the annotated training dataset further comprises an edge attribute dataset defining, for each element of the edge index indicating the physical interaction, at least one characteristic of the physical interaction between the pair of position elements.

15 . The method of claim 13 , wherein:

the training of the neural network comprises adjusting parameters of the neural network to decrease, or minimize, the cross entropy loss between:

(i) amino acid values of outputted amino acid sequences as determined by the neural network applied to the input subset and

(ii) the amino acid values of corresponding amino acid sequences of the target subset.

16 . The method of claim 1 , wherein the method scores mutations of specific residues, said mutations comprising substitutions, additions, and/or deletions of residues.

17 . The method of claim 1 , wherein the method generates a protein sequence based on one or more given protein folds.

18 . A computer program product comprising a non-transitory storage medium having stored thereon computer-readable instructions, wherein the instructions when executed on a processor causes the processor to carry out the method according to claim 1 .

19 . A protein sequence generated by the method of claim 1 .

20 . A system comprising:

at least one memory; and

at least one processor coupled to the memory, and being configured for performing:

receiving an empty or partially filled sequence of node elements, each node element of the sequence representing an amino acid position of the protein sequence and having an amino acid value chosen from one of a predetermined amino acid value and a missing value;

receiving an edge index of edge elements, each edge element of the edge index being associated to a respective pair of the node elements and having at least a presence value, a positive presence value indicating the presence of a physical interaction between the amino acid positions corresponding to the pair of node elements, the edge index thereby defining a desired fold structure of the protein sequence;

processing, by a computer-implemented neural network, a working copy of the partially filled sequence and the edge index to:

determine, for each given edge element of the edge index having a positive presence value, an enhanced edge attribute value set based on the amino acid values of the pair of node elements of the working copy sequence associated to the given edge element;

determine, for each given node element of the working copy sequence, an enhanced amino acid value based on the enhanced edge attribute value sets for every edge element having a positive presence value associated to the given node element; and

for each given node element of the partially filled sequence having the missing value, determining a generated amino acid value based on the enhanced amino acid value corresponding to the given node element.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2022
From: KIM, PHILIP M.; STROKACH, ALEXEY V.; ROMERO, DAVID B.; CORBI VERGE, CARLOS; PEREZ RIBA, ALBERTO
To: THE GOVERNING COUNCIL OF THE UNIVERSITY OF TORONTO
Reel/Frame 060888/0962 →
Continuity (2)
Provisional Application 62944648 · Dec 6, 2019
Related Publication 20220372068A1 · Nov 24, 2022
References Cited (53)
US 20030130797A1 · Skolnick · 2003 [cited by examiner]
US 20030157487A1 · Gold · 2003 [cited by examiner]
US 20110301331A1 · Glaser · 2011 [cited by examiner]
US 20130303387A1 · Sander · 2013 [cited by examiner]
US 20140221216A1 · Cope · 2014 [cited by examiner]
US 20180165412A1 · Frey · 2018 [cited by examiner]
US 20190018019A1 · Shan · 2019 [cited by examiner]
US 20200234788A1 · AlQuraishi · 2020 [cited by examiner]
US 20220230710A1 · Amimeur · 2022 [cited by examiner]
US 20220348999A1 · Tori · 2022 [cited by examiner]
WO WO2020225693A1 · 2020 [cited by examiner]
Bleicher L, Lemke N, Garratt RC. Using amino acid correlation and community detection algorithms to identify functional determinants in protein families. PLoS One. 2011;6(12):e27786. doi: 10.1371/journal.pone.0027786. E… [cited by examiner]
Y. Du and C. Yan, “An Efficient Method for Discovering Functionally Important Motifs in a Group of Protein Structures,” 2018 International Conference on Computational Science and Computational Intelligence (CSCI), Las V… [cited by examiner]
Creswell et al. Generative Adversarial Networks IEEE Signal Processing Magazine pp. 53-65 (Year: 2018) (Year: 2018). [cited by examiner]
L. Liu et al., “Integrating Sequence and Network Information to Enhance Protein-Protein Interaction Prediction Using Graph Convolutional Networks,” 2019 IEEE International Conference on Bioinformatics and Biomedicine (B… [cited by examiner]
D. Zhang and M. R. Kabuka, “Protein Family Classification with Multi-Layer Graph Convolutional Networks,” 2018 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Madrid, Spain, 2018, pp. 2390-2393, … [cited by examiner]
Ingraham et al., “Generative Models for Graph-Based Protein Design” Advances in Neural Information Processing Systems 32 (NeurIPS 2019), Mar. 27, 2019. [cited by applicant]
O'Connell, James et al. “SPIN2: Predicting sequence profiles from protein structures using deep neural networks.” Proteins vol. 86,6 (2018): 629-633. doi:10.1002/prot.25489. [cited by applicant]
Chen et al., “To Improve Protein Sequence Profile Prediction through Image Captioning on Pairwise Residue Distance Map” J. Chem. Inf. Model. 2020, 60, 1, 391-399 Publication Date: Dec. 4, 2019 https://doi.org/10.1021/ac… [cited by applicant]
G. Zhou, J. Wang, X. Zhang and G. Yu, “DeepGOA: Predicting Gene Ontology Annotations of Proteins via Graph Convolutional Network,” 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2019, pp. 1… [cited by applicant]
Strokach et al., “Designing real novel proteins using deep graph neural networks” bioRxiv 868935; Dec. 10, 2019 doi:https://doi.org/10.1101/868935. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/CA2020/051669, mailed Jan. 14, 2021, 9 pages. [cited by applicant]
Bava, K Abdulla et al. “ProTherm, version 4.0: thermodynamic database for proteins and mutants.” Nucleic acids research vol. 32, Database issue (2004): D120-1. doi:10.1093/nar/gkh082. [cited by applicant]
Bommarius, Andreas S, and Marietou F Paye. “Stabilizing biocatalysts.” Chemical Society reviews vol. 42, 15 (2013): 6534-65. doi:10.1039/c3cs60137d. [cited by applicant]
Chevalier, A., Silva, DA., Rocklin, G. et al. Massively parallel de novo protein design for targeted therapeutics. Nature 550, 74-79 (2017). https://doi.org/10.1038/nature23912. [cited by applicant]
Dal Palu, A., Dovier, A. Fogolari, F. Constraint Logic Programming approach to protein structure prediction. BMC Bioinformatics 5, 186 (2004). https://doi.org/10.1186/1471-2105-5-186. [cited by applicant]
Dawson, Natalie L et al. “CATH: an expanded resource to predict protein function through structure and sequence.” Nucleic acids research vol. 45,D1 (2017): D289-D295. doi:10.1093/nar/gkw1098. [cited by applicant]
Deller, Marc C et al. “Protein stability: a crystallographers perspective.” Acta crystallographica. Section F, Structural biology communications vol. 72, Pt 2 (2016): 72-95. [cited by applicant]
Desmet, J., Maeyer, M., Hazes, B. et al. The dead-end elimination theorem and its use in protein side-chain positioning. Nature 356, 539-542 (1992). https://doi.org/10.1038/356539a0. [cited by applicant]
Fischman, Sharon, and Yanay Ofran. “Computational design of antibodies.” Current opinion in structural biology vol. 51 (2018): 156-162. doi:10.1016/j.sbi.2018.04.0070. [cited by applicant]
Khan, Sofia, and Mauno Vihinen. “Performance of protein stability predictors.” Human mutation vol. 31,6 (2010): 675-84. doi:10.1002/humu.21242. [cited by applicant]
Kocić et al., “An End-to-End Deep Neural Network for Autonomous Driving Designed for Embedded Automotive Platforms” Sensors 2019, 19(9), 2064. [cited by applicant]
Kroncke, Brett M et al. “Documentation of an Imperative to Improve Methods for Predicting Membrane Protein Stability.” Biochemistry vol. 55,36 (2016): 5002-9. doi:10.1021/acs.biochem.6b00537. [cited by applicant]
Leach, A R, and A P Lemon. “Exploring the conformational space of protein side chains using dead-end elimination and the A* algorithm.” Proteins vol. 33,2 (1998): 227-39. [cited by applicant]
Lewis, Tony E et al. “Gene3D: Extensive prediction of globular domains in proteins.” Nucleic acids research vol. 46, D1 (2018): D435-D439. doi:10.1093/nar/gkx1069. [cited by applicant]
Li, Zhixiu et al. “Direct prediction of profiles of sequences compatible with a protein structure by neural networks with fragment-based local and energy-based nonlocal profiles.” Proteins vol. 82, 10 (2014): 2565-73. d… [cited by applicant]
McGuffin, L J et al. “The PSIPRED protein structure prediction server.” Bioinformatics (Oxford, England) vol. 16,4 (2000): 404-5. doi:10.1093/bioinformatics/16.4.404. [cited by applicant]
Palm et al., “Recurrent Relational Networks” Part of Advances in Neural Information Processing Systems 31 (NeurIPS 2018). [cited by applicant]
Park et al., “Simultaneous Optimization of Biomolecular Energy Functions on Features from Small Molecules and Macromolecules” J. Chem. Theory Comput. 2016, 12, 12, 6201-6212. [cited by applicant]
Prates et al., “Learning to Solve NP-Complete Problems: A Graph Neural Network for Decision TSP” vol. 33 No. 01: AAAI-19, IAAI-19, EAAI-20—Jul. 17, 2019. [cited by applicant]
Rocklin, Gabriel J et al. “Global analysis of protein folding using massively parallel design, synthesis, and testing.” Science (New York, N.Y.) vol. 357,6347 (2017): 168-175. doi:10.1126/science.aan0693. [cited by applicant]
Schymkowitz, Joost et al. “The FoldX web server: an online force field.” Nucleic acids research vol. 33,Web Server Issue (2005): W382-8. doi:10.1093/nar/gki387. [cited by applicant]
Silver, D., Schrittwieser, J., Simonyan, K. et al. “Mastering the game of Go without human knowledge.” Nature 550, 354-359 (2017). [cited by applicant]
Simonis, H. “Models for Global Constraint Applications.” Constraints 12, 63-92 (2007). [cited by applicant]
Sun MGF, Kim PM (2017) Data driven flexible backbone protein design. PLoS Comput Biol 13(8): e1005722. [cited by applicant]
Sun, Mark G F et al. “Protein engineering by highly parallel screening of computationally designed variants.” Science advances vol. 2,7 e1600692. Jul. 20, 2016, doi:10.1126/sciadv.1600692. [cited by applicant]
Ullah, Abu Dayem, and Kathleen Steinhofel. “A hybrid approach to protein folding problem integrating constraint programming with local search.” BMC bioinformatics vol. 11 Suppl 1,Suppl 1 S39. Jan. 18, 2010, doi:10.1186/… [cited by applicant]
Wainberg et al. “Deep learning in biomedicine.” Nature biotechnology vol. 36,9 (2018): 829-838. doi:10.1038/nbt.4233. [cited by applicant]
Wang, J., Cao, H., Zhang, J.Z.H. et al. Computational Protein Design with Deep Learning Neural Networks. Sci Rep 8, 6349 (2018). https://doi.org/10.1038/s41598-018-24760-x. [cited by applicant]
Witvliet et al., “ELASPIC web-server: proteome-wide structure-based prediction of mutation effects on protein stability and binding affinity.” Bioinformatics, Vol. [cited by applicant]
Xu, Dong, and Yang Zhang. “Ab initio protein structure assembly using continuous structure fragments and optimized knowledge-based force field.” Proteins vol. 80,7 (2012): 1715-35. doi:10.1002/prot.24065. [cited by applicant]
D. Shultis et al., Changing the Apoptosis Pathway through Evolutionary Protein Design. J. Mol. Biol. 431, 825-841 (2019), Elsevier. [cited by applicant]
D. K. Witvliet et al., ELASPIC web-server: proteome-wide structure-based prediction of mutation effects on protein stability and binding affinity. Bioinforma. Oxf. Engl. 32, 1589-1591 (2016), Oxford University Press. [cited by applicant]