System and method for generating a protein sequence
A method and system for generating a protein sequence is implemented using a computer-implemented neural network. An empty or partially filed sequence of node elements, representing amino acid positions of the protein sequence, and an edge index, having edge elements defining physical interaction between amino acid positions, are received. The computer-implemented neural network operates to determine enhanced edge attribute values for edge elements of the edge index and enhanced amino acid values for node elements of the sequence. Amino acid values are generated for elements of the partially filed sequence having missing values.
1 . A method for generating a protein sequence, the method comprising:
receiving an empty or partially filled sequence of node elements, each node element of the sequence representing an amino acid position of the protein sequence and having an amino acid value chosen from one of a predetermined amino acid value and a missing value;
receiving an edge index of edge elements, each edge element of the edge index being associated to a respective pair of the node elements and having at least a presence value, a positive presence value indicating the presence of a physical interaction between the amino acid positions corresponding to the pair of node elements, the edge index thereby defining a desired fold structure of the protein sequence;
processing, by a computer-implemented neural network, a working copy of the partially filled sequence and the edge index to:
determine, for each given edge element of the edge index having a positive presence value, an enhanced edge attribute value set based on the amino acid values of the pair of node elements of the working copy sequence associated to the given edge element;
determine, for each given node element of the working copy sequence, an enhanced amino acid value based on the enhanced edge attribute value sets for every edge element having a positive presence value associated to the given node element; and
for each given node element of the partially filled sequence having the missing value, determining a generated amino acid value based on the enhanced amino acid value corresponding to the given node element.
2 . The method of claim 1 , further comprising:
receiving an edge attribute dataset defining, for each edge element having a positive presence value, an edge attribute value set representing at least one characteristic of the physical interaction between the amino acid positions corresponding to the pair of node elements associated to the given edge element; and
wherein the enhanced edge attribute value set for each given edge element of the edge index is initially determined based on the edge attribute value set for the given edge element in combination with the amino acid values of the pair of node elements of the working copy sequence associated to the given edge element.
3 . The method of claim 2 , wherein:
the edge attribute value set for each given edge element defines a Euclidian distance between the amino acid positions corresponding to the pair of node elements associated to the given edge element.
4 . The method of claim 1 , wherein:
the enhanced edge attribute value set for each given edge element is determined using one or more trained multi-layer perceptrons.
5 . The method of claim 4 , wherein:
the enhanced amino acid value for each given node element of the working copy sequence is determined in a pooling module from summing the enhanced edge attribute value sets for the edge elements having the positive presence value associated to the given node element.
6 . The method of claim 5 , wherein:
the neural network comprises a plurality of repeated residual blocks each encompassing a respective one of the one or more trained multi-layer perceptrons and a respective summing block; and
wherein the enhanced edge attribute value sets and the enhanced amino acid values determined in a preceding one of the residual blocks are provided as inputs to a subsequent one of the residual blocks.
7 . The method of claim 6 , wherein:
in the subsequent one of the residual blocks, the enhanced edge attribute value set for each given edge element of the edge index is determined based on the enhanced edge attribute value set for the given edge element determined in the preceding one of the residual blocks in combination with the enhanced amino acid values of the pair of node elements associated to the given edge element determined in the preceding one of the residual blocks.
8 . The method of claim 4 , wherein:
the enhanced amino acid value for each given node element of the working copy sequence is determined in a pooling module from applying a weighted sum of the enhanced edge attribute value sets for the edge elements having the positive presence value associated to the given node element.
9 . The method of claim 1 , wherein:
the working copy sequence is mapped into an embedding space; and
the processing by the neural network is carried out on the working copy sequence as mapped into the embedding space.
10 . The method of claim 1 , wherein:
the enhanced edge attribute value sets for the edge elements of the edge index are mapped into an embedding space; and
wherein the processing by the neural network is carried out on the edge attribute value sets as mapped into the embedding space.
11 . The method of claim 1 , further comprising:
for each given node element of the partially filled sequence having the missing value, replacing the missing value by the generated amino acid value corresponding to the given node element; and
whereby the partially filled sequence formed of the node elements having predetermined amino acid values and node elements having missing values replaced by the generated amino acid values represents a completed protein sequence.
12 . The method of claim 1 , wherein:
determining the generated amino acid value for each given node element of the partially filled sequence having the missing value comprises:
determining, for each of a plurality of possible amino acid values, a probability metric based on the enhanced amino acid value of the corresponding node element; and
selecting the possible amino acid value having the highest probability metric as the generated amino acid value for the given node element of the partially filed sequence.
13 . The method of claim 1 , wherein:
an annotated training dataset for training the neural network comprises a plurality of protein sequence entries, each entry comprising:
an amino acid sequence of position elements each having a respective amino acid value;
an edge index indicating pairs of position elements of the amino acid sequence having a physical interaction; and
the annotated training dataset further comprises an input subset and a target subset, wherein:
in the input subset, the amino acid values of a portion of the position elements of the amino acid sequences are masked; and
in the target subset, the amino acid values of the position elements of amino acid sequences corresponding to the sequences of the input subset are non-masked.
14 . The method of claim 13 , wherein:
each entry of the annotated training dataset further comprises an edge attribute dataset defining, for each element of the edge index indicating the physical interaction, at least one characteristic of the physical interaction between the pair of position elements.
15 . The method of claim 13 , wherein:
the training of the neural network comprises adjusting parameters of the neural network to decrease, or minimize, the cross entropy loss between:
(i) amino acid values of outputted amino acid sequences as determined by the neural network applied to the input subset and
(ii) the amino acid values of corresponding amino acid sequences of the target subset.
16 . The method of claim 1 , wherein the method scores mutations of specific residues, said mutations comprising substitutions, additions, and/or deletions of residues.
17 . The method of claim 1 , wherein the method generates a protein sequence based on one or more given protein folds.
18 . A computer program product comprising a non-transitory storage medium having stored thereon computer-readable instructions, wherein the instructions when executed on a processor causes the processor to carry out the method according to claim 1 .
19 . A protein sequence generated by the method of claim 1 .
20 . A system comprising:
at least one memory; and
at least one processor coupled to the memory, and being configured for performing:
receiving an empty or partially filled sequence of node elements, each node element of the sequence representing an amino acid position of the protein sequence and having an amino acid value chosen from one of a predetermined amino acid value and a missing value;
receiving an edge index of edge elements, each edge element of the edge index being associated to a respective pair of the node elements and having at least a presence value, a positive presence value indicating the presence of a physical interaction between the amino acid positions corresponding to the pair of node elements, the edge index thereby defining a desired fold structure of the protein sequence;
processing, by a computer-implemented neural network, a working copy of the partially filled sequence and the edge index to:
determine, for each given edge element of the edge index having a positive presence value, an enhanced edge attribute value set based on the amino acid values of the pair of node elements of the working copy sequence associated to the given edge element;
determine, for each given node element of the working copy sequence, an enhanced amino acid value based on the enhanced edge attribute value sets for every edge element having a positive presence value associated to the given node element; and
for each given node element of the partially filled sequence having the missing value, determining a generated amino acid value based on the enhanced amino acid value corresponding to the given node element.