SYSTEMS AND METHODS FOR IDENTIFYING NOVEL PORE-FORMING TOXINS
The present disclosure relates to the field of biotechnology, and, more specifically, to systems and methods for identifying novel pore-forming toxins (PFTs) based on protein structures and sequences.
1 . A method for identifying novel pore-forming toxins (PFTs) based on sequence and structure data, comprising:
identifying, in a dataset comprising PFT information, a plurality of proteins with known sequences and structures;
determining a plurality of protein clusters based on pairwise structural similarity values of the plurality of proteins;
for each respective protein cluster of the plurality of protein clusters, identifying a respective group of proteins that have a pairwise sequence identity a) lower than a threshold pairwise sequence identity, b) above a threshold pairwise sequence identity, or c) within a predetermined pairwise sequence identity range;
generating a graphical model trained using sequence and structure data of proteins from each respective group of proteins, wherein the graphical model is configured to generate a structural segmentation of an input protein based on a sequence of the input protein;
calculating a segment interaction score for the generated structural segmentation of the input protein, wherein the segment interaction score compares the generated structural segmentation of the input protein with structures of the proteins from each respective group of proteins; and
in response to determining that the segment interaction score is greater than a threshold segment interaction score, classifying the input protein as a potential novel PFT.
2 . The method of claim 1 , further comprising:
determining whether the sequence of the input protein is classified as a PFT sequence using a machine learning model configured to classify sequences as a PFT sequence or a non-PFT sequence; and
in response to determining that the sequence is classified as a PFT sequence, identifying the input protein as a novel PFT.
3 . The method of claim 2 , further comprising:
receiving confirmation that the input protein is not a novel PFT; and
re-training the machine learning model such that the sequence of the input protein is identified as a non-PFT sequence.
4 . The method of claim 1 , further comprising generating, for output on a computing device, an indication that the input protein is classified as a potential novel PFT and that the input protein shares a functionality of a particular protein cluster from the plurality of protein clusters.
5 . The method of claim 1 , wherein determining the plurality of protein clusters further comprises:
mapping a structural representation of each protein from the plurality of proteins from a high-dimensional space to a two-dimensional space that preserves structural correlations among the plurality of proteins; and
executing a clustering algorithm on the two-dimensional space to determine the plurality of protein clusters.
6 . The method of claim 5 , wherein the clustering algorithm is a K-means clustering algorithm.
7 . The method of claim 1 , wherein generating the graphical model further comprises:
aligning structures of the proteins from each respective group of proteins to identify common structural regions using an iterative pairwise alignment algorithm; and
identifying consensus and non-consensus secondary structure segments in the aligned structures, wherein training the graphical model comprises maximizing a probability of the identified consensus and non-consensus secondary structure segments in the structural segmentation of the input protein.
8 . The method of claim 1 , wherein the graphical model is a semi-Markov conditional fields (semi-CRFs) model.
9 . The method of claim 1 , wherein proteins in a respective group of proteins share functionality and have a low sequence identity, optionally less than 60, 50, 40, 30, 20, or 10% full length sequence identity.
10 . A system for identifying novel pore-forming toxins (PFTs) based on sequence and structure data, comprising:
a processor, and
memory having stored therein instructions that when executed by the processor cause the processor to:
identify, in a dataset comprising PFT information, a plurality of proteins with known sequences and structures;
determine a plurality of protein clusters based on pairwise structural similarity values of the plurality of proteins;
for each respective protein cluster of the plurality of protein clusters, identify a respective group of proteins that have a pairwise sequence identity a) lower than a threshold pairwise sequence identity, b) above a threshold pairwise sequence identity, or c) within a predetermined pairwise sequence identity range;
generate a graphical model trained using sequence and structure data of proteins from each respective group of proteins, wherein the graphical model is configured to generate a structural segmentation of an input protein based on a sequence of the input protein;
calculate a segment interaction score for the generated structural segmentation of the input protein, wherein the segment interaction score compares the generated structural segmentation of the input protein with structures of the proteins from each respective group of proteins; and
in response to determining that the segment interaction score is greater than a threshold segment interaction score, classify the input protein as a potential novel PFT.
11 . The system of claim 10 , wherein the memory further comprises instructions for:
determining whether the sequence of the input protein is classified as a PFT sequence using a machine learning model configured to classify sequences as a PFT sequence or a non-PFT sequence; and
in response to determining that the sequence is classified as a PFT sequence, identifying the input protein as a novel PFT.
12 . The system of claim 11 , wherein the memory further comprises instructions for:
receiving confirmation that the input protein is not a novel PFT; and
re-training the machine learning model such that the sequence of the input protein is identified as a non-PFT sequence.
13 . The system of claim 10 , wherein the memory further comprises instructions for:
generating, for output on a computing device, an indication that the input protein is classified as a potential novel PFT and that the input protein shares a functionality of a particular protein cluster from the plurality of protein clusters.
14 . The system of claim 10 , wherein determining the plurality of protein clusters further comprises:
mapping a structural representation of each protein from the plurality of proteins from a high-dimensional space to a two-dimensional space that preserves structural correlations among the plurality of proteins; and
executing a clustering algorithm on the two-dimensional space to determine the plurality of protein clusters.
15 . The system of claim 14 , wherein the clustering algorithm is a K-means clustering algorithm.
16 . The system of claim 10 , wherein generating the graphical model further comprises:
aligning structures of the proteins from each respective group of proteins to identify common structural regions using an iterative pairwise alignment algorithm; and
identifying consensus and non-consensus secondary structure segments in the aligned structures, wherein training the graphical model comprises maximizing a probability of the identified consensus and non-consensus secondary structure segments in the structural segmentation of the input protein.
17 . The system of claim 10 , wherein the graphical model is a semi-Markov conditional fields (semi-CRFs) model.
18 . The system of claim 10 , wherein proteins in a respective group of proteins share functionality and have low sequence identity.