IP Library Patent Application 18558275
Patent Application
App. No. 18/558,275

SYSTEMS AND METHODS FOR IDENTIFYING NOVEL PORE-FORMING TOXINS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/558,275
Abstract

The present disclosure relates to the field of biotechnology, and, more specifically, to systems and methods for identifying novel pore-forming toxins (PFTs) based on protein structures and sequences.

Claims (49)

1 . A method for identifying novel pore-forming toxins (PFTs) based on sequence and structure data, comprising:

identifying, in a dataset comprising PFT information, a plurality of proteins with known sequences and structures;

determining a plurality of protein clusters based on pairwise structural similarity values of the plurality of proteins;

for each respective protein cluster of the plurality of protein clusters, identifying a respective group of proteins that have a pairwise sequence identity a) lower than a threshold pairwise sequence identity, b) above a threshold pairwise sequence identity, or c) within a predetermined pairwise sequence identity range;

generating a graphical model trained using sequence and structure data of proteins from each respective group of proteins, wherein the graphical model is configured to generate a structural segmentation of an input protein based on a sequence of the input protein;

calculating a segment interaction score for the generated structural segmentation of the input protein, wherein the segment interaction score compares the generated structural segmentation of the input protein with structures of the proteins from each respective group of proteins; and

in response to determining that the segment interaction score is greater than a threshold segment interaction score, classifying the input protein as a potential novel PFT.

2 . The method of claim 1 , further comprising:

determining whether the sequence of the input protein is classified as a PFT sequence using a machine learning model configured to classify sequences as a PFT sequence or a non-PFT sequence; and

in response to determining that the sequence is classified as a PFT sequence, identifying the input protein as a novel PFT.

3 . The method of claim 2 , further comprising:

receiving confirmation that the input protein is not a novel PFT; and

re-training the machine learning model such that the sequence of the input protein is identified as a non-PFT sequence.

4 . The method of claim 1 , further comprising generating, for output on a computing device, an indication that the input protein is classified as a potential novel PFT and that the input protein shares a functionality of a particular protein cluster from the plurality of protein clusters.

5 . The method of claim 1 , wherein determining the plurality of protein clusters further comprises:

mapping a structural representation of each protein from the plurality of proteins from a high-dimensional space to a two-dimensional space that preserves structural correlations among the plurality of proteins; and

executing a clustering algorithm on the two-dimensional space to determine the plurality of protein clusters.

6 . The method of claim 5 , wherein the clustering algorithm is a K-means clustering algorithm.

7 . The method of claim 1 , wherein generating the graphical model further comprises:

aligning structures of the proteins from each respective group of proteins to identify common structural regions using an iterative pairwise alignment algorithm; and

identifying consensus and non-consensus secondary structure segments in the aligned structures, wherein training the graphical model comprises maximizing a probability of the identified consensus and non-consensus secondary structure segments in the structural segmentation of the input protein.

8 . The method of claim 1 , wherein the graphical model is a semi-Markov conditional fields (semi-CRFs) model.

9 . The method of claim 1 , wherein proteins in a respective group of proteins share functionality and have a low sequence identity, optionally less than 60, 50, 40, 30, 20, or 10% full length sequence identity.

10 . A system for identifying novel pore-forming toxins (PFTs) based on sequence and structure data, comprising:

a processor, and

memory having stored therein instructions that when executed by the processor cause the processor to:

identify, in a dataset comprising PFT information, a plurality of proteins with known sequences and structures;

determine a plurality of protein clusters based on pairwise structural similarity values of the plurality of proteins;

for each respective protein cluster of the plurality of protein clusters, identify a respective group of proteins that have a pairwise sequence identity a) lower than a threshold pairwise sequence identity, b) above a threshold pairwise sequence identity, or c) within a predetermined pairwise sequence identity range;

generate a graphical model trained using sequence and structure data of proteins from each respective group of proteins, wherein the graphical model is configured to generate a structural segmentation of an input protein based on a sequence of the input protein;

calculate a segment interaction score for the generated structural segmentation of the input protein, wherein the segment interaction score compares the generated structural segmentation of the input protein with structures of the proteins from each respective group of proteins; and

in response to determining that the segment interaction score is greater than a threshold segment interaction score, classify the input protein as a potential novel PFT.

11 . The system of claim 10 , wherein the memory further comprises instructions for:

determining whether the sequence of the input protein is classified as a PFT sequence using a machine learning model configured to classify sequences as a PFT sequence or a non-PFT sequence; and

in response to determining that the sequence is classified as a PFT sequence, identifying the input protein as a novel PFT.

12 . The system of claim 11 , wherein the memory further comprises instructions for:

receiving confirmation that the input protein is not a novel PFT; and

re-training the machine learning model such that the sequence of the input protein is identified as a non-PFT sequence.

13 . The system of claim 10 , wherein the memory further comprises instructions for:

generating, for output on a computing device, an indication that the input protein is classified as a potential novel PFT and that the input protein shares a functionality of a particular protein cluster from the plurality of protein clusters.

14 . The system of claim 10 , wherein determining the plurality of protein clusters further comprises:

mapping a structural representation of each protein from the plurality of proteins from a high-dimensional space to a two-dimensional space that preserves structural correlations among the plurality of proteins; and

executing a clustering algorithm on the two-dimensional space to determine the plurality of protein clusters.

15 . The system of claim 14 , wherein the clustering algorithm is a K-means clustering algorithm.

16 . The system of claim 10 , wherein generating the graphical model further comprises:

aligning structures of the proteins from each respective group of proteins to identify common structural regions using an iterative pairwise alignment algorithm; and

identifying consensus and non-consensus secondary structure segments in the aligned structures, wherein training the graphical model comprises maximizing a probability of the identified consensus and non-consensus secondary structure segments in the structural segmentation of the input protein.

17 . The system of claim 10 , wherein the graphical model is a semi-Markov conditional fields (semi-CRFs) model.

18 . The system of claim 10 , wherein proteins in a respective group of proteins share functionality and have low sequence identity.

Assignments (5)
CHANGE OF NAME Recorded Apr 27, 2026
From: BASF AGRICULTURAL SOLUTIONS SEED US LLC
To: BASF AGRICULTURAL SOLUTIONS US LLC
Reel/Frame 075477/0435 →
CHANGE OF NAME Recorded Jan 24, 2025
From: BASF AGRICULTURAL SOLUTIONS SEED US LLC
To: BASF AGRICULTURAL SOLUTIONS US LLC
Reel/Frame 070005/0447 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2023
From: JACOB, THEJU; KAHN, THEODORE; BASF CORPORATION
To: BASF CORPORATION; BASF SE
Reel/Frame 065477/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2023
From: BASF SE
To: BASF AGRICULTURAL SOLUTIONS SEED US LLC
Reel/Frame 065477/0643 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2023
From: LIU, YAN; XU, NAN
To: UNIVERSITY OF SOUTHERN CALIFORNIA
Reel/Frame 065477/0667 →