IP Library Patent Application 18148474
Patent Application
App. No. 18/148,474

METHOD AND SYSTEM FOR PREDICTING A BINDING AFFINITY OF PROTEIN STRUCTURES BASED ON DEEP LEARNING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/148,474
Abstract

A system and a method for predicting a binding affinity of protein structures based on deep learning is disclosed. The method includes capturing, using a data capture module, a multi-dimensional structure of a plurality of protein-protein complexes from a protein sequence data set. The method also includes performing featurization, using a featurization module, of the multi-dimensional structure of the plurality of protein-protein complexes by creating an adjacency matrix for successive spherical shells centered around each type of a plurality of amino acid sequences. The method further includes predicting, using a prediction module, a binding affinity from text sequence of the plurality of amino acid sequences using a pre-trained artificial intelligence model based on the adjacency matrix.

Claims (50)

1 . A processor-implemented method of predicting a binding affinity of protein structures based on deep learning, the method comprising:

capturing, using a data capture module, a multi-dimensional structure of a plurality of protein-protein complexes from a protein sequence data set;

performing featurization, using a featurization module, of the multi-dimensional structure of the plurality of protein-protein complexes by creating an adjacency matrix for successive spherical shells centered around each type of a plurality of amino acid sequences; and

predicting, using a prediction module, a binding affinity from text sequence of the plurality of amino acid sequences using a pre-trained artificial intelligence model based on the adjacency matrix, by generating.

2 . The processor-implemented method of claim 1 , wherein the multi-dimensional structure of the plurality of protein-protein complexes comprises three-dimensional coordinates of the atoms in the molecule along with corresponding chain name.

3 . The processor-implemented method of claim 1 , wherein the artificial intelligence model comprises a convolutional neural network (CNN) model.

4 . The processor-implemented method of claim 1 , wherein performing featurization comprises:

calculating a shell feature of each amino-acid pair comprising distances between plurality of amino acid sequences in protein molecules; and

creating a plurality of feature vectors based on the number of amino acid pairs that fit in the interatomic distances.

5 . The processor-implemented method of claim 3 , wherein calculating the shell feature comprises:

calculating a Euclidean distance between each atomic pair by calculating at least one of a minimum distance and a maximum distance between the plurality of amino acid sequences in a shell of a predetermined radius and a predetermined delta value;

determining the shell feature based on the Euclidean distance for a predetermined inner sphere radius and shell thickness; and

assigning a value of 1 to the feature upon the Euclidean distance being between the predetermined inner sphere and sum of the predetermined inner sphere and a delta value and assigning a value 0 to the feature upon the Euclidean distance being beyond the predetermined inner sphere and sum of the predetermined inner sphere and the delta value.

6 . The method of claim 1 , wherein predicting the binding affinity comprises:

generating one or more PKA values indicative of the binding affinity by the pre-trained machine learning model based on the adjacency matrix, wherein the PKA values comprises one of a numerical value or a floating-point value.

7 . The processor-implemented method of claim 1 , wherein the adjacency matrix comprises intra and inter molecular distance values.

8 . The processor-implemented method of claim 1 , wherein predicting the binding affinity comprises:

determining a plurality of parent structures of the plurality of amino acid sequences using an artificial intelligence-based model;

generating a plurality of protein 3D structures form the plurality of amino acid sequences by performing a homology modelling of the features of the amino acid sequences based on the parent structures and using the convolution neural network model;

subjecting the multi-dimensional PDB structures of antibodies and antigens to a docking process;

generating a PDB complex; and

predicting the binding affinity of the plurality of amino acid sequences.

9 . A processor-implemented method of training an artificial intelligence model for predicting a binding affinity of protein structures, the method comprising:

extracting a plurality of feature vectors from a protein sequence data set;

generating a training set for the artificial intelligence model based on the plurality of feature vectors and importing the training set into the artificial intelligence model;

training and evaluating the artificial intelligence model using the training set for predicting the binding affinity of protein structures.

10 . A system for predicting a binding affinity of protein structures based on deep learning, the system comprising:

a non-transitory memory configured to store a protein sequence data set and one or more executable modules; and

a processor configured to execute the one or more executable modules for predicting a binding affinity of protein structures, wherein the one or more executable modules comprises:

a data capture module configured to capture a multi-dimensional structure of a plurality of protein-protein complexes from a protein sequence data set;

a featurization module configured to perform featurization of the multi-dimensional structure of the plurality of protein-protein complexes by creating an adjacency matrix for successive spherical shells centered around each type of amino acid sequences; and

a prediction module configured to predict a binding affinity from text sequence of the amino acid sequences using a pre-trained artificial intelligence model based on the adjacency matrix.

11 . The system of claim 10 , wherein the multi-dimensional structure of plurality of protein-protein complexes comprises three-dimensional coordinates of the atoms in the molecule along with corresponding chain name.

12 . The system of claim 10 , wherein the artificial intelligence model comprises a convolutional neural network (CNN) model.

13 . The system of claim 10 , wherein the featurization module is further configured to:

calculate a shell feature of each amino-acid pair comprising distances between a plurality of amino acids in the molecules; and

create a plurality of feature vectors based on the number of amino acid pairs that fit in the interatomic distances.

14 . The system of claim 10 , wherein the featurization module is further configured to:

calculate a Euclidean distance between each atomic pair by calculating at least one of a minimum distance and a maximum distance between a plurality of amino acid sequences in a shell of a predetermined radius and a predetermined delta value;

determine the shell feature based on the Euclidean distance for a predetermined inner sphere radius and shell thickness; and

assign a value of 1 to the feature upon the Euclidean distance being between the predetermined inner sphere and sum of the predetermined inner sphere and a delta value and assigning a value 0 to the feature upon the Euclidean distance being beyond the predetermined inner sphere and sum of the predetermined inner sphere and the delta value.

15 . The system of claim 10 , wherein the prediction module is further configured to:

generate one or more PKA values indicative of the binding affinity by the pre-trained machine learning model based on the adjacency matrix, wherein the PKA values comprises one of a numerical value or a floating-point value.

16 . The system of claim 10 , wherein the adjacency matrix comprises intra and inter molecular distance values.

17 . The system of claim 10 , wherein the prediction module is further configured to:

determine parent structures of the plurality of amino acid sequences using an artificial intelligence-based model;

generate a plurality of multi-dimensional protein structures from the plurality of amino acid sequences by performing a homology modeling of the features of the plurality of amino acid sequences based on the parent structures and using the convolution neural network model;

subject the multi-dimensional protein structures of antibodies and antigens to a docking process;

generate a protein data bank complex; and

predict the binding affinity of the plurality of amino acid sequences.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2023
From: INNOPLEXUS CONSULTING SERVICES PVT. LTD.
To: INNOPLEXUS AG
Reel/Frame 063203/0232 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2022
From: KUMAR, SUDHANSHU; JOSEPH, JOEL; GUPTA, ANSH
To: INNOPLEXUS CONSULTING SERVICES PVT. LTD.
Reel/Frame 062242/0510 →