IP Library › Granted Patent US 11,031,094
Granted Patent B2
US 11,031,094 · App. 15/737,231 · Granted Jun 8, 2021

Protein structure prediction system

Inventors: Frederick R. Blattner (Madison, WI); Steven J. Darnell (Madison, WI); Matthew R. Larson (Madison, WI); Amanda E. Mitchell (Madison, WI); John L. Schroeder (Madison, WI)
Assignee: DNASTAR, INC.
G16B15/00G06N7/08G16B5/00G16B30/00G16B99/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,031,094
App. No.
15/737,231
Granted
Jun 8, 2021
Kind
B2
Abstract

The present invention is an accelerated conformational sampling method for predicting target peptide and protein structures comprising a process of determining energy minimized synthetic templates using a simple system for modeling individual molecular bonds within the subject peptide or protein. Use of these synthetic templates greatly reduces the computational resources necessary for optimally determining structural features of the target peptide or protein. The present invention also provides methods for rapid and efficient analysis of the effect of mutations on target peptides and proteins.

Claims (18)

1. A conformational sampling method for predicting the structure of an amino acid sequence, the amino acid sequence comprising a plurality of residues, comprising: a) creating a sequence profile matrix of the amino acid sequence; b) determining an alignment of each respective residue of the amino acid sequence against the sequence profile matrix; c) identifying internal residue contacts for one or more residues in the plurality of residues using the alignment; d) collecting the features of the sequence profile matrix and internal residue contacts for each residue in the plurality of residues into a collected feature matrix; e) aligning the collected feature matrix using one or more threading models with a structural feature database of original templates; f) selecting a plurality of optimally aligned original templates from the aligning e); g) calculating normal modes of motion for each original template in the plurality of optimally aligned original templates; h) perturbing each respective original template in the plurality of optimally aligned original templates, for each pair of calculated normal modes, thereby collectively creating a plurality of synthetic templates; i) scoring the energy difference between each original template and the corresponding synthetic template; j) selecting a subset of synthetic templates from the plurality of synthetic templates based on satisfaction of a predetermined cut-off criterion; k) replacing or supplementing the original templates in the plurality of optimally aligned original templates with the corresponding selected subset of synthetic templates to generate a plurality of modeling templates; l) calculating distance and contact restraints within modeling templates of the plurality of modeling templates; m) performing, with modeling templates in the plurality of modeling templates, Markov Chain Monte Carlo simulations, thereby obtaining simulation results; n) clustering the simulation results, wherein the clustering comprises a plurality of clusters, each cluster in the plurality of clusters representing models from the performing m); o) selecting representative models of each cluster in the plurality of clusters; p) refining the representative models of the selecting o) by energy minimization; and q) selecting the lowest energy refined representative model as the predicted structure of the amino acid sequence.

2. The method of claim 1 , wherein creating the sequence profile matrix comprises predicting the secondary structure of the amino acid sequence.

3. The method of claim 1 , wherein creating the sequence profile matrix comprises predicting solvent accessibility of the amino acid sequence.

4. The method of claim 1 , wherein the predetermined cut-off criterion is a predetermined percentile score and the predetermined percentile score is the 65th percentile.

5. The method of claim 1 , wherein the original templates of k) are supplemented with the corresponding subset of selected synthetic templates.

6. The method of claim 1 , wherein the original templates of k) are replaced with the corresponding subset of selected synthetic templates.

7. A computer system for predicting the structure of an amino acid sequence, the amino acid sequence comprising a plurality of residues, the computer system comprising at least one processor and memory storing at least one program for execution by the at least one processor, the at least one program comprising instructions for: a) creating a sequence profile matrix of the amino acid sequence; b) determining an alignment of each respective residue of the amino acid sequence against the sequence profile matrix; c) identifying internal residue contacts for one or more residues in the plurality of residues using the alignment; d) collecting the features of the sequence profile matrix and internal residue contacts for each residue in the plurality of residues into a collected feature matrix; e) aligning the collected feature matrix using one or more threading models with a structural feature database of original templates; f) selecting a plurality of optimally aligned original templates from the aligning e); g) calculating normal modes of motion for each original template in the plurality of optimally aligned original templates; h) perturbing each respective original template in the plurality of optimally aligned original templates, for each pair of calculated normal modes, thereby collectively creating a plurality of synthetic templates; i) scoring the energy difference between each original template and the corresponding synthetic template; j) selecting a subset of synthetic templates from the plurality of synthetic templates based on satisfaction of a predetermined cut-off criterion; k) replacing or supplementing the original templates in the plurality of optimally aligned original templates with the corresponding selected subset of synthetic templates to generate a plurality of modeling templates; l) calculating distance and contact restraints within modeling templates of the plurality of modeling templates; m) performing, with modeling templates in the plurality of modeling templates, Markov Chain Monte Carlo simulations, thereby obtaining simulation results; n) clustering the simulation results, wherein the clustering comprises a plurality of clusters, each cluster in the plurality of clusters representing models from the performing m); o) selecting representative models of each cluster in the plurality of clusters; p) refining the representative models of the selecting o) by energy minimization; and q) selecting the lowest energy refined representative model as the predicted structure of the amino acid sequence.

8. The computer system of claim 7 , wherein creating the sequence profile matrix comprises predicting the secondary structure of the amino acid sequence.

9. The computer system of claim 7 , wherein creating the sequence profile matrix comprises predicting solvent accessibility of the amino acid sequence.

10. The computer system of claim 7 , wherein the predetermined cut-off criterion is a predetermined percentile score and the predetermined percentile score is the 65th percentile.

11. The computer system of claim 7 , wherein the original templates of k) are supplemented with the corresponding subset of selected synthetic templates.

12. The computer system of claim 7 , wherein the original templates of k) are replaced with the corresponding subset of selected synthetic templates.

13. A non-transitory computer readable storage medium storing a computational module for predicting the structure of an amino acid sequence, the amino acid sequence comprising a plurality of residues, the computational module comprising instructions for: a) creating a sequence profile matrix of the amino acid sequence; b) determining an alignment of each respective residue of the amino acid sequence against the sequence profile matrix; c) identifying internal residue contacts for one or more residues in the plurality of residues using the alignment; d) collecting the features of the sequence profile matrix and internal residue contacts for each residue in the plurality of residues into a collected feature matrix; e) aligning the collected feature matrix using one or more threading models with a structural feature database of original templates; f) selecting a plurality of optimally aligned original templates from the aligning e); g) calculating normal modes of motion for each original template in the plurality of optimally aligned original templates; h) perturbing each respective original template in the plurality of optimally aligned original templates, for each pair of calculated normal modes, thereby collectively creating a plurality of synthetic templates; i) scoring the energy difference between each original template and the corresponding synthetic template; j) selecting a subset of synthetic templates from the plurality of synthetic templates based on satisfaction of a predetermined cut-off criterion; k) replacing or supplementing the original templates in the plurality of optimally aligned original templates with the corresponding selected subset of synthetic templates to generate a plurality of modeling templates; l) calculating distance and contact restraints within modeling templates of the plurality of modeling templates; m) performing, with modeling templates in the plurality of modeling templates, Markov Chain Monte Carlo simulations, thereby obtaining simulation results; n) clustering the simulation results, wherein the clustering comprises a plurality of clusters, each cluster in the plurality of clusters representing models from the performing m); o) selecting representative models of each cluster in the plurality of clusters; p) refining the representative models of the selecting o) by energy minimization; and q) selecting the lowest energy refined representative model as the predicted structure of the amino acid sequence.

14. The non-transitory computer readable storage medium of claim 13 , wherein creating the sequence profile matrix comprises predicting the secondary structure of the amino acid sequence.

15. The non-transitory computer readable storage medium of claim 13 , wherein creating the sequence profile matrix comprises predicting solvent accessibility of the amino acid sequence.

16. The non-transitory computer readable storage medium of claim 13 , wherein the predetermined cut-off criterion is a predetermined percentile score and the predetermined percentile score is the 65th percentile.

17. The non-transitory computer readable storage medium of claim 13 , wherein the original templates of k) are supplemented with the corresponding subset of selected synthetic templates.

18. The non-transitory computer readable storage medium of claim 13 , wherein the original templates of k) are replaced with the corresponding subset of selected synthetic templates.

Continuity (2)
Provisional Application 62193225 · Jul 16, 2015
Related Publication 20180260517A1 · Sep 13, 2018
Cited By (2)
US 12,249,406 US 12,542,197