IP Library Patent Application 11818075
Patent Application
App. No. 11/818,075

Methods and systems of common motif and countermeasure discovery

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
11/818,075
Abstract

Disclosed are computational methods, and associated hardware and software products for identifying a set of target pockets for broad-spectrum drug development based on a provided set of protein motifs. A method of identifying the provided set protein motifs based on a plurality of protein motifs in also disclosed herein. Additional methods for generating a plurality of protein motifs based on both aligned protein structure and sequences are disclosed herein.

Claims (62)

1 . A method of identifying a set of target pockets for broad-spectrum drug development, wherein each target pocket comprises a three-dimensional concavity in a protein structure, the method comprising:

providing a set of protein motifs, wherein each protein motif of the set of motifs comprises a first plurality of conserved residues;

identifying a plurality of pockets based on the set of protein motifs, wherein each pocket comprises a second plurality of conserved residues that define the three dimensional concavity on the surface of a protein structure, wherein the second plurality of conserved residues correspond at least in part with the first plurality of conserved residues;

generating a plurality of binding profiles in association with the plurality of pockets, wherein each binding profile specifies at least one calculated binding activity between each pocket and at least one test molecule;

generating a plurality of pocket similarity values based on the plurality of binding profiles, wherein each pocket similarity value is based on binding profiles associated with at least two pockets of the plurality of pockets;

identifying a set of target pockets based on the plurality of pocket similarity values; and

storing the set of target pockets.

2 . The method of claim 1 , wherein identifying a set of target pockets based on the plurality of pocket similarity values comprises:

generating a cluster based on the plurality of pocket similarity values; and

selecting a set of target pockets based on the cluster.

3 . The method of claim 1 , wherein the binding profiles further specify a plurality of calculated binding activities between each pocket and a plurality of test molecules.

4 . The method of claim 1 , wherein the second plurality of conserved residues is less than 4 residues.

5 . The method of claim 1 , wherein each pocket similarity value is based on each binding profile associated with each pocket and a representative binding profile associated with a representative pocket.

6 . The method of claim 3 , further comprising:

generating a plurality of molecule similarity values based on the plurality of binding profiles, wherein each molecule similarity value is based on binding profiles associated with at least two pockets of the plurality of pockets; and

identifying a set of target molecules based on the plurality of molecule similarity values.

7 . The method of claim 1 , wherein the first plurality of conserved residues are conserved in a set of three-dimensional protein structures.

8 . The method of claim 1 , wherein the first plurality of conserved residues are conserved in a set of protein sequences.

9 . The method of claim 1 , wherein providing a set of protein motifs further comprises:

identifying a plurality of protein motifs, wherein each protein motif of the plurality of protein motifs comprises a plurality of conserved residues and is associated with a protein sequence and protein structure;

generating a plurality of protein motif similarity values, wherein each protein motif similarity value is based on the sequence of a reference protein comprising the protein motif, the structure of the reference protein comprising the protein motif or any combination thereof; and

identifying a set of protein motifs based on the plurality of protein motif similarity values.

10 . The method of claim 9 , further comprising identifying the plurality of protein motifs, wherein identifying the plurality of protein motifs comprises:

providing a plurality of sets of aligned three-dimensional protein structures;

identifying a plurality of spans from the plurality of sets of aligned three-dimensional protein structures, wherein each span is comprised of a plurality of residue positions and each residue position is comprised of a one-to-one set of corresponding residues from the aligned three-dimensional structures whose positions differ by less than a pre-determined distance;

generating a plurality of conservation scores for a plurality of residue positions in the span, wherein each conservation score is generated based on a similarity metric and a one-to-one set of corresponding residues; and

identifying a protein plurality of motifs, wherein each motif is based on the generated plurality of conservation scores.

11 . The method of claim 10 , wherein the pre-determined distance is less than 5 Angstroms.

12 . The method of claim 10 , further comprising generating the plurality of sets of aligned three-dimensional structures wherein generating each set of aligned three dimensional structures comprises

identifying a set of homologous three-dimensional protein structures; and

determining the set of aligned three-dimensional protein structures based on a local alignment of the homologous three-dimensional protein structures, a global alignment of the homologous three-dimensional protein structures or any combination thereof.

13 . The method of claim 12 , wherein the set of homologous three-dimensional structures comprises a structure obtained using x-ray crystallography, electron crystallography, nuclear magnetic resonance, computational protein structure modeling, or combinations thereof.

14 . The method of claim 1 , wherein identifying a plurality of pockets based on the set of motifs further comprises:

generating a set of three-dimensional spheres associated with co-ordinates in three-dimensional space, wherein the set of three-dimensional spheres represent a negative image of the surface of the protein structure;

determining a subset of the set of spheres that fall within a second pre-determined distance from the second set of conserved residues in the protein structure based on the co-ordinates of the set of spheres and the co-ordinates of the second set of conserved residues in the protein structure; and

determining that the second set of conserved residues form a three-dimensional concavity on the surface of the protein structure based on the co-ordinates of the subset of the set of spheres.

15 . The method of claim 14 , wherein the second pre-determined distance is less than 8 Angstroms.

16 . The method of claim 1 , further comprising generating the calculated binding activity between each pocket and each test molecule based on generating a docking between the test molecule and the pocket based on computational protein-ligand docking.

17 . The method of claim 9 , further comprising identifying the plurality of protein motifs, wherein identifying the plurality of protein motifs comprises:

providing a set protein sequences;

generating a sequence alignment of the protein sequences based on a multiple sequence alignment, a pair-wise sequence alignment or any combination thereof;

identifying a plurality of conserved residues in each protein sequence based at least in part on the sequence alignment; and

identifying a plurality of protein motifs, wherein each protein motif comprises the plurality of conserved residues in each protein sequence.

18 . The method of claim 17 , further comprising:

identifying a plurality of test molecules based on the motifs;

determining a plurality of quantitative structural relationships, wherein each quantitative structural relationships is based on the binding activity between the set of motifs and the plurality of test molecules.

19 . A computer-readable storage medium comprising computer program code for identifying a set of target pockets for broad-spectrum drug development, wherein each target pocket comprises a three-dimensional concavity in a protein structure, the computer program code for:

providing a set of protein motifs, wherein each protein motif of the set of motifs comprises a first plurality of conserved residues;

identifying a plurality of pockets based on the set of protein motifs, wherein each pocket comprises a second plurality of conserved residues form a three dimensional concavity on the surface of a protein structure, wherein the second plurality of conserved residues correspond at least in part with the first plurality of conserved residues;

generating a plurality of binding profiles in association with the plurality of pockets, wherein each binding profile specifies at least one calculated binding activity between each pocket and at least one test molecule;

generating a plurality of pocket similarity values based on the plurality of binding profiles, wherein each pocket similarity value is based on binding profiles associated with at least two pockets of the plurality of pockets;

identifying a set of target pockets based on the plurality of pocket similarity values; and

storing the set of target pockets.

20 . The computer-readable storage medium of claim 19 , wherein providing a set of motifs further comprises:

identifying a plurality of protein motifs, wherein each protein motif of the plurality of protein motifs comprises a plurality of conserved residues and is associated with a protein sequence and protein structure;

generating a plurality of protein motif similarity values, wherein each protein motif similarity value is based on the sequence of a reference protein comprising the protein motif, the structure of the reference protein comprising the protein motif or any combination thereof; and

identifying a set of protein motifs based on the plurality of protein motif similarity values.

21 . The computer-readable storage medium of claim 19 , further comprising generating the plurality of motifs, wherein generating the plurality of motifs comprises:

providing a plurality of sets of aligned three-dimensional protein structures;

identifying a plurality of spans from the plurality of sets of aligned three-dimensional protein structures, wherein each span is comprised of a plurality of residue positions and each residue position is comprised of a one-to-one set of corresponding residues from the aligned three-dimensional structures whose positions differ by less than a pre-determined distance;

generating a plurality of conservation scores for a plurality of residue positions in the span, wherein each conservation score is generated based on a similarity metric and a one-to-one set of corresponding residues; and

identifying a protein plurality of motifs, wherein each motif is based on the generated plurality of conservation scores.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2008
From: ZHOU, CAROL L. ECALE; ZEMLA, ADAM T.
To: THE REGENTS OF THE UNIVERSITY OF CALIFIORNIA
Reel/Frame 021848/0420 →
CONFIRMATORY LICENSE Recorded Nov 6, 2007
From: CALIFORNIA, THE REGENTS OF THE UNIVERSITY OF
To: U.S. DEPARTMENT OF ENERGY
Reel/Frame 020077/0139 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2007
From: ROE, DIANA C.; SCHOENIGER, JOSEPH S.
To: SANDIA NATIONAL LABORATORIES
Reel/Frame 019964/0881 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2007
From: REGENTS OF THE UNIVERSITY OF CALIFORNIA, THE
To: LAWRENCE LIVERMORE NATIONAL SECURITY, LLC
Reel/Frame 020012/0032 →