IP Library › Granted Patent US 12,211,588
Granted Patent B1
US 12,211,588 · App. 18/479,524 · Granted Jan 28, 2025

Methods and systems for using deep learning to predict protein-ligand complex structures in protein folding and protein-ligand co-folding

Inventors: Lakshyaditya Singh Aithani (Oxford, GB); Liam Atkinson (Padua, IT)
Assignee: Charm Therapeutics Limited
G16B15/30G16B15/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,211,588
App. No.
18/479,524
Granted
Jan 28, 2025
Kind
B1
Abstract

A method for using machine learning to predict protein-ligand complex structures in protein folding and/or protein-ligand co-folding is presented. The method includes receiving, at a processor, a predicted set of coordinates for (1) a plurality of protein atoms and (2) a plurality of ligand atoms, included in a protein-ligand structure. At least one adjustment to be applied to the predicted set of coordinates is determined based on a plurality of loss functions associated with the plurality of protein atoms and/or the plurality of ligand atoms. A machine learning model is trained to predict geometric information of an unseen protein-ligand structure based on the at least one adjustment.

Claims (58)

1. A method, comprising:

receiving, at a processor, a predicted spatial representation for each atom from a plurality of atoms included in a first protein-ligand structure, the plurality of atoms including (1) a plurality of protein atoms and (2) a plurality of ligand atoms;

receiving, at the processor, a ground-truth spatial representation for each atom from the plurality of atoms included in the first protein-ligand structure;

for at least one protein atom from the plurality of protein atoms:

generating, via the processor and for each protein atom from the at least one protein atom, (1) a predicted protein-centric frame based on the predicted spatial representation of the protein atom and (2) a ground-truth protein-centric frame based on the ground-truth spatial representation of the protein atom, and

for each ligand atom that is from the plurality of ligand atoms and that is associated with at least one of the predicted protein-centric frame or the ground-truth protein-centric frame:

generating, via the processor, a protein-centric translation loss based on (1) the predicted spatial representation for the ligand atom relative to the predicted protein-centric frame and (2) the ground-truth spatial representation for the ligand atom relative to the ground-truth protein-centric frame, to produce a set of protein-centric translation losses;

for each central ligand atom included in the plurality of ligand atoms:

generating, via the processor, (1) a predicted ligand-centric frame based on the predicted spatial representation of the central ligand atom and (2) a ground-truth ligand-centric frame based on the ground-truth spatial representation of the central ligand atom, and

for each protein atom from at least one protein atom associated with at least one of the predicted ligand-centric frame or the ground-truth ligand-centric frame:

generating, via the processor, a ligand-centric translation loss based on (1) the predicted spatial representation for the protein atom relative to the predicted ligand-centric frame and (2) the ground-truth spatial representation for the ligand atom relative to the ground-truth ligand-centric frame, to produce a set of ligand-centric translation losses; and

training a machine learning model to produce a trained machine learning model that is configured to predict a plurality of coordinates associated with a second protein-ligand structure based on the set of protein-centric translation losses and the set of ligand-centric translation losses.

2. The method of claim 1 , wherein the at least one protein atom includes at least one alpha-carbon protein atom and excludes any sidechain protein atoms.

3. The method of claim 1 , wherein the at least one protein atom includes at least one sidechain protein atom.

4. The method of claim 1 , further comprising:

generating a ligand-only loss based on the predicted spatial representation for each ligand atom and the ground-truth spatial representation for each ligand atom, the training the machine learning model being further based, at least in part, on the ligand-only loss.

5. The method of claim 4 , wherein the ligand-only loss includes a N-hop distance loss configured to determine a plausibility of a bond between at least two ligand atoms from the plurality of ligand atoms.

6. The method of claim 4 , wherein the ligand-only loss includes a bond angle loss that is determined for at least one bond between at least two ligand atoms from the plurality of ligand atoms.

7. The method of claim 4 , wherein the ligand-only loss includes a dihedral loss that is associated with at least one bond torsion for at least one bond between at least two ligand atoms from the plurality of ligand atoms.

8. A non-transitory processor-readable medium storing code representing instructions to be executed by one or more processors, the instructions comprising code to cause the one or more processors to:

receive a predicted spatial representation for each atom from a plurality of atoms included in a first protein-ligand structure, the plurality of atoms including (1) a plurality of protein atoms and (2) a plurality of ligand atoms;

receive a ground-truth spatial representation for each atom from the plurality of atoms included in the protein-ligand structure;

for at least one protein atom from the plurality of protein atoms:

generate, for each protein atom from the at least one protein atom, (1) a predicted protein-centric frame including a predicted spatial representation of the protein atom based on the predicted spatial representation of the protein atom and (2) a ground-truth protein-centric frame including a ground truth spatial representation of the protein atom based on the ground-truth spatial representation of the protein atom, and

for each ligand atom from the plurality of ligand atoms and associated with at least one of the predicted protein-centric frame or the ground-truth protein-centric frame:

generate a protein-centric translation loss based on (1) the predicted spatial representation for the ligand atom relative to the predicted protein-centric frame and (2) the ground-truth spatial representation for the ligand atom relative to the ground-truth protein-centric frame, to produce a set of protein-centric translation losses;

for each central ligand atom included in the plurality of ligand atoms:

generate (1) a predicted ligand-centric frame based on the predicted spatial representation of the central ligand atom and (2) a ground-truth ligand-centric frame based on the ground-truth spatial representation of the central ligand atom, and

for each protein atom from at least one protein atom associated with at least one of the predicted ligand-centric frame or the ground-truth ligand-centric frame:

generate a ligand-centric translation loss based on (1) the predicted spatial representation for the protein atom relative to the predicted ligand-centric frame and (2) the ground-truth spatial representation for the ligand atom relative to the ground-truth ligand-centric frame, to produce a set of ligand-centric translation losses; and

train a machine learning model to produce a trained machine learning model that is configured to predict a plurality of coordinates associated with a second protein-ligand structure based on the set of protein-centric translation losses and the set of ligand-centric translation losses.

9. The non-transitory processor-readable medium of claim 8 , wherein the at least one protein atom includes at least one alpha-carbon protein atom and excludes any sidechain protein atoms.

10. The non-transitory processor-readable medium of claim 8 , wherein the at least one protein atom includes at least one sidechain protein atom.

11. The non-transitory processor-readable medium of claim 8 , wherein the instructions further comprise code to cause the one or more processors to:

generate a ligand-only loss based on the predicted spatial representation for each ligand atom and the ground-truth spatial representation for each ligand atom, the training the machine learning model being further based, at least in part, on the ligand-only loss.

12. The non-transitory processor-readable medium of claim 11 , wherein the ligand-only loss includes a N-hop distance loss configured to determine a plausibility of a bond between at least two ligand atoms from the plurality of ligand atoms.

13. The non-transitory processor-readable medium of claim 11 , wherein the ligand-only loss includes a bond angle loss that is determined for at least one bond between at least two ligand atoms from the plurality of ligand atoms.

14. The non-transitory processor-readable medium of claim 11 , wherein the ligand-only loss includes a dihedral loss that is associated with at least one bond torsion for at least one bond between at least two ligand atoms from the plurality of ligand atoms.

15. An apparatus, comprising:

a processor; and

a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to:

receive a predicted spatial representation for each atom from a plurality of atoms included in a first protein-ligand structure, the plurality of atoms including (1) a plurality of protein atoms and (2) a plurality of ligand atoms;

receive a ground-truth spatial representation for each atom from the plurality of atoms included in the protein-ligand structure;

for at least one protein atom from the plurality of protein atoms:

generate, for each protein atom from the at least one protein atom, (1) a predicted protein-centric frame based on the predicted spatial representation of the protein atom and (2) a ground-truth protein-centric frame based on the ground-truth spatial representation of the protein atom, and

for each ligand atom that is from the plurality of ligand atoms and that is associated with at least one of the predicted protein-centric frame or the ground-truth protein-centric frame:

generate a protein-centric translation loss based on (1) the predicted spatial representation for the ligand atom relative to the predicted protein-centric frame and (2) the ground-truth spatial representation for the ligand atom relative to the ground-truth protein-centric frame, to produce a set of protein-centric translation losses;

for each central ligand atom included in the plurality of ligand atoms:

generate (1) a predicted ligand-centric frame including a predicted spatial representation of the central ligand atom and based on the predicted spatial representation of the central ligand atom and (2) a ground-truth ligand-centric frame including a ground truth spatial representation of the central ligand atom based on the ground-truth spatial representation of the central ligand atom, and

for each protein atom from at least one protein atom associated with at least one of the predicted ligand-centric frame or the ground-truth ligand-centric frame:

generate a ligand-centric translation loss based on (1) the predicted spatial representation for the protein atom relative to the predicted ligand-centric frame and (2) the ground-truth spatial representation for the ligand atom relative to the ground-truth ligand-centric frame, to produce a set of ligand-centric translation losses; and

train a machine learning model to predict a plurality of coordinates associated with a second protein-ligand structure based on the set of protein-centric translation losses and the set of ligand-centric translation losses.

16. The apparatus of claim 15 , wherein the at least one protein atom includes at least one alpha-carbon protein atom and excludes any sidechain protein atoms.

17. The apparatus of claim 15 , wherein the at least one protein atom includes at least one sidechain protein atom.

18. The apparatus of claim 15 , wherein the memory storing further instructions that, when executed by the processor, cause the processor to:

generate a ligand-only loss based on the predicted spatial representation for each ligand atom and the ground-truth spatial representation for each ligand atom, the training the machine learning model being further based, at least in part, on the ligand-only loss.

19. The apparatus of claim 18 , wherein the ligand-only loss includes a N-hop distance loss configured to determine a plausibility of a bond between at least two ligand atoms from the plurality of ligand atoms.

20. The apparatus of claim 18 , wherein the ligand-only loss includes a bond angle loss that is determined for at least one bond between at least two ligand atoms from the plurality of ligand atoms.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2023
From: AITHANI, LAKSHYADITYA SINGH; ATKINSON, LIAM
To: CHARM THERAPEUTICS LIMITED
Reel/Frame 065794/0340 →
Continuity (1)
Provisional Application 63411676 · Sep 30, 2022
References Cited (24)
US 20210166779A1 · Jumper et al. · 2021 [cited by applicant]
WO WO2020058174A1 · 2020 [cited by applicant]
WO WO2020058176A1 · 2020 [cited by applicant]
WO WO2020058177A1 · 2020 [cited by applicant]
WO WO2020210591A1 · 2020 [cited by applicant]
WO WO2021110730A1 · 2021 [cited by applicant]
Yan, Xu, et al. “PointSite: a point cloud segmentation tool for identification of protein ligand binding atoms.” Journal of Chemical Information and Modeling 62.11 (2022): 2835-2845. [cited by examiner]
Ahdritz, et al., aqlaboratory/openfold: OpenFold v1.0.1, Nov. 12, 2021. [cited by applicant]
Alcaide, et al., MP-NeRF: A massively parallel method for accelerating protein structure reconstruction from internal coordinates, Journal of Computational Chemistry, Jan. 2022, pp. 74-78. [cited by applicant]
Berman, et al., The protein data bank, Nucleic Acids Research, 2000, pp. 235-242. [cited by applicant]
Eberhardt, et al., AutoDock Vina 1.2. 0: New docking methods, expanded force field, and python bindings, Journal of Chemical Information and Modeling, Jul. 2021, pp. 3891-3898. [cited by applicant]
Du, et al., Insights into protein-ligand interactions: mechanisms, models, and methods, International Journal of Molecular Sciences, Jan. 2016, 34 pages. [cited by applicant]
Frimurer, et al., Ligand-induced conformational changes: improved predictions of ligand binding conformations and affinities, Biophysical Journal, Apr. 2003, pp. 2273-2281. [cited by applicant]
Ganea, et al., Geomol: Torsional geometric generation of molecular 3d conformer ensembles, Advances in Neural Information Processing Systems, Dec. 2021, pp. 13757-13769. [cited by applicant]
Grover, Use of allosteric targets in the discovery of safer drugs, Medical Principles and Practice, Sep. 2013, pp. 418-426. [cited by applicant]
Hassan, et al., Protein-ligand blind docking using QuickVina-W with inter-process spatio-temporal integration, Scientific Reports, Nov. 2017, 13 pages. [cited by applicant]
Jumper, et al., Highly accurate protein structure prediction with AlphaFold, Nature, Aug. 2021, pp. 583-589. [cited by applicant]
Kabsch, et al., A solution for the best rotation to relate two sets of vectors, Acta Crystallographica Section A: Crystal Physics, Diffraction, Theoretical and General Crystallography, Sep. 1976, pp. 922-923. [cited by applicant]
Koes, et al., Lessons learned in empirical scoring with smina from the CSAR 2011 benchmarking exercise, Journal of Chemical Information and Modeling, Aug. 2013, pp. 1893-1904. [cited by applicant]
McNutt, et al., GNINA 1.0: molecular docking with deep learning, Journal of Cheminformatics, Dec. 2021, 43 pages. [cited by applicant]
Stark, et al., EquiBind: Geometric deep learning for drug binding structure prediction, International Conference on Machine Learning, Jun. 2022, pp. pp. 20503-20521. [cited by applicant]
Sverrisson, et al., Fast end-to-end learning on protein surfaces, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15272-15281. [cited by applicant]
Trott, et al., AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading, Journal of Computational Chemistry, Jan. 2010, pp. 455-461. [cited by applicant]
Wang, et al., The PDBbind database: Collection of binding affinities for protein-ligand complexes with known three-dimensional structures, Journal of Medicinal Chemistry, Jun. 2004, pp. 2977-2980. [cited by applicant]
Cited By (4)
US 12,367,329 US 12,632,731 US 12,694,946 US 12,712,051