IP Library › Granted Patent US 12,462,903
Granted Patent B2
US 12,462,903 · App. 18/505,728 · Granted Nov 4, 2025

Utilizing compound-protein machine learning representations to generate bioactivity predictions

Inventors: Seyed Ali Madani Tonekaboni (Toronto, CA); Daniella Fiora Lato (Toronto, CA); Stephen Scott MacKinnon (Burlington, CA)
Assignee: Recursion Pharmaceuticals, Inc.
G16C20/70G16B40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,903
App. No.
18/505,728
Granted
Nov 4, 2025
Kind
B2
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods that utilizing compound-protein machine learning representations to generate target results. For example, the disclosed systems can utilize a compound-protein interaction machine learning model to generate a compound-protein machine learning representation for compound protein pairs. The disclosed systems can utilize the compound-protein machine learning representation to train and utilize other target machine learning models in generating predicted bioactivity results. For example, the disclosed systems train a target machine learning model from compound-protein machine learning representations to generate ADMET predictions and/or biological perturbation program predictions. Furthermore, the disclosed systems can utilize one or more explainability models in conjunction with target machine learning models trained based on compound-protein machine learning representations to identify proteins that contribute to predicted bioactivity results.

Claims (97)

1 . A computer-implemented method comprising:

identifying a plurality of compound-protein pairs comprising a training compound matched to a plurality of proteins;

generating, utilizing a compound-protein interaction machine learning model from the plurality of compound-protein pairs, a plurality of binding scores between the training compound and the plurality of proteins;

generating a compound-protein machine learning representation comprising the plurality of binding scores between the training compound and the plurality of proteins;

generating, utilizing the compound-protein interaction machine learning model, an additional compound-protein machine learning representation comprising an additional plurality of binding scores between an additional training compound and the plurality of proteins; and

iteratively training a target machine learning model to improve accuracy of the target machine learning model by,

for a first training iteration:

inputting the compound-protein machine learning representation comprising the plurality of binding scores between the training compound and the plurality of proteins into the target machine learning model to generate a first predicted bioactivity result for the training compound; and

modifying parameters of the target machine learning model by comparing the first predicted bioactivity result to a ground truth bioactivity result corresponding to the training compound; and

for a second training iteration:

inputting the additional compound-protein machine learning representation comprising the additional plurality of binding scores between the additional training compound and the plurality of proteins into the target machine learning model to generate an additional predicted bioactivity result for the additional training compound; and

modifying the parameters of the target machine learning model by comparing the additional predicted bioactivity result to an additional ground truth bioactivity result corresponding to the additional training compound.

2 . The computer-implemented method of claim 1 , further comprising:

determining, utilizing a loss function, a measure of loss based on comparing the first predicted bioactivity result to the ground truth bioactivity result;

determining, utilizing the loss function, an additional measure of loss based on comparing the additional predicted bioactivity result to the additional ground truth bioactivity result; and

modifying the parameters of the target machine learning model to reduce the measure of loss and the additional measure of loss.

3 . The computer-implemented method of claim 1 , wherein the compound-protein interaction machine learning model comprises a compound-protein interaction neural network having parameters trained to generate binding scores indicating probabilities that compounds will bind to protein pockets of proteins.

4 . The computer-implemented method of claim 1 , further comprising:

determining that the compound-protein interaction machine learning model will perform below a threshold confidence accuracy based on applying a protein confidence filter to the compound-protein machine learning representation; and

in response to determining that the compound-protein interaction machine learning model will perform below a threshold confidence accuracy, generating a refined compound-protein machine learning representation by removing one or more binding scores from the compound-protein machine learning representation.

5 . The computer-implemented method of claim 4 , wherein the first predicted bioactivity result comprises a biological perturbation program prediction and training the target machine learning model comprises training the target machine learning model to generate biological perturbation program predictions utilizing a training dataset by, for a biological perturbation program corresponding to identifying compounds demonstrating a target biological activity:

generating, from the compound-protein machine learning representation utilizing the target machine learning model, the biological perturbation program prediction for the training compound; and

modifying the parameters of the target machine learning model by comparing the biological perturbation program prediction with a ground truth perturbation program result.

6 . The computer-implemented method of claim 5 , further comprising generating the training dataset by:

generating measures of similarity between a target gene of the biological perturbation program and datapoints of the training dataset, wherein the measures of similarity are based on at least one of phenomic data, transcriptomic data, metabolomic data, or proteomic data; and

generating the training dataset by filtering datapoints based on the measures of similarity between the datapoints and the target gene.

7 . The computer-implemented method of claim 6 , wherein the phenomic data comprises phenomic image embeddings and further comprising generating the training dataset by:

identifying phenomic digital images of cell perturbations;

generating, utilizing a machine learning model, phenomic image embeddings from the phenomic digital images; and

generating the training dataset by filtering datapoints based on pheno-similarity measures of the phenomic image embeddings relative to a target gene.

8 . The computer-implemented method of claim 5 , further comprising generating the training dataset by:

generating, utilizing a clustering model, clusters from a dataset utilizing chemical fingerprints of compounds; and

splitting the dataset into the training dataset and a testing data set based on the clusters.

9 . The computer-implemented method of claim 1 , wherein:

the first predicted bioactivity result comprises an absorption, distribution, metabolism, excretion, or toxicity (ADMET) prediction; and

training the target machine learning model comprises training the target machine learning model to generate ADMET predictions for the first training iteration by:

generating, from the compound-protein machine learning representation utilizing the target machine learning model, the ADMET prediction for the training compound; and

modifying the parameters of the target machine learning model by comparing the ADMET prediction to the ground truth bioactivity result, the ground truth bioactivity result comprising a measured ADMET result for the training compound.

10 . A system comprising:

at least one processor; and

at least one non-transitory computer-readable storage medium storing instructions that, when executed by the at least one processor, cause the system to:

identifying a plurality of compound-protein pairs comprising a training compound matched to a plurality of proteins;

generate, utilizing a compound-protein interaction machine learning model from the plurality of compound-protein pairs, a plurality of binding scores between the training compound and the plurality of proteins;

generate a compound-protein machine learning representation comprising the plurality of binding scores between the training compound and the plurality of proteins;

generate, utilizing the compound-protein interaction machine learning model, an additional compound-protein machine learning representation comprising an additional plurality of binding scores between an additional training compound and the plurality of proteins; and

iteratively train a target machine learning model to improve accuracy of the target machine learning model by,

for a first training iteration:

inputting the compound-protein machine learning representation comprising the plurality of binding scores between the training compound and the plurality of proteins into the target machine learning model to generate a first predicted bioactivity result for the training compound; and

modifying parameters of the target machine learning model by comparing the first predicted bioactivity result to a ground truth bioactivity result corresponding to the training compound; and

for a second training iteration:

inputting the additional compound-protein machine learning representation comprising the additional plurality of binding scores between the additional training compound and the plurality of proteins into the target machine learning model to generate an additional predicted bioactivity result for the additional training compound; and

modifying the parameters of the target machine learning model by comparing the additional predicted bioactivity result to an additional ground truth bioactivity result corresponding to the additional training compound.

11 . The system of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the system to:

determine, utilizing a loss function, a measure of loss based on comparing the first predicted bioactivity result to the ground truth bioactivity result;

determine, utilizing the loss function, an additional measure of loss based on comparing the additional predicted bioactivity result to the additional ground truth bioactivity result; and

modify the parameters of the target machine learning model to reduce the measure of loss and the additional measure of loss.

12 . The system of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the system to generate the compound-protein machine learning representation by:

determining machine learning protein confidence scores indicating a measure of confidence of the compound-protein interaction machine learning model in generating binding predictions for the plurality of proteins; and

filtering one or more features based on the machine learning protein confidence scores to generate the compound-protein machine learning representation.

13 . The system of claim 10 , further comprising instructions that, when executed by the at least one processor, cause the system to:

determine that the compound-protein interaction machine learning model will perform below a threshold confidence accuracy based on applying a protein confidence filter to the compound-protein machine learning representation; and

in response to determining that the compound-protein interaction machine learning model will perform below a threshold confidence accuracy, generate a refined compound-protein machine learning representation by removing one or more binding scores from the compound-protein machine learning representation.

14 . The system of claim 13 , wherein the first predicted bioactivity result comprises a biological perturbation program prediction and further comprising instructions that, when executed by the at least one processor, cause the system to train the target machine learning model to generate biological perturbation program predictions utilizing a training dataset by, for a biological perturbation program corresponding to identifying compounds demonstrating a target biological activity:

generating, from the compound-protein machine learning representation utilizing the target machine learning model, the biological perturbation program prediction for the training compound; and

modifying the parameters of the target machine learning model by comparing the biological perturbation program prediction with a ground truth biological perturbation program result.

15 . The system of claim 14 , further comprising instructions that, when executed by the at least one processor, cause the system to generate the training dataset by:

identifying phenomic digital images of cell perturbations;

generating, utilizing a machine learning model, phenomic image embeddings from the phenomic digital images; and

generating the training dataset by filtering datapoints based on a measure of similarity of the phenomic image embeddings relative to the target biological activity.

16 . The system of claim 10 :

wherein the first predicted bioactivity result comprises an absorption, distribution, metabolism, excretion, or toxicity (ADMET) prediction; and

further comprising instructions that, when executed by the at least one processor, cause the system to train the target machine learning model to generate ADMET predictions for the first training iteration by:

generating, from the compound-protein machine learning representation utilizing the target machine learning model, the ADMET prediction for the training compound; and

modifying the parameters of the target machine learning model by comparing the ADMET prediction to the ground truth bioactivity result, the ground truth bioactivity result comprising a measured ADMET result for the training compound.

17 . A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to:

identify a plurality of compound-protein pairs comprising a training compound matched to a plurality of proteins;

generate, utilizing a compound-protein interaction machine learning model from the plurality of compound-protein pairs, a plurality of binding scores between the training compound and the plurality of proteins;

generate a compound-protein machine learning representation comprising the plurality of binding scores between the training compound and the plurality of proteins;

generate, utilizing the compound-protein interaction machine learning model, an additional compound-protein machine learning representation comprising an additional plurality of binding scores between an additional training compound and the plurality of proteins; and

iteratively train a target machine learning model to improve accuracy of the target machine learning model by,

for a first training iteration:

inputting the compound-protein machine learning representation comprising the plurality of binding scores between the training compound and the plurality of proteins into the target machine learning model to generate a first predicted bioactivity result for the training compound; and

modifying parameters of the target machine learning model by comparing the first predicted bioactivity result to a ground truth bioactivity result corresponding to the training compound; and

for a second training iteration:

inputting the additional compound-protein machine learning representation comprising the additional plurality of binding scores between the additional training compound and the plurality of proteins into the target machine learning model to generate an additional predicted bioactivity result for the additional training compound; and

modifying the parameters of the target machine learning model by comparing the additional predicted bioactivity result to an additional ground truth bioactivity result corresponding to the additional training compound.

18 . The non-transitory computer-readable medium of claim 17 , further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the compound-protein machine learning representation by:

determining machine learning protein confidence scores indicating a measure of confidence of the compound-protein interaction machine learning model in generating binding predictions for the plurality of proteins; and

filtering one or more features based on the machine learning protein confidence scores to generate the compound-protein machine learning representation.

19 . The non-transitory computer-readable medium of claim 17 :

wherein the first predicted bioactivity result comprises an absorption, distribution, metabolism, excretion, or toxicity (ADMET) prediction; and

further comprising instructions that, when executed by the at least one processor, cause the computing device to train the target machine learning model to generate ADMET predictions for the first training iteration by:

generating, from the compound-protein machine learning representation utilizing the target machine learning model, the ADMET prediction for the training compound; and

modifying the parameters of the target machine learning model by comparing the ADMET prediction to the ground truth bioactivity result, the ground truth bioactivity result comprising a measured ADMET result for the training compound.

20 . The non-transitory computer-readable medium of claim 17 , wherein the first predicted bioactivity result comprises a biological perturbation program prediction and further comprising instructions that, when executed by the at least one processor, cause the computing device to train the target machine learning model to generate biological perturbation program predictions utilizing a training dataset by, for a biological perturbation program corresponding to identifying compounds demonstrating a target biological activity:

generating, from the compound-protein machine learning representation utilizing the target machine learning model, the biological perturbation program prediction for a training compound; and

modifying the parameters of the target machine learning model by comparing the biological perturbation program prediction with a ground truth biological perturbation program result.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: MADANI TONEKABONI, SEYED ALI; MACKINNON, STEPHEN SCOTT; LATO, DANIELLA FIORA
To: RECURSION PHARMACEUTICALS, INC.
Reel/Frame 065514/0680 →
Continuity (1)
Related Publication 20250157595A1 · May 15, 2025
References Cited (26)
US 20050053999A1 · Gough et al. · 2005 [cited by applicant]
US 20190130290A1 · Kerber et al. · 2019 [cited by applicant]
US 20220076787A1 · Ellington et al. · 2022 [cited by applicant]
US 20220108766A1 · Mackinnon et al. · 2022 [cited by applicant]
US 20220230710A1 · Amimeur et al. · 2022 [cited by applicant]
US 20230063188A1 · Choi et al. · 2023 [cited by applicant]
WO 2020140156A1 · 2020 [cited by applicant]
WO 2022226327A1 · 2022 [cited by applicant]
Lagorce, David, et al. “Computational analysis of calculated physicochemical and ADMET properties of protein-protein interaction inhibitors.” Scientific reports 7.1 (2017): 46277. [cited by examiner]
Atas Guvenilir, Heval, and Tunca Dogan. “How to approach machine learning-based prediction of drug/compound-target interactions.” Journal of Cheminformatics 15.1 (2023): 1-36. [cited by applicant]
Karimi, Mostafa, et al. “DeepAffinity: interpretable deep learning of compound-protein affinity through unified recurrent and convolutional neural networks.” Bioinformatics 35.18 (2019): 3329-3338. [cited by applicant]
U.S. Appl. No. 18/505,754, filed Feb. 14, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/505,754, filed Apr. 26, 2024, Office Action. [cited by applicant]
Hodos et al. Computational approaches to drug repurposing and pharmacology. Willey Interdiscip Rev Syst Biol Med, vol. 8, 46 pages. (Year:2017). [cited by applicant]
Li, Shuya, et al. “MONN: a multi-objective neural network for predicting compound-protein interactions and affinities.” Cell Systems 10.4 (2020): 308-322. (Year: 2020). [cited by applicant]
Mayr, Andreas, et al. “DeepTox: toxicity prediction using deep learning.” Frontiers in Environmental Science 3 (2016): 80. (Year: 2016). [cited by applicant]
Wang, Erniu, et al. “A graph convolutional network-based method for chemical-protein interaction extraction: algorithm development.” JMIR Medical Informatics 8.5 (2020): e17643. (Year: 2020). [cited by applicant]
Wang, Xun, et al. “SSGraphCPI: a novel model for predicting compound-protein interactions based on deep learning.” International Journal of Molecular Sciences 23.7 (2022): 3780. (Year: 2022). [cited by applicant]
U.S. Appl. No. 18/505,748, filed Sep. 23, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/505,748, filed Dec. 20, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/505,748, filed Mar. 19, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/505,754, filed Dec. 3, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/505,754, filed Mar. 31, 2025, Office Action. [cited by applicant]
Hodgson, E., Mailman, R. B., & Chambers, J. E. (Eds.). (1999). Inhibitory concentration (IC). In Macmillan Dictionary of Toxicology (2nd ed.). Macmillan Publishers Ltd. [cited by applicant]
U.S. Appl. No. 18/505,748, Aug. 20, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/505,754, Jul. 16, 2025, Notice of Allowance. [cited by applicant]
Cited By (1)
US 12,694,946