IP Library › Granted Patent US 12,367,397
Granted Patent B2
US 12,367,397 · App. 17/016,640 · Granted Jul 22, 2025

Query-based molecule optimization and applications to functional molecule discovery

Inventors: Samuel Chung Hoffman (New York, NY); Enara C Vijil (Westchester, NY); Pin-Yu Chen (White Plains, NY); Payel Das (Yorktown Heights, NY); Kahini Wadhawan (Delhi, IN)
Assignee: International Business Machines Corporation
G06N3/126G06F16/245G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,397
App. No.
17/016,640
Filed
Sep 10, 2020
Granted
Jul 22, 2025
Kind
B2
Examiner
TRIEU, EM N
Art Unit
2128
USPC
706/13
Abstract

A query-based generic end-to-end molecular optimization (“QMO”) system framework, method and computer program product for optimizing molecules, such as for accelerating drug discovery. The QMO framework decouples representation learning and guided search and applies to any plug-in encoder-decoder with continuous latent representations. QMO framework directly incorporates evaluations based on chemical modeling, analysis packages, and pre-trained machine-learned prediction models for efficient molecule optimization using a query-based guided search method based on zeroth order optimization. The QMO features efficient guided search with molecular property evaluations and constraints obtained using the predictive models and chemical modeling and analysis packages. QMO tasks include optimizing drug-likeness and penalized log P scores with similarity constraints and improving the target binding affinity of existing drugs to pathogens such as the SARS-CoV-2 main protease protein while preserving the desired drug properties. QMO tasks further improves optimizing antimicrobial peptides toward lower toxicity.

Claims (74)

1. An end-to-end query-based molecule optimization method comprising:

receiving, at a pre-trained plug-in encoder-decoder model running on at least one hardware processor of a computer system, a 1-dimensional string representation of an original molecule having a plurality of properties to be optimized;

encoding, using the trained plug-in encoder-decoder model running on the at least one hardware processor, the 1-dimensional string representation of the original molecule to be optimized into a latent vector representation of a latent vector space of the trained encoder-decoder model;

modifying, by the at least one hardware processor, the latent vector representation in the latent vector space;

decoding, by the trained plug-in encoder-decoder model running on the at least one hardware processor, the latent vector representation to obtain a decoded candidate molecule sequence structure;

inputting, by the at least one hardware processor, the decoded candidate molecule sequence structure to each of a plurality of independent, pre-trained machine learned molecular property prediction models, the plug-in encoder-decoder model being detached from the plurality of independent, pre-trained, machine-learned property prediction modules, and running, by the at least one hardware processor, the plurality of machine learned molecular property prediction models to predict for said decoded candidate molecule sequence structure a respective plurality of molecular properties used to evaluate the candidate molecule sequence structure;

obtaining, using the at least one hardware processor, a loss function using a combination of the multiple molecular property predictions, said one or more said respective plurality of molecular property predictions for optimizing molecular properties of said original molecule while satisfying one or more constraints, the obtained loss function comprising: a first function term quantifying a molecular constraint loss to be minimized and a second function term quantifying a molecular property score to be maximized;

performing a pseudo-gradient estimation of said loss function using the at least one hardware processor, said pseudo-gradient estimation comprising a zeroth-order optimization over the latent representation vector space of the trained encoder-decoder model that includes iteratively modifying the latent vector by applying one or more random direction queries using said random vector from said latent representation vector space and obtaining, at each iteration, a pseudo-gradient estimation of the loss function and generating corresponding loss values as a measure of differences between respective plurality of molecular property predictions and a corresponding respective plurality of specified threshold constraints, and

using the generated loss values at each iteration to determine a direction in the latent space for further modifying said latent vector in said latent representation vector space to achieve improved predicted molecular properties;

determining by the at least one hardware processor when said each of the plurality of properties predicted for said corresponding further modified latent vector representation or its corresponding decoded candidate molecule sequence structure satisfy all said corresponding respective plurality of the specified threshold constraints; and

outputting to a manufacturing pipeline, by the at least one hardware processor, the corresponding further modified sequence structure as an optimized original molecule to be manufactured when each of the plurality of properties predicted for said further modified sequence structure satisfies all said respective plurality of the specified threshold constraints, wherein the end-to-end query-based molecule optimization is model-agnostic as to how molecule representations are learned and trained.

2. The method as claimed in claim 1 , wherein a sequence structure of said original molecule to be optimized is a 1-dimensional sequence of symbols, said method further comprising:

encoding, by the at least one hardware processor, the 1-dimensional sequence of symbols by mapping said 1-dimensional sequence of symbols to a data vector in the latent representation vector space, said data vector comprising a latent representation of the 1-dimensional sequence of symbols at a latent dimension.

3. The method as claimed in claim 2 , wherein said modifying said latent vector representation comprises:

adding a perturbation to said data vector to obtain a modified data vector, said perturbation comprising a random vector;

said method further comprising:

decoding said modified data vector to obtain a new modified sequence structure corresponding to the original molecule for said evaluation of the respective plurality of predicted properties.

4. The method as claimed in claim 1 , wherein said loss function is formulated to:

optimize a molecular similarity to said original molecule while satisfying desired chemical properties, said loss function comprising: a first function term quantifying a property validity loss to be minimized and a second function term quantifying a molecular similarity score to be maximized.

5. The method as claimed in claim 3 , wherein said loss function is an objective function to be minimized, said method comprises:

performing a zeroth order gradient descent method to solve said objective function; and

obtaining a generated loss value from solving said objective function.

6. The method as claimed in claim 5 , further comprising:

updating an iterate of said objective function as a function of the generated loss value of a current iteration.

7. A non-transitory computer readable medium comprising instructions that, when executed by at least one hardware processor, configure the at least one hardware processor to:

receive, at a pre-trained plug-in encoder-decoder model running on the least one hardware processor, a 1-dimensional string representation of an original molecule having a plurality of properties to be optimized;

encode, using the trained plug-in encoder-decoder model, the 1-dimensional string representation of the original molecule to be optimized into a latent vector representation of a latent vector space of the trained encoder-decoder model;

modify the latent vector representation in the latent vector space;

decode, by the trained plug-in encoder-decoder model, the latent vector representation in the latent vector space to obtain a decoded candidate molecule sequence structure;

input the decoded candidate molecule sequence structure to each of a plurality of machine learned molecular property prediction models, the plug-in encoder-decoder model being detached from the plurality of independent, pre-trained, machine-learned property prediction modules, and run the plurality of machine learned molecular property prediction models to predict for said decoded candidate molecule sequence structure a respective plurality of molecular properties used to evaluate the candidate molecule sequence structure;

obtain a loss function using a combination of the multiple molecular property predictions, said one or more said respective plurality of molecular property predictions for optimizing molecular properties of said original molecule while satisfying one or more constraints, the obtained loss function comprising: a first function term quantifying a molecular constraint loss to be minimized and a second function term quantifying a molecular property score to be maximized;

perform a pseudo-gradient estimation of said loss function, said pseudo-gradient estimation comprising a zeroth-order optimization over the latent representation vector space of the trained encoder-decoder model that includes iteratively modifying the latent vector by applying one or more random direction queries using said random vector from said latent representation vector space and obtain, at each iteration, a pseudo-gradient estimation of the loss function and generate corresponding loss values as a measure of differences between a respective plurality of molecular property predictions and a corresponding respective plurality of specified threshold constraints, and

use the generated loss values at each iteration to determine a direction in the latent space for further modifying said latent vector representation in said latent vector space to achieve improved predicted properties of the original molecule to be optimized; and

determine when said each of the plurality of properties predicted for the corresponding further modified latent vector representation or its corresponding decoded candidate molecule sequence structure satisfy all said corresponding respective plurality of the specified threshold constraints; and

output to a manufacturing pipeline the corresponding further modified sequence structure as an optimized original molecule to be manufactured when each of the plurality of properties predicted for said further modified sequence structure satisfies all said respective plurality of the specified threshold constraints, wherein an end-to-end query-based molecule optimization is provided that is model-agnostic as to how the molecule representations are learned and trained.

8. The non-transitory computer readable medium as claimed in claim 7 , wherein a sequence structure of said original molecule to be optimized is a 1-dimensional sequence of symbols, said instructions further configure the at least one hardware processor to:

encode the 1-dimensional sequence of symbols by mapping said 1-dimensional sequence of symbols to a data vector in the latent representation vector space, said data vector comprising a latent representation of the 1-dimensional sequence of symbols at a latent dimension.

9. The non-transitory computer readable medium as claimed in claim 8 , wherein to modify said latent vector representation, said instructions further configure the at least one hardware processor to:

add a perturbation to said data vector to obtain a modified data vector, said perturbation comprising a random vector; and

said instructions further configuring the at least one hardware processor to:

decode said modified data vector to obtain a new modified sequence structure corresponding to the original molecule for said evaluation of the respective plurality of predicted properties.

10. The non-transitory computer readable medium as claimed in claim 7 , wherein said loss function is formulated to:

optimize a molecular similarity to said original molecule while satisfying desired chemical properties, said loss function comprising: a first function term quantifying a property validity loss to be minimized and a second function term quantifying a molecular similarity score to be maximized.

11. The non-transitory computer readable medium as claimed in claim 9 , wherein the loss function is an objective function to be minimized, the instructions further configure the at least one hardware processor to:

perform a zeroth order gradient descent method to solve said objective function; and

obtain a generated loss value from solving said objective function.

12. The non-transitory computer readable medium as claimed in claim 11 , wherein the instructions further configure the at least one hardware processor to:

update an iterate of said objective function as a function of the generated loss value of a current iteration.

13. An end-to-end computer-implemented query-based molecule optimization system comprising:

a memory storage device; and

a hardware processor coupled to said memory storage device and configured to perform a method to:

receive, at a pre-trained plug-in encoder-decoder model running on the least one hardware processor, a 1-dimensional string representation of an original molecule having a plurality of properties to be optimized;

encode, using the trained plug-in encoder-decoder model, the 1-dimensional string representation of the original molecule to be optimized into a latent vector representation of a latent vector space of the trained encoder-decoder model;

modify the latent vector representation in the latent vector space;

decode, by the trained plug-in encoder-decoder model, the latent vector representation in the latent vector space to obtain a decoded candidate molecule sequence structure;

input the decoded candidate molecule sequence structure to each of a plurality of independent, pre-trained machine learned molecular property prediction models, the plug-in encoder-decoder model being detached from the plurality of independent, pre-trained, machine-learned property prediction modules, and run the plurality of machine learned molecular property prediction models to predict for said decoded candidate molecule sequence structure a respective plurality of molecular properties used to evaluate the candidate molecule sequence structure;

obtain a loss function using a combination of the multiple molecular property predictions, said one or more said respective plurality of molecular property predictions for optimizing molecular properties of said original molecule while satisfying one or more constraints, the obtained loss function comprising: a first function term quantifying a molecular constraint loss to be minimized and a second function term quantifying a molecular property score to be maximized;

perform a pseudo-gradient estimation of said loss function, said pseudo-gradient estimation comprising a zeroth-order optimization over the latent representation vector space of the trained encoder-decoder model that includes iteratively modifying the latent vector by applying one or more random direction queries using said random vector from said latent representation vector space and obtain, at each iteration, a pseudo-gradient estimation of the loss function and generate corresponding loss values as a measure of differences between a respective plurality of molecular property predictions and a corresponding respective plurality of specified threshold constraints, and

use the generated loss values at each iteration to determine a direction in the latent space for further modifying said latent vector representation in said latent vector space to achieve improved predicted properties of the original molecule to be optimized;

determine when said each of the plurality of properties predicted for the corresponding further modified latent vector representation or its corresponding decoded candidate molecule sequence structure satisfy all said corresponding respective plurality of the specified threshold constraints; and

output to a manufacturing pipeline the corresponding further modified sequence structure as an optimized original molecule to be manufactured when each of the plurality of properties predicted for said further modified sequence structure satisfies all said respective plurality of the specified threshold constraints, wherein the end-to-end query-based molecule optimization is model-agnostic as to how molecule representations are learned and trained.

14. The computer-implemented system as claimed in claim 13 , wherein a sequence structure of said original molecule to be optimized is a 1-dimensional sequence of symbols, said hardware processor further configured to:

encode the 1-dimensional sequence of symbols by mapping said 1-dimensional sequence of symbols to a data vector in the latent representation vector space, said data vector comprising a latent representation of the 1-dimensional sequence of symbols at a latent dimension.

15. The computer-implemented system as claimed in claim 14 , wherein said modifying said latent vector representation comprises:

adding a perturbation to said data vector to obtain a modified data vector, said perturbation comprising a random vector; and

said hardware processor is further configured to:

decode said further modified data vector to obtain a new modified sequence structure corresponding to the original molecule for said evaluation of the respective plurality of predicted properties.

16. The computer-implemented system as claimed in claim 15 , wherein said loss function is formulated to:

optimize a molecular similarity to said original molecule while satisfying desired chemical properties, said loss function comprising: a first function term quantifying a property validity loss to be minimized and a second function term quantifying a molecular similarity score to be maximized.

17. The computer-implemented system as claimed in claim 16 , wherein the loss function is an objective function to be minimized, the hardware processor is further configured to:

perform a zeroth order gradient descent method to solve said objective function; and

obtain a generated loss value from solving said objective function.

18. The computer-implemented system as claimed in claim 17 , wherein the hardware processor is further configured to:

update an iterate of said objective function as a function of the generated loss value of a current iteration.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2020
From: HOFFMAN, SAMUEL CHUNG; VIJIL, ENARA C; CHEN, PIN-YU; DAS, PAYEL; WADHAWAN, KAHINI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 053732/0567 →
Continuity (1)
Related Publication 20220076137A1 · Mar 10, 2022
References Cited (69)
US 11049590B1 · Lee · 2021 [cited by examiner]
US 20070192037A1 · Jojic · 2007 [cited by examiner]
US 20200082916A1 · Polykovskiy et al. · 2020 [cited by applicant]
US 20200090049A1 · Aliper et al. · 2020 [cited by applicant]
WO WO2109231624A2 · 2019 [cited by applicant]
WO WO2020243440A1 · 2020 [cited by examiner]
Deep-learning-based inverse design model for intelligent (Year: 2018). [cited by examiner]
Deep-learning-based inverse design model for intelligent discovery of organic molecules (Year: 2018). [cited by examiner]
ZOO: Zeroth Order Optimization Based Black-box Attacks toDeep Neural Networks without Training Substitute Models (Year: 2017). [cited by examiner]
AutoZOOM: Autoencoder-based Zeroth Order Optimization Method for Attacking Black-box Neural Networks (Year: 2019). [cited by examiner]
Generative Network Complex for the Automated Generation of Drug-like Molecules (Year: 2020). [cited by examiner]
Chenthamarakshan et al., “CogMol: Target-Specific and Selective Drug Design for COVID-19 Using Deep Generative Models,” [Submitted on Apr. 2, 2020 (v1), last revised Jun. 24, 2020 (this version, v2)] https://arxiv.org/a… [cited by applicant]
Liu, “A Primer on Zeroth-Order Optimization in Signal Processing and Machine Learning.” [Submitted on Jun. 11, 2020 (v1), last revised Jun. 21, 2020 (this version, v2)], https://arxiv.org/abs/2006.06224, 13 pages. [cited by applicant]
Winter et al. “Efficient multi-objective molecular optimization in a continuous latent space.” Chemical Science, vol. 10, 34, pp. 8016-8024, Jul. 8, 2019. [cited by applicant]
Wang et al., “Zeroth Order Optimization by a Mixture of Evolution Strategies,” Sep. 25, 2019 (modified: Dec. 24, 2019), ICLR 2020 Conference Blind Submission, https://openreview.net/forum?id=SKIE_CNFPr, pp. 1-11. [cited by applicant]
Gómez-Bombarelli et al., “Automatic chemical design using a data-driven continuous representation of molecules,” [Submitted on Oct. 7, 2016 (v1), last rev. Dec. 5, 2017 (this version, v3)], https://arxiv.org/abs/1610.02… [cited by applicant]
Griffiths et al., “Constrained Bayesian optimization for automatic chemical design using variational autoencoders,” Chem. Sci., 11, pp. 577-586, 2020.Downloaded on Aug. 7, 2020. [cited by applicant]
Zhao et al., “On the Design of Black-box Adversarial Examples by Leveraging Gradient-free Optimization and Operator Splitting Method”, IEEE International Conference on Computer Vision, 2019, pp. 121-130. [cited by applicant]
Zhou et al., “Optimization of Molecules via Deep Reinforcement Learning”, Scientific reports, vol. 9, No. 1, 2019, pp. 1-10. [cited by applicant]
Altschul et al., “Basic Local Alignment Search Tool”, Journal of molecular biology, vol. 215, No. 3, 1990, pp. 403-410. [cited by applicant]
Bahdanau et al., “Neural machine translation by jointly learning to align and translate”, in International Conference on Learning Representations, 2015, 15 pages. [cited by applicant]
Bickerton et al., “Quantifying the chemical beauty of drugs”, Nature chemistry, vol. 4, No. 2, 2012, 27 pages. [cited by applicant]
Bohacek et al., “The art and practice of structure-based drug design: A molecular modeling perspective”, Medicinal research reviews, vol. 16, No. 1, 1996, 3 pages (copy of table of contents attached). [cited by applicant]
Chen et al., “ZO-AdaMM: Zeroth-Order Adaptive Momentum Method for Black-Box Optimization”, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 12 pages. [cited by applicant]
Cheng et al., “Query-Efficient Hard-label Black-box Attack: An Optimization-based Approach”, arXiv, Jul. 12, 2018, 12 pages. [cited by applicant]
Coates et al., “Novel classes of antibiotics or more of the same?”, British journal of pharmacology, vol. 163, No. 1, 2011, pp. 184-194. [cited by applicant]
Dalke et al., “mmpdb: An Open-Source Matched Molecular Pair Platform for Large Multiproperty Data Sets”, Journal of Chemical Information and Modeling, vol. 58, No. 5, 2018, pp. 902-910. [cited by applicant]
Das et al., “Accelerating Antimicrobial Discovery with Controllable Deep Generative Models and Molecular Dynamics”, arXiv, Feb. 26, 2021, 64 pages. [cited by applicant]
Dossetter et al., “Matched Molecular Pair Analysis in drug discovery”, Drug Discovery Today, vol. 18, No. 15-16, 2013, pp. 724-731. [cited by applicant]
Ertl et al., “Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions”, Journal of Cheminformatics, vol. 1, Jun. 2009, 11 pages. [cited by applicant]
Fu et al., “CORE: Automatic Molecule Optimization Using Copy & Refine Strategy”, The Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI-20), 2020, 8 pages. [cited by applicant]
Ghadimi et al., “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,” SIAM Journal on Optimization, vol. 23, No. 4, 2013, 25 pages. [cited by applicant]
Golovin et al., “Gradientless Descent: High-Dimensional Zeroth-Order Optimization”, arXiv, May 18, 2020, 26 pages. [cited by applicant]
Griffen et al., “Matched Molecular Pairs as a Medicinal Chemistry Tool: Miniperspective”, Journal of medicinal chemistry, vol. 54, No. 22, 2011, pp. 7739-7750. [cited by applicant]
Guimaraes et al., “Objective-Reinforced Generative Adversarial Networks (ORGAN) for Sequence Generation Models”, arXiv, Feb. 7, 2018, 7 pages. [cited by applicant]
Gupta et al., “In Silico Approach for Predicting Toxicity of Peptides and Proteins”, PloS one, vol. 8, No. 9, 2013, 10 pages. [cited by applicant]
Hasan et al., “HLPpred-Fuse: improved and robust prediction of hemolytic peptide and its activity by fusing multiple feature representation”, Bioinformatics, vol. 36, Apr. 2020, pp. 3350-3356. [cited by applicant]
Huynh et al., “In Silico Exploration of the Molecular Mechanism of Clinically Oriented Drugs for Possibly Inhibiting SARS-COV-2's Main Protease”, The Journal of Physical Chemistry Letters, vol. 11, 2020, pp. 4413-4420. [cited by applicant]
Ilyas et al., “Black-box Adversarial Attacks with Limited Queries and Information”, International Coference on International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Jang et al., “Categorical reparameterization with gumbel-softmax”, Published as a conference paper at ICLR (International Conference on Learning Representations), 2017, 12 pages. [cited by applicant]
Jeon et al., “Identification of Antiviral Drug Candidates against SARS-COV-2 from FDA-Approved Drugs”, Antimicrobial Agents and Chemotherapy, vol. 64, No. 7, Jul. 2020, 9 pages. [cited by applicant]
Jiménez-Luna et al., “DeltaDelta neural networks for lead optimization of small molecule potency”, Chemical Science, vol. 10, No. 47, 2019, pp. 10911-10918. [cited by applicant]
Jin et al., “Hierarchical graph-to-graph translation for molecules”, arXiv, Oct. 18, 2019, 14 pages. [cited by applicant]
Jin et al., “Junction Tree Variational Autoencoder for Molecular Graph Generation”, Proceedings of the 35 th International Conference on Machine Learning, Stockholm, Sweden, PMLR 80, 2018, 10 pages. [cited by applicant]
Jin et al., “Learning multimodal graph-to-graph translation for molecule optimization”, International Conference on Learning Representations, 2019, 13 pages. [cited by applicant]
Jin et al., “Structure of Mpro from SARS-COV-2 and discovery of its inhibitors”, Nature, vol. 582, Jun. 2020, 24 pages. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization”, International Conference on Learning Representations, 2015, 15 pages. [cited by applicant]
Korovina et al., “ChemBO: Bayesian Optimization of Small Organic Molecules with Synthesizable Recommendations”, Proceedings of the 23rdInternational Conference on Artificial Intelligence and Statistics (AISTATS) 2020, P… [cited by applicant]
Landrum et al., “rdkit/rdkit: 2020_09_5 (Q3 2020) Release”, https://doi.org/10.5281/zenodo.4570805, Mar. 1, 2021, 8 pages. [cited by applicant]
Liu et al., “signSGD via Zeroth-order Oracle”, Published as a conference paper at ICLR (International Conference on Learning Representations), 2019, 24 pages. [cited by applicant]
Liu et al., “Zeroth-Order Online Alternating Direction Method of Multipliers: Convergence Analysis and Applications”, in International Conference on Artificial Intelligence and Statistics, vol. 84, Apr. 9-11, 2018, pp. … [cited by applicant]
Liu et al., “Zeroth-Order Stochastic Variance Reduction for Nonconvex Optimization”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), 2018, 11 pages. [cited by applicant]
Maragakis, et al., A Deep-Learning View of Chemical Space Designed to Facilitate Drug Discovery, Journal of Chemical Information and Modeling, vol. 60, 2020, pp. 4487-4496. [cited by applicant]
Olivecrona et al., “Molecular de-novo design through deep reinforcement learning”, Journal of cheminformatics, vol. 9, No. 1, 2017, 14 pages. [cited by applicant]
Polykovskiy et al., “Molecular Sets (Moses): A Benchmarking Platform for Molecular Generation Models”, arXiv, Oct. 28, 2020, 19 pages. [cited by applicant]
Qin et al., “Artificial intelligence method to design and fold alphahelical structural proteins from the primary amino acid sequence”, Extreme Mechanics Letters, vol. 36, 2020, 23 pages. [cited by applicant]
Reutlinger et al., “Multi-Objective Molecular De Novo Design by Adaptive Fragment Prioritization”, Angewandte Chemie International Edition, vol. 53, No. 16, 2014, pp. 4244-4248. [cited by applicant]
Reymond et al., “The enumeration of chemical space,” Wiley Interdisciplinary Reviews: Computational Molecular Science, vol. 2, No. 5, 2012, 2 pages (abstract only). [cited by applicant]
Rogers et al., “Extended-Connectivity Fingerprints”, Journal of chemical information and modeling, vol. 50, No. 5, 2010, pp. 742-754. [cited by applicant]
Sanchez-Lengeling et al., “Optimizing distributions over molecular space. An Objective-Reinforced Generative Adversarial Network for Inverse-design Chemistry (Organic)”, chemrxiv preprint chemrxiv.5309668, 2017, 18 page… [cited by applicant]
Skalic et al., “Shape-Based Generative Modeling for de Novo Drug Design”, Journal of chemical information and modeling, vol. 59, No. 3, 2019, pp. 1205-1214. [cited by applicant]
Sterling et al., “ZINC 15 - Ligand Discovery for Everyone”, Journal of chemical information and modeling, vol. 55, No. 11, 2015, pp. 2324-2337. [cited by applicant]
Trott et al., “AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization and multithreading”, Journal of Computational Chemistry, vol. 31, No. 2, 2010, pp. 455-461. [cited by applicant]
Weininger, “Smiles, a Chemical Language and Information System. 1. Introduction to Methodology and Encoding Rules”, Journal of chemical information and computer sciences, vol. 28, No. 1, 1988, pp. 31-36. [cited by applicant]
Winter et al., “Learning continuous and data-driven molecular descriptors by translating equivalent chemical representations”, Chemical science, vol. 10, No. 6, 2019, pp. 1692-1701. [cited by applicant]
Xiao et al., “iAMP-2L: A two-level multi-label classifier for identifying antimicrobial peptides and their functional types”, Analytical Biochemistry, vol. 436, 2013, pp. 168-177. [cited by applicant]
Yang et al., “Improving Molecular Design by Stochastic Iterative Target Augmentation”, Proceedings of the 37 th International Conference on Machine Learning, Online, PMLR 119, 2020, 2020, 11 pages. [cited by applicant]
You et al., “Graph Convolutional Policy Network for Goal-Directed Molecular Graph Generation”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), 12 pages. [cited by applicant]
Yuan et al., LigBuilder 2: A Practical de Novo Drug Design Approach, Journal of chemical information and modeling, vol. 51, No. 5, 2011, pp. 1083-1091. [cited by applicant]
Cited By (2)
US 12,694,954 US 12,718,904