IP Library › Granted Patent US 12,327,616
Granted Patent B2
US 12,327,616 · App. 17/708,384 · Granted Jun 10, 2025

Pre-training molecule embedding GNNs using contrastive learning based on scaffolding

Inventors: Mohammad Reza Sarshogh (Seattle, WA); Robin Abraham (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G16C20/30G06N3/08G16C20/40G16C20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,616
App. No.
17/708,384
Granted
Jun 10, 2025
Kind
B2
Abstract

Systems and methods are provided for generating a training dataset for training a molecule embedding module using contrastive learning, wherein the definition of similarity is based on molecular scaffold similarity. For example, systems access a molecular dataset and separate the molecular dataset into positive samples and negative samples. Systems then generate a training dataset comprising the positive samples and negative samples. Systems and methods are also provided for using the trained molecule embedding module to generate molecule embeddings and for building an end-to-end machine learning model configured to perform molecular embedding analysis and molecular property prediction, the model comprising the trained molecule embedding module and a property prediction module.

Claims (51)

1. A method for generating training batches for training a molecule embedding module configured to be used in an end-to-end model configured to predict molecular properties, the method executed by one or more computer processors comprising:

accessing a molecular dataset comprising a plurality of molecules; and

separating the molecular dataset into a plurality of positive samples and a plurality of negative samples by at least:

accessing a molecular scaffold dataset comprising a plurality of molecular scaffolds, each molecular scaffold defining a core structure of a particular molecule and one or more underlying characteristics of the particular molecule;

determining that at least two molecules included in the plurality of molecules are similar based on the at least two molecules corresponding to a same molecular scaffold;

labeling the at least two molecules as a positive sample;

determining that at least two other molecules included in the plurality of molecules are dissimilar based on the at least two other molecules corresponding to different molecular scaffolds;

labeling the at least two other molecules as a negative sample; and

generating a training batch comprising the plurality of positive samples corresponding to pairs of similar molecules and the plurality of negative samples corresponding to pairs of dissimilar molecules, the training batch configured to train the molecule embedding module using contrastive learning;

applying the molecule embedding module to a training dataset to train the molecule embedding module to generate molecule embeddings using the contrastive learning;

applying the molecule embedding module to an input molecular dataset comprising a plurality of molecules; and

generating a plurality of molecule embeddings, each molecule embedding corresponding to a molecule included in the input molecular dataset.

2. The method of claim 1 , further comprising:

identifying a target molecule to be added to the training batch;

identifying a particular molecular scaffold associated with the target molecule;

determining that no other molecule included in the training batch is associated with the particular molecular scaffold; and

in response to determining that no other molecule included in the training batch is associated with the particular molecular scaffold, adding the target molecule to the training batch.

3. The method of claim 1 , further comprising:

identifying a target molecule to be added to the training batch;

identifying a particular molecular scaffold associated with the target molecule;

determining that at least one other molecule included in the training batch is associated with the particular molecular scaffold; and

in response to determining that at least one other molecule included in the training batch is associated with the particular molecular scaffold, refraining from adding the target molecule to the training batch.

4. The method of claim 1 , wherein the molecule embedding module comprises one or more graph neural network layers.

5. A method for identifying positive samples from a set of molecular in order to map similar molecules close to each other in a molecular embedding space, the method executed by one or more computer processors comprising:

(i) accessing a molecule embedding module that has been previously trained to map similar molecules close to each other, and dissimilar molecules far apart, the molecule embedding module having been previously trained by a process that includes at least:

accessing a molecular dataset comprising a plurality of molecules; and

separating the molecular dataset into a plurality of positive samples and a plurality of negative samples by at least:

accessing a molecular scaffold dataset comprising a plurality of molecular scaffolds, each molecular scaffold defining a core structure of a particular molecule and one or more underlying characteristics of the particular molecule;

determining that at least two molecules are similar molecules;

labeling the at least two molecules as a positive sample;

determining that at least two other molecules are dissimilar molecules and labeling the at least two other molecules as a negative sample; and

generating a training batch comprising the plurality of positive samples corresponding to pairs molecules corresponding to a same molecular scaffold and the plurality of negative samples corresponding to pairs of dissimilar molecules which correspond to different molecular scaffolds; and

(ii) applying the molecule embedding module to an input molecular dataset to generate a plurality of molecule embeddings, each molecule embedding corresponding to at least one molecule included in the input molecular dataset, wherein the input molecular dataset comprises a plurality of molecular graphs; and

identifying positive molecular samples and negative molecular samples from the plurality of molecule embeddings, wherein each positive molecular sample is identified based on a similar molecular scaffold and each negative molecular sample is identified based on dissimilar molecular scaffold.

6. A method for building an end-to-end machine learning model configured to predict molecular properties, the method executed by one or more computer processors comprising:

(i) accessing a molecule embedding module that has been previously trained to map similar molecules close to each other and dissimilar molecules far apart in an embedding space, by applying the molecule embedding module to a training dataset, the training dataset being generated by a process that includes at least:

accessing a molecular dataset comprising a plurality of molecules;

separating the molecular dataset into a plurality of positive samples and a plurality of negative samples by at least:

accessing a molecular scaffold dataset comprising a plurality of molecular scaffolds, each molecular scaffold defining a core structure of a particular molecule and one or more underlying characteristics of the particular molecule,

determining that at least two molecules are similar molecules,

labeling the at least two molecules as a positive sample,

determining that at least two other molecules are dissimilar molecules and labeling the at least two other molecules as a negative sample; and

generating a training batch comprising the plurality of positive samples corresponding to pairs of molecules corresponding to a same molecular scaffold and the plurality of negative samples corresponding to pairs of dissimilar molecules which correspond to different molecular scaffolds; and

(ii) accessing a property prediction module configured to predict one or more molecular properties; and

(iii) generating an end-to-end machine learning model configured to predict molecular properties by compiling the molecular embedding module with the property prediction module;

applying the molecule embedding module to an input molecular dataset comprising a plurality of molecules;

generating a plurality of molecule embeddings, each molecule embedding corresponding to a molecule included in the input molecular dataset;

identifying a particular molecular property associated with one or more candidate molecules; and

using the property prediction module to predict that a target molecule is associated with the particular molecular property based on a similarity score to the one or more candidate molecules.

7. The method of claim 6 , wherein the molecule embedding dataset comprises a plurality of molecular graphs.

8. The method of claim 6 , wherein the molecule embedding graph comprises one or more graph neural networks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2022
From: SARSHOGH, MOHAMMAD REZA; ABRAHAM, ROBIN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 059479/0329 →
Continuity (2)
Provisional Application 63301177 · Jan 20, 2022
Related Publication 20230230662A1 · Jul 20, 2023
References Cited (37)
US 8595154B2 · Breckenridge · 2013 [cited by examiner]
US 11127488B1 · Bajpai · 2021 [cited by examiner]
US 11144790B2 · Li · 2021 [cited by examiner]
US 11288575B2 · Tomioka · 2022 [cited by examiner]
US 11500905B2 · Baughman · 2022 [cited by examiner]
US 11720813B2 · Babu · 2023 [cited by examiner]
US 11762635B2 · Brown · 2023 [cited by examiner]
US 11763599B2 · Wang · 2023 [cited by examiner]
US 20200105375A1 · Pan · 2020 [cited by examiner]
US 20200388401A1 · Spiro · 2020 [cited by examiner]
US 20210004670A1 · Tripathi · 2021 [cited by examiner]
US 20210081804A1 · Stojevic · 2021 [cited by examiner]
US 20220198286A1 · Prat · 2022 [cited by examiner]
“Losses-PyTorch Metric Learning”, Retrieved from: https://web.archive.org/web/20210330140241/https://kevinmusgrave.github.io/pytorch-metric-learning/losses/, Mar. 30, 2021, 30 Pages. [cited by applicant]
“RDKit: Open-Source Cheminformatics Software”, Retrieved from: https://web.archive.org/web/20220316172251/https://www.rdkit.org/, Mar. 16, 2022, 2 Pages. [cited by applicant]
“SMILES Tutorial”, Retrieved from: https://web.archive.org/web/20170127075822/https://archive.epa.gov/med/med_archive_03/web/html/smiles.html, Jan. 27, 2017, 3 Pages. [cited by applicant]
Brody, et al., “How Attentive are Graph Attention Networks?”, In repository of arXiv:2105.14491v3, Jan. 31, 2022, 26 Pages. [cited by applicant]
Chen, et al., “A Simple Framework for Contrastive Learning of Visual Representations”, In Proceedings of the International Conference on Machine Learning, Nov. 21, 2020, 11 Pages. [cited by applicant]
Cho, et al., “Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation”, In Repository of arXiv:1406.1078v3, Sep. 3, 2014, 15 Pages. [cited by applicant]
Corso, et al., “Principal Neighbourhood Aggregation for Graph Nets”, In Journal of Advances in Neural Information Processing Systems, 2020, 12 Pages. [cited by applicant]
Gilmer, et al., “Neural Message Passing for Quantum Chemistry”, In Proceedings of the 34th International Conference on Machine Learning, Jul. 17, 2017, 10 Pages. [cited by applicant]
He, et al., “Momentum Contrast for Unsupervised Visual Representation Learning”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 9726-9735. [cited by applicant]
Hu, et al., “Strategies for Pre-Training Graph Neural Networks”, In repository of arXiv:1905.12265v3, Feb. 18, 2020, 22 Pages. [cited by applicant]
Jiménez-Luna, et al., “Artificial intelligence in drug discovery: recent advances and future perspectives”, In journal of Expert opinion on drug discovery, vol. 16, Issue 9, Sep. 2, 2021, pp. 949-959. [cited by applicant]
Kipf, et al., “Semi-Supervised Classification with Graph Convolutional Networks”, In Repository of arXiv:1609.02907, Feb. 22, 2017, 14 Pages. [cited by applicant]
Li, et al., “Gated Graph Sequence Neural Networks”, In repository of arXiv:1511.05493v4, Sep. 22, 2017, 20 Pages. [cited by applicant]
Li, et al., “Learn molecular representations from large-scale unlabeled molecules for drug discovery”, In repository of arXiv:2012.11175v1, Dec. 21, 2020, 20 Pages. [cited by applicant]
Liu, et al., “N-Gram Graph: Simple Unsupervised Representation for Graphs, with Applications to Molecules”, In Journal of Advances in Neural Information Processing Systems, 2019, 13 Pages. [cited by applicant]
Oord, et al., “Representation Learning with Contrastive Predictive Coding”, In Repository of arXiv:1807.03748v1, Jul. 10, 2018, 13 Pages. [cited by applicant]
Rong, et al., “Self-supervised Graph Transformer on Large-scale Molecular Data”, In Proceedings of the 34th International Conference on Neural Information Processing Systems, Dec. 2020, 13 Pages. [cited by applicant]
Sutskever, et al., “Sequence to Sequence Learning with Neural Networks”, In Proceedings of Advances in neural information processing systems, Sep. 10, 2014, 9 Pages. [cited by applicant]
Wang, et al., “Molecular Contrastive Learning of Representations via Graph Neural Networks”, In repository of arXiv:2102.10056v2, Jan. 17, 2022, 19 Pages. [cited by applicant]
Zhu, “Dual-view Molecule Pre-training”, In repository of arXiv:2106.10234v2, Oct. 13, 2021, 15 Pages. [cited by applicant]
“ImageNet”, Retrieved from: https://web.archive.org/web/20211021000659/https://www.image-net.org/, Oct. 21, 2021, 1 Page. [cited by applicant]
Gawehn, et al., “Deep Learning in Drug Discovery”, In Journal of Molecular Informatics, vol. 35, Issue 1, Jan. 2016, pp. 3-14. [cited by applicant]
U.S. Appl. No. 63/301,177, filed Jan. 20, 2022. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/052916”, Mailed Date: Apr. 19, 2023, 11 Pages. [cited by applicant]