IP Library Granted Patent US 11,587,646
Granted Patent B2
US 11,587,646 · App. 16/702,119 · Granted Feb 21, 2023

Method for simultaneous characterization and expansion of reference libraries for small molecule identification

Inventors: Sean M. Colby (Richland, WA); Ryan S. Renslow (Richland, WA)
Assignee: Battelle Memorial Institute
G16C20/20G06N3/08G06N20/10G16C20/30G16C20/60G16C20/70G16C60/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,587,646
App. No.
16/702,119
Granted
Feb 21, 2023
Kind
B2
Abstract

A variational autoencoder (VAE) has been developed to learn a continuous numerical, or latent, representation of molecular structure to expand reference libraries for small molecule identification. The VAE has been extended to include a chemical property decoder, trained as a multitask network, to shape the latent representation such that it assembles according to desired chemical properties. The approach is unique in its application to metabolomics and small molecule identification, focused on properties that are obtained from experimental measurements (m/z, CCS) paired with its training paradigm, which involves a cascade of transfer learning iterations. First, molecular representation is learned from a large dataset of structures with m/z labels. Next, in silico property values are used to continue training. Finally, the network is further refined by being trained with the experimental data. The trained network is used to predict chemical properties directly from structure and generate candidate structures with desired chemical properties. The network is extensible to other training data and molecular representations, and for use with other analytical platforms, for both chemical property and feature prediction as well as molecular structure generation.

Claims (47)

1. A method of simultaneous characterization and expansion of reference libraries for small molecule identification comprising:

extending a variational autoencoder (VAE) to include a chemical property decoder, so as to form an extended VAE,

training the extended VAE as a multi-task network, to shape a latent representation that self-assembles according to preselected criteria, and

finding correlations embedded in the latent representation,

wherein the correlations embedded in the latent representation enable prediction of chemical properties from structures generated along manifolds defined by chemical property analogues.

2. The method of claim 1 wherein training includes processing a cascade of transfer learning iterations comprising: a first dataset of unlabeled structures, a second dataset of properties calculated in silico and a third dataset of limited experimental data for fine tuning.

3. The method of claim 2 wherein the first dataset is larger than the second dataset, and the second dataset is larger than the third dataset.

4. The method of claim 2 wherein an output of processing the cascade of transfer learning iterations comprises a set of weights.

5. The method of claim 1 wherein the correlations embedded in the latent representation enable prediction of the structures from the chemical properties.

6. The method of claim 1 further comprising implementing character embedding to learn relationships among different characters or symbols that are used to represent molecular structure.

7. The method of claim 1 further comprising implementing a beam search to enable the vector representation of a structure to be converted back into multiple, most-probable full structures.

8. The method of claim 1 further comprising implementing a global attention mechanism which informs a user why certain chemical bonds are giving certain properties.

9. The method of claim 1 further comprising implementing a molecular structure discriminator which is an added deep learning network that determines whether a structure is both syntactically and chemically valid.

10. The method of claim 1 further comprising implementing a molecular structure discriminator which is used to modify a training loss function during training of the VAE, such that the latent space would be penalized for and therefore be trained to avoid creating structures that have physically/chemically invalid bonding types, molecular topologies, and/or nonbonding interactions.

11. The method of claim 1 further comprising implementing a molecular structure discriminator that is trained simultaneous to the VAE with an external structure discriminator, such that the latent space is forced to produce increasingly valid structures and that the structure discriminator becomes fine-tuned to reject the increasingly smaller space of invalid structures that the VAE may produce, which therefore iteratively improves the percent of valid structures that the VAE produces and increase the discrimination power of the molecular structure discriminator.

12. An apparatus comprising:

a memory configured for storing an application, the application configured for:

extending a variational autoencoder (VAE) to include a chemical property decoder, so as to form an extended VAE,

training the extended VAE as a multi-task network, to shape a latent representation that self-assembles according to preselected criteria, and

finding correlations embedded in the latent representation, wherein the correlations embedded in the latent representation enable prediction of chemical properties from structures generated along manifolds defined by chemical property analogues; and

a processor configured for processing the application.

13. The apparatus of claim 12 wherein training includes processing a cascade of transfer learning iterations comprising: a first dataset of unlabeled structures, a second dataset of properties calculated in silico and a third dataset of limited experimental data for fine tuning.

14. The apparatus of claim 13 wherein the first dataset is larger than the second dataset, and the second dataset is larger than the third dataset.

15. The apparatus of claim 13 wherein an output of processing the cascade of transfer learning iterations comprises a set of weights.

16. The apparatus of claim 12 wherein the correlations embedded in the latent representation enable prediction of the structures from the chemical properties.

17. The apparatus of claim 12 wherein the application is further configured for implementing character embedding to learn relationships among different characters or symbols that are used to represent molecular structure.

18. The apparatus of claim 12 wherein the application is further configured for implementing a beam search to enable the vector representation of a structure to be converted back into multiple, most-probable full structures.

19. The apparatus of claim 12 wherein the application is further configured for implementing a global attention mechanism which informs a user why certain chemical bonds are giving certain properties.

20. The apparatus of claim 12 wherein the application is further configured for implementing a molecular structure discriminator which is an added deep learning network that determines whether a structure is both syntactically and chemically valid.

21. The apparatus of claim 12 wherein the application is further configured for implementing a molecular structure discriminator which is used to modify a training loss function during training of the VAE, such that the latent space would be penalized for and therefore be trained to avoid creating structures that have physically/chemically invalid bonding types, molecular topologies, and/or nonbonding interactions.

22. The apparatus of claim 12 wherein the application is further configured for implementing a molecular structure discriminator that is trained simultaneous to the VAE with an external structure discriminator, such that the latent space is forced to produce increasingly valid structures and that the structure discriminator becomes fine-tuned to reject the increasingly smaller space of invalid structures that the VAE may produce, which therefore iteratively improves the percent of valid structures that the VAE produces and increase the discrimination power of the molecular structure discriminator.

23. A system comprising:

a first device for storing training data; and

a second device configured for:

extending a variational autoencoder (VAE) to include a chemical property decoder, so as to form an extended VAE,

training the extended VAE as a multi-task network using the training data, to shape a latent representation that self-assembles according to preselected criteria, and

finding correlations embedded in the latent representation, wherein the correlations embedded in the latent representation enable prediction of chemical properties from structures generated along manifolds defined by chemical property analogues.

24. The system of claim 23 wherein training includes processing a cascade of transfer learning iterations comprising: a first dataset of unlabeled structures, a second dataset of properties calculated in silico and a third dataset of limited experimental data for fine tuning.

25. The system of claim 24 wherein the first dataset is larger than the second dataset, and the second dataset is larger than the third dataset.

26. The system of claim 24 wherein an output of processing the cascade of transfer learning iterations comprises a set of weights.

27. The system of claim 23 wherein the correlations embedded in the latent representation enable prediction of the structures from the chemical properties.

28. The system of claim 23 wherein the second device is further configured for implementing character embedding to learn relationships among different characters or symbols that are used to represent molecular structure.

29. The system of claim 23 wherein the second device is further configured for implementing a beam search to enable the vector representation of a structure to be converted back into multiple, most-probable full structures.

30. The system of claim 23 wherein the second device is further configured for implementing a global attention mechanism which informs a user why certain chemical bonds are giving certain properties.

31. The system of claim 23 wherein the second device is further configured for implementing a molecular structure discriminator which is an added deep learning network that determines whether a structure is both syntactically and chemically valid.

32. The system of claim 23 wherein the second device is further configured for implementing a molecular structure discriminator which is used to modify a training loss function during training of the VAE, such that the latent space would be penalized for and therefore be trained to avoid creating structures that have physically/chemically invalid bonding types, molecular topologies, and/or nonbonding interactions.

33. The system of claim 23 wherein the second device is further configured for implementing a molecular structure discriminator that is trained simultaneous to the VAE with an external structure discriminator, such that the latent space is forced to produce increasingly valid structures and that the structure discriminator becomes fine-tuned to reject the increasingly smaller space of invalid structures that the VAE may produce, which therefore iteratively improves the percent of valid structures that the VAE produces and increase the discrimination power of the molecular structure discriminator.

Assignments (2)
CONFIRMATORY LICENSE Recorded Jan 29, 2020
From: BATTELLE MEMORIAL INSTITUTE, PACIFIC NORTHWEST DIVISION
To: U.S. DEPARTMENT OF ENERGY
Reel/Frame 051752/0274 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2019
From: COLBY, SEAN M.; RENSLOW, RYAN S.
To: BATTELLE MEMORIAL INSTITUTE
Reel/Frame 051165/0607 →
Continuity (2)
Provisional Application 62774663 · Dec 3, 2018
Related Publication 20200176087A1 · Jun 4, 2020
Cited By (6)
US 12,190,236 US 12,368,503 US 12,587,274 US 12,603,701 US 12,627,372 US 12,664,813