IP Library Granted Patent US 12712052
Granted Patent B2
US 12712052 · App. 18/076,280 · Granted Aug 18, 2026

Method and system to identify natural products from mass spectrometry and genomics data

Inventors: Bahar Behsaz (Pasadena, CA); Liu Cao (Pittsburgh, PA); Mustafa Guler (Pittsburgh, PA); Yi-Yuan Lee (Ithaca, NY); Hosein Mohimani (Pittsburgh, PA); Mihir Mongia (Pittsburg, PA); Donghui Yan (Pittburgh, PA)
Assignee: Carnegie Mellon University
G16B40/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12712052
App. No.
18/076,280
Filed
Dec 6, 2022
Granted
Aug 18, 2026
Kind
B2
Art Unit
2852
USPC
702/27
Abstract

A method and system is for receiving data representing gene clusters, the gene clusters including one or more genes configured to encode one or more polypeptides or other small molecules; accessing a machine learning model, the machine learning model being trained with a training dataset that associates the gene clusters to structures of one or more small molecules represented in the data; applying the machine learning model to the data representing the gene clusters; identifying, based on applying the machine learning model, one or more monomers associated with at least one gene cluster represented in the data; and determining a structure for a natural product including the one or more monomers.

Claims (63)

1 . A method comprising:

receiving data representing gene clusters, the gene clusters including one or more genes configured to encode one or more polypeptides or other small molecules;

accessing a machine learning model, the machine learning model being trained with a training dataset that associates the gene clusters to structures of one or more small molecules represented in the data;

wherein the machine learning model is trained by performing operations comprising:

accessing a set of hypothetical structures for natural products including the structure for the natural product;

generating a set of random structures of molecules, the random structures including small molecules;

testing, using mass spectrometry data representing known structures and the set of random structures, the set of hypothetical structures for the natural products including the structure for the natural product;

generating a score for the structure, the score indicating a match between the structure and a known structure represented in the mass spectrometry data;

filtering, based on the score, one or more hypothetical structures from the set of hypothetical structures to generate a filtered set of hypothetical structures that includes the structure for the natural product; and

generating the training dataset for training the machine learning model, the training dataset including the filtered set of hypothetical structures;

applying the machine learning model to the data representing the gene clusters;

identifying, based on applying the machine learning model, one or more monomers associated with at least one gene cluster represented in the data; and

determining a structure for a natural product including the one or more monomers.

2 . The method of claim 1 , further comprising:

determining a class associated with the gene clusters; and

accessing, based on the class, a set of training data that is specific to the class associated with the gene clusters.

3 . The method of claim 1 , further comprising predicting a biological activity of an identified natural product based on the machine learning model that is trained with the training dataset.

4 . The method of claim 3 , further comprising generating, based on predicting the activity, a data library comprising data that associates a gene cluster with a respective biological activity.

5 . The method of claim 1 , further comprising purifying the natural product based on the determined structure.

6 . The method of claim 1 , wherein determining the structure for the natural product including the one or more monomers comprises:

predicting, based on the one or more monomers that are identified, a core molecule that is assembled by combining a group of monomers;

determining one or more particular gene clusters represented in the data that cause a change of a structure of the core molecule; and

identifying an enzyme associated with one or more particular gene clusters that cause the change to the structure of the core molecule.

7 . The method of claim 6 , wherein the core molecule includes a peptide, and wherein the change comprises an addition of an amino acid.

8 . The method of claim 6 , wherein the change comprises an addition of a lipid tail to the core molecule.

9 . The method of claim 6 , wherein the change comprises an addition of a monomer to the core molecule.

10 . The method of claim 1 , wherein the data representing the gene clusters comprises one or more data signatures, wherein data signatures comprise a location of a gene cluster with respect to one or more other gene clusters; and

wherein determining the structure for a natural product including the one or more monomers is based on the data signatures.

11 . A system for searching a database to identify structures of molecular compounds from mass spectrometry data, the system comprising:

at least one processor; and

a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

receiving data representing gene clusters, the gene clusters including one or more genes configured to encode one or more polypeptides or proteins;

accessing a machine learning model, the machine learning model being trained with a training dataset that associates structures of small molecules to one or more of the gene clusters represented in the data;

wherein the machine learning model is trained by performing operations comprising:

accessing a set of hypothetical structures for natural products including the structure for the natural product;

generating a set of random structures of molecules, the random structures including small molecules;

testing, using mass spectrometry data representing known structures and the set of random structures, the set of hypothetical structures for the natural products including the structure for the natural product;

generating a score for the structure, the score indicating a match between the structure and a known structure represented in the mass spectrometry data;

filtering, based on the score, one or more hypothetical structures from the set of hypothetical structures to generate a filtered set of hypothetical structures that includes the structure for the natural product; and

generating the training dataset for training the machine learning model, the training dataset including the filtered set of hypothetical structures;

applying the machine learning model to the data representing the gene clusters;

identifying, based on applying the machine learning model, one or more monomers associated with at least one gene cluster represented in the data; and

determining a structure for a natural product including the one or more monomers.

12 . The system of claim 11 , the operations further comprising:

determining a class associated with the gene clusters; and

accessing, based on the class, a set of training data that is specific to the class associated with the gene clusters.

13 . The system of claim 11 , the operations further comprising predicting a biological activity of an identified natural product based on the machine learning model that is trained with the training dataset.

14 . The system of claim 13 , the operations further comprising generating, based on predicting the activity, a data library comprising data that associates a gene cluster with a respective biological activity.

15 . The system of claim 11 , the operations further comprising purifying the natural product based on the determined structure.

16 . The system of claim 11 , wherein determining the structure for the natural product including the one or more monomers comprises:

predicting, based on the one or more monomers that are identified, a core molecule that is assembled by combining a group of monomers;

determining one or more particular gene clusters represented in the data that cause a change of a structure of the core molecule; and

identifying an enzyme associated with one or more particular gene clusters that cause the change to the structure of the core molecule.

17 . The system of claim 16 , wherein the core molecule includes a peptide, and wherein the change comprises an addition of an amino acid.

18 . The system of claim 16 , wherein the core molecule includes a non-ribosomal peptide.

19 . The system of claim 16 , wherein the core molecule includes a ribosomally synthesized and post-translationally modified peptide.

20 . The system of claim 16 , wherein the core molecule includes a polyketide.

21 . The system of claim 16 , wherein the core molecule includes a saccharide or aminoglycoside.

22 . The system of claim 16 , wherein the change comprises an addition of a lipid tail to the core molecule.

23 . The system of claim 16 , wherein the core molecule includes a hybrid of non-ribosomal peptide and/or ribosomally synthesized and post-translationally modified peptide and/or a polyketide and/or a saccharide or aminoglycoside, and wherein the change comprises an addition of a monomer.

24 . The system of claim 16 , wherein the data representing the gene clusters comprises one or more data signatures, wherein data signatures comprise a location of a gene cluster with respect to one or more other gene clusters; and

wherein determining the structure for a natural product including the one or more monomers is based on the data signatures.

25 . The system of claim 16 , wherein structures of predicted molecules are stored in a computer format that allows for accelerated search against mass spectra.