IP Library Patent Application 18654581
Patent Application
App. No. 18/654,581

IDENTIFYING BIOSYNTHETIC GENE CLUSTERS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/654,581
Abstract

A biosynthetic gene cluster (BGC) prediction system identifies candidate BGCs within genomes using an iteratively trained machine-learned model. The system identifies, in a genome sequence, a set of domains, each identified domain corresponding to a set of domain identifiers. The set of domain identifiers corresponds to a set of vectors. The iteratively trained model is applied to the set of vectors to produce a BGC class score for each domain. The system selects candidate BGCs by averaging GBC class scores across genes within a domain and comparing the average BGC class scores to a threshold. The system predicts a molecular activity of biosynthetic products derived from the selected BGCs, and provides for display, on a user interface, the candidate BGCs and predicted molecular activity.

Claims (63)

1 . A method comprising:

identifying, in a genome sequence, a set of domains, each identified domain corresponding to a set of domain identifiers, the set of domain identifiers corresponding to a set of vectors;

applying an interactively trained model to the set of vectors to produce a biosynthetic gene cluster (BGC) class score for each domain, wherein the model was trained by:

identifying a set of positive vectors representing known BGCs;

synthesizing a set of negative vectors unlikely to represent BGCs;

applying the model to the positive and negative sets of vectors to generate predictions of whether each vector is a positive or negative vector; and

updating weights of the model based on the predictions;

selecting candidate BGCs by averaging BGC class scores across genes within a domain and comparing the average BGC class scores to a threshold;

predicting a molecular activity of biosynthetic products derived from the selected BGCs; and

providing for display, on a user interface, the candidate BGCs and predicted molecular activity.

2 . The method of claim 1 , further comprising:

processing candidate BGCs, wherein processing includes merging and filtering candidate BGCs based on at least one of: a presence of known BGCs, a cluster length, or a distance between candidate BGCs.

3 . The method of claim 1 , further comprising:

merging consecutive candidate BGC genes that are adjacent in the genome sequence.

4 . The method of claim 1 , wherein the model is a bi-directional long short-term memory (LSTM) block.

5 . The method of claim 1 , wherein the domain identifiers are maintained in genomic order.

6 . The method of claim 1 , wherein each vector in the set of vectors comprises one hundred elements, each element a real number, each element representing a property of the domain based on its genomic context.

7 . The method of claim 1 , further comprising:

predicting, for each candidate BGC, with a classifier, a secondary metabolite class based on a biosynthetic product and molecular activity of the candidate BGC.

8 . The method of claim 7 , wherein the classifier is a random forest classifier.

9 . The method of claim 1 , wherein the set of negative vectors are synthesized by:

retrieving a genome sequence with known BGCs;

modifying the genome sequence by replacing a portion of the genes within the known BGCs with random genes of similar length;

generating a set of identifiers for each domain in the modified genome sequence; and

applying a shallow neural network block to each domain in the modified genome sequence to produce a negative set of vectors.

10 . The method of claim 1 , wherein applying the model to the set of vectors further comprises applying a sigmoid activation function.

11 . A non-transitory computer-readable storage medium containing computer program code comprising instructions that, when executed by a processor, causes the processor to:

identify, in a genome sequence, a set of domains, each identified domain corresponding to a set of domain identifiers, the set of domain identifiers corresponding to a set of vectors;

apply an interactively trained model to the set of vectors to produce a biosynthetic gene cluster (BGC) class score for each domain, wherein the model was trained by:

identifying a set of positive vectors representing known BGCs;

synthesizing a set of negative vectors unlikely to represent BGCs;

applying the model to the positive and negative sets of vectors to generate predictions of whether each vector is a positive or negative vector; and

updating weights of the model based on the predictions;

select candidate BGCs by averaging BGC class scores across genes within a domain and comparing the average BGC class scores to a threshold;

predict a molecular activity of biosynthetic products derived from the selected BGCs; and

provide for display, on a user interface, the candidate BGCs and predicted molecular activity.

12 . The non-transitory computer-readable storage medium of claim 11 , wherein the instructions, when executed by the processor, further cause the processor to:

process candidate BGCs by merging and filtering candidate BGCs based on at least one of: a presence of known BGCs, a cluster length, or a distance between candidate BGCs.

13 . The non-transitory computer-readable storage medium of claim 11 , wherein the instructions, when executed by the processor, cause the processor to:

merge consecutive candidate BGC genes.

14 . The non-transitory computer-readable storage medium of claim 11 , wherein the model is a bi-directional long short-term memory (LSTM) block.

15 . The non-transitory computer-readable storage medium of claim 11 , wherein the domain identifiers are maintained in genomic order.

16 . The non-transitory computer-readable storage medium of claim 11 , wherein each vector in the set of vectors comprises one hundred elements, each element a real number, each element representing a property of the domain based on its genomic context.

17 . The non-transitory computer-readable storage medium of claim 11 , wherein the instructions, when executed by the processor, further cause the processor to:

predict, for each candidate BGC, with a classifier, a secondary metabolite class based on a biosynthetic product and molecular activity of the candidate BGC.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the classifier is a random forest classifier.

19 . The non-transitory computer-readable storage medium of claim 17 , wherein the set of negative vectors are synthesized by:

retrieving a genome sequence with known BGCs;

modifying the genome sequence by replacing a portion of the genes within the known BGCs with random genes of similar length;

generating a set of identifiers for each domain in the modified genome sequence; and

applying a shallow neural network block to each domain in the modified genome sequence to produce a negative set of vectors.

20 . A computer system, comprising:

one or more processors; and

a non-transitory computer-readable storage medium storing instructions that, when executed by the one or more processor, causes the one or more processor to:

identify, in a genome sequence, a set of domains, each identified domain corresponding to a set of domain identifiers, the set of domain identifiers corresponding to a set of vectors;

apply an interactively trained model to the set of vectors to produce a biosynthetic gene cluster (BGC) class score for each domain, wherein the model was trained by:

identifying a set of positive vectors representing known BGCs;

synthesizing a set of negative vectors unlikely to represent BGCs;

applying the model to the positive and negative sets of vectors to generate predictions of whether each vector is a positive or negative vector; and

updating weights of the model based on the predictions;

select candidate BGCs by averaging BGC class scores across genes within a domain and comparing the average BGC class scores to a threshold;

predict a molecular activity of biosynthetic products derived from the selected BGCs; and

provide for display, on a user interface, the candidate BGCs and predicted molecular activity.

Assignments (3)
MERGER Recorded Jun 18, 2024
From: MERCK SHARP & DOHME CORP.
To: MERCK SHARP & DOHME LLC
Reel/Frame 067753/0791 →
CHANGE OF NAME Recorded Jun 18, 2024
From: MSD IT GLOBAL INNOVATION CENTER S.R.O.
To: MSD CZECH REPUBLIC S.R.O.
Reel/Frame 067774/0843 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: HANNIGAN, GEOFFREY D.; PRIHODA, DAVID; SOUKUP, JINDRICH; WOELK, CHRISTOPHER HARRON; BITTON, DANNY A.
To: MERCK SHARP & DOHME CORP.; MSD IT GLOBAL INNOVATION CENTER S. R. O.
Reel/Frame 067737/0794 →