IP Library Granted Patent US 12,651,642
Granted Patent B2
US 12,651,642 · App. 19/193,962 · Granted Jun 9, 2026

Early fusion of natural and protein language models for generative AI-based protein and drug design

Inventor: Stephen Gbejule Odaibo (Sugar Land, TX)
Assignee: Deep EigenMatics, Inc.
G16B15/30G06N3/04G16B40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,642
App. No.
19/193,962
Granted
Jun 9, 2026
Kind
B2
Abstract

Methods and apparatus using a mixture of representation modalities including natural language, protein sequence, protein structure, property-vector, and small molecule drug representations to jointly train a neural network which accepts mixed modality queries as input and produces mixed modality output responses including representations of proteins for synthesis and of small molecule drugs for manufacture. In one embodiment of the invention, multicapitate transformers wherein each decoder head has a distinct loss function and represents a distinct modality, are used. Modality-specific embeddings are implemented for the mixed modality input query, and an autoregressive process yields the output protein for synthesis or small molecule drug for manufacture.

Claims (69)

1 . A method; comprising:

a) receiving, at a processor, a dataset consisting of a plurality of representations of features of a plurality of proteins:

i) wherein respective modalities of the representations include:

(1) natural language representation modality; and

(2) sequence representation modality; and

ii) wherein the represented features include one or more of sequence, structure, function, interactions, interactors, binding partners, attributes, and properties;

b) using the dataset to train a neural network:

i) wherein the neural network is configured to accept as input data, a query consisting of one or more of the respective modalities, and to yield as output data, a response to the query, wherein the response comprises of one or more of the respective modalities;

ii) wherein the neural network has multiple heads, each with its own loss function; and

iii) wherein the neural network heads include one head for natural language representation output and a different head for protein sequence representation output; and

c) using the trained neural network to generate a representation of an output protein, in response to a query specifying conditions on the protein; and

d) synthesizing the protein.

2 . The method of claim 1 , further comprising synthesizing the protein.

3 . The method of claim 1 , wherein in addition to the natural language head and the protein sequence head, the neural network also has a different head for a protein structure representation modality.

4 . The method of claim 3 , wherein the neural network is a transformer.

5 . The method of claim 4 , wherein the transformer is autoregressive.

6 . The method of claim 5 , wherein for each respective head of the transformer, final output is a probability distribution over a set of possible values at that head.

7 . The method of claim 6 , wherein a query specifies a target receptor and requests a peptide ligand of the receptor; and wherein a representation of a peptide ligand of the specified target receptor is generated by the trained neural network.

8 . The method of claim 7 , further comprising:

a) generating, using a given query a representation of a peptide ligand by randomly sampling an output probability distribution of an active head, at each iteration of an autoregression process;

b) storing the resultant peptide ligand representation in memory;

c) using the given query, repeating the random-sampling-based generation process a plurality of times, each yielding a candidate peptide ligand;

d) assessing interaction, efficacy, and properties of each candidate ligand with the target receptor;

e) selecting, based on the assessment, a candidate ligand; and

f) synthesizing the selected ligand.

9 . The method of claim 8 , wherein for each of the following modalities: natural language, protein sequence, and protein structure, an input embedding used for input data of each respective modality is distinct from an input embedding used for input data of any of the other modalities.

10 . A method; comprising:

a) receiving, at a processor, a dataset consisting of a plurality of representations of features of a plurality of proteins:

i) wherein respective modalities of the representations include:

(1) natural language representation modality,

(2) sequence representation modality,

(3) structure representation modality, and

(4) small molecule drug representation modality; and

ii) wherein the represented features include one or more of sequence, structure, function, interactions, interactors, binding partners, attributes, and properties;

b) using the dataset to train a neural network:

i) wherein the neural network is configured to accept as input data, a query consisting of one or more of the respective modalities, and to yield as output data, a response to the query, wherein the response comprises of one or more of the respective modalities,

ii) wherein the neural network has multiple heads, each with its own loss function, and

iii) wherein the neural network heads include one head for natural language representation output, a different head for protein sequence representation output, a different head for protein structure representation output, and a different head for small molecule drug representation output;

c) using the trained neural network to generate a representation of an output small molecule drug, in response to a query specifying conditions on the small molecule drug; and

d) manufacturing the small molecule drug.

11 . The method of claim 10 , further comprising manufacturing the small molecule drug.

12 . The method of claim 11 , wherein the neural network is a transformer.

13 . The method of claim 12 , wherein the transformer is autoregressive.

14 . The method of claim 13 , wherein for each respective head of the transformer, final output is a probability distribution over a set of possible values at that head.

15 . The method of claim 14 , wherein a query specifies a target receptor and requests a small molecule drug ligand of the receptor; and wherein a representation of a small molecule drug ligand of the specified target receptor is generated by the trained neural network.

16 . The method of claim 15 , further comprising:

a) generating, using a given query a representation of a small molecule drug ligand by randomly sampling an output probability distribution of an active head, at each iteration of an autoregression process;

b) storing the resultant small molecule drug ligand representation in memory;

c) using the given query, repeating the random-sampling-based generation process a plurality of times, each yielding a candidate small molecule drug ligand;

d) assessing interaction, efficacy, and properties of each candidate ligand with the target receptor;

e) selecting, based on the assessment, a candidate ligand; and

f) synthesizing the selected ligand.

17 . The method of claim 16 , wherein for each of the following modalities: natural language, protein sequence, protein structure, and small molecule drug, an input embedding used for input data of each respective modality is distinct from an input embedding used for input data of any of the other modalities.

18 . An apparatus, comprising a processor and an associated memory; wherein the memory stores instructions that when executed by the processor, are configured to cause the processor to:

a) receive a dataset consisting of a plurality of representations of features of a plurality of proteins:

i) wherein respective modalities of the representations include:

(1) natural language representation modality,

(2) sequence representation modality,

(3) structure representation modality, and

(4) small molecule drug representation modality; and

ii) wherein the represented features include one or more of sequence, structure, function, interactions, interactors, binding partners, attributes, and properties;

b) use the dataset to train a neural network:

i) wherein the neural network is configured to accept as input data, a query consisting of one or more of the respective modalities, and to yield as output data, a response to the query, wherein the response comprises of one or more of the respective modalities,

ii) wherein the neural network has a single heads, and

iii) wherein final output is a probability distribution over possible values of each modality including over auxiliary tokens;

c) use the trained neural network to generate a representation of an output ligand, in response to a query specifying conditions on the ligand; and

d) synthesize the ligand.

19 . The apparatus of claim 18 , wherein the neural network is an autoregressive transformer.

20 . The apparatus of claim 19 , wherein the input query requests a peptide ligand of a specified target receptor; and wherein the synthesized ligand is a ligand of the specified target receptor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2026
From: DEEP EIGENMATICS LLC
To: DEEP EIGENMATICS, INC.
Reel/Frame 073363/0835 →
Continuity (1)
Related Publication 20250279161A1 · Sep 4, 2025
References Cited (5)
CN 117912545A · 2024 [cited by examiner]
Wei, Jason, et al. “Chain-of-thought prompting elicits reasoning in large language models.” Advances in neural information processing systems 35 (2022): 24824-24837. [cited by examiner]
English machine translation of Zhou CN-117912545-A. (Year: 2025). [cited by examiner]
Gao, Wenhao, and Connor W. Coley. “The synthesizability of molecules proposed by generative models.” Journal of chemical information and modeling 60.12 (2020): 5714-5723. [cited by examiner]
Merrifield, Robert B. “Solid phase peptide synthesis. I. The synthesis of a tetrapeptide.” Journal of the American Chemical Society 85.14 (1963): 2149-2154. [cited by examiner]