IP Library › Patent Application 17589623
Patent Application
App. No. 17/589,623

SYSTEMS AND METHODS FOR FEW-SHOT PROTEIN FITNESS PREDICTION WITH GENERATIVE MODELS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/589,623
Abstract

Embodiments are directed to finetuning a pre-trained language model using generative fitness finetuning. The generative fitness finetuning reuses a probability distribution learned during unsupervised training of the pre-trained language model to finetune and assay labeled data. The generative fitness finetuning trains the language model to classify a relative fitness of protein sequence pairs based on the corresponding probability of the protein sequences in the pairs. The generative fitness finetuning identifies protein sequences in the pairs with a higher probability as also having higher fitness. The trained and finetuned language model identifies fitness of a protein sequence.

Claims (56)

1 . A method for training a language model to predict a protein fitness score for a protein sequence, the method comprising:

training the language model on a training dataset of unlabeled protein sequences to learn a probability distribution over protein sequences in the training dataset; and

finetuning the language model by training the probability distribution learned during training as a pairwise classifier that classifies a relative fitness of protein sequence pairs until a loss function is minimized, wherein the finetuning further comprises:

receiving a protein sequence pair of the protein sequence pairs, the protein sequence pair including a first protein sequence and a second protein sequence;

determining, using the probability distribution, a first probability of the first protein sequence and a second probability of the second protein sequence;

selecting a protein sequence from the first protein sequence and the second protein sequence that corresponds to a higher probability between the first probability and the second probability as the protein sequence with a higher fitness; and

determining a value of the loss function based on the selecting the protein sequence.

2 . The method of claim 1 , wherein the language model is a generative language model.

3 . The method of claim 1 , wherein the value of the loss function includes the first probability, the second probability, and a score based on a first fitness value of the first protein sequence and a second fitness value of the second protein sequence.

4 . The method of claim 3 , wherein determining the score further comprising:

determining the first fitness value using the first probability of the first protein sequence;

determining the second fitness value using the second probability of the second protein sequence; and

determining the score based on first fitness value and the second fitness value.

5 . The method of claim 1 , wherein the training dataset further includes fitness labels.

6 . The method of claim 1 , further comprising:

generating a fitness score for a new protein sequence using the trained and finetuned language model.

7 . The method of claim 1 , further comprising:

selecting the protein sequence pairs from a few-shot protein dataset, wherein sequences in the few-shot protein dataset are variants of protein sequences that exist in nature.

8 . A system for training a language model to predict a protein fitness score for a protein sequence, the system comprising:

a memory configured to store the language model and a protein fitness prediction module; and

a processor coupled to the memory and configured to cause the protein fitness prediction module to perform operations, the operations comprising:

training the language model on a training dataset of unlabeled protein sequences to learn a probability distribution over protein sequences in the training dataset; and

finetuning the language model by training the probability distribution learned during training as a pairwise classifier that classifies a relative fitness of protein sequence pairs until a loss function is minimized, wherein the finetuning further comprises:

receiving a protein sequence pair of the protein sequence pairs, the protein sequence pair including a first protein sequence and a second protein sequence;

determining, using the probability distribution, a first probability of the first protein sequence and a second probability of the second protein sequence;

selecting a protein sequence from the first protein sequence and the second protein sequence that corresponds to a higher probability between the first probability and the second probability as the protein sequence with a higher fitness; and

determining a value of the loss function based on the selecting the protein sequence.

9 . The system of claim 8 , wherein the language model is a generative language model.

10 . The system of claim 8 , wherein the value of the loss function includes the first probability, the second probability, and a score based on a first fitness value of the first protein sequence and a second fitness value of the second protein sequence.

11 . The system of claim 10 , wherein the operations for determining the score further comprise:

determining the first fitness value using the first probability of the first protein sequence;

determining the second fitness value using the second probability of the second protein sequence; and

determining the score based on first fitness value and the second fitness value.

12 . The system of claim 8 , wherein the operations further comprise:

selecting the protein sequence pairs from a few-shot protein dataset, wherein sequences in the few-shot protein dataset are variants of protein sequences that exist in nature.

13 . The system of claim 8 , further comprising:

selecting the protein sequence pairs from a few-shot protein dataset, wherein sequences in the few-shot protein dataset are variants of protein sequences that exist in nature.

14 . The system of claim 8 , further comprising:

generating a fitness score for a new protein sequence using the trained and finetuned language model.

15 . A non-transitory computer readable medium having instructions stored thereon, that when executed by a processor causes the processor to perform operations, the operations comprising:

training a language model on a training dataset of unlabeled protein sequences to learn a probability distribution over protein sequences in a training dataset; and

finetuning the language model by training the probability distribution learned during training as a pairwise classifier that classifies a relative fitness of protein sequence pairs, wherein the finetuning further comprises:

receiving a protein sequence pair of the protein sequence pairs, the protein sequence pair including a first protein sequence and a second protein sequence;

determining, using the probability distribution, a first probability of the first protein sequence and a second probability of the second protein sequence; and

selecting a protein sequence from the first protein sequence and the second protein sequence that corresponds to a higher probability between the first probability and the second probability as the protein sequence with a higher fitness.

16 . The non-transitory computer readable medium of claim 15 , further comprising:

finetuning, using a loss function, the language model, wherein the loss function includes the first probability, the second probability, and a score based on a first fitness value of the first protein sequence and a second fitness value of the second protein sequence.

17 . The non-transitory computer readable medium of claim 16 , wherein determining the score further comprises:

determining the first fitness value using the first probability of the first protein sequence;

determining the second fitness value using the second probability of the second protein sequence; and

determining the score based on first fitness value and the second fitness value.

18 . The non-transitory computer readable medium of claim 15 , further comprising:

generating a fitness score for a new protein sequence using the trained and finetuned language model.

19 . The non-transitory computer readable medium of claim 15 , wherein the training dataset further includes fitness labels.

20 . The non-transitory computer readable medium of claim 15 , further comprising:

selecting the protein sequence pairs from a few-shot protein dataset, wherein sequences in the few-shot protein dataset are variants of protein sequences that exist in nature.

Assignments (2)
CHANGE OF NAME Recorded Aug 4, 2026
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 076118/0548 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2022
From: KRAUSE, BEN; MADANI, ALI
To: SALESFORCE.COM, INC.
Reel/Frame 060720/0199 →