IP Library Patent Application 17221405
Patent Application
App. No. 17/221,405

DECODING APPROACHES FOR PROTEIN IDENTIFICATION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/221,405
Abstract

Methods and systems are provided for accurate and efficient identification and quantification of proteins. In an aspect, disclosed herein is a method for identifying a protein in a sample of unknown proteins, comprising receiving information of a plurality of empirical measurements performed on the unknown proteins; comparing the information of empirical measurements against a database comprising a plurality of protein sequences, each protein sequence corresponding to a candidate protein among a plurality of candidate proteins; and for each of one or more of the plurality of candidate proteins, generating a probability that the candidate protein generates the information of empirical measurements, a probability that the plurality of empirical measurements is not observed given that the candidate protein is present in the sample, or a probability that the candidate protein is present in the sample; based on the comparison of the information of empirical measurements against the database.

Claims (106)

1 . A method for training a probabilistic computational model, comprising:

(a) obtaining, by a computer, the probabilistic computational model, which probabilistic computational model comprises a set of binding probabilities corresponding to a set of affinity reagents configured to bind to a set of amino acids of a protein;

(b) contacting a training set of unknown proteins with the set of affinity reagents;

(c) obtaining a plurality of empirical measurements of the training set of unknown proteins, which plurality of empirical measurements comprises binding measurements of the set of affinity reagents to at least one of the training set of unknown proteins;

(d) determining, by the computer, an updated binding probability for at least one affinity reagent of the set of affinity reagents, based at least in part on empirical measurements for the at least one affinity reagent; and

(e) repeating at least one iteration of (b), (c), and (d) to iteratively optimize the set of binding probabilities.

2 . The method of claim 1 , wherein (e) comprises using additional training sets of unknown proteins or sets of affinity reagents.

3 . The method of claim 1 , further comprising:

(f) identifying a test protein in a test sample of unknown proteins using the trained probabilistic computational model.

4 . The method of claim 3 , wherein (f) comprises:

(i) assaying the test sample to obtain a second plurality of empirical measurements comprising binding measurements of a test set of affinity reagents,

(ii) comparing at least a portion of the second plurality of empirical measurements against a computer database comprising protein sequences of a plurality of candidate proteins,

(iii) based in the comparing in (ii), for a candidate protein in the plurality of candidate proteins, determining a probability that the candidate protein is present in the test sample, at least in part by applying the trained probabilistic computational model to the plurality of empirical measurements, and

(iv) identifying the test protein in the test sample of unknown proteins based on the probability determined in (iii).

5 . The method of claim 4 , wherein (iii) comprises calculating, for each of the plurality of empirical measurements, a value given by P(measurement outcome protein) which equals a probability that a measurement outcome comprising the empirical measurement is observed given that the candidate protein is present in the sample.

6 . The method of claim 4 , wherein (iii) comprises:

for each of a set of candidate proteins, calculating a probability that a measurement outcome set comprising a plurality of N empirical measurements is observed given that the candidate protein is present in the test sample, based on a product of a plurality of N probabilities that each of the plurality of N empirical measurements is observed given that the given candidate protein is present in the test sample, thereby generating a set of probabilities that the measurement outcome set is observed for each of the set of candidate proteins; and

generating the probability that the candidate protein is present in the test sample using the expression:

Pr

(

outcome

set

|

candidate

protein

)

i

=

1

P

Pr

(

outcome

set

|

protein

i

)

,

wherein Σ i=1 P Pr(outcome set |protein i ) is a sum of the set of probabilities.

7 . The method of claim 4 , further comprising generating a confidence level that the candidate protein matches one of the unknown proteins in the test sample.

8 . The method of claim 4 , further comprising generating a sensitivity of identifying the test protein with a pre-determined threshold.

9 . The method of claim 4 , wherein the test protein in the test sample is truncated or degraded, or does not originate from a protein terminus.

10 . The method of claim 4 , further comprising calculating a probability that the measurement outcome comprising the empirical measurement is observed given that the candidate protein is present in the sample.

11 . The method of claim 10 , further comprising using the expression:

P

(

measurement

outcome

|

candidate

protein

)

=

1

σ

2

π

exp

(

-

u

2

2

)

,

wherein σ=|CV*expected outcome value |,

wherein CV is a coefficient of variation of the empirical measurement,

wherein u=(measured outcome value—expected outcome value)/σ,

wherein the measured outcome is at least one of a measured length, a measured hydrophobicity, and a measured isoelectric point of at least one of the unknown proteins in the test sample, and wherein the expected outcome value is at least one of an expected length, an expected hydrophobicity, and an expected isoelectric point of the at least one of the unknown proteins in the test sample.

12 . The method of claim 11 , wherein the expected length of the at least one of the unknown proteins is a length of a protein sequence of the at least one of the unknown proteins in the test sample.

13 . The method of claim 11 , wherein the expected hydrophobicity of the at least one of the unknown proteins is a grand average of hydropathy (gravy) score determined based on a protein sequence of the at least one of the unknown proteins in the test sample.

14 . The method of claim 11 , wherein the expected isoelectric point of the at least one of the unknown proteins is determined based on a protein sequence of the at least one of the unknown proteins in the test sample.

15 . The method of claim 4 , wherein the computer database comprises protein sequences corresponding to at least 10 different candidate proteins.

16 . The method of claim 4 , further comprising, for each of the plurality of candidate proteins, generating a probability that the candidate protein is present in the sample; and identifying a given candidate protein as matching the test protein when the given candidate protein has a largest value of P(measurement outcomelprotein) among the plurality of candidate proteins.

17 . The method of claim 1 , wherein the plurality of empirical measurements comprises at least one of length, hydrophobicity, and isoelectric point of one or more of the unknown proteins in the training set.

18 . The method of claim 17 , wherein assaying the sample to obtain the plurality of empirical measurements comprises fractionating one or more of the unknown proteins based on at least one of length, hydrophobicity, and isoelectric point to produce fractionated proteins; and obtaining the plurality of empirical measurements from the fractionated proteins.

19 . The method of claim 18 , wherein assaying the sample comprises fractionating at least one of the unknown proteins based on at least one of: the length by gel filtration or size exclusion chromatography, the hydrophobicity by hydrophobic interaction chromatography, and the isoelectric point by ion exchange chromatography.

20 . The method of claim 1 , wherein the plurality of empirical measurements comprises binding of the set of affinity reagents or non-specific binding of the set of affinity reagents.

21 . The method of claim 20 , wherein the pre-determined threshold is less than a 1% false identification rate.

22 . The method of claim 1 , wherein the plurality of empirical measurements comprises measurements performed on mixtures of antibodies.

23 . The method of claim 1 , wherein the plurality of empirical measurements comprises measurements performed on the training set of unknown proteins in presence of single amino acid variants (SAVs) caused by non-synonymous single-nucleotide polymorphisms (SNPs).

24 . The method of claim 1 , wherein (a) comprises initializing the set of binding probabilities with an initial binding probability.

25 . The method of claim 1 , wherein (d) comprises determining the updated binding probability for the at least one affinity reagent using a proportion of unknown proteins in the training set containing a binding site recognized by the at least one affinity reagent that are bound to the at least one affinity reagent.

26 . The method of claim 1 , wherein (e) comprises performing an expectation maximization algorithm on the plurality of empirical measurements.

27 . The method of claim 1 , wherein the set of affinity reagents comprises at least 10 different affinity reagents.

28 . The method of claim 1 , wherein (e) comprises performing at least 10 iterations of (b), (c), and (d).

29 . The method of claim 1 , wherein (e) comprises performing a number of iterations of (b), (c), and (d) sufficient to achieve a sensitivity of protein identification of at least about 90%.

Assignments (2)
CHANGE OF NAME Recorded Apr 13, 2023
From: NAUTILUS BIOTECHNOLOGY, INC.
To: NAUTILUS SUBSIDIARY, INC.
Reel/Frame 063325/0533 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2021
From: PATEL, SUJAL M.; MALLICK, PARAG; EGERTSON, JARRETT D.
To: NAUTILUS BIOTECHNOLOGY, INC.
Reel/Frame 057028/0450 →