IP Library Granted Patent US 11,875,880
Granted Patent B2
US 11,875,880 · App. 16/679,155 · Granted Jan 16, 2024

Systems and methods for calculating protein confidence values

Inventor: Ignat V. Shilov (Palo Alto, CA)
Assignee: DH Technologies Development Pte Ltd.
G16C20/20G16B15/00G16B20/00G16B40/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,875,880
App. No.
16/679,155
Granted
Jan 16, 2024
Kind
B2
Abstract

A method is disclosed for calculating and recalculating protein confidence values by recalculating peptide confidence values in proteomic analysis. A plurality of scans are performed of a sample that is proteolytically digested into surrogate peptide analytes. A plurality of spectra are obtained, and a plurality of peptides are identified. A protein database is searched for proteins matching peptides from the plurality of peptides. Peptide confidence values are determined for the set of peptides. A protein confidence value is calculated for each protein in the set of proteins. A protein from the set of proteins with a largest protein confidence value is selected, the largest protein confidence value for the protein is saved, the protein from the set of proteins is removed, and peptides corresponding to the protein is removed from the set of peptides. The protein confidence value is recalculated for each protein in the set of proteins.

Claims (85)

1. A method for calculating and recalculating protein confidence values by recalculating peptide confidence values in proteomic analysis in order to distinguish proteins found in a sample from random or false positive results, comprising:

(a) performing a plurality of scans of a sample that is proteolytically digested into surrogate peptide analytes producing a plurality of spectra using one or more mass spectrometers;

(b) obtaining the plurality of spectra from the mass spectrometer using a processor;

(c) identifying a plurality of peptides from the plurality of spectra using the processor;

(d) searching a protein database for proteins matching peptides from the plurality of peptides producing a set of proteins and a corresponding set of peptides using the processor;

(e) determining peptide confidence values for the set of peptides using the processor, wherein a peptide confidence, Ci, of an i th peptide of the set of peptides is a probability that the ith peptide is identified from the plurality of spectra;

(f) calculating a protein confidence value for each protein in the set of proteins based on one or more peptide confidence values of one or more corresponding peptides from the set of peptides using the processor, wherein a protein confidence value P p is a probability calculated according to

P p =1−Π(1 −Ci ),

where Π(1−Ci) is the product of one or more peptide confidence values;

(g) selecting a protein from the set of proteins with a largest protein confidence value, saving the largest protein confidence value for the protein, removing the protein from the set of proteins, and removing one or more peptides corresponding to the protein from the set of peptides using the processor;

(h) recalculating the protein confidence value, P p , for each protein in the set of proteins based on one or more peptide posterior probability values of one or more corresponding peptides from the set of peptides according to

P p =1−Π(1−P(+|B) i ) using the processor,

wherein a posterior probability of the i th peptide, P(+|B), of the set of peptides is calculated using

P

(

+

|

B

)

=

P

(

B

|

+

)

·

P

(

+

)

P

(

B

)

where P(B|+) is the peptide confidence value of the i th peptide, C i , P(B) is the marginal probability of observing the peptide with a given confidence, and P(+) is the prior probability of randomly selecting a true positive and P(+) is calculated from all confidence values of peptides currently in the set of peptides and all confidence values of peptides removed from the set of peptides to account for the effect of removing the one or more peptides corresponding to the removed protein from the set of peptides, and

(i) repeating steps (f)-(g) until all proteins are removed from the set of proteins or until a protein confidence value of the selected protein with a largest protein confidence value is below a threshold of interest and identifies proteins found in the sample as the proteins with the saved largest protein confidence values using the processor.

2. The method of claim 1 , wherein determining peptide confidence values for the set of peptides comprises using a target-decoy method and the protein database.

3. The method of claim 1 , wherein determining peptide confidence values comprises using a heuristic.

4. A computer program product, comprising a tangible computer-readable storage medium whose contents include a program with instructions being executed on a processor so as to perform a method for calculating and recalculating protein confidence values by recalculating peptide confidence values in proteomic analysis in order to distinguish proteins found in a sample from random or false positive results, the method comprising:

(a) providing a system, wherein the system comprises distinct software modules, and wherein the distinct software modules comprise a measurement module and an analysis module;

(b) obtaining a plurality of spectra from one or more mass spectrometers that perform a plurality of scans of a sample that is proteolytically digested into surrogate peptide analytes producing a plurality of spectra using the measurement module;

(c) identifying a plurality of peptides from the plurality of spectra using the analysis module;

(d) searching a protein database for proteins matching peptides from the plurality of peptides producing a set of proteins and a corresponding set of peptides using the analysis module;

(e) determining peptide confidence values for the set of peptides using the analysis module, wherein a peptide confidence, Ci, of an i th peptide of the set of peptides is a probability that the ith peptide is identified from the plurality of spectra;

(f) calculating a protein confidence value for each protein in the set of proteins based on one or more peptide confidence values of one or more corresponding peptides from the set of peptides using the analysis module, wherein a protein confidence value P p is a probability calculated according to

P p =1−Π(1 −Ci ),

where Π(1−Ci) is the product of one or more peptide confidence values;

(g) selecting a protein from the set of proteins with a largest protein confidence value, saving the largest protein confidence value for the protein, removing the protein from the set of proteins, and removing one or more peptides corresponding to the protein from the set of peptides using the analysis module;

(h) recalculating the protein confidence value, P p , for each protein in the set of proteins based on one or more peptide posterior probability values of one or more corresponding peptides from the set of peptides according to

P p =1−Π(1−P(+|B) i ) using the analysis module, wherein a posterior probability of the i th peptide, P(+|B), of the set of peptides is calculated using

P

(

+

|

B

)

=

P

(

B

|

+

)

·

P

(

+

)

P

(

B

)

where P(B|+) is the peptide confidence value of the i th peptide, C i , P(B) is the marginal probability of observing the peptide with a given confidence, and P(+) is the prior probability of randomly selecting a true positive and P(+) is calculated from all confidence values of peptides currently in the set of peptides and all confidence values of peptides removed from the set of peptides to account for the effect of removing the one or more peptides corresponding to the removed protein from the set of peptides, and

(i) repeating steps (f)-(g) until all proteins are removed from the set of proteins or until a protein confidence value of the selected protein with a largest protein confidence value is below a threshold of interest and identifies proteins found in the sample as the proteins with the saved largest protein confidence values using the analysis module.

5. The computer program product of claim 4 , wherein the analysis module determines peptide confidence values for the set of peptides using a target-decoy method and the protein database.

6. The computer program product of claim 4 , wherein the analysis module determines peptide confidence values using a heuristic.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2019
From: SHILOV, IGNAT V.
To: DH TECHNOLOGIES DEVELOPMENT PTE. LTD.
Reel/Frame 050963/0274 →
CHANGE OF ADDRESS Recorded Nov 9, 2019
From: DH TECHNOLOGIES DEVELOPMENT PTE. LTD.
To: DH TECHNOLOGIES DEVELOPMENT PTE. LTD.
Reel/Frame 050977/0433 →
Continuity (3)
Division 13697707
Provisional Application 61334763 · May 14, 2010
Related Publication 20200075132A1 · Mar 5, 2020