IP Library Granted Patent US 12,334,189
Granted Patent B2
US 12,334,189 · App. 18/943,185 · Granted Jun 17, 2025

System and method for processing experimental data

Inventors: Maximilien Burq (San Diego, CA); Jure Zbontar (San Diego, CA); Peter Cimermancic (San Diego, CA)
Assignee: Tesorai, Inc.
G16B15/00G06N3/0464G16B40/10G16B40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,189
App. No.
18/943,185
Granted
Jun 17, 2025
Kind
B2
Abstract

The method for processing experimental data can include: determining experimental data (e.g., mass spectrometry spectra) and processing the experimental data. In variants, processing the experimental data can include: identifying one or more molecules, comparing experimental samples, determining a quantification, evaluating a quality of the experimental data, and/or otherwise processing the experimental data. The method can optionally include determining supplemental information, determining a set of candidate molecules, training a model, and/any other suitable steps.

Claims (44)

1. A method for molecule identification, comprising:

training a first encoder using a known match between a training sample and a set of training spectra, the training sample comprising a set of training molecules, wherein training the first encoder comprises:

using the first encoder, determining a set of molecule predictions based on the set of training spectra, wherein a number of true-positive molecule predictions and a number of true-negative molecule predictions are determined based on a comparison between the set of molecule predictions and the set of training molecules;

determining a loss for the set of molecule predictions, the loss comprising an accuracy metric determined based on the number of true-positive molecule predictions and the number of true-negative molecule predictions; and

training the first encoder based on the loss;

determining a mass spectrometry spectrum for a molecule;

determining an embedding for the mass spectrometry spectrum, using the first encoder;

for each candidate molecule in a set of candidate molecules:

determining an embedding for the candidate molecule based on a sequence for the candidate molecule, using a second encoder; and

using a scoring model, determining a score for the candidate molecule based on the embedding for the mass spectrometry spectrum and the embedding for the candidate molecule; and

selecting a candidate molecule from the set of candidate molecules based on the scores.

2. The method of claim 1 , wherein determining the score for the candidate molecule comprises: determining a similarity metric based on the embedding for the mass spectrometry spectrum and the embedding for the candidate molecule; and determining the score based on the similarity metric.

3. The method of claim 2 , wherein the similarity metric comprises a distance between the embedding for the mass spectrometry spectrum and the embedding for the candidate molecule.

4. The method of claim 1 , wherein the first encoder comprises a CNN.

5. The method of claim 4 , wherein the first encoder further comprises a transformer.

6. The method of claim 1 , wherein the second encoder comprises a combination of an RNN and a transformer.

7. The method of claim 1 , wherein the mass spectrometry spectrum comprises an experimental dataset containing intensities across at least two dimensions, wherein a first dimension of the experimental dataset is mass to charge ratio (m/z) and a second dimension of the experimental dataset is retention time.

8. The method of claim 7 , wherein the experimental dataset contains intensities across a third dimension, wherein the third dimension is ion mobility.

9. The method of claim 1 , wherein the first encoder is configured to accommodate receiving as input experimental datasets with a variable number of dimensions.

10. The method of claim 1 , further comprising determining a location of a post-translational modification within the candidate molecule based on the embedding for the mass spectrometry spectrum and the embedding for the candidate molecule.

11. The method of claim 1 , wherein individual matches between each training spectrum in the set of training spectra and a corresponding training molecule in the set of training molecules is not known, wherein the set of molecule predictions is determined using the first encoder, the second encoder, and the scoring model.

12. The method of claim 11 , wherein the loss for the set of molecule predictions is further determined based on a second set of molecule predictions determined using at least two first-generation search engines.

13. The method of claim 1 , wherein the first encoder, the second encoder, and the scoring model are not trained using decoy molecules.

14. The method of claim 1 , wherein the embedding for the mass spectrometry spectrum is determined based on the mass spectrometry spectrum and supplemental experimental information, wherein the supplemental experimental information comprises at least one of: instrument type or fragmentation method.

15. A method for molecule identification, comprising:

training an encoder using a known match between a training sample and a set of training spectra, the training sample comprising a set of training molecules, wherein training the encoder comprises:

using the encoder, determining a set of molecule predictions based on the set of training spectra, wherein a number of true-positive molecule predictions and a number of true-negative molecule predictions are determined based on a comparison between the set of molecule predictions and the set of training molecules;

determining a loss for the set of molecule predictions, the loss comprising an accuracy metric determined based on the number of true-positive molecule predictions and the number of true-negative molecule predictions; and

training the encoder based on the loss;

determining a first set of mass spectrometry spectra for a first set of molecules in a first sample;

determining a second set of mass spectrometry spectra for a second set of molecules in a second sample;

determining an embedding for each mass spectrometry spectrum in the first set of mass spectrometry spectra, using the encoder;

determining an embedding for each mass spectrometry spectrum in the second set of mass spectrometry spectra, using the encoder;

clustering the embeddings for each mass spectrometry spectrum in the first and second sets of mass spectrometry spectra into a set of clusters;

identifying an embedding of interest based on the set of clusters; and

identifying a molecule based on the embedding of interest, wherein the molecule is present in the second sample and is not present in the first sample.

16. The method of claim 15 , wherein the encoder encodes each mass spectrometry spectrum into an embedding space, wherein the embeddings are clustered within the embedding space.

17. The method of claim 15 , wherein each cluster in the set of clusters has a cluster purity of at least 50%.

18. The method of claim 15 , wherein identifying the molecule comprises:

for each candidate molecule in a set of candidate molecules:

determining an embedding for the candidate molecule based on a sequence for the candidate molecule, using a second encoder; and

using a scoring model, determining a score for the candidate molecule based on the embedding of interest and the embedding for the candidate molecule; and

selecting a candidate molecule from the set of candidate molecules based on the scores.

19. The method of claim 15 , wherein the first sample and the second sample comprise samples from a multiple-arm experiment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2025
From: BURQ, MAXIMILIEN; ZBONTAR, JURE; CIMERMANCIC, PETER
To: TESORAI, INC.
Reel/Frame 070104/0330 →
Continuity (3)
Provisional Application 63682215 · Aug 12, 2024
Provisional Application 63597505 · Nov 9, 2023
Related Publication 20250157568A1 · May 15, 2025
References Cited (51)
US 10705100B1 · Hansen · 2020 [cited by examiner]
US 11629384B2 · Batenchuk et al. · 2023 [cited by applicant]
US 11862298B1 · Palaniappan et al. · 2024 [cited by applicant]
US 20190302054A1 · Bleiholder · 2019 [cited by examiner]
US 20200392178A1 · Manica · 2020 [cited by examiner]
US 20210215685A1 · Xia · 2021 [cited by examiner]
US 20210233640A1 · Banchereau · 2021 [cited by examiner]
US 20220208540A1 · Behsaz · 2022 [cited by examiner]
US 20230080329A1 · Liu · 2023 [cited by examiner]
US 20230245439A1 · Batenchuk et al. · 2023 [cited by applicant]
US 20240207626A1 · Tiwary · 2024 [cited by examiner]
CN 118155746A · 2024 [cited by applicant]
Modeling Lower-Order Statistics to Enable Decoy-Free FDR Estimation in Proteomics Madej et al Published: Mar. 24, 2023 (Year: 2023). [cited by examiner]
Some recommendations for multi-arm multi-stage trials Wason et al Statistical Methods in Medical Research 2016, vol. 25(2) 716-727 (Year: 2016). [cited by examiner]
Litsa, et al., “An End-to-End Deep Learning Framework for Translating Mass Spectra to De-Novo Molecules”, Communications Chemistry, Jun. 23, 2023, vol. 6, A1ticle 132, [retrieved online Jan. 4, 2025]. <See Entire Docume… [cited by applicant]
Zheng, et al., “A BERT-Based Pretraining Model for Extracting Molecular Structural Information from a SMILES Sequence”, ournal of Cheminformatics, Jun. 19, 2024, vol. 16, Article 71, [retrieved online Feb. 26, 2025]. Re… [cited by applicant]
Zhou, et al., “S-MolSearch: 3D Semi-supervised Contrastive Learning for Bioactive Molecule Search”, In: The Thirty-Eighth Annual Conference on Neural Information Processing Systems. Vancouver, Canada: NeurIPS 2024, Sep.… [cited by applicant]
Altenburg, et al., “yHydra: Deep Learning enables an Ultra Fast Open Search by Jointly Embedding MS/MS Spectra and Peptides of Mass Spectrometry-based Proteomics”, bioRxiv, https://www.biorxiv.org/content/10.1101/2021.1… [cited by applicant]
Ananth, et al., “A learned score function improves the power of mass spectrometry database search”, bioRxiv, https://www.biorxiv.org/content/10.1101/2024.01.26.577425v2, Feb. 7, 2024. [cited by applicant]
Bassani-Sternberg, et al., “Direct identification of clinically relevant neoepitopes presented on native human melanoma tissue by mass spectrometry”, Nature Communications, Nov. 21, 2016. [cited by applicant]
Bekker-Jense, et al., “An Optimized Shotgun Strategy for the Rapid Generation of Comprehensive Human Proteomes”, Cell Systems, 4, pp. 587-599, Jun. 28, 2017. [cited by applicant]
Bittremieux, et al., “A learned embedding for efficient joint analysis of millions of mass spectra”, Nat. Methods, 19(6) pp. 675-678, Jun. 2022. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners”, arXiv:2005.14165, Jul. 22, 2020. [cited by applicant]
Butler, et al., “MS2Mol: A transformer model for illuminating dark chemical space from mass spectra”, ChemRxiv, Sep. 5, 2023, Version 4. [cited by applicant]
Cox, et al., “MaxQuant enables high peptide identification rates, individualized p.p.b.-range mass accuracies and proteome-wide protein quantification”, Nature Biotechnology, vol. 26, No. 12, Dec. 2008. [cited by applicant]
Eng, et al., “An Approach to Correlate Tandem Mass Spectral Data of Peptides with Amino Acid Sequences in a Protein Database”, Journal of the American Society for Mass Spectrometry, vol. 5, Issue 11, Nov. 1, 1994. [cited by applicant]
Freestone, et al., “Re-investigating the correctness of decoy-based false discovery rate control in proteomics tandem mass spectrometry”, bioRxiv, https://doi.org/10.1101/2023.06.21.546013, Jun. 24, 2023. [cited by applicant]
Freestone, et al., “Reinvestigating the Correctness of Decoy-Based False Discovery Rate Control in Proteomics Tandem Mass Spectrometry”, Journal of Proteome Research, 2024, 23, 1907-1914. [cited by applicant]
Freestone, et al., “Semi-supervised learning while controlling the FDR with an application to tandem mass spectrometry analysis”, bioRxiv, https://www.biorxiv.org/content/10.1101/2023.10.26.564068v3, Jan. 26, 2024. [cited by applicant]
Gabriel, et al., “Prosit-TMT: Deep Learning Boosts Identification of TMT-Labeled Peptides”, Anal. Chem., 2022, 94, 7181-7190. [cited by applicant]
Gessulat, et al., “Prosit: proteome-wide prediction of peptide tandem mass spectra by deep learning”, Nature Methods, vol. 16, pp. 509-518, Jun. 2019. [cited by applicant]
Griss, Johannes, “Recognizing millions of consistently unidentified spectra across hundreds of shotgun proteomics datasets”, Nature Methods, vol. 13, No. 8, Aug. 2016. [cited by applicant]
Jiang, et al., “The Future of Proteomics is Up in the Air: Can Ion Mobility Replace Liquid Chromatography for High Through put Proteomics?”, Journal of Proteome Research, 2024, 23, 1871-1882. [cited by applicant]
Kall, et al., “Semi-supervised learning for peptide identification from shotgun proteomics datasets”, Nature Methods, vol. 4, No. 11, Nov. 2007. [cited by applicant]
Keller, et al., “Empirical Statistical Model to Estimate the Accuracy of Peptide Identifications Made by MS/MS and Database Search”, Analytical Chemistry, vol. 74, No. 20, Oct. 15, 2002. [cited by applicant]
Kim, et al., “MS-GF + makes progress towards a universal database search tool for proteomics”, Nature Communications, Oct. 31, 2014. [cited by applicant]
Kleikamp, et al., “Metaproteomics, metagenomics and 16S rRNA sequencing provide different perspectives on the aerobic granular sludge microbiome”, Water Research, 246 (2023) 120700. [cited by applicant]
Klimek, et al., “The Standard Protein Mix Database: A Diverse Dataset to Assist in the Production of Improved Peptide and Protein Identification Software Tools”, J Proteome Res. Jan. 2008; 7(1): 96-103. [cited by applicant]
Kong, et al., “MSFragger: ultrafast and comprehensive peptide identification in shotgun proteomics”, Nat Methods. May 2017; 14(5): 513-520. [cited by applicant]
Lazear, Michaelr. , “Sage: An Open-Source Tool for Fast Proteomics Searching and Quantification at Scale”, Journal of Proteome, 2023, 22, 3652-3659. [cited by applicant]
Meier, et al., “Online Parallel Accumulation-Serial Fragmentation (PASEF) with a Novel Trapped Ion Mobility Mass Spectrometer”, Molecular & Cellular Proteomics, 17, 2534-2545, Dec. 2018. [cited by applicant]
Radford, et al., “Improving Language Understanding by Generative Pre-Training”, Preprint 2018. [cited by applicant]
Radford, et al., “Language Models are Unsupervised Multitask Learners”, OpenAI, https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners. pdf, 2019. [cited by applicant]
Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, arXiv:2103.00020v1, Feb. 26, 2021. [cited by applicant]
Radford, et al., “Robust Speech Recognition via Large-Scale Weak Supervision”, arXiv:2212.04356v1, Dec. 6, 2022. [cited by applicant]
Schoenholz, et al., “Peptide-Spectra Matching from Weak Supervision”, arXiv:1808.06576v2, Aug. 22, 2018. [cited by applicant]
Tiwary, et al., “High-quality MS/MS spectrum prediction for data-dependent and data-independent acquisition data analysis”, Nature Methods, vol. 16, pp. 519-525, Jun. 2019. [cited by applicant]
Wang, et al., “Assembling the Community-Scale Discoverable Human Proteome”, Cell Systems, 7, 412-421, Oct. 24, 2018. [cited by applicant]
Wen, et al., “Assessment of false discovery rate control in tandem mass spectrometry analysis using entrapment”, bioRxiv, Jun. 4, 2024. [cited by applicant]
Wilhelm, et al., “Deep learning boosts sensitivity of mass spectrometry-based immunopeptidomics”, Nature Communications, 12, Article No. 3346, Jun. 23, 2021. [cited by applicant]
Williams, et al., “Automated Coupling of Nanodroplet Sample Preparation with Liquid Chromatography-Mass Spectrometry for High-Throughput Single-Cell Proteomics”, Anal Chem., Aug. 4, 2020; 92(15): 10588-10596. [cited by applicant]