IP Library Granted Patent US 12,249,401
Granted Patent B2
US 12,249,401 · App. 16/631,778 · Granted Mar 11, 2025

Systems and methods for analyzing mixed cell populations

Inventors: Aaron M. Newman (San Mateo, CA); Arash Ash Alizadeh (San Mateo, CA)
Assignee: The Board of Trustees of the Leland Stanford Junior University
G16B30/00C12Q1/6881G16B20/00G16B40/20G16B50/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,249,401
App. No.
16/631,778
Granted
Mar 11, 2025
Kind
B2
Abstract

The present disclosure provides systems and methods for analyzing a mixed population of cells. In particular, the present disclosure provides systems and methods for digital cytometry of a biological sample, digital analysis of a biological sample, digital purification of a biological sample, evaluation of a disease in an individual, and prediction of a clinical outcome of a disease therapy.

Claims (34)

1. A method comprising:

(a) obtaining a biological sample comprising a plurality of biological macromolecules from a plurality of distinct cell types from a subject having cancer;

(b) processing said biological sample to generate a feature profile of said plurality of biological macromolecules, wherein said feature profile comprises a plurality of features associated with said plurality of distinct cell types;

(c) computer processing said feature profile, using a deconvolution module, to (1) quantify an abundance of at least one of said plurality of distinct cell types in said biological sample, wherein quantifying said abundance comprises applying a batch correction procedure to remove technical variation in said abundance, and (2) generate one or more differential gene expression profiles that are differential across subtypes of said at least one of said plurality of distinct cell types,

wherein said batch correction procedure removes said technical variation between a reference matrix B of a plurality of feature signatures and said feature profile,

wherein said batch correction procedure is applied in a single cell reference mode (S-mode) or a bulk reference mode (B-mode),

(i) wherein said applying said batch correction procedure in said S-mode comprises removing technical differences between said reference matrix B derived from a set of single cell reference profiles and an input set of mixture samples M by:

(1) obtaining a plurality of estimates of a plurality of cell frequencies F* within said input set of mixture samples M, given said reference matrix B and said set of single cell reference profiles R, and

(2) refining said plurality of estimates of said plurality of cell frequencies F* by performing said batch correction procedure on said reference matrix B to obtain an adjusted reference matrix, and applying said adjusted reference matrix to said input set of mixture samples M, and

(ii) wherein said applying said batch correction procedure in said B-mode comprises removing said technical differences between said reference matrix B derived from bulk reference profiles and an input set of mixture samples M by:

(1) generating a plurality of mixture samples M* comprising a linear combination of a plurality of imputed cell type proportions in said input set of mixture samples M and corresponding profiles in said reference matrix B, and

(2) performing said batch correction on said input set of mixture samples M to eliminate said batch effects between said input set of mixture samples M and said plurality of mixture samples M*;

(d) predicting a clinical outcome of a cancer therapy on said subject for said cancer, based at least in part on said abundance and said one or more differential gene expression profiles of said at least one of said plurality of distinct cell types in said biological sample, wherein said at least one of said plurality of distinct cell types in said biological sample comprises a type of cancer cell; and

(e) administering said cancer therapy to said subject based on said predicted clinical outcome of said cancer therapy, wherein said cancer therapy is selected from the group consisting of a chemotherapy, an immunotherapy, and an immunochemotherapy.

2. The method of claim 1 , wherein obtaining said biological sample does not comprise physical isolation of cells from said biological sample.

3. The method of claim 1 , wherein said feature profile is generated from single-cell gene expression measurements of a plurality of cells of each of said plurality of distinct cell types.

4. The method of claim 3 , wherein said single-cell gene expression measurements are generated by single-cell RNA sequencing (scRNA-Seq).

5. The method of claim 1 , wherein said abundance is a fractional abundance of said at least one of said plurality of distinct cell types in said biological sample.

6. The method of claim 1 , wherein quantifying said abundance comprises optimizing a regression between said feature profile and said reference matrix B of feature signatures for a second plurality of distinct cell types, wherein said feature profile is modeled as a linear combination of said reference matrix B.

7. The method of claim 6 , wherein optimizing said regression comprises solving for a set of regression coefficients of said regression, wherein said solution minimizes a linear loss function and an L 2 -norm penalty function.

8. The method of claim 7 , wherein said linear loss function is a linear E-insensitive loss function.

9. The method of claim 6 , wherein optimizing said regression comprises using a support vector regression (SVR) or a non-negative matrix factorization (NMF).

10. The method of claim 9 , wherein said SVR is ε-SVR or v (nu)-SVR.

11. The method of claim 1 , wherein said biological sample comprises bulk tissue, a formalin-fixed, paraffin-embedded (FFPE) tissue, a frozen tissue, a blood sample, a sample derived from a solid tissue sample, or a tumor sample.

12. The method of claim 1 , wherein said biological macromolecules comprise nucleic acids, proteins, metabolites, carbohydrates, sugars, lipids, or a combination thereof.

13. The method of claim 1 , wherein said batch correction procedure is applied in said single cell reference mode (S-mode).

14. The method of claim 1 , further comprising generating a cell-type-specific state for said at least one of said plurality of distinct cell types.

15. The method of claim 14 , wherein said cell-type-specific state is a sample-level cell-type-specific state.

16. The method of claim 14 , wherein said cell-type-specific state is a group-level cell-type-specific state.

17. The method of claim 1 , wherein said batch correction procedure is applied in said bulk reference mode (B-mode).

18. The method of claim 1 , wherein said one or more differential gene expression profiles processing is generated at least in part by:

(a) stratifying one or more gene expression profiles of said plurality of biological macromolecules into two or more subtypes based on bulk expression; and

(b) determining one or more cell type-specific gene expression coefficients for each of said two or more subtypes.

19. The method of claim 18 , wherein said plurality of biological samples are obtained from a plurality of conditions of interest such that said cell type-specific gene expression coefficients are differential across said plurality of conditions of interest.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2020
From: NEWMAN, AARON M.; ALIZADEH, ARASH ASH
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERISTY
Reel/Frame 052262/0902 →
Continuity (2)
Provisional Application 62535645 · Jul 21, 2017
Related Publication 20200176080A1 · Jun 4, 2020
References Cited (73)
US 6661004B2 · Aumond et al. · 2003 [cited by applicant]
US 10167514B2 · Newman · 2019 [cited by examiner]
US 11802314B2 · Newman · 2023 [cited by examiner]
US 12031183B2 · Newman · 2024 [cited by examiner]
US 20050108753A1 · Saidi et al. · 2005 [cited by applicant]
US 20130110756A1 · Zhang et al. · 2013 [cited by applicant]
US 20160217253A1 · Newman · 2016 [cited by examiner]
US 20160341731A1 · Sood et al. · 2016 [cited by applicant]
US 20210040442A1 · Rajagopal · 2021 [cited by examiner]
CN 103217411 · 2013 [cited by applicant]
JP T2014517954 · 2014 [cited by applicant]
WO WO2007035613 · 2007 [cited by applicant]
Chen, Ziyi, et al. “Inference of immune cell composition on the expression profiles of mouse tissue.” Scientific reports 7.1 (2017): 1-11. [cited by examiner]
Ju, Wenjun, et al. “Defining cell-type specificity at the transcriptional level in human disease.” Genome research 23.11 (2013): 1862-1873. [cited by examiner]
Sun, Bing-Yu, et al. “Combined feature selection and cancer prognosis using support vector machine regression.” IEEE/ACM transactions on computational biology and bioinformatics 8.6 (2010): 1671-1677. [cited by examiner]
Johnson, W. Evan, Cheng Li, and Ariel Rabinovic. “Adjusting batch effects in microarray expression data using empirical Bayes methods.” Biostatistics 8.1 (2007): 118-127. (Year: 2007). [cited by examiner]
Wilhelm-Benartzi, Charlotte S., et al. “Review of processing and analysis methods for DNA methylation array data.” British journal of cancer 109.6 (2013): 1394-1402. (Year: 2013). [cited by examiner]
Wagner, A. et al. (2016) Revealing the vectors of cellular identity with single cell genomics. Nature Biotechnology vol. 14, No. 11, p. 1145-1168. (Year: 2016). [cited by examiner]
Tung P-Y, et al. (Jan. 2017) Batch effects and the effective design of single cell gene expression studies. Scientific reports, volum3 7, e39921, 15 pages. (Year: 2017). [cited by examiner]
Goh, W. W-B, et al (Jun. 2017) Why batch effects matter in Omics data and how to avoid them. Trends in Biotechnology, vol. 25, No. 6, p. 498-507. (Year: 2017). [cited by examiner]
Chen, C. et al. (2011) Removing batch effects in analysis of expression microarray data: an evaluation of six batch adjustment methods. PLOS One vol. 6, issue 2, e17238, 10 pages. (Year: 2011). [cited by examiner]
Caicedo, J. C. et al. (Aug. 11, 2017) Data-analysis strategies for image-based cell profiling. Nature Methods, vol. 14, No. 9, p. 849-863. (Year: 2017). [cited by examiner]
Definition of normal distribution, wikipedia.com downloaded Jan. 2024 (Year: 2024). [cited by examiner]
Definition of sampling, and random sampling, wikipedia.com, downloaded Jan. 2024 (Year: 2024). [cited by examiner]
Definition of simple random sampling, wikipedia.com downloade Jan. 2024 (Year: 2024). [cited by examiner]
Meng, et al (2013) scalable simple random sampling and stratified sampling, Proceedings of the 30th int conf on machine learning, atlanta georgia, vol. 28. (Year: 2013). [cited by examiner]
Le et al. (2020) “A Review of Digital Cytometry Methods: Estimating the Relative Abundance of Cell Types in a Bulk of Cells”, Briefings in Bioinformatics: 1-12. [cited by applicant]
Li et al. (2016) “Comprehensive Analyses of Tumor Immunity: Implications for Cancer Immunotherapy”, Genome Biology, 2016, vol. 17, No. 1: 1-16. [cited by applicant]
Newman et al. (2017) “Data Normalization Considerations for Digital Tumor Dissection”, Genome Biology, 2017, vol. 18, No. 1: 1-6. [cited by applicant]
Newman et al. (2019) “Determining Cell Type Abundance and Expression From Bulk Tissues with Digital Cytometry”, Nature Biotechnology, vol. 37, No. 7: 773-782. [cited by applicant]
Steen et al. (2020) “Profiling Cell Type Abundance and Expression in Bulk Tissues with CIBERSORTx”, Methods Mol Bio, vol. 2117: 135-157. [cited by applicant]
Cobos et al., (2018) “Computational deconvolution of transcriptomics data from mixed cell populations”, Bioinformatics, vol. 34(11), pp. 1969-1779. [cited by applicant]
Krishnan et al., (2011) “Quantitative Analysis of Sub-Epithelial Connective Tissue Cell Population of Oral Submucous Fibrosis Using Support Vector Machine”, Journal of Medical Imaging and Health Informatics, vol. 1, No.… [cited by applicant]
Newman et al., (2015) “Robust enumeration of cell subsets from tissue expression profiles”, Nature Methods, vol. 12, No. 5, pp. 453-457. [cited by applicant]
Abbas et al., “Immune response in silico (IRIS): immune-specific genes identified from a compendium of microarray expression data” Genes and Immunity, 2005, pp. 319-331, vol. 6, No. 4. [cited by applicant]
Abbas et al., “Deconvolution of Blood Microarray Data Identifies Cellular Activation Patterns in Systematic Kupus Erythematosus”, PLOS One, Jul. 2009, p. e6098, vol. 4, No. 7. [cited by applicant]
Ahn et al., “DeMix: deconvolution for mixed cancer transcriptomes using raw measured data” Bioinformatics, 2013, pp. 1865-1871, vol. 29, No. 15. [cited by applicant]
Benita et al., “Gene enrichment profiles reveal T-cell development, differentiation, and lineage-specific transcription factors including ZBTB25 as a novel NF-AT repressor”, Blood, 2010, pp. 5376-5384, vol. 115. [cited by applicant]
Burdick and Murray, “Deconvolution of gene expression from cell populations across the C. elegans lineage”, BMC Bioinformatics, Jun. 22, 2012, p. 204, vol. 14. [cited by applicant]
Cherlassky et al., “Practical selection of SVM parameters and noise estimation for SVM regression”, Neural Netk, 2004, pp. 113-126, vol. 17. [cited by applicant]
Coussens et al., “Neutralizing tumor-promoting chronic inflammation: a magic bullet?”, Science, 2013, pp. 286-291, vol. 339. [cited by applicant]
Drucker et al., “Support Vector Regression Machines”, MIT Press, 1997, pp. 155-161, vol. 9. [cited by applicant]
Farrar et al., “Multicollinearity in Regression Analysis: The Problem Revisited”, R. R. Rev. Econ. Stat., 1967, pp. 92-107, vol. 49. [cited by applicant]
Gaujoux and Seoighe, “CellMix: a comprehensive toolbox for gene expression deconvolution”, Bioinformatics, 2013, pp. 2211-2212, vol. 29, No. 17. [cited by applicant]
Gong and Szutakowski, “DeconRNASeq: a statistical framework for deconvolution of heterogeneous tissue samples based on mRNA-Seq data”, Bioinformatics, 2013, pp. 1083-1085, vol. 29, No. 8. [cited by applicant]
Gong et al., “Optimal Deconvolution of Transcriptional Profiling Data Using Quadratic Programming with Application to Complex Clinical Blood Samples”, PLOS One, Nov. 2011, p. e27156. [cited by applicant]
Hanahan et al., “Hallmarks of Cancer: The Next Generation”, Cell, 2011, pp. 646-674, vol. 144. [cited by applicant]
Kuhn et al., “Population-specific expression analysis (PSEA) reveals molecular changes in diseased brain”, Nat Methods, 2011, pp. 945-947, vol. 8. [cited by applicant]
Levy et al., “Active Idiotypic Vaccination Versus Control Immunotherapy for Folicular Lymphoma”, J Clin. Oncol., 2014, pp. 1797-1803, vol. 32. [cited by applicant]
Liebner et al., “MMAD: microarray microdissection with analysis of difference is a computational tool for deconvoluting cell type=speficic contributions from tissue samples”, Bioinformatics, 2014, pp. 682-689, vol. 30, … [cited by applicant]
Lu et al., “Expression deconvolution: A reinterpretation of DNA microarray data reveals dynamic changes in cell populations”, PNAS, 2003, pp. 10370-10375, vol. 100, No. 18. [cited by applicant]
Lukk et al., “A global map of human gene expression”, NAt. Biotechnol, 2010, pp. 322-324, vol. 28. [cited by applicant]
Mackey et al. (2011) “Divide-and-Conquer Matrix Factorization” Advances in Neural Information Processing Systems 24, edited by J. Shawe-Taylor et al. Proceedings from the conference, “Neural information Processing Syste… [cited by applicant]
Newman et al. (2014) “An ultrasensitive method for quantitating circulating tumor DNA with broad patient coverage” Nature Medicine 20: 548-554. [cited by applicant]
Qiao et al., “PERT: A Method for Expression Deconvolution of Human Blood Samples from Varied Microenvironmental and Developmental Conditions”, PLOS Comput. Biol., 2012, p. e1002838, vol. 8. [cited by applicant]
Scholkopf et al., “New Support Vector Algorithms”, Neural Comput., 2000, pp. 1207-1245, vol. 12. [cited by applicant]
Shen-Orr and Gajoux, “Computational deconvolution: extracting cell type-specific information from heterogeneous samples”, Curr. Opin. Immunol., 2013, pp. 571-578, vol. 25. [cited by applicant]
Shen-Orr et al., “Cell type-specific gene expression differences in complex tissues”, Nat. Methods., 2010, pp. 287-289, vol. 7. [cited by applicant]
Storey et al., “Statistical significance for genomewide studies”, Proc. Natl. Acad. Sci. U. S. A., 2003, pp. 9440-9445, vol. 100. [cited by applicant]
Wang et al., “The doubly regularized support vector machine”, Statistica Sinica, 2006, pp. 589-615, vol. 16, No. 2. [cited by applicant]
Wang et al. (2013) “Non-negative matrix factorization by maximizing correntropy for cancer clustering” BMC Bioinformatics 14(107): 107 (pp. 1-11). [cited by applicant]
Yoshihara et al., “Inferring tumour purity and stromal and immune cell a dmixture from expression data”, Nat. Commun., 2013, p. 2612, vol. 4. [cited by applicant]
Zhong and Liu, “Gene expression deconvolution in linear space”, Nat. Methods., 2012, pp. 8-9, vol. 9. [cited by applicant]
Zhong et al., “Digital sorting of complex tissues for cell type-specific gene expression profiles”, BMC Bioinformatics, 2013, p. 89, vol. 14. [cited by applicant]
Zuckerman (2013) PLOS Computational Biology 9:e1003189. [cited by applicant]
Gaiteri et al., (2013) “Beyond modules and hubs: the potential of gene co-expression networks for investigating molecular mechanisms of complex brain disorders.”, Genes, Brain and Behavior, 13: 13-24. [cited by applicant]
Newmen et al., (2014) “Identifying stem cell gene expression patterns and phenotypic networks with AutoSOME.”, Methods in Molecular Biology, 1150: 115-130. [cited by applicant]
Rivenbark et al., (2013) “Molecular and cellular heterogeneity in breast cancer: challenges for personalized medicine”, American Journal of Pathology, 183 (4): 1113-1124. [cited by applicant]
Smola, (2004) “A tutorial on support vector regression.”, Statistics and computing, 14: 199-222. [cited by applicant]
Sotiriou et al., (2009) “Gene-expression patterns in Breast Cancer.”, The New England Journal of Medicine, 360 (8): 790-800. [cited by applicant]
Thiebaut, (2002) “Optimization issues in blind deconvolution algorithms”, Proc. SPIE 4847, Astronomical Data Analysis II, 1-9. [cited by applicant]
Yin, (2013) “Identification of differential gene pathways with sparse principal component analysis”, Georgia State University, 1-26. [cited by applicant]
Zheng et al., (2014) “Deconvolution of High Dimensional Mixtures via Boosting, with Application to Diffusion-Weighted MRI of Human Brain”, Advances in Neural Information Processing Systems, 27: 2699-2707. [cited by applicant]
Cited By (5)
US 12,562,239 US 12,571,054 US 12,584,180 US 12,674,208 US 12,692,551