IP Library Granted Patent US 12,562,239
Granted Patent B2
US 12,562,239 · App. 19/057,786 · Granted Feb 24, 2026

Systems and methods for analyzing mixed cell populations

Inventors: Aaron M. Newman (Palo Alto, CA); Arash Ash Alizadeh (San Mateo, CA)
Assignee: The Board of Trustees of the Leland Stanford Junior University
G16B30/00C12Q1/6881G16B20/00G16B40/20G16B50/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,239
App. No.
19/057,786
Granted
Feb 24, 2026
Kind
B2
Abstract

The present disclosure provides systems and methods for analyzing a mixed population of cells. In particular, the present disclosure provides systems and methods for digital cytometry of a biological sample, digital analysis of a biological sample, digital purification of a biological sample, evaluation of a disease in an individual, and prediction of a clinical outcome of a disease therapy.

Claims (33)

1 . A method comprising:

(a) obtaining a biological sample from a subject, wherein said biological sample comprises a tumor sample comprising a plurality of distinct cell types;

(b) extracting ribonucleic acid (RNA) molecules from said biological sample;

(c) assaying said RNA molecules to generate a feature profile,

wherein said assaying comprises RNA sequencing, wherein said RNA sequencing comprises (i) reverse transcribing said RNA molecules to produce complementary deoxyribonucleic acid (cDNA) molecules and (ii) amplifying said cDNA molecules,

wherein said feature profile comprises a plurality of features associated with said plurality of distinct cell types,

wherein said feature profile comprises a gene expression profile of cells in said biological sample, wherein said gene expression profile represents an RNA transcriptome of said cells in said biological sample; and

(d) computer processing said feature profile to (1) quantify an abundance of at least one of said plurality of distinct cell types in said biological sample, wherein quantifying said abundance comprises applying a batch correction procedure to remove technical variation in said abundance, and (2) generate one or more differential gene expression profiles in said RNA transcriptome of said cells in said biological sample, wherein said one or more differential gene expression profiles are differential across subtypes of said at least one of said plurality of distinct cell types,

wherein quantifying said abundance comprises optimizing a regression between said feature profile and said reference matrix B of feature signatures for a second plurality of distinct cell types,

wherein said feature profile is modeled as a linear combination of said reference matrix B,

wherein optimizing said regression comprises solving for a set of regression coefficients of said regression, wherein said solution minimizes a linear loss function and an L 2 -norm penalty function,

wherein said batch correction procedure removes said technical variation between a reference matrix B of a plurality of feature signatures and said feature profile,

wherein said batch correction procedure is applied in a single cell reference mode (S-mode) or a bulk reference mode (B-mode),

(i) wherein said applying said batch correction procedure in said S-mode comprises removing technical differences between said reference matrix B derived from a set of single cell reference profiles and an input set of mixture samples M by:

(1) obtaining a plurality of estimates of a plurality of cell frequencies F* within said input set of mixture samples M, given said reference matrix B and said set of single cell reference profiles R, and

(2) refining said plurality of estimates of said plurality of cell frequencies F* by performing said batch correction procedure on said reference matrix B to obtain an adjusted reference matrix, and applying said adjusted reference matrix to said input set of mixture samples M, and

(ii) wherein said applying said batch correction procedure in said B-mode comprises removing said technical differences between said reference matrix B derived from bulk reference profiles and an input set of mixture samples M by:

(1) generating a plurality of mixture samples M* comprising a linear combination of a plurality of imputed cell type proportions in said input set of mixture samples M and corresponding profiles in said reference matrix B, and

(2) performing said batch correction on said input set of mixture samples M to eliminate said batch effects between said input set of mixture samples M and said plurality of mixture samples M*.

2 . The method of claim 1 , wherein obtaining said biological sample does not comprise physical isolation of cells from said biological sample.

3 . The method of claim 1 , further comprising enriching said biological sample for at least one distinct cell type of said plurality of distinct cell types.

4 . The method of claim 1 , wherein said feature profile is generated from single-cell gene expression measurements of a plurality of cells of each of said plurality of distinct cell types.

5 . The method of claim 4 , wherein said RNA sequencing comprises single-cell RNA sequencing (scRNA-Seq).

6 . The method of claim 1 , wherein said abundance is a fractional abundance of said at least one of said plurality of distinct cell types in said biological sample.

7 . The method of claim 1 , wherein said linear loss function is a linear ε-insensitive loss function.

8 . The method of claim 1 , wherein optimizing said regression comprises using a support vector regression (SVR) or a non-negative matrix factorization (NMF).

9 . The method of claim 8 , wherein said SVR is ε-SVR.

10 . The method of claim 8 , wherein said SVR is v(nu)-SVR.

11 . The method of claim 1 , wherein said batch correction procedure is applied in said single cell reference mode (S-mode).

12 . The method of claim 1 , wherein said batch correction procedure is applied in said bulk reference mode (B-mode).

13 . The method of claim 1 , further comprising generating a cell-type-specific state for said at least one of said plurality of distinct cell types.

14 . The method of claim 13 , wherein said cell-type-specific state is a sample-level cell-type-specific state.

15 . The method of claim 13 , wherein said cell-type-specific state is a group-level cell-type-specific state.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2025
From: NEWMAN, AARON M.; ALIZADEH, ARASH ASH
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
Reel/Frame 070276/0304 →
Continuity (3)
Continuation 16631778
Provisional Application 62535645 · Jul 21, 2017
Related Publication 20250322909A1 · Oct 16, 2025
References Cited (97)
US 6661004B2 · Aumond et al. · 2003 [cited by applicant]
US 10167514B2 · Newman · 2019 [cited by examiner]
US 11225689B2 · Shekhar · 2022 [cited by examiner]
US 11756651B2 · Givechian · 2023 [cited by applicant]
US 11802314B2 · Newman · 2023 [cited by examiner]
US 12031183B2 · Newman · 2024 [cited by applicant]
US 12249401B2 · Newman · 2025 [cited by examiner]
US 20050108753A1 · Saidi et al. · 2005 [cited by applicant]
US 20130110756A1 · Zhang et al. · 2013 [cited by applicant]
US 20160217253A1 · Newman et al. · 2016 [cited by applicant]
US 20160341731A1 · Sood et al. · 2016 [cited by applicant]
US 20190233898A1 · Newman et al. · 2019 [cited by applicant]
US 20190338364A1 · Newman et al. · 2019 [cited by applicant]
US 20200075169A1 · Lau · 2020 [cited by applicant]
US 20210040442A1 · Rajagopal · 2021 [cited by examiner]
US 20250129425A1 · Newman · 2025 [cited by applicant]
US 20250243551A1 · Newman · 2025 [cited by applicant]
US 20250322909A1 · Newman · 2025 [cited by examiner]
CN 103217411 · 2013 [cited by applicant]
JP T2014517954 · 2014 [cited by applicant]
WO WO2007035613 · 2007 [cited by applicant]
WO WO2015109234A1 · 2015 [cited by examiner]
WO WO2018191558A1 · 2018 [cited by examiner]
Abbas et al., “Deconvolution of Blood Microarray Data Identifies Cellular Activation Patterns in Systematic Kupus Erythematosus”, PLOS One, Jul. 2009, p. e6098, vol. 4, No. 7. [cited by applicant]
Abbas et al., “Immune response in silico (IRIS): immune-specific genes identified from a compendium of microarray expression data” Genes and Immunity, 2005, pp. 319-331, vol. 6, No. 4. [cited by applicant]
Ahn et al., “DeMix: deconvolution for mixed cancer transcriptomes using raw measured data” Bioinformatics, 2013, pp. 1865-1871, vol. 29, No. 15. [cited by applicant]
Benita et al., “Gene enrichment profiles reveal T-cell development, differentiation, and lineage-specific transcription factors including ZBTB25 as a novel NF-AT repressor”, Blood, 2010, pp. 5376-5384, vol. 115. [cited by applicant]
Burdick and Murray, “Deconvolution of gene expression from cell populations across the C. elegans lineage”, BMC Bioinformatics, Jun. 22, 2012, p. 204, vol. 14. [cited by applicant]
Caicedo, J.C. et al., (2017) Data-analysis strategies for image-based cell profiling. Nature Methods, vol. 14, No. 9, p. 849-863. (Aug. 11, 2017). [cited by applicant]
Chen, C. et al, (2011) Removing batch effects in analysis of expression microarray data: an evaluation of six batch adjustment methods. PLOS ONE vol. 6, issue 2, e17238, 10 pages. [cited by applicant]
Chen et al. (2017) “Inference of immune cell composition on the expression profiles of mouse tissue.” Scientific reports 7: 1-11. [cited by applicant]
Cherlassky et al., “Practical selection of SVM parameters and noise estimation for SVM regression”, Neural Netk, 2004, pp. 113-126, vol. 17. [cited by applicant]
Cobos et al., (2018) “Computational deconvolution of transcriptomics data from mixed cell populations”, Bioinformatics, 34(11), pp. 1969-1779. [cited by applicant]
Coussens et al., “Neutralizing tumor-promoting chronic inflammation: a magic bullet?”, Science, 2013, pp. 286-291, vol. 339. [cited by applicant]
Definition of normal distribution, Wikipedia.com downloaded Jan. 2024 (Year: 2024). [cited by applicant]
Definition of sampling, and random sampling, Wikipedia.com, downloaded Jan. 2024 (Year: 2024). [cited by applicant]
Definition of simple random sampling, Wikipedia.com downloaded Jan. 2024 (Year: 2024). [cited by applicant]
Drucker et al., “Support Vector Regression Machines”, MIT Press, 1997, pp. 155-161, vol. 9. [cited by applicant]
Farrar et al., “Multicollinearity in Regression Analysis: The Problem Revisited”, R. R. Rev. Econ. Stat., 1967, pp. 92-107, vol. 49. [cited by applicant]
Gaiteri et al., (2013) “Beyond modules and hubs: the potential of gene co-expression networks for investigating molecular mechanisms of complex brain disorders.”, Genes, Brain and Behavior, 13: 13-24. [cited by applicant]
Gaujoux and Seoighe, “CellMix: a comprehensive toolbox for gene expression deconvolution”, Bioinformatics, 2013, pp. 2211-2212, vol. 29, No. 17. [cited by applicant]
Goh et al., (2017) Why batch effects matter in Omics data and how to avoid them. Trends in Biotechnology, vol. 25, No. 6, p. 498-507 (Jun. 2017). [cited by applicant]
Gong and Szutakowski, “DeconRNASeq: a statistical framework for deconvolution of heterogeneous tissue samples based on mRNA-Seq data”, Bioinformatics, 2013, pp. 1083-1085, vol. 29, No. 8. [cited by applicant]
Gong et al., “Optimal Deconvolution of Transcriptional Profiling Data Using Quadratic Programming with Application to Complex Clinical Blood Samples”, PLOS One, Nov. 2011, p. e27156. [cited by applicant]
Hanahan et al., “Hallmarks of Cancer: The Next Generation”, Cell, 2011, pp. 646-674, vol. 144. [cited by applicant]
Johnson et al. (2007) “Adjusting batch effects in microarray expression data using empirical Bayes methods” Biostatistics 8(1): 118-127. [cited by applicant]
Ju et al. (2013) “Defining cell-type specificity at the transcriptional level in human disease.” Genome research 23: 1862-1873. [cited by applicant]
Krishnan et al. (2011) “Quantitative Analysis of Sub-Epithelial Connective Tissue Cell Population of Oral Submucous Fibrosis Using Support Vector Machine”, Journal of Medical Imaging and Health Informatics, vol. 1, No. … [cited by applicant]
Kuhn et al., “Population-specific expression analysis (PSEA) reveals molecular changes in diseased brain”, Nat Methods, 2011, pp. 945-947, vol. 8. [cited by applicant]
Le et al. (2020) “A Review of Digital Cytometry Methods: Estimating the Relative Abundance of Cell Types in a Bulk of Cells”, Briefings in Bioinformatics: 1-12. [cited by applicant]
Levy et al., “Active Idiotypic Vaccination Versus Control Immunotherapy for Folicular Lymphoma”, J Clin. Oncol., 2014, pp. 1797-1803, vol. 32. [cited by applicant]
Li et al. (2016) “Comprehensive Analyses of Tumor Immunity: Implications for Cancer Immunotherapy”, Genome Biology, 2016, vol. 17, No. 1: 1-16. [cited by applicant]
Liebner et al., “MMAD: microarray microdissection with analysis of difference is a computational tool for deconvoluting cell type=speficic contributions from tissue samples”, Bioinformatics, 2014, pp. 682-689, vol. 30, … [cited by applicant]
Lu et al., “Expression deconvolution: A reinterpretation of DNA microarray data reveals dynamic changes in cell populations”, PNAS, 2003, pp. 10370-10375, vol. 100, No. 18. [cited by applicant]
Lukk et al., “A global map of human gene expression”, NAt. Biotechnol, 2010, pp. 322-324, vol. 28. [cited by applicant]
Mackey et al. (2011) “Divide-and-Conquer Matrix Factorization” Advances in Neural Information Processing Systems 24, edited by J. Shawe-Taylor et al. Proceedings from the conference, “Neural information Processing Syste… [cited by applicant]
Meng et al., (2013) Scalable simple random sampling and stratified sampling, Proceedings of the 30 [cited by applicant]
Newman et al. (2014) “An ultrasensitive method for quantitating circulating tumor DNA with broad patient coverage” Nature Medicine 20: 548-554. [cited by applicant]
Newman et al., (2014) “Identifying stem cell gene expression patterns and phenotypic networks with AutoSOME.”, Methods in Molecular Biology, 1150: 115-130. [cited by applicant]
Newman et al. (2015) “Robust enumeration of cell subsets from tissue expression profiles”, Nature Methods, vol. 12, No. 5, pp. 453-457. [cited by applicant]
Newman et al. (2017) “Data Normalization Considerations for Digital Tumor Dissection”, Genome Biology, 2017, vol. 18, No. 1: 1-6. [cited by applicant]
Newman et al. (2019) “Determining Cell Type Abundance and Expression From Bulk Tissues with Digital Cytometry”, Nature Biotechnology, vol. 37, No. 7: 773-782. [cited by applicant]
Qiao et al., “PERT: A Method for Expression Deconvolution of Human Blood Samples from Varied Microenvironmental and Developmental Conditions”, PLOS Comput. Biol., 2012, p. e1002838, vol. 8. [cited by applicant]
Rivenbark et al., (2013) “Molecular and cellular heterogeneity in breast cancer: challenges for personalized medicine”, American Journal of Pathology, 183 (4): 1113-1124. [cited by applicant]
Scholkopf et al., “New Support Vector Algorithms”, Neural Comput., 2000, pp. 1207-1245, vol. 12. [cited by applicant]
Shen-Orr and Gajoux, “Computational deconvolution: extracting cell type-specific information from heterogeneous samples”, Curr. Opin. Immunol., 2013, pp. 571-578, vol. 25. [cited by applicant]
Shen-Orr et al., “Cell type-specific gene expression differences in complex tissues”, Nat. Methods., 2010, pp. 287-289, vol. 7. [cited by applicant]
Smola (2004) “A tutorial on support vector regression.”, Statistics and computing, 14: 199-222. [cited by applicant]
Sotiriou et al., (2009) “Gene-expression patterns in Breast Cancer.”, The New England Journal of Medicine, 360 (8): 790-800. [cited by applicant]
Steen et al. (2020) “Profiling Cell Type Abundance and Expression in Bulk Tissues with CIBERSORTx”, Methods Mol Bio, vol. 2117: 135-157. [cited by applicant]
Storey et al., “Statistical significance for genomewide studies”, Proc. Natl. Acad. Sci. U. S. A., 2003, pp. 9440-9445, vol. 100. [cited by applicant]
Sun et al. (2010) “Combined feature selection and cancer prognosis using support vector machine regression.” IEEE/ACM transactions on computational biology and bioinformatics 8(6): 1671-1677. [cited by applicant]
Thiebaut (2002) “Optimization issues in blind deconvolution algorithms”, Proc. SPIE 4847, Astronomical Data Analysis II, 1-9. [cited by applicant]
Tung et al., (2017) Batch effects and the effective design of single cell gene expression studies. Scientific reports, vol. 7, e39921, 15 pages (Jan. 2017). [cited by applicant]
Wagner et al., (2016) Revealing the vectors of cellular identify with single cell genomics. Nature Biotechnology vol. 14, No. 11, p. 1145-1168. [cited by applicant]
Wang et al. (2013) “Non-negative matrix factorization by maximizing correntropy for cancer clustering” BMC Bioinformatics 14(107): 107 (pp. 1-11). [cited by applicant]
Wang et al., “The doubly regularized support vector machine”, Statistica Sinica, 2006, pp. 589-615, vol. 16, No. 2. [cited by applicant]
Wilhelm-Benartzi et al. (2013) “Review of processing and analysis methods for DNA methylation array data” British J of Cancer 109(6): 1394-1402. [cited by applicant]
Yin, (2013) “Identification of differential gene pathways with sparse principal component analysis”, Georgia State University, 1-26. [cited by applicant]
Yoshihara et al., “Inferring tumour purity and stromal and immune cell a dmixture from expression data”, Nat. Commun., 2013, p. 2612, vol. 4. [cited by applicant]
Zheng et al., (2014) “Deconvolution of High Dimensional Mixtures via Boosting, with Application to Diffusion-Weighted MRI of Human Brain”, Advances in Neural Information Processing Systems, 27:2699-2707. [cited by applicant]
Zhong and Liu, “Gene expression deconvolution in linear space”, Nat. Methods., 2012, pp. 8-9, vol. 9. [cited by applicant]
Zhong et al., “Digital sorting of complex tissues for cell type-specific gene expression profiles”, BMC Bioinformatics, 2013, p. 89, vol. 14. [cited by applicant]
Zuckerman (2013) PLOS Computational Biology 9:e1003189. [cited by applicant]
Abraham; Scalable approaches for analysis of human genome-wide expression and genetic variation data. PhD thesis, Dept. of Computing and Information Systems, The University of Melbourne, 313 pages (2012). [cited by applicant]
Arieshanti et al., “Analysis of SELDI-TOF-MS Using E-Support Vector Regression for Ovarian Cancer Identification”, The 15 [cited by applicant]
Boardman, Extrinsic regularization in parameter optimization for support vector machines. Master of Computer Science, Dalhousie University, 132 pages. (2006). [cited by applicant]
Chiu, Using support vector regression to model the correlation between the clinical metasases time and gene expression profile for breast cancer. Artificial Intelligence in Medicine, vol. 44, p. 221-231. (2008). [cited by applicant]
Chuang et al., “Dimension Reduction with Support Vector Regression for Ovarian Cancer Microarray Data”, 2005 IEEE International conference on systems, man, and cybernetics. p. 1-5. DOI: 10.1109/ICSMC.2005.1571284 (2005). [cited by applicant]
Felton, Identification of carcinoma cells in peripheral blood samples of patients with advanced breast carcinoma using RT-PCT amplification of CK7 and MUC1. The Breast. vol. 13, p. 35-41. (2004). [cited by applicant]
Greene et al., “Big Data Bioinformatics”, Journal of Cellular Physiology, vol. 229, p. 1896-1900. (2014). [cited by applicant]
Jimenez et al., “Feasibility of gene expression signature analysis in prostate cancer biopsy specimens to predict outcomes following radiation therap.” Radiation Oncology, vol. 87, Issue 2, supplement, S669 conference a… [cited by applicant]
Karlik et al., “Personalized Cancer Treatment by Using Naïve Bayes Classifier”, International Journal of Machine Learning and Computing, vol. 2, No. 3, p. 339-344. (2012). [cited by applicant]
Mahmoodian et al., “Using support vector regression in gene selection and fuzzy rule generation for relapse time prediction of breast cancer”, Biocybernetrics and biomedical engineering, vol. 36, p. 468-472. (2016). [cited by applicant]
Mohammadi et al. (2017) A Critical Survey of Deconvolution Methods for Separating Cell Types in Complex Tissues, Proceedings of the IEEE, 105(2): 340-366. [cited by applicant]
Wei, RNA-seq accurately identifies cancer biomarker signatures to distinguish tissue of origin. Neoplasia, vol. 16, No. 11, p. 918-927. (2014). [cited by applicant]
Wilson et al. (2015) “Outcomes and endpoints in trials of cancer treatment: the past, present and future,” Lancet Oncology, 16: e32-e42. [cited by applicant]