IP Library Granted Patent US 12,626,780
Granted Patent B2
US 12,626,780 · App. 16/352,739 · Granted May 12, 2026

Method and system for selecting, managing, and analyzing data of high dimensionality

Inventors: Darya Filippova (Sunnyvale, CA); Anton Valouev (La Canada, CA); Virgil Nicula (Cupertino, CA); Karthik Jagadeesh (San Francisco, CA); M. Cyrus Maher (San Mateo, CA); Matthew H. Larson (San Francisco, CA); Monica Portela dos Santos Pimentel (San Jose, CA); Robert Abe Paine Calef (Redwood City, CA)
Assignee: GRAIL, Inc.
G16B30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,780
App. No.
16/352,739
Granted
May 12, 2026
Kind
B2
Abstract

A system, method and computer program product for analyzing data of high dimensionality (e.g., sequence reads of nucleic acid samples in connection with a disease condition) are provided.

Claims (58)

1 . A method of analyzing sequence reads of nucleic acid samples in connection with a cancer condition, comprising:

receiving, at a data collection component of a computer system, a first set of sequence reads of the nucleic acid samples from each healthy subject in a reference group of healthy subjects, wherein the nucleic acid samples comprise a set of cell-free DNA (cfDNA) fragments and wherein the first set of sequence reads is derived from a DNA sequencing process and wherein the first set of sequence reads are not known to contain one or more mutations indicative of the cancer condition;

aligning, using a processor of the computer system, each sequence read in the first set of sequence reads to regions in a reference genome;

establishing, using the processor and subsequent to the aligning, a low variability filter, wherein the establishing comprises:

identifying, using the processor, numbers of sequence reads aligned to each of the regions in the reference genome across each healthy subject in the reference group of healthy subjects;

deriving, using the processor and based on the identifying, quantity data associated with each of the regions based on the numbers of sequence reads;

calibrating, using the processor, the quantity data for each of the regions;

computing, using the processor, reference quantities for each of the regions based on the calibrated quantity data, wherein the reference quantities comprise at least a first reference quantity and a second reference quantity, wherein the first reference quantity corresponds to an average of the calibrated quantity data and wherein the second reference quantity is a standard deviation of the calibrated quantity data;

determining, using the processor, a difference between the first reference quantity and the second reference quantity for each of the regions;

classifying each of the regions as a high variability region or a low variability region by comparing the difference against a predetermined threshold, wherein the high variability region is defined by the difference being greater than the predetermined threshold and wherein the low variability region is defined by the difference between less than the predetermined threshold;

receiving, at the data collection component, a training set of sequence reads from subjects in a training group, wherein the training set includes a second set of sequence reads of nucleic acid samples from healthy subjects and a third set of sequence reads of nucleic acid samples from cancer subjects who are known to have the cancer condition, wherein the second set of sequence reads of nucleic acid samples are not known to contain the one or more mutations indicative of the cancer condition and wherein the third set of sequence reads of nucleic acid samples do contain the one or more mutations indicative of the cancer condition and wherein the training set of sequence reads from the subjects in the training group comprise cfDNA fragments of varying length;

applying, using the processor, the established low variability filter to the training set of sequence reads;

generating, using the processor and based on the applying, a filtered training set of sequence reads, wherein the generating the filtered training set of sequence reads comprises discarding one or more sequence reads in the training set of sequence reads that are aligned to the high variability region in the reference genome training, using the processor, a machine learning model on the filtered training set of sequence reads;

applying, using the processor, a predictive capability of the trained machine learning model to identify differences between the second set of sequence reads of nucleic acid samples from the healthy subjects and the third set of sequence reads of nucleic acid samples from the cancer subjects.

2 . The method of claim 1 , wherein the cancer condition is a cancer type selected from the group consisting of lung cancer, ovarian cancer, kidney cancer, bladder cancer, hepato-biliary cancer, pancreatic cancer, upper gastrointestinal cancer, sarcoma, breast cancer, liver cancer, prostate cancer, brain cancer, and combinations thereof.

3 . The method of claim 1 , further comprising: performing initial data processing of the first set of sequence reads of nucleic acid samples from each healthy subject in the reference group of healthy subjects based on a fourth set of sequence reads of nucleic acid samples from a baseline group of healthy subjects, wherein the reference group and the baseline group do not overlap, and wherein the initial data processing comprises correction of GC biases or normalization of numbers of sequence reads that align to regions of the reference genome.

4 . The method of claim 1 , further comprising: performing initial data processing of the sequence reads of nucleic acid samples from each subject in the training group based on a fourth set of sequence reads of nucleic acid samples from a baseline group of healthy subjects, wherein the baseline group and the training group do not overlap, and wherein the initial data processing comprises correction of GC biases or normalization of numbers of sequence reads aligned to regions of the reference genome.

5 . The method of claim 1 , wherein the quantity data consists of one quantity corresponding to a total number of sequence reads that align to the low variability region.

6 . The method of claim 1 , wherein the quantity data comprises multiple quantities each corresponding to a subset of the sequence reads that align to the low variability region, wherein each sequence read within a same subset corresponds to nucleic acid samples having a same predetermined fragment size or size range, wherein sequence reads in different subsets correspond to nucleic acid samples having a different fragment size or size range.

7 . The method of claim 1 , wherein the one or more parameters are determined by principal component analysis (PCA).

8 . The method of claim 1 , further comprising: refining the one or more parameters in a multi-fold cross-validation process by dividing the second filtered training set of sequence reads into a filtered training subset and a filtered validation subset.

9 . The method of claim 8 , wherein the filtered training and validation subsets in one fold of the multi-fold cross-validation process are different from another filtered training and validation subset in another fold of the multi-fold cross-validation process.

10 . The method of claim 1 , wherein the high variability region in the reference genome corresponds to a plurality of regions in the reference genome that exhibit variability above the predetermined threshold and wherein each of the plurality of regions has the same size.

11 . The method of claim 1 , wherein the high variability region in the reference genome includes a plurality of high variability regions that correspond to a plurality of regions in the reference genome that exhibit variability above the predetermined threshold and wherein each of the plurality of regions do not have the same size.

12 . The method of claim 1 , wherein the one or more parameters are determined based on a subset of the training set of sequence reads.

13 . The method of claim 1 , wherein the nucleic acid samples from the subjects in the training group comprise cfDNA fragments that are longer than the predetermined threshold length, wherein the predetermined threshold length is less than 160 nucleotides.

14 . The method of claim 13 , wherein the predetermined threshold length is 140 nucleotides or less.

15 . The method of claim 13 , wherein the sequence reads in the training set includes sequence reads of cfDNA fragments in the nucleic acid samples from the subjects in the training group having a length falling between a second threshold length and a third threshold length, wherein: the second threshold length is from 240 to 260 nucleotides, and the third threshold length is from 290 nucleotides to 310 nucleotides.

16 . A computer system comprising: one or more processors; and a non-transitory computer-readable medium including one or more sequences of instructions that, when executed by the one or more processors, cause the processors to

receive, at a data collection component of the computer system, a first set of sequence reads of nucleic acid samples from each healthy subject in a reference group of healthy subjects, wherein the nucleic acid samples comprise cell-free DNA (cfDNA) fragments and wherein the first set of sequence reads is derived from a DNA sequencing process and wherein the first set of sequence reads are not known to contain one or more mutations indicative of a cancer condition;

align, using the one or more processors, each sequence read in the first set of sequence reads to regions in a reference genome;

establish, using the one or more processors and subsequent to the aligning, a low variability filter, wherein the instructions to establish comprise instructions to:

identify, using the one or more processors, numbers of sequence reads aligned to each of the regions in the reference genome across each healthy subject in the reference group of healthy subjects;

derive, using the one or more processors and based on the identifying, quantity data associated with each of the regions based on the numbers of sequence reads;

calibrate, using the one or more processors, the quantity data for each of the regions;

compute, using the one or more processors, reference quantities for each of the regions based on the calibrated quantity data, wherein the reference quantities comprise at least a first reference quantity and a second reference quantity, wherein the first reference quantity corresponds to an average of the calibrated quantity data and wherein the second reference quantity is a standard deviation of the calibrated quantity data;

determine, using the one or more processors, a difference between the first reference quantity and the second reference quantity for each of the regions;

classify, using the one or more processors, each of the regions as a high variability region or a low variability region by comparing the difference against a predetermined threshold, wherein the high variability region is defined by the difference being greater than the predetermined threshold and wherein the low variability region is defined by the difference between less than the predetermined threshold;

receive, at the data collection component, a training set of sequence reads from subjects in a training group, wherein the training set includes a second set of sequence reads of nucleic acid samples from healthy subjects and a third set of sequence reads of nucleic acid samples from cancer subjects who are known to have a cancer condition, wherein the second set of sequence reads of nucleic acid samples are not known to contain the one or more mutations indicative of the cancer condition and wherein the third set of sequence reads of nucleic acid samples do contain the one or more mutations indicative of the cancer condition and wherein the training set of sequence reads from the subjects in the training group comprise cfDNA fragments of varying length;

apply, using the one or more processors, the established low variability filter to the training set of sequence reads;

generate, using the one or more processors and based on the applying, a filtered training set of sequence reads, wherein the instructions to generate the filtered training set of sequence reads comprise instructions to discard one or more sequence reads in the filtered training set of sequence reads that are aligned to the high variability region in the reference genome;

training, using the processor, a machine learning model on the filtered training set of sequence reads;

applying, using the processor, a predictive capability of the trained machine learning model to identify differences between the second set of sequence reads of nucleic acid samples from the healthy subjects and the third set of sequence reads of nucleic acid samples from the cancer subjects.

17 . A non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor of a computer system, cause the computer system to perform a method comprising:

receiving, at a data collection component of a computer system, a first set of sequence reads of nucleic acid samples from each healthy subject in a reference group of healthy subjects, wherein the nucleic acid samples comprise a first set of cell-free DNA (cfDNA) fragments and wherein the first set of sequence reads is derived from a DNA sequencing process and wherein the first set of sequence reads are not known to contain one or more mutations indicative of a cancer condition;

aligning, using a processor of the computer system, each sequence read in the first set of sequence reads to regions in a reference genome;

establishing, using the processor and subsequent to the aligning, a low variability filter, wherein the establishing comprises:

identifying, using the processor, numbers of sequence reads aligned to each of the regions in the reference genome across each healthy subject in the reference group of healthy subjects;

deriving, using the processor and based on the identifying, quantity data associated with each of the regions based on the numbers of sequence reads;

calibrating, using the processor, the quantity data for each of the regions;

computing, using the processor, reference quantities for each of the regions based on the calibrated quantity data, wherein the reference quantities comprise at least a first reference quantity and a second reference quantity, wherein the first reference quantity corresponds to an average of the calibrated quantity data and wherein the second reference quantity is a standard deviation of the calibrated quantity data;

determining, using the processor, a difference between the first reference quantity and the second reference quantity for each of the regions;

classifying each of the regions as a high variability region or a low variability region by comparing the difference against a predetermined threshold, wherein the high variability region is defined by the difference being greater than the predetermined threshold and wherein the low variability region is defined by the difference between less than the predetermined threshold;

receiving, at the data collection component, a training set of sequence reads from subjects in a training group, wherein the training set includes a second set of sequence reads of nucleic acid samples from healthy subjects and a third set of sequence reads of nucleic acid samples from cancer subjects who are known to have a cancer condition, wherein the second set of sequence reads of nucleic acid samples are not known to contain the one or more mutations indicative of the cancer condition and wherein the third set of sequence reads of nucleic acid samples do contain the one or more mutations indicative of the cancer condition and wherein the training set of sequence reads from the subjects in the training group comprise cfDNA fragments of varying length;

applying, using the processor, the established low variability filter to the training set of sequence reads;

generating, using the processor and based on the applying, a filtered training set of sequence reads, wherein the generating the filtered training set of sequence reads comprises discarding one or more sequence reads in the filtered training set of sequence reads that are aligned to the high variability region in the reference genome;

training, using the processor, a machine learning model on the filtered training set of sequence reads;

applying, using the processor, a predictive capability of the trained machine learning model to identify differences between the second set of sequence reads of nucleic acid samples from the healthy subjects and the third set of sequence reads of nucleic acid samples from the cancer subjects.

Assignments (3)
CHANGE OF NAME Recorded Feb 13, 2025
From: GRAIL, LLC
To: GRAIL, INC.
Reel/Frame 070208/0555 →
MERGER AND CHANGE OF NAME Recorded Oct 13, 2021
From: GRAIL, INC.; SDG OPS, LLC
To: GRAIL, LLC
Reel/Frame 057788/0719 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2019
From: FILIPPOVA, DARYA; VALOUEV, ANTON; NICULA, VIRGIL; JAGADEESH, KARTHIK; MAHER, M. CYRUS; LARSON, MATTHEW H.; PORTELA DOS SANTOS PIMENTEL, MONICA; CALEF, ROBERT ABE PAINE
To: GRAIL, INC.
Reel/Frame 048942/0137 →
Continuity (2)
Provisional Application 62642461 · Mar 13, 2018
Related Publication 20190287649A1 · Sep 19, 2019
References Cited (182)
US 6466923B1 · Young · 2002 [cited by applicant]
US 8642349B1 · Yeatman et al. · 2014 [cited by applicant]
US 8706422B2 · Lo et al. · 2014 [cited by applicant]
US 9260745B2 · Rava et al. · 2016 [cited by applicant]
US 9373059B1 · Heifets et al. · 2016 [cited by applicant]
US 10095831B2 · Duenwald et al. · 2018 [cited by applicant]
US 10457995B2 · Talasaz · 2019 [cited by applicant]
US 10741269B2 · Chudova · 2020 [cited by examiner]
US 20030215120A1 · Uppaluri et al. · 2003 [cited by applicant]
US 20100112590A1 · Lo et al. · 2010 [cited by applicant]
US 20100145894A1 · Semizarov et al. · 2010 [cited by applicant]
US 20120053253A1 · Stone et al. · 2012 [cited by applicant]
US 20130034546A1 · Rava et al. · 2013 [cited by applicant]
US 20130325360A1 · Deciu et al. · 2013 [cited by applicant]
US 20140066317A1 · Talasaz · 2014 [cited by applicant]
US 20140080715A1 · Lo et al. · 2014 [cited by applicant]
US 20140100121A1 · Lo et al. · 2014 [cited by applicant]
US 20140242588A1 · Van Den Boom et al. · 2014 [cited by applicant]
US 20140371078A1 · Abdueva · 2014 [cited by applicant]
US 20150102216A1 · Röder et al. · 2015 [cited by applicant]
US 20150220838A1 · Martin et al. · 2015 [cited by applicant]
US 20160002739A1 · Schutz et al. · 2016 [cited by applicant]
US 20160232290A1 · Rava et al. · 2016 [cited by applicant]
US 20160251704A1 · Talasaz et al. · 2016 [cited by applicant]
US 20160364522A1 · Frey et al. · 2016 [cited by applicant]
US 20170107576A1 · Babiarz et al. · 2017 [cited by applicant]
US 20170121767A1 · Dor et al. · 2017 [cited by applicant]
US 20170220735A1 · Duenwald et al. · 2017 [cited by applicant]
US 20170240973A1 · Eltoukhy et al. · 2017 [cited by applicant]
US 20170270245A1 · van Rooyen et al. · 2017 [cited by applicant]
US 20170329893A1 · di Iulio et al. · 2017 [cited by applicant]
US 20170342477A1 · Jensen et al. · 2017 [cited by applicant]
US 20170342500A1 · Marquard et al. · 2017 [cited by applicant]
US 20170362638A1 · Chudova et al. · 2017 [cited by applicant]
US 20180046851A1 · Kienzle et al. · 2018 [cited by applicant]
US 20180144261A1 · Wnuk et al. · 2018 [cited by applicant]
US 20180148777A1 · Kirkizlar et al. · 2018 [cited by applicant]
US 20180173845A1 · Sigurjonsson et al. · 2018 [cited by applicant]
US 20190106737A1 · Underhill · 2019 [cited by applicant]
US 20190131016A1 · Cohen et al. · 2019 [cited by applicant]
US 20190164627A1 · Blocker et al. · 2019 [cited by applicant]
US 20190287646A1 · Hubbell et al. · 2019 [cited by applicant]
US 20190287649A1 · Filippova et al. · 2019 [cited by applicant]
US 20190287652A1 · Gross et al. · 2019 [cited by applicant]
US 20190316209A1 · Hubbell et al. · 2019 [cited by applicant]
US 20200005901A1 · Cohen et al. · 2020 [cited by applicant]
US 20200219587A1 · Hubbell · 2020 [cited by applicant]
US 20200294624A1 · Filippova et al. · 2020 [cited by applicant]
US 20200365229A1 · Fields et al. · 2020 [cited by applicant]
US 20200372296A1 · Maher · 2020 [cited by applicant]
US 20210125683A1 · Zhou · 2021 [cited by examiner]
US 20210324477A1 · Xiang et al. · 2021 [cited by applicant]
AU 2015360298 · 2016 [cited by applicant]
WO 2010037001A2 · 2010 [cited by applicant]
WO WO2010051318 · 2010 [cited by applicant]
WO 2011127136A1 · 2011 [cited by applicant]
WO WO2018071621 · 2012 [cited by applicant]
WO WO2013052907 · 2013 [cited by applicant]
WO WO2014149134 · 2014 [cited by applicant]
WO WO2016094853 · 2015 [cited by applicant]
WO 2016094853A1 · 2016 [cited by applicant]
WO WO2017161175 · 2017 [cited by applicant]
WO WO2017181202 · 2017 [cited by applicant]
WO WO2017212428 · 2017 [cited by applicant]
WO 2018031929A1 · 2018 [cited by applicant]
WO WO2018022890 · 2018 [cited by applicant]
WO WO2018022906 · 2018 [cited by applicant]
WO 2018009723A1 · 2018 [cited by applicant]
WO WO2019084559 · 2019 [cited by applicant]
WO WO2019232435 · 2019 [cited by applicant]
WO WO2021119471 · 2021 [cited by applicant]
Robinson, Peter N., Marten Jager, and Rosario Michael Piro. Computational Exome and Genome Analysis. First edition. Boca Raton, FL: CRC Press, 2017. Web.(2017) (Year: 2017). [cited by examiner]
Clark, Travis A. et al. “Analytical Validation of a Hybrid Capture-Based Next-Generation Sequencing Clinical Assay for Genomic Profiling of Cell-Free Circulating Tumor DNA.” The Journal of molecular diagnostics: JMD 20.… [cited by examiner]
Xia, Ligang et al. “Statistical Analysis of Mutant Allele Frequency Level of Circulating Cell-Free DNA and Blood Cells in Healthy Individuals.” Scientific reports 7.1 (2017): 7526-7. Web and Supplemental (Year: 2017). [cited by examiner]
Gunasegaran, Thineswaran, and Yu.-N Cheah. “Evolutionary Cross Validation.” 2017 8th International Conference on Information Technology (ICIT). IEEE, 2017. 89-95. Web.(2017) (Year: 2017). [cited by examiner]
Odegaard, Justin I et al. “Validation of a Plasma-Based Comprehensive Cancer Genotyping Assay Utilizing Orthogonal Tissue- and Plasma-Based Methodologies.” Clinical cancer research 24.15 (2018): 3539-3549. Web. (Year: 2… [cited by examiner]
Minarik, Gabriel et al. Utilization of Benchtop Next Generation Sequencing Platforms Ion Torrent PGM and MiSeq in Noninvasive Prenatal Testing for Chromosome 21 Trisomy and Testing of Impact of In Silico and Physical Si… [cited by examiner]
Kersaudy-Kerhoas, Maïwenn, and Elodie Sollier. “Micro-Scale Blood Plasma Separation: From Acoustophoresis to Egg-Beaters.” Lab on a chip 13.17 (2013): 3323-3346. Web. (Year: 2013). [cited by examiner]
Ulz, P., Thallinger, G. G., Auer, M., Graf, R., Kashofer, K., Jahn, S. W., Abete, L., Pristauz, G., Petru, E., Geigl, J. B., Heitzer, E., & Speicher, M. R. (2016). Inferring expressed genes by whole-genome sequencing of… [cited by examiner]
Smedley, D., Schubach, M., Jacobsen, . (2016). A Whole-Genome Analysis Framework for Effective Identification of Pathogenic Regulatory Variants in Mendelian Disease. American journal of human genetics, 99(3), 595-606. (… [cited by examiner]
Hudecova, I., & Chiu, R. W. (2017). Non-invasive prenatal diagnosis of thalassemias using maternal plasma cell free DNA. Best practice & research. Clinical obstetrics & gynaecology, 39, 63-73. (Year: 2017). [cited by examiner]
Enyedi, Márton Zsolt et al. “Simultaneous detection of BRCA mutations and large genomic rearrangements in germline DNA and FFPE tumor samples.” Oncotarget 7.38 (2016): 61845-61859. Web. (Year: 2016). [cited by examiner]
Leek and Storey, 2007, “Capturing Heterogeneity in Gene Expression Studies by Surrogate Variable Analysis,” PLoS Genet 3, pp. 1724-1735. [cited by applicant]
Rumelhart et al., 1988, “Neurocomputing: Foundations of research,” ch. Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press. [cited by applicant]
Zeiler, 2012 “ADADELTA: an adaptive learning rate method,” CoRR, vol. abs/212.5701. [cited by applicant]
U.S. Appl. No. 62/679,347, filed Jun. 1, 2018. [cited by applicant]
“U.S. Appl. No. 62/847,223, entitled Model-Based Featurization and Classification,” filed May 13, 2019. [cited by applicant]
U.S. Appl. No. 62/851,486, entitled “Systems and Methods for Determining Whether a Subject Has a Cancer Condition Using Transfer Learning,” filed May 22, 2019. [cited by applicant]
U.S. Appl. No. 62/827,682, entitled “Systems and Methods for Using Fragment Lengths as a Predictor of Cancer,” filed Apr. 1, 2019. [cited by applicant]
Agresti, “An Introduction to Categorical Data Analysis” Chapter 5, John Wiley & Son, New York, 103-144, 1996. [cited by applicant]
Alkan, et al. Nat Genet 41, p. 1061-7, 2009. [cited by applicant]
Angermueller, Christof, et al. “DeepCpG: accurate prediction of single-cell DNA methylation states using deep learning”, Genome Biology, 18(1), 2017. [cited by applicant]
Benjamin, et al., 2012, “Summaraizing and Correcting the GC Content Bias in High-Throughput Sequencing” Nucleic Acids Research, vol. 40, Issue 10, 1-14. [cited by applicant]
Boeva, et al. “Control-Free Calling of Copy Number Alternations in Deep-Sequencing Data using GC-Content Normalization”, Bioinformatics, 27(2), 266-269. [cited by applicant]
Boser, et al., “A training algorithm for optimal margin classifiers”, Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., 142-152, 1992. [cited by applicant]
Casadio, et al. “Urine cell-free DNA integrity as a marker for early bladder cancer diagnosis: preliminary data,” Urol Oncol. 31(8), 1744-1750. [cited by applicant]
Chan, et al. “Clinical Sciences Reviews Committee of the Association of Clinical Biochemists Cell-free nucleic acids in plasma, serum and urine: a new tool in molecular diagnosis”. [cited by applicant]
Chen, et al., “Conflict of CpG density and DNA methylation are proximally and distally involved in gene regulation in human and mouse tissues”, Epgenetics 13(7), 721-741, 2018. [cited by applicant]
De Mattos-Arruda and Caldas, 2016, “Cell-free circulating tumour DNA as a liquid biopsy in breast cancer,” Mol Oncol. 2016;10(3):464-474. [cited by applicant]
Du, et al., “Comparison of Beta-value and M-value methods for quantifying methylation levels by microarray analysis”, BMC Bioinformatics 11, 587, 2010. [cited by applicant]
Duda, “Pattern Classification”, Second Edition, John Wiley & Sons, Inc., 259, 262-265. [cited by applicant]
Fabregat, et al., “The Reactome Pathway Knowledgebase”, Nucleic Acids Res., 46, D649-D655, 2018. [cited by applicant]
Frenel et al., 2015, Serial next-generation sequencing of circulating cell-free DNA evaluating tumor clone response to molecularly targeted drug administration. Clin Cancer Res. 21(20):4586-4596. [cited by applicant]
Furey, et al., “Support vector machine classification and validation of cancer tissue samples using microarray expression data”, Bioinformatics 16, 906-914, 2000. [cited by applicant]
Goessl et al., “Fluorescent methylation-specific polymerase chain reaction for DNA-based detection of prostate cancer in bodily fluids,” Cancer Res. 2000;60(21):5941-5945. [cited by applicant]
Gold, “Softmax to Softassign: Neural Network Algorithms for Combinatorial Optimization”, Journal of Artificial Neural Networks 2, 381-399, 1996. [cited by applicant]
Hao et al., “Circulating cell-free DNA in serum as a biomarker for diagnosis and prognostic prediction of colorectal cancer,” Br J Cancer. 2014; 111(8): 1482-1489. [cited by applicant]
Hassoun, “Fundamentals of Artificial Neural Networks”, Massachusetts Institute of Technology, 1995. [cited by applicant]
Hastie, The Elements of Statistical Learning, Springer, New York, 2001. [cited by applicant]
Heitzer et al., 2013, “Establishment of tumor-specific copy number alterations from plasma DNA of patients with cancer,” Int J Cancer. 133(2):346-356. [cited by applicant]
Heitzer et al., 2015, “Circulating tumor DNA as a liquid biopsy for cancer,” Clin Chem. 61(1):112-123. [cited by applicant]
Hoadley, et al. “Multi-platform analysis of 12 cancer types reveals molecular classification within and across tissues-of-origin”, Cell, 158(4). 929-944, 2014. [cited by applicant]
Jaccard, “Étude comparative de la distribution florale dans une portion des Alpes et des Jura”, Bulletin de la Société vaudoise des sciences naturelles, 37 547-579, 1901. [cited by applicant]
Jaccard, “The Distribution of the flora in the alpine zone” New Phytologist, 11, 37-50, 1912. [cited by applicant]
Jenkinson, et al.,“Potential energy landscapes identify the information-theoretic nature of the epigenome”, Nat. Genet. 49(5), 719-729, 2017. [cited by applicant]
Kanehisa, et al., KEGG: Kyoto Encyclopedia of Genes and Genomes, Nucleic Acids Res., 28(1), 27-30, 2000. [cited by applicant]
Kim et al., 2014, “Circulating cell-free DNA as a promising biomarker in patients with gastric cancer: diagnostic validity and significant reduction of clDNA after surgical resection,” Ann Surg Treat Res. 2014;86(3): 13… [cited by applicant]
Kosub “A note on the triangle inequality for the Jaccard distance” arXiv:1612.02696. [cited by applicant]
Larochelle, et al., “Exploring strategies for training deep neural networks”, J Mach Learn Res 10, 1-40, 2009. [cited by applicant]
Levandowsky, et al., “Distance between sets”, Nature 234(5), 34-35, 1971. [cited by applicant]
Li, et al., “Mapping short DNA sequencing reads and calling variants using mapping quality scores”, Genome Research, 18, 1852-8, 2008. [cited by applicant]
Lipkus, “A proof of the triangle inequality for the Tanimoto distance”, Journal of Mathematical Chemistry, 26 (1-3), 263-265, 1999. [cited by applicant]
Liu, et al., “Bisulfite-free direct detection of 5-methylcytosine and 5-hydroxymethylcytosine at base resolution”, Nature Biotechnology 37, 424-429, 2019. [cited by applicant]
Lo et al., 2010, “Maternal plasma DNA sequencing reveals the genome-wide genetic and mutational profile of the fetus,” Sci Transl Med. 2(61):6lra91. [cited by applicant]
Miller, et al. “ReadDepth: A parallel R package for detecting copy numer alterations from short sequencing reads” PLOS One, vol. 6, 1-7, 2011. [cited by applicant]
Moulton, et al., “Maximally Consistent Sampling and the Jaccard Index of Probability Distributions”, International Conference on Data Mining, Workshop on High Dimensional Data Mining, 2018. [cited by applicant]
Nuel, “Exact distribution of a pattern in a set of random sequences generated by a Markov source: applications to biological data”, Algorithms for Molecular Biology 5(15, 2010. [cited by applicant]
Oxnard, et al. “Simultaneous Multi-cancer Detection and Tissue of Origin (TOO) Localization Using Targeted Bisulfite Sequencing of Plasma Cell-free DNA (cfDNA)”, American Society of Clinical Oncology (ASCO) Breakthrough… [cited by applicant]
Poplin, Ryan, et al. “A universal SNP and small-indel variant caller using deep neural networks” Nature Biotechnology, 36(10), 2018. [cited by applicant]
Qian, et al., “Intelligent Surveillance Systems”, Springer, 161, 2011. [cited by applicant]
Raptis and Menard, 1980, “Quantitation and characterization of plasma DNA in normals and patients with systemic lupus erythematosus,” J Clin Invest. 66(6): 1391-1399. [cited by applicant]
Rogers, et al., “A Computer Program for Classifying Plants”, Science 132(3434), 1115-1118, 1960. [cited by applicant]
Salvi et al., 2016, “Cell-free DNA as a diagnostic marker for cancer: current insights,” Onco Targets Ther. 9:6549-6559. [cited by applicant]
Shao et al. 2015 “Quantitative analysis of cell-free DNA in ovarian cancer,” Oncol Lett. 2015; 10(6):3478-3482. [cited by applicant]
Shapiro et al., 1983, “Determination of circulating DNA levels in patients with benign or malignant gastrointestinal disease,” Cancer. 51(11):2116-2120. [cited by applicant]
Siegel et al., 2015, “Cancer statistics,” CA Cancer J Clin. 65(1):5-29. [cited by applicant]
Sorber L. et al., J Mol Diagn., 19(1): 162-68 (2017). [cited by applicant]
Sozzi et al., 2003 “Quantification of free circulating DNA as a diagnostic marker in lung cancer,” J Clin Oncol. 21(21):3902-3908. [cited by applicant]
Stroun et al., “Neoplastic characteristics of the DNA found in the plasma of cancer patients,” Oncology. 1989;46(5):318-322). [cited by applicant]
Tan, et al., “Introduction to Data Mining”, 2005. [cited by applicant]
Tanimoto, et al., “An Elementary Mathematical theory of Classification and Prediction”, Internal IBM Technical Report, 1958. [cited by applicant]
Terry et al., 2016, “A prospective evaluation of early detection biomarkers for ovarian cancer in the European EPIC cohort,” Clin Cancer Res. Apr. 8, 2016. [cited by applicant]
Vapnik, “Statistical Learning Theory”, Wiley, New York, 1998. [cited by applicant]
Vincent, et al., “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion”, J Mach Learn Res, 11, 3371-3408, 2010. [cited by applicant]
Yang, et al., “DistAl: An Inter-pattern Distance-based Constructive Learning Algorithm”, Intelligent Data Analysis, 3(1), 55-83, 1999. [cited by applicant]
Yoon et al., 2009, Genome Research 19(9): 1586. [cited by applicant]
Yu, et al., “GaussianCpG: a Gaussian model for detection of CpG island in human genome sequences”, BMC Genomics 18(4), 392, 2017. [cited by applicant]
Zhang, et al., “A novel method to quantify local CpG methylation density by regional methylation elongation assay on microarray”, BMC Genomics 9:59, 9 pages, 2008. [cited by applicant]
Zhang, et al. “Tumor markers CA19-9, CA242 and CEA in the diagnosis of pancreatic cancer: a meta-analysis” Int J Clin Exp Med. 8(7), 11683-11691, 2015. [cited by applicant]
Zonta et al., “Assessment of DNA integrity, applications for cancer research,” Adv Clin Chem. 2015;70: 197-246. [cited by applicant]
U.S. Appl. No. 62/642,507, filed Mar. 13, 2018, and entitled “Identifying Copy Number Aberrations” (67 pages). [cited by applicant]
Breiman, “Random Forests—Random Features”, Technical Report 567, Sep. 1999, Statistics Department, U.C. Berkeley, pp. 1-29. [cited by applicant]
Nancy R. Cook, “Use and Misuse of the Receiver Operating Characteristic Curve in Risk Prediction”, vol. 115, Issue 7, Feb. 20, 2007, pp. 928-935. [cited by applicant]
Richard O. Duda et al., “Pattern Classification and Scene Analysis”, John Wiley & Sons, Inc., New York, Published in “A Wiley-Interscience . . . ”, Sep. 1, 1974, Computer Science, 8 pages. [cited by applicant]
Clark et al., Analytical Validation of a Hybrid Capture-Based Next-Generation Sequencing Clinical Assay for Genomic Profiling of Cell-Free Circulating Tumor DNA, The Journal of Molecular Diagnostics, Elsevier, 2018 20(5… [cited by applicant]
Gunasegaran et al., Evolutionary Cross Validation, 2017 8th International Conference on Information Technology (ICIT), IEEE, 2017, pp. 89-95. [cited by applicant]
Kersaudy-Kerhoas et al., Micro-scale blood plasma separation: from acoustophoresis to egg-beaters, RSC Publishing, Lab on a Chip, 2013 13(17), pp. 3323-3346. [cited by applicant]
Minarik et al., Utilization of Benchtop Next Generation Sequencing Platforms Ion Torrent PGM and MiSeq in Noninvasive Prenatal Testing for Chromosome 21 Trisomy and Testing of Impact of In Silico and Physical Size Selec… [cited by applicant]
Odegaard et al., Validation of a Plasma-Based Comprehensive Cancer Genotyping Assay Utilizing Orthogonal Tissue- and Plasma-Based Methodologies, 2018, pp. 3539-3549, American Association for Cancer Research. [cited by applicant]
Robinson et al., Computational Exome and Genome Analysis, First edition, Boca Raton, FL: CRC Press, 2017, 575 pages. [cited by applicant]
Xia, Ligang et al., Statistical Analysis of Mutant Allele Frequency Level of Circulating Cell-Free DNA and Blood Cells in Healthy Individuals, Scientific Reports 7.1, pp. 2017, 7526-7527. [cited by applicant]
Kang et al., “CancerLocator: non-invasive cancer diagnosis and tissue-of-origin prediction using methylation profiles of cell-free DNA,” Genome Biology, 2017, vol. 18, No. 53: pp. 1-12. [cited by applicant]
Changhong Shan, U.S. Appl. No. 62/642,506, filed Mar. 13, 2018, “Single Radio Voice Call Continuity (SRVCC) Handover and Return to Next Generation Radio Access Network (NG-RAN) After SRVCC” (54 pages). [cited by applicant]
I. T. Jolliffe, “Principal Component Analysis”, Springer, New York © 2002, 518 pages. [cited by applicant]
Steven Kothen-Hill et al., “Deep Learning Mutation Prediction Enables Early Stage Lung Cancer Detection in Liquid Biopsy”, Workshop track—ICLR 2018, 24 pages. [cited by applicant]
Aengus S. O'Marcaigh, et al., “Estimating the Predictive Value of a Diagnostic Test: How to Prevent Misleading or Confusing Results, Clinical Pediatrics”, vol. 32, Issue 8, First Published Aug. 1, 1993, 7 pages. [cited by applicant]
Margaret Sullivan Pepe, et al., “Limitations of the Odds Ration in Gauging the Performance of a Diagnostic, Prognostic, or Screening Marker”, American Journal of Epidemiology, American Journal of Epidemiology, vol. 159,… [cited by applicant]
Alkes L. Price, et al., “Principal components analysis corrects for stratification in genome-wide association studies”, vol. 38 No. 8, Aug. 2006, Nature Genetics, pp. 904-909. [cited by applicant]
Swanton, et al., “Phylogenetic ctDNA analysis depicts early stage lung cancer evolution”, Nature. Apr. 26, 2017, Author manuscript; available in PMC Feb. 14, 2018; 545(7655), 446-451, 47 pages. [cited by applicant]
Yuchen Yuan et al., “DeepGene: an advanced cancer type classifier based on deep learning and somatic point mutations”, Dec. 2016, BMC Bioinformatics 17(Suppl 17):476, 15 pages. [cited by applicant]
Zhao, et al., “Detection of fetal subchromosomal abnormalities by sequencing circulating cell-free DNA from maternal plasma”, Clinical Chemistry, vol. 61, Issue 4, Apr. 1, 2015, pp. 608-616. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2019/022139, mailed Jul. 11, 2019, 18 pages. [cited by applicant]
U.S. Appl. No. 62/818,013, filed Mar. 13, 2019, and entitled “Systems and Methods for Enriching for Cancer-Derived Fragments Using Fragment Size”. [cited by applicant]
U.S. Appl. No. 62/679,347, filed Jun. 1, 2018, and entitled “Models for Targeted Sequencing”. [cited by applicant]
Erickson R, “Somatic gene mutation and human disease other than cancer: An update”, 2010, Mutat Res.;705 (2):96-106. [cited by applicant]
Erickson, “Somatic gene mutation and human disease other than cancer”, 2003, Mutat Res., 543(2): pp. 125-136. [cited by applicant]
Trevor Hastie et al., “Chapter 9: Additive Models, Trees, and Related Models”, The Elements of Statistical Learning, Second Edition, Springer Science+Business Media, LLC, 2009, pp. 295-336. [cited by applicant]
Alan Agresti, “Building and Applying Logistic Regression Models”, An Introduction to Categorical Data Analysis, Second Edition, 2006, pp. 137-172. [cited by applicant]
Nello Cristianini et al., “An Introduction to Support Vector Machines and other kernel-based learning methods”, Cambridge University Press, First Published 2000, pp. 93-124. [cited by applicant]
Karen Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, Published as a conference paper at ICLR 2015, pp. 1-14. [cited by applicant]
International Search Report and Written Opinion mailed Sep. 13, 2019 in International Application No. PCT/US2019/034994 (20 pages). [cited by applicant]
International Search Report and Written Opinion mailed Mar. 9, 2021 in International Application No. PCT/US2020/064577 (13 pages). [cited by applicant]