IP Library Granted Patent US 12,191,000
Granted Patent B2
US 12,191,000 · App. 18/151,197 · Granted Jan 7, 2025

Systems and methods for classifying patients with respect to multiple cancer classes

Inventors: M. Cyrus Maher (San Mateo, CA); Anton Valouev (Palo Alto, CA); Darya Filippova (Sunnyvale, CA); Virgil Nicula (Cupertino, CA); Karthik Jagadeesh (San Francisco, CA); Oliver Claude Venn (San Francisco, CA); Samuel S. Gross (Sunnyvale, CA); John F. Beausang (Menlo Park, CA); Robert Abe Paine Calef (Redwood City, CA)
Assignee: GRAIL, INC.
G16B30/00G06N5/04G06N20/00G16B20/20G16B40/00G16H10/40G16H10/60G16H50/20G16H50/70G16H70/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,191,000
App. No.
18/151,197
Granted
Jan 7, 2025
Kind
B2
Abstract

Technical solutions for classifying patients with respect to multiple cancer classes are provided. The classification can be done using cell-free whole genome sequencing information from subjects. A reference set of subjects is used to train classifiers to recognize genomic markers that distinguish such cancer classes. The classifier training includes dividing the reference genome into a set of non-overlapping bins, applying a dimensionality reduction method to obtain a feature set, and using the feature set to train classifiers. For subjects with unknown cancer class, the trained classifiers provide probabilities or likelihoods that the subject has a respective cancer class for each cancer in a set of cancer classes. The present disclosure thus describes methods to improve the screening and detection of cancer class from among several cancer classes. This serves to facilitate early and appropriate treatment for subjects afflicted with cancer.

Claims (84)

1. A method of training an untrained first classifier to classify a test subject of a given species to a cancer class in a plurality of cancer classes using a computer system comprising one or more processors, the method comprising:

obtaining, by the computer system and for each respective reference subject in a first plurality of reference subjects, (i) a cancer class of the respective reference subject and (ii) a sequencing construct for the respective reference subject that includes a first bin count for each respective bin in a plurality of bins that collectively represent all or a portion of a reference genome of the species, wherein each respective first bin count representative of a number of nucleic acid fragments measured from nucleic acids in a biological sample obtained from the respective reference subject that maps onto a different and non-overlapping portion of the reference genome of the species, wherein, for each respective cancer class in the plurality of cancer classes, the first plurality of reference subjects includes at least one reference subject that has the respective cancer class;

improving a computational efficiency of the computer system by collectively subjecting, by the computer system, the first bin count of each bin in the plurality of bins for each reference subject in the first plurality of reference subjects to a dimensionality reduction method thereby obtaining a feature set, wherein the feature set consists of a number of features that is fewer than the number of bins in the plurality of bins;

resampling, using the computer system, the feature set a plurality of times, wherein the resampling comprises forming, for each respective training iteration in a plurality of training iterations, a trained component classifier via:

omitting, from the feature set, a subset of values for features in the feature set for the first plurality of reference subjects; and

forming the trained component classifier by inputting, in conjunction with the cancer class of respective reference subjects in the first plurality of reference subjects as ground truth, remaining values for the features in the feature set as collective input to a respective untrained component classifier; and

constructing, as a result of the resampling and by collectively leveraging output generated by the trained component classifier formed in each of the plurality of training iterations, a trained first classifier having an improved cancer class recognition ability over the untrained first classifier.

2. The method of claim 1 , wherein the sequencing construct for each respective reference subject in the first plurality of reference subjects is obtained by targeted panel or whole genome sequencing.

3. The method of claim 1 , wherein each respective reference subject in the first plurality of reference subjects is human.

4. The method of claim 1 , wherein the plurality of cancer classes is at least two or more cancer classes selected from the group consisting of non-cancer, bladder cancer, brain cancer, breast cancer, colorectal cancer, endometrial cancer, esophageal cancer, head/neck cancer, kidney cancer, liver cancer, hematological cancer, lung cancer, a lymphoma, leukemia, a melanoma, a lymphoma, ovarian cancer, pancreatic cancer, prostate cancer, rectal cancer, renal cancer, thyroid cancer and uterine cancer.

5. The method of claim 1 , wherein:

the biological sample or methylation biological sample obtained from the respective reference subject is a whole blood sample from the respective reference subject, and

the nucleic acids in the biological sample or methylation biological sample obtained from the respective reference subject are genomic DNA.

6. The method of claim 1 , wherein:

the biological sample or methylation biological sample obtained from the respective reference subject is a plasma sample from the respective reference subject.

7. The method of claim 1 , wherein the dimensionality reduction method further comprises application of principal component analysis using the cancer class and the sequencing construct of each reference subject in the first plurality of reference subjects thereby identifying the feature set comprising a plurality of principal components.

8. The method of claim 1 , further comprising regularizing the first plurality of reference subjects set using the cancer class and the sequencing construct of each reference subject in the first plurality of reference subjects.

9. The method of claim 1 , further comprising scaling the first bin count for each respective bin in the plurality of bins for each respective reference subject in the first plurality of reference subjects by:

taking a log transformation of the respective first bin count thereby forming a log transformed first bin count for the respective bin,

subtracting a mean value of the respective log transformed first bin count across the first plurality of reference subjects from the log transformed first bin count of the respective bin thereby forming a first normalized first bin count for the respective bin; and

dividing, subsequent to the subtracting, the respective first normalized first bin count for the respective bin by a standard deviation of the first normalized bin first count across the first plurality of reference subjects thereby scaling the first bin count for each respective bin in the plurality of bins for each respective reference subject in the first plurality of reference subjects.

10. The method of claim 1 , further comprising estimating a performance of the first trained classifier as an average performance of the trained component classifier formed in each respective iteration in the plurality of training iterations.

11. The method of claim 1 , wherein:

the sequencing construct for the respective reference subject further includes a second bin count for each respective bin in the plurality of bins, each respective second bin count representative of a number of nucleic acid fragments that are in a second size range that were measured from nucleic acids in the biological sample obtained from the respective reference subject that maps onto the different and non-overlapping portion of the reference genome,

each respective first bin count representative of a number of nucleic acid fragments that are in a first size range that were measured from nucleic acids in the biological sample obtained from the respective reference subject that maps onto the different and non-overlapping portion of the reference genome,

the collectively subjecting step further provides the second bin count of each bin in each respective plurality of bins across the first plurality of reference subjects to the dimensionality reduction method thereby obtaining the feature set, and

the first size range is different than the second size range.

12. The method of claim 1 , wherein:

the sequencing construct for the respective reference subject includes a respective set of bin counts for each respective bin in the plurality of bins, wherein the respective set of bin counts includes the first bin count, and wherein each respective bin count in the respective set of bin counts is representative of a number of nucleic acid fragments that are in a size range corresponding to the respective bin count that were measured from nucleic acids in the biological sample obtained from the respective reference subject that maps onto the different and non-overlapping portion of the reference genome, and

the collectively subjecting step provides the respective set of bin counts of each bin in the plurality of bins across the first plurality of reference subjects to the dimensionality reduction method thereby obtaining the feature set, and wherein

the respective set of bins includes at least three different bin counts, and wherein each bin count in the respective set of bin counts corresponds to a different size range.

13. A method of classifying a test subject of a given species to a cancer class in a plurality of cancer classes using a computer system comprising one or more processors, the method comprising:

using a trained first classifier to classify the test subject to a cancer class in the plurality of cancer classes using nucleic acid fragments in a biological sample obtained from the test subject, wherein the trained first classifier having been trained via:

obtaining, by the computer system and for each respective reference subject in a first plurality of reference subjects, (i) a cancer class of the respective reference subject and (ii) a sequencing construct for the respective reference subject that includes a first bin count for each respective bin in a plurality of bins that collectively represent all or a portion of a reference genome of the species, wherein each respective first bin count representative of a number of nucleic acid fragments measured from nucleic acids in a biological sample obtained from the respective reference subject that maps onto a different and non-overlapping portion of the reference genome of the species, wherein, for each respective cancer class in the plurality of cancer classes, the first plurality of reference subjects includes at least one reference subject that has the respective cancer class;

improving a computational efficiency of the computer system by collectively subjecting, by the computer system, the first bin count of each bin in the plurality of bins for each reference subject in the first plurality of reference subjects to a dimensionality reduction method thereby obtaining a feature set, wherein the feature set consists of a number of features that is fewer than the number of bins in the plurality of bins;

resampling, using the computer system, the feature set a plurality of times, wherein the resampling comprises forming, for each respective training iteration in a plurality of training iterations, a trained component classifier via:

omitting, from the feature set, a subset of values for features in the feature set for the first plurality of reference subjects; and

forming the trained component classifier by inputting, in conjunction with the cancer class of respective reference subjects in the first plurality of reference subjects as ground truth, remaining values for the features in the feature set as collective input to a respective untrained component classifier; and

constructing, as a result of the resampling and by collectively leveraging output generated by the trained component classifier formed in each of the plurality of training iterations, a trained first classifier having an improved cancer class recognition ability over the untrained first classifier.

14. The method of claim 13 , wherein:

the using step combines a first call on cancer class made by the trained first classifier for the test subject with a second call made by a trained second classifier,

input to the second trained classifier for the second call comprises a methylation pattern measured in a methylation biological sample obtained from the test subject, and

the trained second classifier having been trained using a respective methylation pattern measured in a respective reference methylation biological sample from each reference subject in a second plurality of reference subjects.

15. The method of claim 14 , wherein the test subject is deemed to have a first cancer class in the plurality of cancer classes when both the first call and the second call identify the test subject as having the same cancer class in the plurality of cancer classes.

16. The method of claim 14 , wherein the test subject is deemed to have a first cancer class in the plurality of cancer classes when (i) the trained first classifier calls the first cancer class with a higher probability than all other cancer classes in the plurality of cancer classes and (ii) the second call identifies the test subject as having the first cancer class.

17. The method of claim 14 , wherein the test subject is deemed to have a first cancer class in the plurality of cancer classes when (i) the first trained classifier calls the first cancer class with a call that is among the top two cancer classes in the plurality of cancer classes in terms of probability and (ii) the second call identifies the test subject as having the first cancer class.

18. The method of claim 14 , wherein the test subject is deemed to have a first cancer class in the plurality of cancer classes when (i) the first trained classifier calls the first cancer class with a higher probability than all other cancer classes in the plurality of cancer classes and (ii) the second trained classifier calls the first cancer class with a higher probability than all other cancer classes in the plurality of cancer classes.

19. The method of claim 14 , wherein the test subject is deemed to have a first cancer class in the plurality of cancer classes when (i) the first trained classifier calls the first cancer class with a call that is among the top two cancer classes in the plurality of cancer classes in terms of probability and (ii) the second trained classifier calls the first cancer class with a call that is among the top two cancer classes in the plurality of cancer classes in terms of probability.

20. The method of claim 14 , wherein:

the biological sample obtained from the test subject is a plasma sample from the test subject, and

the biological sample or the methylation biological sample obtained from the respective reference subject during the training was a plasma sample from the respective reference subject.

21. The method of claim 14 , wherein:

the classifying uses the trained first classifier to classify the test subject to the cancer class in the plurality of cancer classes and an aggressiveness of the cancer class using (i) nucleic acid fragments of cell-free nucleic acids in the biological sample or the methylation biological sample obtained from the test subject and (ii) an indication of whether a predetermined genetic marker is absent or present in the test subject, and

the trained first classifier having been further trained such that:

the obtaining step further comprises, for each respective reference subject in the first plurality of reference subjects, an indication of whether the predetermined genetic marker is absent or present in the respective reference subject, and

the improving step further uses the indication of whether the predetermined genetic marker is absent or present in each respective reference subject in the first plurality of reference subjects as ground truth.

22. The method of claim 13 , wherein:

the trained first classifier is a multinomial classifier that provides a plurality of likelihoods responsive to the nucleic acid fragments obtained from cell-free nucleic acids from the test subject, wherein each respective likelihood in the plurality of likelihoods is a likelihood that the test subject has a corresponding cancer class in the plurality of cancer classes.

23. The method of claim 13 , further comprising:

identifying a recommended treatment to the test subject based upon the cancer class of the test subject determined by the first trained classifier.

24. The method of claim 13 , further comprising:

identifying a recommended first treatment to the test subject when the first trained classifier determines that the test subject has a first cancer, and

identifying a recommended second treatment to the test subject when the first trained classifier determines that the test subject has a second cancer, wherein

the first cancer and the second cancer are different cancers.

25. The method of claim 24 , wherein the first cancer and the second cancer are each independently selected from the group consisting of bladder cancer, brain cancer, breast cancer, colorectal cancer, endometrial cancer, esophageal cancer, head/neck cancer, kidney cancer, liver cancer, hematological cancer, lung cancer, a lymphoma, leukemia, a melanoma, a lymphoma, ovarian cancer, pancreatic cancer, prostate cancer, rectal cancer, renal cancer, thyroid cancer and uterine cancer.

26. A computer system comprising:

one or more processors; and

a non-transitory computer-readable medium including one or more sequences of instructions that, when executed by the one or more processors, cause the one or more processors to perform a method of training an untrained first classifier to classify a test subject of a given species to a cancer class in a plurality of cancer classes, the method comprising:

obtaining, by the computer system and for each respective reference subject in a first plurality of reference subjects, (i) a cancer class of the respective reference subject and (ii) a sequencing construct for the respective reference subject that includes a first bin count for each respective bin in a plurality of bins that collectively represent all or a portion of a reference genome of the species, wherein each respective first bin count representative of a number of nucleic acid fragments measured from nucleic acids in a biological sample obtained from the respective reference subject that maps onto a different and non-overlapping portion of the reference genome of the species, wherein, for each respective cancer class in the plurality of cancer classes, the first plurality of reference subjects includes at least one reference subject that has the respective cancer class;

improving a computational efficiency of the computer system by collectively subjecting, by the computer system, the first bin count of each bin in the plurality of bins for each reference subject in the first plurality of reference subjects to a dimensionality reduction method thereby obtaining a feature set, wherein the feature set consists of a number of features that is fewer than the number of bins in the plurality of bins;

resampling, using the computer system, the feature set a plurality of times, wherein the resampling comprises forming, for each respective training iteration in a plurality of training iterations, a trained component classifier via:

omitting, from the feature set, a subset of values for features in the feature set for the first plurality of reference subjects; and

forming the trained component classifier by inputting, in conjunction with the cancer class of respective reference subjects in the first plurality of reference subjects as ground truth, remaining values for the features in the feature set as collective input to a respective untrained component classifier; and

constructing, as a result of the resampling and by collectively leveraging output generated by the trained component classifier formed in each of the plurality of training iterations, a trained first classifier having an improved cancer class recognition ability over the untrained first classifier.

27. A computer system, comprising:

one or more processors; and

a non-transitory computer-readable medium including one or more sequences of instructions that, when executed by the one or more processors, cause the one or more processors to perform a method of classifying a test subject of a given species to a cancer class in a plurality of cancer classes, the method comprising:

using a trained first classifier to classify the test subject to a cancer class in the plurality of cancer classes using nucleic acid fragments in a biological sample obtained from the test subject, the trained first classifier having been trained via:

obtaining, by the computer system and for each respective reference subject in a first plurality of reference subjects, (i) a cancer class of the respective reference subject and (ii) a sequencing construct for the respective reference subject that includes a first bin count for each respective bin in a plurality of bins that collectively represent all or a portion of a reference genome of the species, wherein each respective first bin count representative of a number of nucleic acid fragments measured from nucleic acids in a biological sample obtained from the respective reference subject that maps onto a different and non-overlapping portion of the reference genome of the species, wherein, for each respective cancer class in the plurality of cancer classes, the first plurality of reference subjects includes at least one reference subject that has the respective cancer class;

improving a computational efficiency of the computer system by collectively subjecting, by the computer system, the first bin count of each bin in the plurality of bins for each reference subject in the first plurality of reference subjects to a dimensionality reduction method thereby obtaining a feature set, wherein the feature set consists of a number of features that is fewer than the number of bins in the plurality of bins;

resampling, using the computer system, the feature set a plurality of times, wherein the resampling comprises forming, for each respective training iteration in a plurality of training iterations, a trained component classifier via:

omitting, from the feature set, a subset of values for features in the feature set for the first plurality of reference subjects; and

forming the trained component classifier by inputting, in conjunction with the cancer class of respective reference subjects in the first plurality of reference subjects as ground truth, remaining values for the features in the feature set as collective input to a respective untrained component classifier; and

constructing, as a result of the resampling and by collectively leveraging output generated by the trained component classifier formed in each of the plurality of training iterations, a trained first classifier having an improved cancer class recognition ability over the untrained first classifier.

Assignments (4)
CHANGE OF NAME Recorded Feb 7, 2025
From: GRAIL, LLC
To: GRAIL, INC.
Reel/Frame 070154/0635 →
CHANGE OF NAME Recorded Sep 25, 2024
From: GRAIL, LLC
To: GRAIL, INC.
Reel/Frame 069046/0910 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2024
From: MAHER, M. CYRUS; VALOUEV, ANTON; FILIPPOVA, DARYA; NICULA, VIRGIL; JAGADEESH, KARTHIK; VENN, OLIVER CLAUDE; GROSS, SAMUEL S.; BEAUSANG, JOHN F.; CALEF, ROBERT ABE PAINE
To: GRAIL, INC.
Reel/Frame 068696/0762 →
MERGER AND CHANGE OF NAME Recorded Sep 25, 2024
From: GRAIL, INC.; SDG OPS, LLC
To: GRAIL, LLC
Reel/Frame 068696/0829 →
Continuity (3)
Continuation 16709537 · Dec 10, 2019
Provisional Application 62777693 · Dec 10, 2018
Related Publication 20230170048A1 · Jun 1, 2023
References Cited (158)
US 8642349B1 · Yeatman et al. · 2014 [cited by applicant]
US 8706422B2 · Lo et al. · 2014 [cited by applicant]
US 9260745B2 · Rava et al. · 2016 [cited by applicant]
US 9373059B1 · Heifets et al. · 2016 [cited by applicant]
US 10457995B2 · Talasaz · 2019 [cited by applicant]
US 20100112590A1 · Lo et al. · 2010 [cited by applicant]
US 20100145894A1 · Semizarov et al. · 2010 [cited by applicant]
US 20120053253A1 · Stone et al. · 2012 [cited by applicant]
US 20130034546A1 · Rava et al. · 2013 [cited by applicant]
US 20130325360A1 · Deciu et al. · 2013 [cited by applicant]
US 20140066317A1 · Talasaz · 2014 [cited by applicant]
US 20140080715A1 · Lo et al. · 2014 [cited by applicant]
US 20140100121A1 · Lo et al. · 2014 [cited by applicant]
US 20140242588A1 · Boom et al. · 2014 [cited by applicant]
US 20140371078A1 · Abdueva · 2014 [cited by applicant]
US 20150102216A1 · Roder · 2015 [cited by examiner]
US 20160002739A1 · Schütz et al. · 2016 [cited by applicant]
US 20160232290A1 · Rava et al. · 2016 [cited by applicant]
US 20160251704A1 · Talasaz et al. · 2016 [cited by applicant]
US 20160364522A1 · Frey et al. · 2016 [cited by applicant]
US 20170107576A1 · Babiarz et al. · 2017 [cited by applicant]
US 20170121767A1 · Dor et al. · 2017 [cited by applicant]
US 20170220735A1 · Duenwald et al. · 2017 [cited by applicant]
US 20170240973A1 · Eltoukhy et al. · 2017 [cited by applicant]
US 20170270245A1 · van Rooyen · 2017 [cited by examiner]
US 20170329893A1 · Julio et al. · 2017 [cited by applicant]
US 20170342477A1 · Jensen et al. · 2017 [cited by applicant]
US 20170342500A1 · Marquard et al. · 2017 [cited by applicant]
US 20170362638A1 · Chudova et al. · 2017 [cited by applicant]
US 20180046851A1 · Kienzle et al. · 2018 [cited by applicant]
US 20180144261A1 · Wnuk et al. · 2018 [cited by applicant]
US 20180148777A1 · Kirkizlar et al. · 2018 [cited by applicant]
US 20180173845A1 · Sigurjonsson et al. · 2018 [cited by applicant]
US 20190106737A1 · Underhill · 2019 [cited by applicant]
US 20190164627A1 · Blocker et al. · 2019 [cited by applicant]
US 20190287646A1 · Hubbell · 2019 [cited by applicant]
US 20190287649A1 · Filippova et al. · 2019 [cited by applicant]
US 20190287652A1 · Gross et al. · 2019 [cited by applicant]
US 20190316209A1 · Hubbell et al. · 2019 [cited by applicant]
US 20200005901A1 · Cohen · 2020 [cited by examiner]
US 20200219587A1 · Hubbell · 2020 [cited by applicant]
US 20200294624A1 · Filippova et al. · 2020 [cited by applicant]
US 20200365229A1 · Fields et al. · 2020 [cited by applicant]
US 20200372296A1 · Maher · 2020 [cited by applicant]
US 20210324477A1 · Xiang · 2021 [cited by examiner]
CN 106795562A · 2017 [cited by applicant]
CN 107430588A · 2017 [cited by applicant]
CN 108064314A · 2018 [cited by applicant]
WO 2010037001A2 · 2010 [cited by applicant]
WO 2010051318A2 · 2010 [cited by applicant]
WO 2011127136A1 · 2011 [cited by applicant]
WO 2012071621A1 · 2012 [cited by applicant]
WO 2013052907A2 · 2013 [cited by applicant]
WO 2014149134A2 · 2014 [cited by applicant]
WO 2016094853A1 · 2016 [cited by applicant]
WO 2017062382A1 · 2017 [cited by applicant]
WO 2017161175A1 · 2017 [cited by applicant]
WO 2017212428A1 · 2017 [cited by applicant]
WO 2017181202A3 · 2018 [cited by applicant]
WO 2018009723A1 · 2018 [cited by applicant]
WO 2018022890A1 · 2018 [cited by applicant]
WO 2018022906A1 · 2018 [cited by applicant]
WO 2018031929A1 · 2018 [cited by applicant]
WO 2019084559A1 · 2019 [cited by applicant]
WO 2019232435A1 · 2019 [cited by applicant]
WO 2021119471A1 · 2021 [cited by applicant]
“International Search Report and Written Opinion” for International Application No. PCT/US2019/034994, Sep. 13, 2019, 28 pages. [cited by applicant]
“International Search Report and Written Opinion,” for International Application No. PCT/US2019/22139, Jul. 11, 2019, 18 pages. [cited by applicant]
“Patent Application,” U.S. Appl. No. 62/679,347, entitled “Models for Targeted Sequencing”, filed Jun. 1, 2018. [cited by applicant]
“Patent Application,” U.S. Appl. No. 62/818,013, entitled Systems and Methods for Enriching for Cancer-Derived Fragments Using Fragment Size, filed Mar. 13, 2019, Mar. 13, 2019. [cited by applicant]
“Patent Application,” U.S. Appl. No. 62/827,682, entitled “Systems and Methods for Using Fragment Lengths as a Predictor of Cancer,” filed Apr. 1, 2019, Apr. 10, 2019. [cited by applicant]
“Patent Application,” U.S. Appl. No. 62/851,486, entitled “Systems and Methods for Determining Whether a Subject Has a Cancer Condition Using Transfer Learning,” filed May 22, 2019, May 22, 2019. [cited by applicant]
“Patent Application,” U.S. Appl. No. 62/847,223, entitled “Model-Based Featurization and Classification,” filed May 13, 2019, May 13, 2019. [cited by applicant]
Aengus S. O'Marcaigh, et al., “Estimating the Predictive Value of a Diagnostic Test: How to Prevent Misleading or Confusing Results, Clinical Pediatrics”, vol. 32, Issue 8, First Published Aug. 1, 1993, 7 pages. [cited by applicant]
Alan Agresti, “Building and Applying Logistic Regression Models”, An Introduction to Categorical Data Analysis, Second Edition, Aug. 7, 2006, pp. 137-172. [cited by applicant]
Alan Agresti, “Introduction to Categorical Data Analysis; Chapter 5: Logist Analysis”, Wiley-Interscience, A John Wiley & Sons Inc., Publication, Copyright 1996, 44 pages. [cited by applicant]
Alan H. Lipkus, “A proof of the triangle inequality for the Tanimoto distance”, Journal of Mathematical Chemistry, vol. 26, Published Oct. 1, 1999, pp. 263-265. [cited by applicant]
Alkan et al., “Personalized copy number and segmental duplication maps using next-generation sequencing,” 2009, Nat Genet 41:1061-7. [cited by applicant]
Alkes L. Price, et al., “Principal components analysis corrects for stratification in genome-wide association studies”, vol. 38 No. 8, Aug. 2006, Nature Genetics, pp. 904-909. [cited by applicant]
Antonio Fabregat et al., “The Reactome Pathway Knowledgebase”, Nucleic Acids Research, 2018, vol. 46, Database Issue D649-D655, Published online Nov. 14, 2017, 7 pp. [cited by applicant]
Benjamini et al., “Summaraizing and Correcting the GC Content Bias in High-Throughput Sequencing,” 2012, Nucleic Acids Research 40(10): e72. [cited by applicant]
Bernhard E. Boser et al., “A Training Algorithm for Optimal Margin Classifiers”, Proceedings of the fifth annual workshop on Computational learning theory, Jul. 1992, pp. 144-152. [cited by applicant]
Boeva et al., “Control-free calling of copy number alterations indeep-sequencing data using GC-content normalization”, 2011, Bioinformatics 27(2), p. 268-9. [cited by applicant]
Breiman, “Random Forests—Random Features”, Technical Report 567, Sep. 1999, Statistics Department, U.C. Berkeley, pp. 1-29. [cited by applicant]
Casadio et al., 2013, “Urine cell-free DNA integrity as a marker for early bladder cancer diagnosis: preliminary data,” Urol Oncol. 31(8):1744-1750. [cited by applicant]
Chan et al., 2003, “Clinical Sciences Reviews Committee of the Association of ClinicalBiochemists Cell-free nucleic acids in plasma, serum and urine: a new tool in molecular diagnosis,” Ann Clin Biochem. 40(Pt 2):122-13… [cited by applicant]
Changhong Shan, U.S. Appl. No. 62/642,506, filed Mar. 13, 2018, “Single Radio Voice Call Continuity (SRVCC) Handover and Return to Next Generation Radio Access Network (NG-RAN) After SRVCC” (54 pages). [cited by applicant]
Christof Angermueller et al., “DeepCpG: accurate prediction of single-cell DNA methylation states using deep learning”, Genome Biology, 18:67, Apr. 11, 2017, pp. 1-13. [cited by applicant]
David J. Rogers et al., “A Computer Program for Classifying Plants”, vol. 132, Issue 3434, Oct. 21, 1960. [cited by applicant]
De Mattos-Arruda and Caldas, 2016, “Cell-free circulating tumour DNA as a liquid biopsy in breast cancer,” Mol Oncol. 10(3):464-474. [cited by applicant]
Dingdong Zhang et al., “A novel method to quantify local CpG methylation density by regional methylation elongation assay on microarray”, Published Jan. 31, 2008, BMC Genomics 9, Article No. 59, BioMed Central, 9 pp. [cited by applicant]
Duda, “Support Vector Machines”, Pattern Classification (Chapter 5—Linear Discriminant Functions), 2001, pp. 259 and 262-265. [cited by applicant]
Erickson, “Somatic gene mutation and human disease other than cancer: An update,” Mutat Res 705(2), 2010, 96-106. [cited by applicant]
Erickson, et al., “Somatic gene mutation and human disease other than cancer,” Mutat Res 543(2), 2003, 125-136. [cited by applicant]
Frenel et al., 2015, “Serial next-generation sequencing of circulating cell-free DNA evaluating tumor clone response to molecularly targeted drug administration,” Clin Cancer Res.21(20):4586-4596. [cited by applicant]
Fushun Chen et al., “Conflicts of CpG density and DNA methylation are proximally and distally involved in gene regulation in human and mouse tissues”, Epigenetics, ISSN: 1559-2294 (Print) 1559-2308 (Online), 2018, vol. … [cited by applicant]
G.R. Oxnard et al., “Simultaneous multi-cancer detection and tissue of origin (TOO) localization using targeted bisulfite sequencing of plasma cell-free DNA (cfDNA)”, New Diagnostic Tools, ESMO, vol. 30, Supplement 5, O… [cited by applicant]
Garrett Jenkinson et al., “Potential energy landscapes identify the information-theoretic nature of the epigenome”, vol. 49, No. 5, May 2017, Nature Genetics, pp. 719-732. [cited by applicant]
Goessl et al., 2000, “Fluorescent methylation-specific polymerase chain reaction for DNA-based detection of prostate cancer in bodily fluids,” Cancer Res. 60(21):5941-5945. [cited by applicant]
Gold, 1996, “Softmax to Softassign: Neural Network Algorithms for Combinatorial Optimization,” Journal of Artificial Neural Networks 2, 381-399. [cited by applicant]
Gregory Nuel et al., “Exact distribution of a pattern in a set of random sequences generated by a Markov source: applications to biological data”, Algorithms for Molecular Biology 5, Article No. 15, (2010), pp. 1-18. [cited by applicant]
Hao et al., 2014, “Circulating cell-free DNA in serum as a biomarker for diagnosis and prognostic prediction of colorectal cancer,” Br J Cancer 111(8):1482-1489. [cited by applicant]
Heitzer et al., 2013, “Establishment of tumorspecific copy No. alterations from plasma DNA of patients with cancer,” Int J Cancer. 133(2):346-356). [cited by applicant]
Heitzer et al., 2015, “Circulating tumor DNA as a liquid biopsy for cancer,” Clin Chem. 61(1):112-123. [cited by applicant]
Heng Li et al., “Mapping short DNA sequencing reads and calling variants using mapping quality scores”, Genome Research, ISSN 1088-9051/08, accepted in revised form Aug. 13, 2008, pp. 1851-1858. [cited by applicant]
Hoadley, et al., “Multi-platform analysis of 12 cancer types reveals molecular classification within and across tissues-of-origin,” Cell 158(4), 2014, 929-944. [cited by applicant]
Hugo Larochelle et al., “Exploring Strategies for Training Deep Neural Networks”, Journal of Machine Learning Research 1 (2009), pp. 1-40. [cited by applicant]
Huihuan Qian et al., “Intelligent Surveillance Systems”, International Series on Intelligent Systems, Control and Automation: Science and Engineering, vol. 51, © Springer Science+Business Media B.V. 2011, 187 pp. [cited by applicant]
I. T. Jolliffe, “Principal Component Analysis”, Springer, New York © 2002, 518 pages. [cited by applicant]
International Search Report and Written Opinion mailed Mar. 9, 2021 in International Application No. PCT/US2020/064577 (13 pages). [cited by applicant]
Jihoon Yang et al., “DistAI: An inter-pattern distance-based constructive learning algorithm”, Intelligent Data Analysis, vol. 3, No. 1, Published Jan. 1, 1999, pp. 55-73. [cited by applicant]
Karen Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, Published as a conference paper at ICLR 2015, Apr. 10, 2015, 14 pp. [cited by applicant]
Kim et al., 2014, “Circulating cell-free DNA as a promising biomarker in patients with gastriccancer: diagnostic validity and significant reduction of cfDNA after surgical resection,” Ann Surg Treat Res 86(3):136-1420. [cited by applicant]
Laure Sorber et al., “A Comparison of Cell-Free DNA Isolation Kits: Isolation and Quantification of Cell-Free DNA in Plasma”, The Journal of Molecular Diagnostics, vol. 19, No. 1, Jan. 2017, pp. 162-168. [cited by applicant]
Levandowsky et al., “Distance between Sets”, Nature vol. 234, Nov. 5, 1971, pp. 34-35. [cited by applicant]
Iu et al., 2019, “Bisulfite-free direct detection of 5-methylcytosine and 5-hydroxymethylcytosine at base resolution,” Nature Biotechnology 37, pp. 424-429. [cited by applicant]
Lo et al., 2010, “Maternal plasma DNA sequencing reveals the genome-wide genetic and mutational profile of the jetus,” Sci Transl Med. 2(61):61ra91. [cited by applicant]
Margaret Sullivan Pepe, et al., “Limitations of the Odds Ration in Gauging the Performance of a Diagnostic, Prognostic, or Screening Marker”, American Journal of Epidemiology, American Journal of Epidemiology, vol. 159,… [cited by applicant]
Miller et al., “ReadDepth: A Parallel R Package for Detecting Copy Number Alterations from Short Sequencing Reads”, 2011, PLoS ONE 6(1), p. e16327. [cited by applicant]
Minoru Kanehisa et al., “KEGG: Kyoto Encyclopedia of Genes and Genomes”, Nucleic Acids Research, 2000, vol. 28, No. 1, pp. 27-30. [cited by applicant]
Mohamad H. Hassoun, “Fundamentals of Artificial Neural Networks”, Adaptive Multilayer Nerual Networks I, pp. 197-283. [cited by applicant]
Nancy R. Cook, “Use and Misuse of the Receiver Operating Characteristic Curve in Risk Prediction”, vol. 115, Issue 7, Feb. 20, 2007, pp. 928-935. [cited by applicant]
Nello Cristianini et al., “An Introduction to Support Vector Machines and other kernel-based learning methods”, Support Vector Machines, © Cambridge University Press 2000, pp. 93-124. [cited by applicant]
Ning Yu et al., “GaussianCpG: a Gaussian model for detection of CpG island in human genome sequences”, The Author(s) BMC Genomics 2017, 18(Suppl 4):392, 9 pp. [cited by applicant]
Pan Du et al., “Comparison of Beta-value and M-value methods for quantifying methylation levels by microarray analysis”, Du et al. BMC Bioinformatics 2010, 11:587, 9 pp. [cited by applicant]
Pang-Ning Tan et al., “Instructor's Solution Manual”, © 2006 Pearson Addison-Wesley, 169 pp. [cited by applicant]
Pascal Vincent et al., “Stacked Denoising Autoencoders: Learning Useful Representations in a Deep Network with a Local Denoising Criterion”, Journal of Machine Learning Research, 11(110), © 2010, pp. 3371-3408. [cited by applicant]
Paul Jaccard, “The Distribution of the Flora in the Alpine Zone”, The New Phytologist, vol. 11, No. 2, Feb. 29, 1912, pp. 37-50. [cited by applicant]
Raptis et al., 1980, “Quantitation and characterization of plasma DNA in normals and patients with systemic lupus erythematosus,” J Clin Invest. 66(6):1391-1399. [cited by applicant]
Richard O. Duda et al., “Pattern Classification and Scene Analysis”, © 1973 John Wiley & Sons, Inc., 496 pp. [cited by applicant]
Richard O. Duda et al., “Pattern Classification and Scene Analysis”, John Wiley & Sons, Inc., New York, A Wiley-Interscience Publication, Copyright 1973, 496 pages. [cited by applicant]
Ryan Moulton et al., “Maximally Consistent Sampling and the Jaccard Index of Probability Distributions”, Sep. 11, 2018, 11 pp. [cited by applicant]
Ryan Poplin et al., “A universal SNP and small-indel variant caller using deep neural networks”, Nature Biotechnology, vol. 36, No. 10, Oct. 2018, 9 pp. [cited by applicant]
Salvi et al., 2016, “Cell-free DNA as a diagnostic marker for cancer: current insights,” Onco Targets Ther. 9:6549-6559. [cited by applicant]
Shao et al., 2015 “Quantitative analysis of cell-free DNA in ovarian cancer,” Oncol Lett 10(6):3478-3482). [cited by applicant]
Shapiro et al., 1983, “Determination of circulating DNA levels in patients with benign or malignant gastrointestinal disease,” Cancer. 51(11):2116-2120. [cited by applicant]
Siegel et al., 2015, “Cancer statistics,” CA Cancer J Clin. 65(1):5-29. [cited by applicant]
Sozzi et al., “Quantification of free circulating DNA as a diagnostic marker in lung cancer,” J Clin Oncol. 21(21):3902-3908. [cited by applicant]
Steven Kothen-Hill et al., “Deep Learning Mutation Prediction Enables Early Stage Lung Cancer Detection in Liquid Biopsy”, Workshop track—ICLR 2018, 24 pages. [cited by applicant]
Stroun et al., 1989, “Neoplastic characteristics of the DNA found in the plasma of cancer patients,” Oncology 46(5):318-322. [cited by applicant]
Sven Kosub, “A note on the triangle inequality for the Jaccard distance”, Dec. 9, 2016, pp. 1-5. [cited by applicant]
Swanton, et al., “Phylogenetic ctDNA analysis depicts early stage lung cancer evolution”, Nature. Apr. 26, 2017, Author manuscript; available in PMC Feb. 14, 2018; 545(7655), 446-451, 47 pages. [cited by applicant]
T. Hastie et al., “Additive Models, Trees, and Related Methods”, The Elements of Statistical Learning, Second Edition, © Springer Science+Business Media, LLC 2009, pp. 295-336. [cited by applicant]
T.T. Tanimoto, “An Elementary Mathematical Theory of Classification and Prediction”, Nov. 17, 1958, 11 pp. [cited by applicant]
Terrence S. Furey et al., “Support vector machine classification and validation of cancer tissue samples using microarray expression data”, Bioinformatics, vol. 16, No. 10, 2000, pp. 906-914. [cited by applicant]
Terry et al., 2016, “A prospective evaluation of early detection biomarkers for ovarian cancer in the European EPIC cohort,” Clin Cancer Res. Apr. 8, 2016; Epub. [cited by applicant]
U.S. Appl. No. 62/642,507, filed Mar. 13, 2018, entitled, “Identifying Copy Number Aberrations” (67 pages). [cited by applicant]
Vladimir N. Vapnik, “Statistical Learning Theory”, © 1998 by John Wiley & Sons Inc., 28 pp. [cited by applicant]
Yoon et al., “Sensitive and Accurate Detection of Copy Number Variants Using Read Depth of Coverage”, 2009, Genome Research 19(9):1586. [cited by applicant]
Yuchen Yuan et al., “DeepGene: an advanced cancer type classifier based on deep learning and somatic point mutations”, Dec. 2016, BMC Bioinformatics 17(Suppl 17):476, 15 pages. [cited by applicant]
Zhang et al., 2015, “Tumor markers CA19-9, CA242 and CEA in the diagnosis of pancreatic cancer: a meta-analysis,” Int J Clin Exp Med. 8(7):11683-11691. [cited by applicant]
Zhao, et al., “Detection of fetal subchromosomal abnormalities by sequencing circulating cell-free DNA from maternal plasma”, Clinical Chemistry, vol. 61, Issue 4, Apr. 1, 2015, pp. 608-616. [cited by applicant]
Zonta et al., 2015, “Assessment of DNA integrity, applications for cancer research,” Adv Clin Chem 70:197-246. [cited by applicant]
Hudecova, I., & Chiu, R. W. (2017). Non-invasive prenatal diagnosis of thalassemias using maternal plasma cell free DNA. Best practice & research. Clinical obstetrics & gynaecology, 39, 63-73. (Year: 2017). [cited by applicant]
Smedley, D., Schubach, M., Jacobsen, J. O. B., Kohler, S., Zemotjel, T., Spielmann, M., Jager, M., Hochheiser, H., Washington, N.L., McMurry, J. A., Haendel, M.A., Mungall, C.J., Lewis, S.E., Groza, T., Valentini, G., R… [cited by applicant]
Ulz, P., Thallinger, G. G., Auer, M., Graf, R., Kashofer, K., Jahn, S. W., Abete, L., Pristauz, G., Petru, E., Geigl, J. B., Heitzer, E., & Speicher, M. R. (2016). Inferring expressed genes by whole-genome sequencing of… [cited by applicant]
Jaccard, Paul, Comparative Study of the Floral distribution in a portion of the Als and the Jura, Bull.Soc.Vaud.Sci Nat. XXXVII, 142, pp. 547-579, 1901, English translation. [cited by applicant]
Jaccard, Étude comparative de la distribution florale dans une portion des Alpes et des Jura, 1901, pp. 547-579, Bulletin de la Société vaudoise des sciences naturelles, 37. [cited by applicant]