Population based treatment recommender using cell free dna
Systems and methods are disclosed for generating a therapeutic response predict or detecting a disease, by: using a genetic analyzer to generate genetic information; receiving into computer memory a training dataset comprising, for each of a plurality of individuals having a disease, (1) genetic information from the individual generated at first time point and (2) treatment response of the individual to one or more therapeutic interventions determined at a second, later, time point; and implementing a machine learning algorithm using the dataset to generate at least one computer implemented classification algorithm, wherein the classification algorithm, based on genetic information from a subject, predicts therapeutic response of the subject to a therapeutic intervention.
1 . A method, comprising:
obtaining or having obtained a cell-free biological sample comprising cell-free DNA (cfDNA) molecules from a subject from at least two different time points;
ligating molecular barcodes to a plurality of the cfDNA molecules, wherein the molecular barcodes uniquely tag polynucleotides in the cfDNA molecules to generate tagged cfDNA molecules;
generating a plurality of sequencing reads from the tagged cfDNA molecules;
grouping sequence reads into families, each family generated from a uniquely tagged parent polynucleotide;
identifying, in the plurality of sequencing reads, genetic variants comprising sequence variants and copy number variants;
quantifying the frequency of genetic variants;
storing genetic information comprising the genetic variants and frequencies of the genetic variants and clinical information of the subject in at least one database;
obtaining genetic information from each of a plurality of subjects in the at least one database;
operably connecting at least one extractor to the at least one database, wherein the extractor is configured to extract one or more features from the genetic information and the clinical information stored in the at least one database;
performing a training process of one or more classifiers that are included in at least one recommender, wherein the training process includes:
determining a plurality of classes of subjects by identifying individual sets of subjects having a shared characteristic;
generating a training data set using a multi-parametric model, the training data set indicating characteristics of polynucleotides of samples of subjects that belong to each class of the plurality of classes; and
training a machine learning algorithm using the training data set to generate one or more trained classifiers, wherein individual trained classifiers of the one or more trained classifiers determines one or more classes of the plurality of classes for a test sample and
operably connecting the at least one recommender to the at least one extractor, thereby generating the system;
wherein the at least one recommender is configured to:
implement the one or more trained classifiers to select a treatment from among a plurality of treatment options for a cancer of at least one test subject of the plurality of subjects, based at least in part on an analysis of the one or more features at a first time point of the two or more time points; and
implement the one or more trained classifiers to select a later treatment from among a plurality of treatment options for a cancer of at least one test subject of the plurality of subjects, wherein selection of the treatment, based on at least an analysis of the:
one or more features comprising genetic information for the given subject from a second time point of the two or more time points; and
amount of time between at least two or more time points.
2 . The method of claim 1 , wherein the plurality of subjects comprises subjects having cancer.
3 . The method of claim 1 , wherein the plurality of subjects comprises subjects in which cancer is not detected.
4 . The method of claim 1 , further comprising operably connecting at least one genetic analyzer to the system, wherein the genetic analyzer is configured to analyze the genetic information.
5 . The method of claim 1 , further comprising operably connecting at least one report generator to the at least one recommender, wherein the at least one report generator is configured to generate a report comprising information corresponding to the selected treatment.
6 . The method of claim 1 , wherein the genetic information comprises unstructured text data.
7 . The method of claim 1 , further comprising storing the clinical information in at least one database array; and wherein the clinical information comprises patient information from physicians and laboratory test results.
8 . The method of claim 1 , comprising pre-processing the training data set by transforming at least a portion of the clinical information obtained from at least a portion of the plurality of subjects into class-conditional probabilities.
9 . The method of claim 1 , wherein the clinical information comprises computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, bone scans, positron emission tomography (PET) scans, bone marrow tests, X-rays, endoscopies, lymphangiograms, intravenous urograms (IVU), intravenous pyelograms (IVP), lumbar punctures, cystoscopies, immunological tests, histology reports, or cancer marker tests.
10 . The method of claim 1 , further comprising using the system to predict a course of the treatment for the at least one test subject having cancer.
11 . The method of claim 1 , further comprising operably connecting at least one classifier to the at least one extractor and to the at least one recommender, wherein the at least one classifier is configured to classify the one or more features extracted from the genetic information and the clinical information.
12 . The method of claim 11 , wherein the one or more features comprise treatment responses to therapeutic interventions, and wherein the at least one classifier is configured to classify the one or more features to generate one or more treatment classifications.
13 . The method of claim 12 , wherein the one or more treatment classifications comprise responsive to treatment, non-responsive to treatment, or a level of responsiveness to treatment.
14 . The method of claim 11 , further comprising operably connecting at least one inference unit to the at least one classifier and to the at least one recommender.
15 . The method of claim 14 , wherein an output of the at least one classifier is provided to the at least one inference unit.
16 . The method of claim 1 , wherein a sample of the at least one test subject includes at least one of cell-free polynucleotides or polynucleotides derived from a tumor sample.
17 . The method of claim 1 , wherein the plurality of classes is selected from the group consisting of: subjects in which cancer is not detected, subjects having breast cancer, subjects having colon cancer, subjects having lung cancer, subjects having pancreatic cancer, subjects having prostate cancer, subjects having ovarian cancer, subjects having melanoma, and subjects having liver cancer.
18 . The method of claim 1 , wherein the one or more features comprise an actionable tumor-specific genomic alteration comprising single nucleotide variants (SNVs) or insertions or deletions (indels).
19 . The method of claim 1 , wherein the treatment and later treatment are not the same.
20 . The method of claim 1 , wherein the tagged cfDNA molecules are amplified prior to sequencing.
21 . The method of claim 20 , wherein the comparison further comprises a relative excess or relative deficit to total number of sequence reads at test and control loci.
22 . The method of claim 1 , comprising estimating a tumor burden based on a comparison of the relative number of sequence reads bearing a variant, to the total number of sequence reads generated from the sample.