Systems and methods for cancer condition determination using autoencoders
A method for discriminating a cancer state is provided. A first dataset is obtained for a plurality of subjects having a first cancer state. Each subject has a plurality of nucleic acid methylation fragments with methylation patterns comprising CpG site methylation states. An autoencoder including an encoder and decoder is trained by evaluating the error in the autoencoder reconstruction of the methylation pattern and nucleic acid sequence of each nucleic acid methylation fragment in the first dataset. A second dataset is obtained for a plurality of subjects having a second cancer state. A plurality of features is identified by inputting the methylation pattern and nucleic acid sequence of each nucleic acid methylation fragment in the second dataset into the trained autoencoder and computing a score determined by the autoencoder reconstruction of the methylation pattern. The plurality of features is used to train a supervised model that discriminates a cancer state.
1 . A method comprising:
A) obtaining a test dataset, in electronic form, wherein:
the test dataset comprises, for a test subject, a corresponding methylation pattern and a corresponding nucleic acid sequence of each respective nucleic acid methylation fragment in a plurality of nucleic acid methylation fragments determined by a methylation sequencing of nucleic acids in a biological sample obtained from the test subject, and
the corresponding methylation pattern comprises a methylation state of each respective CpG site in a corresponding plurality of CpG sites in the respective nucleic acid methylation fragment; and
B) training an autoencoder by a training process that comprises two stages, and the two stages comprises:
using a first stage of supervised training to train an untrained autoencoder using a first methylation pattern dataset that comprises training samples of a first cancer state indicating an absence of cancer, wherein the first stage of supervised training determines similarities of output reconstructions compared to the first methylation pattern dataset,
selecting a second methylation pattern dataset based on one or more selection criteria, the second methylation pattern dataset comprising training samples of the first cancer state and training samples of a second cancer state indicating a presence of cancer according to the one or more selection criteria, and
using a second stage of supervised training to train the trained autoencoder using the second methylation pattern dataset to generate reconstruction scores;
C) applying the trained autoencoder to the corresponding methylation pattern and the corresponding nucleic acid sequence of each respective nucleic acid methylation fragment in all or a portion of the plurality of nucleic acid methylation fragments to determine whether the test subject has a cancer state, wherein the trained autoencoder includes 1000 or more weights, and wherein the C) applying comprises:
for each respective nucleic acid methylation fragment in the all or the portion of the plurality of nucleic acid methylation fragments:
(i) using the trained autoencoder to reconstruct a corresponding methylation pattern based on the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment;
(ii) computing a corresponding score determined at least in part by the reconstructed methylation pattern; and
(ii) determining, as an output, whether the test subject has the cancer state using each corresponding score;
D) determining, based on the output, that the cancer state of the test subject is the second cancer state; and
E) causing an administration of a treatment of the cancer state to the test subject based on determining that the cancer state is the second cancer state, wherein the treatment includes administering a dosage of Lenalidomid, Pembrolizumab, Trastuzumab, Bevacizumab, Rituximab, Ibrutinib, Human Papillomavirus Quadrivalent (Types 6, 11, 16, and 18) Vaccine, Pertuzumab, Pemetrexed, Nilotinib, Nilotinib, Denosumab, Abiraterone acetate, Promacta, Imatinib, Everolimus, Palbociclib, Erlotinib, Bortezomib, Bortezomib, or a generic equivalent thereof.
2 . The method of claim 1 , wherein the corresponding score of the respective nucleic acid methylation fragment:
is determined by a correctness of the reconstruction of the corresponding methylation pattern of the respective nucleic acid methylation fragment by the trained autoencoder, and
is independent of a correctness of the reconstruction of the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment by the trained autoencoder.
3 . The method of claim 1 , wherein the corresponding score of the respective nucleic acid methylation fragment:
is determined by a correctness of the reconstruction of the corresponding methylation pattern of the respective nucleic acid methylation fragment by the trained autoencoder, and
is further determined by the correctness of the reconstruction of the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment by the trained autoencoder.
4 . The method of claim 2 , wherein the correctness of the reconstruction of the corresponding methylation pattern of the respective nucleic acid methylation fragment by the trained autoencoder is determined, at least in part, by a Hamming distance between the reconstruction of the corresponding methylation pattern of the respective nucleic acid methylation fragment and the actual methylation pattern of the respective nucleic acid methylation fragment.
5 . The method of claim 1 , wherein the plurality of nucleic acid methylation fragments comprises one thousand or more, ten thousand or more, 100 thousand or more, one million or more, ten million or more, 100 million or more, 500 million or more, one billion or more, two billion or more, three billion or more, four billion or more, five billion or more, six billion or more, seven billion or more, eight billion or more, nine billion or more, or 10 billion or more nucleic acid methylation fragments.
6 . The method of claim 1 , wherein after the A) obtaining and prior to the C) applying:
filtering the plurality of nucleic acid methylation fragments by removing, from the plurality of nucleic acid methylation fragments, each respective nucleic acid methylation fragment that fails to satisfy one or more selection criteria.
7 . The method of claim 6 , wherein:
the respective nucleic acid methylation fragment fails to satisfy a selection criterion in the one or more selection criteria when the methylation pattern of the respective nucleic acid methylation fragment has an output p-value that fails to satisfy a p-value threshold, and
the output p-value of the respective nucleic acid methylation fragment is determined, at least in part, based upon a comparison of the methylation pattern of the respective nucleic acid methylation fragment over a plurality of CpG sites of the respective nucleic acid methylation fragment to a corresponding distribution of methylation patterns of those nucleic acid methylation fragments in a training dataset that have the corresponding plurality of CpG sites.
8 . The method of claim 6 , wherein:
the respective nucleic acid methylation fragment fails to satisfy a selection criterion in the one or more selection criteria when an output p-value provided by a trained Markov model, responsive to input of the methylation pattern of the respective nucleic acid methylation fragment, fails the selection criterion, and
the trained Markov model is trained, at least in part, based upon evaluation of a methylation state of each CpG site in a plurality of CpG sites of the respective nucleic acid methylation fragment across those nucleic acid methylation fragments in a training dataset that have the corresponding plurality of CpG sites.
9 . The method of claim 6 , wherein the respective nucleic acid methylation fragment fails to satisfy a selection criterion in the one or more selection criteria when the respective nucleic acid methylation fragment has less than a threshold number of CpG sites.
10 . The method of claim 9 , wherein the threshold number of CpG sites is 4, 5, 6, 7, 8, 9, or 10.
11 . The method of claim 6 , wherein the respective nucleic acid methylation fragment fails to satisfy a selection criterion in the one or more selection criteria when the respective nucleic acid methylation fragment has less than a threshold number of residues.
12 . The method of claim 11 , wherein the threshold number of residues is a fixed value between 20 and 90.
13 . The method of claim 6 , wherein the filtering removes a nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments that has the same corresponding methylation pattern and the same corresponding nucleic acid sequence as another nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments.
14 . The method of claim 1 , wherein the trained autoencoder is a variational autoencoder, a stacked denoising deep autoencoder, a deep recurrent autoencoder, a convolutional autoencoder, or a transformer network.
15 . The method of claim 1 , wherein the trained autoencoder is a deep recurrent autoencoder and the B) applying, for a respective nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments:
feeds a first track of the deep recurrent autoencoder the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment broken up into a plurality of k-mers, and
feeds a second track of the deep recurrent autoencoder the corresponding methylation pattern of the respective nucleic acid methylation fragment.
16 . The method of claim 1 , wherein the trained autoencoder is a deep recurrent autoencoder and the C) applying the trained autoencoder:
feeds a first track of the deep recurrent autoencoder the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment on a residue basis, and
feeds a second track of the deep recurrent autoencoder the corresponding methylation pattern of the respective nucleic acid methylation fragment.
17 . The method of claim 1 , wherein the trained autoencoder comprises:
an encoder that encodes the corresponding methylation pattern and the corresponding nucleic acid sequence of the corresponding nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments thereby forming a plurality of latent features; and
a decoder that decodes the plurality of latent features into a reconstruction of the corresponding methylation pattern and the corresponding nucleic acid sequence of the corresponding nucleic acid methylation fragment.
18 . The method of claim 1 , wherein the methylation state of a respective CpG site in the plurality of CpG sites in the respective nucleic acid methylation fragment is:
methylated when the respective CpG site is determined by the methylation sequencing to be methylated,
unmethylated when the respective CpG site is determined by the methylation sequencing to not be methylated, and
flagged as “other” when the methylation sequencing is unable to call the methylation state of the respective CpG site as methylation or unmethylated.
19 . The method of claim 1 , wherein the methylation sequencing is i) whole genome methylation sequencing or ii) targeted DNA methylation sequencing using a plurality of nucleic acid probes.
20 . The method of claim 1 , wherein the second cancer state is a stage of a specified cancer.
21 . The method of claim 1 , wherein the methylation sequencing of nucleic acids in a biological sample obtained from the respective subject is methylation sequencing of cell-free nucleic acids in the biological sample.
22 . The method of claim 1 , wherein the biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the test subject.
23 . The method of claim 1 , wherein the test dataset comprises:
a first corresponding nucleic acid sequence of a first nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments determined by the methylation sequencing of nucleic acids in the biological sample obtained from the test subject wherein the first corresponding nucleic acid sequence is from a forward strand or a reverse strand of the first nucleic acid methylation fragment or wherein the first corresponding nucleic acid sequence is a reverse strand of the first nucleic acid methylation fragment and is in reverse complement form or is flagged as being reverse strand of the first nucleic acid methylation fragment.