IP Library Granted Patent US 12,646,621
Granted Patent B2
US 12,646,621 · App. 17/191,914 · Granted Jun 2, 2026

Systems and methods for cancer condition determination using autoencoders

Inventors: Virgil Nicula (Cupertino, CA); Joshua Newman (Mountain View, CA)
Assignee: Grail, Inc.
G16H50/20G16B20/00G16B40/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,621
App. No.
17/191,914
Granted
Jun 2, 2026
Kind
B2
Abstract

A method for discriminating a cancer state is provided. A first dataset is obtained for a plurality of subjects having a first cancer state. Each subject has a plurality of nucleic acid methylation fragments with methylation patterns comprising CpG site methylation states. An autoencoder including an encoder and decoder is trained by evaluating the error in the autoencoder reconstruction of the methylation pattern and nucleic acid sequence of each nucleic acid methylation fragment in the first dataset. A second dataset is obtained for a plurality of subjects having a second cancer state. A plurality of features is identified by inputting the methylation pattern and nucleic acid sequence of each nucleic acid methylation fragment in the second dataset into the trained autoencoder and computing a score determined by the autoencoder reconstruction of the methylation pattern. The plurality of features is used to train a supervised model that discriminates a cancer state.

Claims (56)

1 . A method comprising:

A) obtaining a test dataset, in electronic form, wherein:

the test dataset comprises, for a test subject, a corresponding methylation pattern and a corresponding nucleic acid sequence of each respective nucleic acid methylation fragment in a plurality of nucleic acid methylation fragments determined by a methylation sequencing of nucleic acids in a biological sample obtained from the test subject, and

the corresponding methylation pattern comprises a methylation state of each respective CpG site in a corresponding plurality of CpG sites in the respective nucleic acid methylation fragment; and

B) training an autoencoder by a training process that comprises two stages, and the two stages comprises:

using a first stage of supervised training to train an untrained autoencoder using a first methylation pattern dataset that comprises training samples of a first cancer state indicating an absence of cancer, wherein the first stage of supervised training determines similarities of output reconstructions compared to the first methylation pattern dataset,

selecting a second methylation pattern dataset based on one or more selection criteria, the second methylation pattern dataset comprising training samples of the first cancer state and training samples of a second cancer state indicating a presence of cancer according to the one or more selection criteria, and

using a second stage of supervised training to train the trained autoencoder using the second methylation pattern dataset to generate reconstruction scores;

C) applying the trained autoencoder to the corresponding methylation pattern and the corresponding nucleic acid sequence of each respective nucleic acid methylation fragment in all or a portion of the plurality of nucleic acid methylation fragments to determine whether the test subject has a cancer state, wherein the trained autoencoder includes 1000 or more weights, and wherein the C) applying comprises:

for each respective nucleic acid methylation fragment in the all or the portion of the plurality of nucleic acid methylation fragments:

(i) using the trained autoencoder to reconstruct a corresponding methylation pattern based on the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment;

(ii) computing a corresponding score determined at least in part by the reconstructed methylation pattern; and

(ii) determining, as an output, whether the test subject has the cancer state using each corresponding score;

D) determining, based on the output, that the cancer state of the test subject is the second cancer state; and

E) causing an administration of a treatment of the cancer state to the test subject based on determining that the cancer state is the second cancer state, wherein the treatment includes administering a dosage of Lenalidomid, Pembrolizumab, Trastuzumab, Bevacizumab, Rituximab, Ibrutinib, Human Papillomavirus Quadrivalent (Types 6, 11, 16, and 18) Vaccine, Pertuzumab, Pemetrexed, Nilotinib, Nilotinib, Denosumab, Abiraterone acetate, Promacta, Imatinib, Everolimus, Palbociclib, Erlotinib, Bortezomib, Bortezomib, or a generic equivalent thereof.

2 . The method of claim 1 , wherein the corresponding score of the respective nucleic acid methylation fragment:

is determined by a correctness of the reconstruction of the corresponding methylation pattern of the respective nucleic acid methylation fragment by the trained autoencoder, and

is independent of a correctness of the reconstruction of the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment by the trained autoencoder.

3 . The method of claim 1 , wherein the corresponding score of the respective nucleic acid methylation fragment:

is determined by a correctness of the reconstruction of the corresponding methylation pattern of the respective nucleic acid methylation fragment by the trained autoencoder, and

is further determined by the correctness of the reconstruction of the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment by the trained autoencoder.

4 . The method of claim 2 , wherein the correctness of the reconstruction of the corresponding methylation pattern of the respective nucleic acid methylation fragment by the trained autoencoder is determined, at least in part, by a Hamming distance between the reconstruction of the corresponding methylation pattern of the respective nucleic acid methylation fragment and the actual methylation pattern of the respective nucleic acid methylation fragment.

5 . The method of claim 1 , wherein the plurality of nucleic acid methylation fragments comprises one thousand or more, ten thousand or more, 100 thousand or more, one million or more, ten million or more, 100 million or more, 500 million or more, one billion or more, two billion or more, three billion or more, four billion or more, five billion or more, six billion or more, seven billion or more, eight billion or more, nine billion or more, or 10 billion or more nucleic acid methylation fragments.

6 . The method of claim 1 , wherein after the A) obtaining and prior to the C) applying:

filtering the plurality of nucleic acid methylation fragments by removing, from the plurality of nucleic acid methylation fragments, each respective nucleic acid methylation fragment that fails to satisfy one or more selection criteria.

7 . The method of claim 6 , wherein:

the respective nucleic acid methylation fragment fails to satisfy a selection criterion in the one or more selection criteria when the methylation pattern of the respective nucleic acid methylation fragment has an output p-value that fails to satisfy a p-value threshold, and

the output p-value of the respective nucleic acid methylation fragment is determined, at least in part, based upon a comparison of the methylation pattern of the respective nucleic acid methylation fragment over a plurality of CpG sites of the respective nucleic acid methylation fragment to a corresponding distribution of methylation patterns of those nucleic acid methylation fragments in a training dataset that have the corresponding plurality of CpG sites.

8 . The method of claim 6 , wherein:

the respective nucleic acid methylation fragment fails to satisfy a selection criterion in the one or more selection criteria when an output p-value provided by a trained Markov model, responsive to input of the methylation pattern of the respective nucleic acid methylation fragment, fails the selection criterion, and

the trained Markov model is trained, at least in part, based upon evaluation of a methylation state of each CpG site in a plurality of CpG sites of the respective nucleic acid methylation fragment across those nucleic acid methylation fragments in a training dataset that have the corresponding plurality of CpG sites.

9 . The method of claim 6 , wherein the respective nucleic acid methylation fragment fails to satisfy a selection criterion in the one or more selection criteria when the respective nucleic acid methylation fragment has less than a threshold number of CpG sites.

10 . The method of claim 9 , wherein the threshold number of CpG sites is 4, 5, 6, 7, 8, 9, or 10.

11 . The method of claim 6 , wherein the respective nucleic acid methylation fragment fails to satisfy a selection criterion in the one or more selection criteria when the respective nucleic acid methylation fragment has less than a threshold number of residues.

12 . The method of claim 11 , wherein the threshold number of residues is a fixed value between 20 and 90.

13 . The method of claim 6 , wherein the filtering removes a nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments that has the same corresponding methylation pattern and the same corresponding nucleic acid sequence as another nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments.

14 . The method of claim 1 , wherein the trained autoencoder is a variational autoencoder, a stacked denoising deep autoencoder, a deep recurrent autoencoder, a convolutional autoencoder, or a transformer network.

15 . The method of claim 1 , wherein the trained autoencoder is a deep recurrent autoencoder and the B) applying, for a respective nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments:

feeds a first track of the deep recurrent autoencoder the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment broken up into a plurality of k-mers, and

feeds a second track of the deep recurrent autoencoder the corresponding methylation pattern of the respective nucleic acid methylation fragment.

16 . The method of claim 1 , wherein the trained autoencoder is a deep recurrent autoencoder and the C) applying the trained autoencoder:

feeds a first track of the deep recurrent autoencoder the corresponding nucleic acid sequence of the respective nucleic acid methylation fragment on a residue basis, and

feeds a second track of the deep recurrent autoencoder the corresponding methylation pattern of the respective nucleic acid methylation fragment.

17 . The method of claim 1 , wherein the trained autoencoder comprises:

an encoder that encodes the corresponding methylation pattern and the corresponding nucleic acid sequence of the corresponding nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments thereby forming a plurality of latent features; and

a decoder that decodes the plurality of latent features into a reconstruction of the corresponding methylation pattern and the corresponding nucleic acid sequence of the corresponding nucleic acid methylation fragment.

18 . The method of claim 1 , wherein the methylation state of a respective CpG site in the plurality of CpG sites in the respective nucleic acid methylation fragment is:

methylated when the respective CpG site is determined by the methylation sequencing to be methylated,

unmethylated when the respective CpG site is determined by the methylation sequencing to not be methylated, and

flagged as “other” when the methylation sequencing is unable to call the methylation state of the respective CpG site as methylation or unmethylated.

19 . The method of claim 1 , wherein the methylation sequencing is i) whole genome methylation sequencing or ii) targeted DNA methylation sequencing using a plurality of nucleic acid probes.

20 . The method of claim 1 , wherein the second cancer state is a stage of a specified cancer.

21 . The method of claim 1 , wherein the methylation sequencing of nucleic acids in a biological sample obtained from the respective subject is methylation sequencing of cell-free nucleic acids in the biological sample.

22 . The method of claim 1 , wherein the biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the test subject.

23 . The method of claim 1 , wherein the test dataset comprises:

a first corresponding nucleic acid sequence of a first nucleic acid methylation fragment in the plurality of nucleic acid methylation fragments determined by the methylation sequencing of nucleic acids in the biological sample obtained from the test subject wherein the first corresponding nucleic acid sequence is from a forward strand or a reverse strand of the first nucleic acid methylation fragment or wherein the first corresponding nucleic acid sequence is a reverse strand of the first nucleic acid methylation fragment and is in reverse complement form or is flagged as being reverse strand of the first nucleic acid methylation fragment.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Oct 13, 2021
From: GRAIL, INC.; SDG OPS, LLC
To: GRAIL, LLC
Reel/Frame 057788/0719 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 16, 2021
From: NICULA, VIRGIL; NEWMAN, JOSHUA
To: GRAIL, INC.
Reel/Frame 056881/0521 →
Continuity (2)
Provisional Application 62985258 · Mar 4, 2020
Related Publication 20210358626A1 · Nov 18, 2021
References Cited (55)
US 20180237863A1 · Namsaraev et al. · 2018 [cited by applicant]
US 20190287652A1 · Gross et al. · 2019 [cited by applicant]
US 20200239964A1 · Gross et al. · 2020 [cited by applicant]
US 20200239965A1 · Fields et al. · 2020 [cited by applicant]
US 20200365229A1 · Fields et al. · 2020 [cited by applicant]
US 20200372296A1 · Maher · 2020 [cited by applicant]
US 20200385813A1 · Venn · 2020 [cited by applicant]
WO WO2018081130 · 2018 [cited by applicant]
WO WO2018163051A1 · 2018 [cited by examiner]
WO 2019195268A2 · 2019 [cited by applicant]
WO WO2019200410A1 · 2019 [cited by examiner]
WO WO2019204360A1 · 2019 [cited by examiner]
WO WO2020069350 · 2020 [cited by applicant]
WO WO2020154682 · 2020 [cited by applicant]
Jelinek, Jaroslav, et al. “Conserved DNA methylation patterns in healthy blood cells and extensive changes in leukemia measured by a new quantitative technique.” Epigenetics 7.12 (2012): 1368-1378. (Year: 2012). [cited by examiner]
Crowgey, Erin L., et al. “Epigenetic machine learning: utilizing DNA methylation patterns to predict spastic cerebral palsy.” BMC bioinformatics 19 (2018): 1-10. (Year: 2018). [cited by examiner]
Angermueller, Christof, et al. “DeepCpG: accurate prediction of single-cell DNA methylation states using deep learning.” Genome biology 18.1 (2017): 1-13. (Year: 2017). [cited by examiner]
Zhang, Xiaoyu, et al. “Integrated Multi-omics Analysis Using Variational Autoencoders: Application to Pan-cancer Classification.” arXiv preprint arXiv:1908.06278 (2019). (Year: 2019). [cited by examiner]
Liu, Qian, and Pingzhao Hu. “Association analysis of deep genomic features extracted by denoising autoencoders in breast cancer.” Cancers 11.4 (2019): 494. (Year: 2019). [cited by examiner]
Titus, Alexander J., et al. “Unsupervised deep learning with variational autoencoders applied to breast tumor genome-wide DNA methylation data with biologic feature extraction.” BioRxiv 433763 (2018). (Year: 2018). [cited by examiner]
U.S. Appl. No. 17/119,606, entitled “Cancer classification using patch convolutional neural networks,” filed Dec. 11, 2020. [cited by applicant]
Agresti, [cited by applicant]
Ameniya et al. 2019, “The Encode Blacklist: Identification of Problematic Regions of the Genome,” Scientific Reports 9, article No. 9354. [cited by applicant]
Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5 [cited by applicant]
Breiman, 1999, “Random Forests—Random Features,” Technical Report 567, Statistics Department, U.C. Berkeley, Sep. 1999. [cited by applicant]
Doersch, 2016, “Tutorial on variational autoencoders.” arXiv preprint arXiv: 1606.05908. [cited by applicant]
Du et al., 2010, BMC Bioinformatics 11 :587, doi: 10.1186/1471-2105-11-587. [cited by applicant]
Duda, [cited by applicant]
Feng et al., 2020, “Soft Gradient Boosting Machine,” arXiv:2006.04059. [cited by applicant]
Fernandes et al., 2017, “Transfer Learning with Partial Observability Applied to Cervical Cancer Screening,” Pattern Recognition and Image Analysis: 8 [cited by applicant]
Furey et al., 2000, [cited by applicant]
Grunau et al., 2001, “MethDB—a public database for DNA methylation data,” Nucleic Acids Research 29(1), 270-274. [cited by applicant]
Hachiya et al., 2017, “Genomewide identification of inter-individually variable DNA methylation sites improves the efficacy of epigenetic association studies,” NPJ Genom Med. 2017. 2: 11. [cited by applicant]
Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology. [cited by applicant]
Hastie, 2001, [cited by applicant]
Huang et al., 2021, “MethHC 2.0: information repository of DNA methylation and gene expression in human cancer,” Nucleic Acids Research 49(Dl), Dl268-Dl275. [cited by applicant]
Jones, 2002, Oncogene 21 :5358-5360. [cited by applicant]
Kingma and Max, 2019, [cited by applicant]
Klein et al., 2018, “Development of a comprehensive cell-free DNA (cIDNA) assay for early detection of multiple tumor types: The Circulating Cell-free Genome Atlas (CCGA) study,” J. Clin. Oncology 36(15), 12021-12021; d… [cited by applicant]
Kristiadi, 2016, “Variational Autoencoder: Intuition and Implementation,”. [cited by applicant]
Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40. [cited by applicant]
Liu et al., “Bisulfite-free direct detection of 5-methylcytosine and 5-hydroxymethylcytosine at base resolution,” Nat Biotechnol, doi: 10.1038/s41587-019-0041-2. [cited by applicant]
Liu et al., 2019, “Genome-wide cell-free DNA (cIDNA) methylation signatures and effect on tissue of origin (TOO) performance,” J. Clin. Oncology 37(15), 3049-3049; doi: 10.1200/JCO.2019.37.15 suppl.3049. [cited by applicant]
Ongenaert et al., “PubMeth: a cancer methylation database combining text-mining and expert annotation,” Nucleic Acids Research: doi: 10.1093/nar/gkm788. [cited by applicant]
Paska and Hudler, 2015, Biochemia Medica 25(2): 161-176. [cited by applicant]
Schliep et al., 2003, Bioinformatics 19(1): i255-i263. [cited by applicant]
Taheri and Mammadov, “Learning the naive Bayes classifier with optimization models,” International Journal of Applied Mathematics and Computer Science 23( 4), 787-795. [cited by applicant]
Vapnik, 1998, [cited by applicant]
Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408. [cited by applicant]
Warton and Samimi, 2015, Front Mol Biosci, 2(13) doi: 10.3389/fmolb.2015.00013. [cited by applicant]
Yoon, 2009, “Hidden Markov Models and their Applications in Biological Sequence Analysis,” Curr. Genomics. Sep; 10(6): 402-415, doi: 10.2174/138920209789177575. [cited by applicant]
Ziller et al., 2015, “Coverage recommendations for methylation analysis by whole-genome bisulfite sequencing,” Nature Methods. 12(3):230-232, doi: 10.1038/nmeth.3152. [cited by applicant]
International Search Report and Written Opinion for PCT/US2021/020787; dated Jun. 16, 2021; 20 pages. [cited by applicant]
Mohammed Khwaja et. al; “A Deep Autoencoder System for Differentiation of Cancer Types Based on DNA Methylation State”; Oct. 5, 2018; 8 pages. [cited by applicant]
Masser, D.R. et al., “Targeted DNA Methylation Analysis by Next-generation Sequencing,” Journal of Visualized Experiments (96), Feb. 24, 2015, pp. 1-11. [cited by applicant]