IP Library Granted Patent US 12,580,051
Granted Patent B2
US 12,580,051 · App. 17/187,319 · Granted Mar 17, 2026

Identifying methylation patterns that discriminate or indicate a cancer condition

Inventors: Collin Melton (Menlo Park, CA); Earl Hubbell (Palo Alto, CA); Oliver Claude Venn (San Francisco, CA)
Assignee: GRAIL, Inc.
G16B40/20C12Q1/6886G06F18/24155G16B20/00G16B45/00C12Q2600/154
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,580,051
App. No.
17/187,319
Granted
Mar 17, 2026
Kind
B2
Abstract

Systems and methods of identifying methylation patterns discriminating or indicating a cancer condition are provided. First and second datasets are obtained. Each dataset comprises a plurality of fragment methylation patterns determined by methylation sequencing of nucleic acids obtained from a first or second set of subjects and comprising a methylation state of each CpG site in a corresponding plurality of CpG sites. Each plurality of subjects has a respective first or second state of the cancer condition. First and second interval maps are generated for each respective dataset, each comprising a plurality of nodes characterized by a start methylation site, an end methylation site, a representation of each different fragment methylation pattern and a count of fragments. The first and second interval maps are scanned for qualifying methylation patterns within a predetermined range of CpG sites, satisfying one or more selection criteria, thereby identifying methylation patterns discriminating a cancer condition.

Claims (96)

1 . A computer-implemented method of identifying a plurality of qualifying methylation patterns that discriminate or indicate a cancer condition, the method comprising:

A) obtaining, at a computer system, a first dataset, in electronic form, wherein the first dataset comprises a corresponding fragment methylation pattern of each respective fragment in a first plurality of fragments, wherein the corresponding fragment methylation pattern of each respective fragment (i) is determined by a methylation sequencing of nucleic acids from a respective biological sample obtained from a corresponding subject in a first set of subjects and (ii) comprises a methylation state of each CpG site in a corresponding plurality of CpG sites in the respective fragment, and wherein the first plurality of fragments comprises more than 1000 fragments;

B) obtaining, at the computer system, a second dataset, in electronic form, wherein the second dataset comprises a corresponding fragment methylation pattern of each respective fragment in a second plurality of fragments, wherein the corresponding fragment methylation pattern of each respective fragment (i) is determined by a methylation sequencing of nucleic acids from a respective biological sample obtained from a corresponding subject in a second set of subjects and (ii) comprises a methylation state of each CpG site in a corresponding plurality of CpG sites in the respective fragment, wherein each subject in the first set of subjects has a first state of the cancer condition and each subject in the second set of subjects has a second state of the cancer condition, and wherein the second plurality of fragments comprises more than 1000 fragments;

C) generating, using one or more processors associated with the computer system, one or more first state interval maps for one or more corresponding genomic regions using the first dataset, wherein:

each first state interval map in the one or more first state interval maps comprises a corresponding independent plurality of nodes, wherein the corresponding independent plurality of nodes comprises more than 50 nodes,

each first state interval map is represented as a hierarchical tree structure, wherein each node in the hierarchical tree structure corresponds to a genomic region, nodes at shallower depths represent larger genomic regions, nodes at deeper depths represent smaller genomic regions, and the hierarchical tree structure is constructed recursively by partitioning the genomic regions into sub-regions, and

each respective node in each corresponding independent plurality of nodes in the one or more first state interval maps is characterized by (1) corresponding start methylation site, (2) corresponding end methylation site and, (3) each different fragment methylation pattern observed across the first plurality of fragments in the first dataset between the corresponding start methylation site and the corresponding end methylation site of the respective node, (i) a representation of the different fragment methylation pattern and (ii) a count of fragments in the first dataset whose fragment methylation pattern begins at the corresponding start methylation site and ends at the corresponding end methylation site and has the different fragment methylation pattern;

D) generating, using the one or more processors, one or more second state interval maps for one or more corresponding genomic regions using the second dataset, wherein:

each second state interval map in the one or more second state interval maps comprises a corresponding independent plurality of nodes, wherein the corresponding independent plurality of nodes comprises more than 50 nodes, and

each respective node in each corresponding independent plurality of nodes in the one or more second state interval maps is characterized by a corresponding start methylation site, a corresponding end methylation site and, for each different fragment methylation pattern observed across the second plurality of fragments in the second dataset between the corresponding start methylation site and the corresponding end methylation site of the respective node, (i) a representation of the different fragment methylation pattern and (ii) a count of fragments in the second dataset whose fragment methylation pattern begins at the corresponding start methylation site and ends at the corresponding end methylation site and has the different fragment methylation pattern;

E) filtering, using the one or more processors, the one or more first state interval maps and the one or more second state interval maps to remove one or more nodes or corresponding genomic sub-regions that satisfy one or more exclusion criteria, wherein the one or more exclusion criteria comprises:

i) belonging to a blacklisted region of a genome;

ii) having a noise level greater than a threshold across non-cancer control samples; or

iii) having a low discriminatory power between cancer and non-cancer states;

F) constructing, using the one or more processors, a filtered data structure comprising the remaining first nodes and second nodes, and fragment methylation pattern representations associated with the remaining first nodes and second nodes, from the filtered one or more first state interval maps and the filtered one or more second state interval maps;

G) scanning, using the one or more processors, the filtered data structure for a plurality of qualifying methylation patterns, wherein each qualifying methylation pattern in the plurality of qualifying methylation patterns:

(i) has a length that is in a predetermined CpG site number range, within the fragment methylation patterns of the one or more first state interval maps and the one or more second state interval maps,

(ii) satisfies one or more selection criteria, and

(iii) spans a corresponding CpG interval between a corresponding initial CpG site and a corresponding final CpG site,

thereby identifying the plurality of qualifying methylation patterns that discriminates or indicates a cancer condition,

H) training, using the one or more processors, a classifier for classifying a state of the cancer condition using the plurality of qualifying methylation patterns and the associated filtered data structure, wherein the training comprises constructing a primary training dataset that comprises the plurality of qualifying methylation patterns as canonical sets of methylation state vectors, in conjunction with cell source labels corresponding to the first set of subjects and the second set of subjects, and wherein the classifier comprises a neural network;

I) applying, using the one or more processors, the trained classifier to a third dataset comprising fragment methylation patterns obtained from a biological sample from a test subject; and

J) receiving, at the computer system, output from the trained classifier indicating a state of the cancer condition in the test subject.

2 . The method of claim 1 , wherein the one or more selection criteria specifies that a methylation pattern:

(i) is represented in the one or more first state interval maps with a first frequency that satisfies a first frequency threshold,

(ii) is represented in the one or more first state interval maps with a coverage that satisfies a first state depth threshold, and

(iii) is represented in the one or more second state interval maps with a second frequency that satisfies a second frequency threshold.

3 . The method of claim 2 , wherein:

(i) the methylation pattern is represented in the one or more first state interval maps with a first frequency that satisfies a first frequency threshold when the frequency of the methylation pattern in the one or more first state interval maps exceeds the first frequency threshold,

(ii) the methylation pattern is represented in the one or more first state interval maps with a coverage that satisfies the first state depth threshold when the coverage of the methylation pattern in the one or more first state interval maps exceeds the first state depth threshold, and

(iii) the methylation pattern is represented in the one or more second state interval maps with a second frequency that satisfies the second frequency threshold when the frequency of the methylation pattern in the one or more second state interval maps is less than the second frequency threshold.

4 . The method of claim 1 , wherein the first and second plurality of fragments are cell-free nucleic acids.

5 . The method of claim 1 , wherein:

the one or more first state interval maps are a plurality of first state interval maps;

the one or more second state interval maps are a plurality of second state interval maps; the one or more corresponding genomic regions are a plurality of genomic regions; and each respective genomic region in the plurality of genomic regions is represented by a first state interval map in the plurality of first state interval maps and a second state interval map in the plurality of second state interval maps and wherein

the plurality of genomic regions is between 10 and 30, or

each genomic region in the plurality of genomic regions is a different human chromosome, or

the plurality of genomic regions consists of between two and 1000 genomic regions, between 500 and 5,000 genomic regions, between 1,000 and 20,000 genomic regions or between 5,000 and 50,000 genomic regions.

6 . The method of claim 5 , wherein the methylation sequencing of the A) obtaining and B) obtaining is targeted sequencing using a plurality of probes and each genomic region in the plurality of genomic regions is associated with at least one probe in the plurality of probes.

7 . The method of claim 1 , wherein the predetermined CpG number range is between 2 and 100 contiguous CpG sites in a human reference genome.

8 . The method of claim 1 , wherein there are more than 10,000 CpG sites, more than 25,000 CpG sites, more than 50,000 CpG sites, or more than 80,000 CpG sites across the one or more corresponding genomic regions.

9 . The method of claim 1 , wherein an average sequence read length of a corresponding plurality of sequence reads obtained by the methylation sequencing for a respective fragment is between 140 and 280 nucleotides.

10 . The method of claim 1 , wherein each genomic region in the one or more corresponding genomic regions represents between 500 base pairs and 10,000 base pairs of a human genome reference sequence.

11 . The method of claim 1 , wherein the methylation sequencing is i) whole-genome methylation sequencing or ii) targeted DNA methylation sequencing using a plurality of nucleic acid probes.

12 . A computer system for identifying a plurality of qualifying methylation patterns that discriminate or indicate a cancer condition, the computer system comprising:

at least one processor; and

a memory storing at least one program for execution by the at least one processor, the at least one program comprising instructions for:

A) obtaining a first dataset, in electronic form, wherein the first dataset comprises a corresponding fragment methylation pattern of each respective fragment in a first plurality of fragments, wherein the corresponding fragment methylation pattern of each respective fragment (i) is determined by a methylation sequencing of nucleic acids from a respective biological sample obtained from a corresponding subject in a first set of subjects and (ii) comprises a methylation state of each CpG site in a corresponding plurality of CpG sites in the respective fragment, and wherein the first plurality of fragments comprises more than 1000 fragments;

B) obtaining a second dataset, in electronic form, wherein the second dataset comprises a corresponding fragment methylation pattern of each respective fragment in a second plurality of fragments, wherein the corresponding fragment methylation pattern of each respective fragment (i) is determined by a methylation sequencing of nucleic acids from a respective biological sample obtained from a corresponding subject in a second set of subjects and (ii) comprises a methylation state of each CpG site in a corresponding plurality of CpG sites in the respective fragment, wherein each subject in the first set of subjects has a first state of the cancer condition and each subject in the second set of subjects has a second state of the cancer condition, and wherein the second plurality of fragments comprises more than 1000 fragments;

C) generating one or more first state interval maps for one or more corresponding genomic regions using the first dataset, wherein:

each first state interval map in the one or more first state interval maps comprises a corresponding independent plurality of nodes, wherein the corresponding independent plurality of nodes comprises more than 50 nodes,

each first state interval map is represented as a hierarchical tree structure, wherein each node in the hierarchical tree structure corresponds to a genomic region, nodes at shallower depths represent larger genomic regions, nodes at deeper depths represent smaller genomic regions, and the hierarchical tree structure is constructed recursively by partitioning the genomic regions into sub-regions, and

each respective node in each corresponding independent plurality of nodes in the one or more first state interval maps is characterized by a corresponding start methylation site, a corresponding end methylation site and, for each different fragment methylation pattern observed across the first plurality of fragments in the first dataset between the corresponding start methylation site and the corresponding end methylation site of the respective node, (i) a representation of the different fragment methylation pattern and (ii) a count of fragments in the first dataset whose fragment methylation pattern begins at the corresponding start methylation site and ends at the corresponding end methylation site and has the different fragment methylation pattern;

D) generating one or more second state interval maps for one or more corresponding genomic regions using the second dataset, wherein:

each second state interval map in the one or more second state interval maps comprises a corresponding independent plurality of nodes, wherein the corresponding independent plurality of nodes comprises more than 50 nodes, and

each respective node in each corresponding independent plurality of nodes in the one or more second state interval maps is characterized by a corresponding start methylation site, a corresponding end methylation site and, for each different fragment methylation pattern observed across the second plurality of fragments in the second dataset between the corresponding start methylation site and the corresponding end methylation site of the respective node, (i) a representation of the different fragment methylation pattern and (ii) a count of fragments in the second dataset whose fragment methylation pattern begins at the corresponding start methylation site and ends at the corresponding end methylation site and has the different fragment methylation pattern;

E) filtering the one more first state interval maps and the one or more second state interval maps to remove one or more nodes or corresponding genomic sub-regions that satisfy one or more exclusion criteria, wherein the one or more exclusion criteria comprises:

i) belonging to a blacklisted region of a genome;

ii) having a noise level greater than a threshold across non-cancer control samples; or

iii) having a low discriminatory power between cancer and non-cancer states:

F) constructing a filtered data structure comprising the remaining first nodes and second nodes, and fragment methylation pattern representations associated with the remaining first nodes and second nodes, from the filtered one or more first state interval maps and the filtered one or more second state interval maps;

G) scanning the filtered data structure for a plurality of qualifying methylation patterns, wherein each qualifying methylation pattern in the plurality of qualifying methylation patterns:

(i) has a length that is in a predetermined CpG site number range, within the fragment methylation patterns of the one or more first state interval maps and the one or more second state interval maps,

(ii) satisfies one or more selection criteria, and

(iii) spans a corresponding CpG interval between a corresponding initial CpG site and a corresponding final CpG site,

thereby identifying the plurality of qualifying methylation patterns that discriminates or indicates a cancer condition,

H) training a classifier for classifying a state of the cancer condition using the plurality of qualifying methylation patterns and the associated filtered data structure, wherein the training comprises constructing a primary training dataset that comprises the plurality of qualifying methylation patterns as canonical sets of methylation state vectors, in conjunction with cell source labels corresponding to the first set of subjects and the second set of subjects, and wherein the classifier comprises a neural network;

I) applying the trained classifier to a third dataset comprising fragment methylation patterns obtained from a biological sample from a test subject; and

J) receiving output from the trained classifier indicating a state of the cancer condition in the test subject.

13 . A non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor associated with a computer system, cause the processor to perform a method for identifying a plurality of qualifying methylation patterns that discriminate or indicate a cancer condition, the method comprising:

A) obtaining, at the computer system, a first dataset, in electronic form, wherein the first dataset comprises a corresponding fragment methylation pattern of each respective fragment in a first plurality of fragments, wherein the corresponding fragment methylation pattern of each respective fragment (i) is determined by a methylation sequencing of nucleic acids from a respective biological sample obtained from a corresponding subject in a first set of subjects and (ii) comprises a methylation state of each CpG site in a corresponding plurality of CpG sites in the respective fragment, and wherein the first plurality of fragments comprises more than 1000 fragments;

B) obtaining, at the computer system, a second dataset, in electronic form, wherein the second dataset comprises a corresponding fragment methylation pattern of each respective fragment in a second plurality of fragments, wherein the corresponding fragment methylation pattern of each respective fragment (i) is determined by a methylation sequencing of nucleic acids from a respective biological sample obtained from a corresponding subject in a second set of subjects and (ii) comprises a methylation state of each CpG site in a corresponding plurality of CpG sites in the respective fragment, wherein each subject in the first set of subjects has a first state of the cancer condition and each subject in the second set of subjects has a second state of the cancer condition, and wherein the second plurality of fragments comprises more than 1000 fragments;

C) generating, using the processor, one or more first state interval maps for one or more corresponding genomic regions using the first dataset, wherein:

each first state interval map in the one or more first state interval maps comprises a corresponding independent plurality of nodes, wherein the corresponding independent plurality of nodes comprises more than 50 nodes,

each first state interval map is represented as a hierarchical tree structure, wherein each node in the hierarchical tree structure corresponds to a genomic region, nodes at shallower depths represent larger genomic regions, nodes at deeper depths represent smaller genomic regions, and the hierarchical tree structure is constructed recursively by partitioning the genomic regions into sub-regions, and

each respective node in each corresponding independent plurality of nodes in the one or more first state interval maps is characterized by (1) a corresponding start methylation site, (2) a corresponding end methylation site and, (3) for each different fragment methylation pattern observed across the first plurality of fragments in the first dataset between the corresponding start methylation site and the corresponding end methylation site of the respective node, (i) a representation of the different fragment methylation pattern and (ii) a count of fragments in the first dataset whose fragment methylation pattern begins at the corresponding start methylation site and ends at the corresponding end methylation site and has the different fragment methylation pattern;

D) generating one or more second state interval maps for one or more corresponding genomic regions using the second dataset, wherein:

each second state interval map in the one or more second state interval maps comprises a corresponding independent plurality of nodes, wherein the corresponding independent plurality of nodes comprises more than 50 nodes, and

each respective node in each corresponding independent plurality of nodes in the one or more second state interval maps is characterized by (1) a corresponding start methylation site, (2) a corresponding end methylation site and, (3) for each different fragment methylation pattern observed across the second plurality of fragments in the second dataset between the corresponding start methylation site and the corresponding end methylation site of the respective node, (i) a representation of the different fragment methylation pattern and (ii) a count of fragments in the second dataset whose fragment methylation pattern begins at the corresponding start methylation site and ends at the corresponding end methylation site and has the different fragment methylation pattern;

E) filtering, using the processor, the one more first state interval maps and the one or more second state interval maps to remove one or more nodes or corresponding genomic sub-regions that satisfy one or more exclusion criteria, wherein the one or more exclusion criteria comprises:

i) belonging to a blacklisted region of a genome;

ii) having a noise level greater than a threshold across non-cancer control samples; or

iii) having a low discriminatory power between cancer and non-cancer states;

E) constructing, using the processor, a filtered data structure comprising the remaining first nodes and second nodes, and fragment methylation pattern representations associated with the remaining first nodes and second nodes, from the filtered one or more first state interval maps and the filtered one or more second state interval maps:

G) scanning the hierarchical tree structures of the one or more first state interval maps and the one or more second state interval maps for a plurality of qualifying methylation patterns, wherein each qualifying methylation pattern in the plurality of qualifying methylation patterns:

(i) has a length that is in a predetermined CpG site number range, within the fragment methylation patterns of the one or more first state interval maps and the one or more second state interval maps,

(ii) satisfies one or more selection criteria, and

(iii) spans a corresponding CpG interval between a corresponding initial CpG site and a corresponding final CpG site,

thereby identifying the plurality of qualifying methylation patterns that discriminates or indicates a cancer condition, and training a classifier for classifying a state of the cancer condition using the plurality of qualifying methylation patterns, wherein the training comprises constructing a primary training dataset that comprises the plurality of qualifying methylation patterns as canonical sets of methylation state vectors, in conjunction with cell source labels corresponding to the first set of subjects and the second set of subjects, and the classifier comprises a neural network, a logistic regression model, a support vector machine, a Naive Bayes model, a boosted trees model, a random forest model, a decision tree model, a multinomial logistic regression model, a linear model, or a linear regression model.

14 . The method of claim 1 , wherein the scanning the hierarchical tree structures of the one or more first state interval maps and the one or more second state interval maps for the plurality of qualifying methylation patterns comprises:

generating a query for a qualifying methylation pattern using hierarchical query propagation;

propagating the query down the hierarchical tree structure, with each node processing a corresponding genomic region independently; and

aggregating results of the qualifying patterns identified from child nodes at corresponding parent nodes as the query response propagates back up the hierarchical tree structure,

wherein the hierarchical query propagation reduces computational overhead by allowing intermediate nodes to aggregate results and discard irrelevant data from corresponding child nodes.

15 . The method of claim 1 , wherein the training of the classifier further comprises:

incorporating coefficients derived from one or more auxiliary training datasets to supplement the primary training dataset using transfer learning, wherein the coefficients are applied via matrix multiplication to obtain an intermediate classifier that is further trained using the primary training dataset.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Oct 13, 2021
From: GRAIL, INC.; SDG OPS, LLC
To: GRAIL, LLC
Reel/Frame 057788/0719 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2021
From: MELTON, COLLIN; HUBBELL, EARL; VENN, OLIVER CLAUDE
To: GRAIL, INC.
Reel/Frame 057696/0562 →
Continuity (2)
Provisional Application 62983443 · Feb 28, 2020
Related Publication 20210292845A1 · Sep 23, 2021
References Cited (119)
US 7601499B2 · Berka et al. · 2009 [cited by applicant]
US 11168356B2 · Lo · 2021 [cited by examiner]
US 11410750B2 · Gross · 2022 [cited by examiner]
US 11482303B2 · Nicula · 2022 [cited by examiner]
US 11581062B2 · Maher · 2023 [cited by examiner]
US 11685958B2 · Gross · 2023 [cited by examiner]
US 11725251B2 · Gross · 2023 [cited by examiner]
US 11773450B2 · Yip · 2023 [cited by examiner]
US 11783915B2 · Nicula · 2023 [cited by examiner]
US 11795513B2 · Gross · 2023 [cited by examiner]
US 12024750B2 · Gross · 2024 [cited by examiner]
US 20160017419A1 · Chiu et al. · 2016 [cited by applicant]
US 20170212428A1 · Roos et al. · 2017 [cited by applicant]
US 20180237863A1 · Namsaraev et al. · 2018 [cited by applicant]
US 20190177792A1 · Liu et al. · 2019 [cited by applicant]
US 20190287649A1 · Filippova et al. · 2019 [cited by applicant]
US 20190287652A1 · Gross et al. · 2019 [cited by applicant]
US 20200005899A1 · Nicula et al. · 2020 [cited by applicant]
US 20200340064A1 · Gross et al. · 2020 [cited by applicant]
US 20200365229A1 · Fields et al. · 2020 [cited by applicant]
US 20200385813A1 · Venn · 2020 [cited by applicant]
US 20210065842A1 · Valouev et al. · 2021 [cited by applicant]
US 20210285042A1 · Singh et al. · 2021 [cited by applicant]
US 20210292845A1 · Melton et al. · 2021 [cited by applicant]
US 20210313006A1 · Gross et al. · 2021 [cited by applicant]
US 20210327534A1 · Nicula et al. · 2021 [cited by applicant]
US 20210358626A1 · Nicula et al. · 2021 [cited by applicant]
WO WO2014043763 · 2014 [cited by applicant]
WO WO2017212428 · 2017 [cited by applicant]
WO WO2018081130A1 · 2018 [cited by applicant]
WO WO2019195268A2 · 2019 [cited by applicant]
WO WO2019204360A1 · 2019 [cited by applicant]
WO WO2019209954 · 2019 [cited by applicant]
WO WO2020069350A1 · 2020 [cited by applicant]
WO WO2020132148A1 · 2020 [cited by applicant]
WO WO2020154682A3 · 2020 [cited by applicant]
WO WO2020237184 · 2020 [cited by applicant]
Doerge et al. (2002) Mapping and analysis of quantitative trait loci in experimental populations. Nature reviews genetics, vol. 3, p. 43-53. (Year: 2002). [cited by examiner]
Bian, 2018, “Comparing the performance of selected variant callers using synthetic data and genome segmentation,” BMC Bioinformatics 19:429. [cited by applicant]
Chaudhary et al., 2017, Journal of Clinical Oncology, 35(15), suppl.e14529. [cited by applicant]
Grunau et al., 2001, “MethDB—a public database for DNA methylation data,” Nucleic Acids Research 29(1), 270-274. [cited by applicant]
Huang et al., 2021, “MethHC 2.0: information repository of DNA methylation and gene expression in human cancer,” Nucleic Acids Research 49(D1), D1268-D1275. [cited by applicant]
Hachiya et al., 2017, “Genome-wide identification of inter-individually variable DNA methylation sites improves the efficacy of epigenetic association studies,” NPJ Genom Med. 2017. 2:11. [cited by applicant]
Karczewski et al., 2019, “Variation across 141,456 human exomes and genomes reveals the spectrum of loss-of-function intolerance across human protein-coding genes,” bioRxiv doi.org/10.1101/531210. [cited by applicant]
Ongenaert et al., “PubMeth: a cancer methylation database combining text-mining and expert annotation,” Nucleic Acids Research: doi:10.1093/nar/gkm788. [cited by applicant]
Sherry et al., 2001, “dbSNP: the NCBI database of genetic variation” Nuc. Acids. Res. 29, 308-311. [cited by applicant]
U.S. Appl. No. 62/847,223, filed May 13, 2019, Fields et al. [cited by applicant]
U.S. Appl. No. 62/985,258, filed Mar. 4, 2020, Nicula et al. [cited by applicant]
U.S. Appl. No. 62/877,755, filed Jul. 23, 2019, Valouev et al. [cited by applicant]
U.S. Appl. No. 63/003,087, filed Mar. 31, 2020, Gross et al. [cited by applicant]
U.S. Appl. No. 16/839,469, filed Apr. 3, 2020, Yip et al. [cited by applicant]
International Search Report and Written Opinion dated May 31, 2021 for PCT Application No. PCT/US2020/066217. 23 pages. [cited by applicant]
Invitation to Pay Additional Fees and Partial International Search Report dated Apr. 8, 2021 for PCT Application No. PCT/US2020/066217. 16 pages. [cited by applicant]
Agresti. An Introduction to Categorical Data Analysis. John Wiley & Son, New York. 1996; Chapter 5, pp. 103-144. [cited by applicant]
Boser, et al. A training algorithm for optimal margin classifiers. Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, PA. 1992; 142-152. [cited by applicant]
Breiman. Random Forests—Random Features. Technical Report 567, Statistics Department, U.C. Berkeley. 1999. 29 pages. [cited by applicant]
Casadio, et al. Urine cell-free DNA integrity as a marker for early bladder cancer diagnosis: preliminary data. Urol Oncol. 2013; 31(8):1744-1750. [cited by applicant]
Cavalcante and Santor. annotatr: genomicBreiman regions in context. Bioinformatics. 2017; 33(15):2381-2383. [cited by applicant]
Chan, et al. Cell-free nucleic acids in plasma, serum and urine: a new tool in molecular diagnosis. Ann Clin Biochem. 2003; 40(Pt 2):122-130. [cited by applicant]
Cristianini and Shawe-Taylor. An Introduction to Support Vector Machines. Chapter 6, Cambridge University Press, Cambridge. 2000; pp. 93-124. [cited by applicant]
De Mattos-Arruda, et al. Cell-free circulating tumour DNA as a liquid biopsy in breast cancer. Mol Oncol. 2016; 10(3):464-474. [cited by applicant]
Duda. Pattern Classification. Second Edition, John Wiley & Sons, Inc. 2001; pp. 259, 262-265. [cited by applicant]
Duda, et al. Pattern Classification and Scene Analysis. John Wiley & Sons, Inc., New York. 1973. [cited by applicant]
Everitt. Cluster analysis (3d ed.). Wiley, New York, N.Y. 1993. [cited by applicant]
Fernandes, et al. Transfer Learning with Partial Observability Applied to Cervical Cancer Screening. Pattern Recognition and Image Analysis: 8th Iberian Conference Proceedings. 2017. pp. 243-250. [cited by applicant]
Frenel, et al. Serial next-generation sequencing of circulating cell-free DNA evaluating tumor clone response to molecularly targeted drug administration. Clin Cancer Res. 2015; 21(20): 4586-4596. [cited by applicant]
Furey, et al. Support vector machine classification and validation of cancer tissue samples using microarray expression data. Bioinformatics. 2000; 16:906-914. [cited by applicant]
Goessl, et al. Fluorescent methylation-specific polymerase chain reaction for DNAbased detection of prostate cancer in bodily fluids. Cancer Res. 2000; 60(21):5941-5845. [cited by applicant]
Hao, et al. Circulating cell-free DNA in serum as a biomarker for diagnosis and prognostic prediction of colorectal cancer. Br J Cancer. 2014; 111(8):1482-1489. [cited by applicant]
Hassoun. Adaptive Multilayer Neural Networks I. Fundamentals of Artificial Neural Networks, Chapter 5, Massachusetts Institute of Technology. 1995. 87 pages. [cited by applicant]
Hastie, et al. Additive Models, Trees, and Related Models. The Elements of Statistical Learning, Second Edition, Springer Science+Business Media, LLC. 2009; 295-336. [cited by applicant]
Heitzer, et al. Circulating tumor DNA as a liquid biopsy for cancer. Clin Chem. 2015; 61(1):112-123. [cited by applicant]
Heitzer, et al. Establishment of tumorspecific copy No. alterations from plasma DNA of patients with cancer. Int J Cancer. 2013; 133(2):346-356. [cited by applicant]
Jensen, et al. High-Throughput Massively Parallel Sequencing for Fetal Aneuploidy Detection from Maternal Plasma. PLoS One. 2013. 8; e57381. [cited by applicant]
Kaufman, et al. Finding Groups in Data: An Introduction to Cluster Analysis. Wiley, New York, N.Y. 1990. [cited by applicant]
Kim, et al. Circulating cell-free DNA as a promising biomarker in patients with gastric cancer: diagnostic validity and significant reduction of cfDNA after surgical resection. Ann Surg Treat Res. 2014; 86(3):136-142. [cited by applicant]
Langmead, et al. Fast gapped-read alignment with Bowtie 2. Nat Methods. 2012; 9:357-359. [cited by applicant]
Larochelle, et al. Exploring strategies for training deep neural networks. J Mach Learn Res. 2009; 10:1-40. [cited by applicant]
Li, et al. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009; 25(14):1754-1760. [cited by applicant]
Li, et al. Mappability and read length. Frontiers In Genetics. 2014; 5(381):1-7. [cited by applicant]
Lo, et al. Maternal Plasma DNA Sequencing Reveals the Genome-Wide Genetic and Mutational Profile of the Fetus. Science Translational Medicine, 2010; vol. 2, Issue 61, 14 pages, [online], retrieved from the internet URL:… [cited by applicant]
Raptis, et al. Quantitation and characterization of plasma DNA ins normals and patients with systemic lupus erythematosu. J Clin Invest. 1980; 66(6): 1391-1399. [cited by applicant]
Salvi, et al. Cell-free DNA as a diagnostic marker for cancer: current insights. Onco Targets Ther. 2016; 9: 6549-6559. [cited by applicant]
Shao, et al. Quantitative analysis of cell-free DNA in ovarian cancer. Oncol Lett. 2015; 10(6):3478-3482. [cited by applicant]
Shapiro, et al. Determination of circulating DNA levels in patients with benign or malignant gastrointestinal disease. Cancer. 1983; 51(11): 2116-2120. [cited by applicant]
Smith, et al. Evaluating alignment and variant-calling software for mutation identification in [cited by applicant]
Song. Comparison of co-expression measures: mutual information, correlation, and model based indices. BMC bioinformatics. 2012; 13.1; 1-21. [cited by applicant]
Sozzi. Quantification of free circulating DNA as a diagnostic marker in lung cancer. J Clin Oncol. 2003; 21(21): 3902-3908. [cited by applicant]
Stroun, et al. Neoplastic characteristics of the DNA found in the plasma of cancer patients. Oncology. 1989; 46(5): 318-322. [cited by applicant]
Swanton, et al. Phylogenetic ctDNA analysis depicts early stage lung cancer evolution. Nature 2017; 545(7655):446-451. [cited by applicant]
Vapnik. Perceptrons and Their Generalizations. Statistical Learning Theory, Chapter 9, Wiley, New York. 1998; 375-400. [cited by applicant]
Vincent, et al. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. J Mach Learn Res. 2010; 11:3371-3408. [cited by applicant]
Yang, et al. DistAI: An Inter-pattern Distance-based Constructive Learning Algorithm. Intelligent Data Analysis. 1999; 3(1):55-73. [cited by applicant]
Yoon. Hidden Markov Models and their Applications in Biological Sequence Analysis. Curr. Genomics. 2009; 10(6):402-415. [cited by applicant]
Zonta, et al. Assessment of DNA integrity, applications for cancer research. Adv Clin Chem. 2015; 70:197-246. [cited by applicant]
Chan, et al. Noninvasive detection of cancer-associated genome-wide hypomethylation and copy number aberrations by plasma DNA bisulfite sequencing. PNAS. Nov. 19, 2013; vol. 110, No. 47, 18761-18768. [cited by applicant]
Fan, et al. Methods for genome-wide DNA methylation analysis in human cancer. Briefings in Functional Genomics. 2016; 15(6): 432-442. [cited by applicant]
Lapin, et al. Fragment size and level of cell-free DNA provide prognostic information in patientes with advanced pancreatic cancer. J Translational Medicine. 2018; 16(300). [cited by applicant]
Lee, et al. Analyzing the cancer methylome through targeted bisulfite sequencing. Cancer Letters. 2013; vol. 340: 171-178. [cited by applicant]
Liu, et al. Bisulfite-free direct detection of 5- methylcytosine and 5-hydroxymethylcytosine at base resolution. Nature Biotechnology. 2019; 37:424-429. [cited by applicant]
International Search Report and Written Opinion dated Jun. 18, 2021 for PCT Application No. PCT/US2021/020012. 16 pages. [cited by applicant]
Anonymous. Intervalmap.pkg.go.dec. Published Nov. 11, 2019. Retrieved from https://pkg.go.dev/github.com/grailbio/[email protected]/intervalmap. 6 pages. [cited by applicant]
Du et al., 2010, BMC Bioinformatics 11:587. [cited by applicant]
Gross, S. et al., U.S. Appl. No. 62/642,480, filed Mar. 13, 2018. [cited by applicant]
Gross, S. et al., U.S. Appl. No. 62/972,375, filed Feb. 10, 2020. [cited by applicant]
Jones, 2002, Oncogene 21:5358-5360. [cited by applicant]
Kang, Shuli et al., “CancerLocator: non-invasive cancer diagnosis and tissue-of-origin prediction using methylation profiles of cell-free DNA”, Genome Biology, vol. 18, No. 1. [cited by applicant]
Klein et al., 2018, “Development of a comprehensive cell-free DNA (cfDNA) assay for early detection of multiple tumor types: The Circulating Cell-free Genome Atlas (CCGA) study,” J. Clin. Oncology 36(15), 12021-12021. [cited by applicant]
Lee et al., 2013, “MatchTree: Flexible, scalable, and fault-tolerant wide-area resource discovery with distributed matchmaking and aggregation,” Fut Gen Comp Sys 29, 1596-1610. [cited by applicant]
Liu et al., 2019, “Genome-wide cell-free DNA (cfDNA) methylation signatures and effect on tissue of origin (TOO) performance,” J. Clin. Oncology 37(15), 3049-3049. [cited by applicant]
Masser et al., 2015, “Targeted DNA Methylation Analysis by Next-generation Sequencing,” J. Vis. Exp. (96). [cited by applicant]
Newman, J., et al., U.S. Appl. No. 62/834,904, filed Apr. 16, 2019. [cited by applicant]
Nicula, V., et al., U.S. Appl. No. 62/948,129, filed Dec. 13, 2019. [cited by applicant]
Paska, et al., 2015, Biochemia Medica 25(2):161-176. [cited by applicant]
Wang et al., 2015, “Syntax-based Deep Matching of Short Texts,” arXiv: 1503.02427v6[cs.CL]. [cited by applicant]
Warton, et al., 2015, Front Mol Biosci, 2(13). [cited by applicant]
Wald, 2007, “On Fast Construction of SAH-based Bounding vol. Hierarchies,” IEEE, doi:10.1109/RT.2007.4342588. [cited by applicant]
Venn, Oliver Claud, U.S. Appl. No. 62/781,549, filed Dec. 18, 2018. [cited by applicant]
Ziller et al., 2015, “Coverage recommendations for methylation analysis by whole-genome bisulfite sequencing,” Nature Methods. 12(3):230-232. [cited by applicant]