Fragmentation for measuring methylation and disease
Fragmentation of cell-free DNA molecules is measured and used for various purposes, including determining methylation, e.g., at a particular site of a DNA molecule, at a particular genomic site in a reference genome for a biological sample (e.g., plasma, serum, urine, saliva) of cell-free DNA of a subject, or for a particular region in the reference genome for the biological sample (also just referred to as a sample). Various types of fragmentation measurements can be used, e.g., end motifs and cleavage profiles. Another purpose is determining a fractional concentration of DNA of a particular tissue type (e.g., clinically-relevant DNA). Another purpose is determining a pathology of a subject using a biological sample including cell-free DNA. The cell-free DNA can be of the subject or of a pathogen (e.g., a virus) in the subject's sample. Sites/regions that are hypermethylated, hypomethylated, 5hmC-enriched, and 5hmC-depleted for a particular tissue type can be used.
1 . A method for measuring methylation of a CpG site in a genome of a subject using cell-free DNA molecules, the method comprising:
analyzing a plurality of cell-free DNA molecules from a biological sample of the subject, wherein the plurality of cell-free DNA molecules is not treated using a process that differentially modifies nor differentially recognizes DNA molecules depending on their methylation status, wherein the analyzing includes at least one selected from a group consisting of massively parallel sequencing and PCR that provides sequence reads of the plurality of cell-free DNA molecules, and wherein analyzing each cell-free DNA molecule of the plurality of cell-free DNA molecules includes:
determining, using one or more of the sequence reads, a genomic position in a reference genome corresponding to at least one end of the cell-free DNA molecule;
determining, using the sequence reads, a first amount of the plurality of cell-free DNA molecules from the biological sample ending at a first position within a window around the CpG site, the first position being between-1 to +1 position of the window, the CpG site having a 0 position being C;
determining one or more other amounts of the plurality of cell-free DNA molecules, and wherein the first amount, the one or more other amounts, or both include one or more cell-free DNA molecules of the plurality of cell-free DNA molecules that end at a position (1) that is in a region including the CpG site and (2) that is other than the 0 position for the CpG site or any other CpG site; and
determining, using the first amount and the one or more other amounts, a classification for methylation of the CpG site in the genome of the subject.
2 . The method of claim 1 , further comprising:
determining a second amount of cell-free DNA molecules ending at a second position within the window around the CpG site, the second position being different than the first position, the one or more other amounts including the second amount.
3 . The method of claim 2 , wherein the first position is the 0 position, and wherein the second position is at +1 or −1 from the CpG site.
4 . The method of claim 2 , wherein the window is at least −2 to +2 from the CpG site.
5 . The method of claim 2 , wherein determining the classification includes:
determining a separation value using the first amount and the second amount; and
comparing the separation value to a calibration value, wherein the calibration value is determined using cell-free DNA molecules that are from one or more calibration samples and that are located at CpG sites having known classifications.
6 . The method of claim 5 , wherein the classification is a quantitative value, and wherein comparing the separation value to the calibration value includes comparing the separation value to a calibration function.
7 . The method of claim 6 , wherein the quantitative value is a range that is 30% or less.
8 . The method of claim 2 , wherein the CpG site is a first CpG site and the classification is a first classification, and wherein the window includes a second CpG site and at least two positions other than the first CpG site and the second CpG site, and the method further comprising:
determining a third amount of cell-free DNA molecules ending at the second CpG site;
for each position of the at least two positions within the window:
determining a respective amount of cell-free DNA molecules ending at the position, thereby determining respective amounts including the second amount, wherein the at least two positions include the first position;
generating a feature vector including the respective amounts, the first amount, and the third amount; and
inputting the feature vector into a machine learning model as part of determining the classification for the first CpG site and determining a second classification of the second CpG site, wherein the machine learning model is trained using cell-free DNA molecules located within windows around CpG sites having known classifications.
9 . The method of claim 1 , wherein the window includes at least two positions other than the first position, the method further comprising:
for each position of the at least two positions within the window:
determining a respective amount of cell-free DNA molecules ending at the position; and
comparing the first amount of cell-free DNA molecules ending at the first position to the respective amount of cell-free DNA molecules ending at the position as part of determining the classification, wherein the one or more other amounts include the respective amount.
10 . The method of claim 1 , wherein the window includes at least two positions other than the first position, the method further comprising:
for each position of the at least two positions within the window:
determining a respective amount of cell-free DNA molecules ending at the position, thereby determining respective amounts, wherein the one or more other amounts include the respective amount;
generating a feature vector including the respective amounts and the first amount; and
inputting the feature vector into a machine learning model as part of determining the classification, wherein the machine learning model is trained using cell-free DNA molecules located within windows around CpG sites having known classifications.
11 . The method of claim 10 , wherein the feature vector forms a matrix with each row corresponding to a base of a strand, and wherein a column includes a non-zero amount in the row corresponding to the base at a respective position.
12 . The method of claim 10 , wherein the machine learning model is a convolutional neural network.
13 . The method of claim 10 , wherein the first position and the at least two positions include all positions within the window, the window being at least +4 to −4 from the CpG site.
14 . The method of claim 10 , wherein the machine learning model uses a sequence context within the window.
15 . The method of claim 14 , wherein the machine learning model is trained for the sequence context within the window.
16 . The method of claim 14 , wherein the feature vector includes the sequence context.
17 . The method of claim 1 , wherein the classification for methylation indicates that the CpG site is in a hypermethylated state or a hypomethylated state for the cell-free DNA molecules at the CpG site, wherein the hypermethylated state indicates a methylation density above a first threshold that is at least 70%, and wherein the hypomethylated state indicates the methylation density below a second threshold that is 30% or less.
18 . The method of claim 1 , wherein the first amount is normalized.
19 . The method of claim 18 , wherein the normalization uses the number of the plurality of cell-free DNA molecules ending within a region including the CpG site, wherein the one or more other amounts include the number of the plurality of cell-free DNA molecules ending within the region including the CpG site.
20 . The method of claim 18 , wherein the normalization uses a number of the plurality of cell-free DNA molecules covering the CpG site, wherein the one or more other amounts include the number of the plurality of cell-free DNA molecules ending within a region including the CpG site.
21 . The method of claim 18 , wherein the normalization uses an average or median depth of the plurality of cell-free DNA molecules in a region including the CpG site, wherein the one or more other amounts includes the average or median depth.
22 . The method of claim 1 , further comprising:
determining, using the first amount, another classification for methylation of a different CpG site in the genome of the subject, the different CpG site being within 600 nucleotides (nt) downstream from the 5′ end of the CpG site.
23 . The method of claim 1 , wherein analyzing the plurality of cell-free DNA molecules includes sequencing the plurality of cell-free DNA molecules.
24 . The method of claim 1 , wherein analyzing the plurality of cell-free DNA molecules includes using PCR.
25 . The method of claim 24 , wherein the PCR targets sequences in a repeat region.
26 . The method of claim 1 , wherein analyzing the plurality of cell-free DNA molecules includes analyzing at least 10,000 cell-free DNA molecules.
27 . The method of claim 1 , wherein determining the genomic position of each of the plurality of cell-free DNA molecules includes aligning, using a computer system, one or more sequence reads of the cell-free DNA molecule to the reference genome.
28 . The method of claim 27 , wherein the subject is a human, and wherein the reference genome is a reference human genome.
29 . The method of claim 1 , further comprising:
outputting, by a computer system, the classification for methylation of the CpG site.