Systems and methods to detect rare mutations and copy number variation
The present disclosure provides a system and method for the detection of rare mutations and copy number variations in cell free polynucleotides. Generally, the systems and methods comprise sample preparation, or the extraction and isolation of cell free polynucleotide sequences from a bodily fluid; subsequent sequencing of cell free polynucleotides by techniques known in the art; and application of bioinformatics tools to detect rare mutations and copy number variations as compared to a reference. The systems and methods also may contain a database or collection of different rare mutations or copy number variation profiles of different diseases, to be used as additional references in aiding detection of rare mutations, copy number variation profiling or general genetic profiling of a disease.
1. A method, comprising:
(a) providing a population of cell-free DNA (“cfDNA”) molecules obtained from a bodily sample from a subject;
(b) converting the population of cfDNA molecules into a population of tagged parent polynucleotides, wherein each of the tagged parent polynucleotides comprises (i) a sequence from a cfDNA molecule of the population of cfDNA molecules, and (ii) an identifier sequence comprising one or more polynucleotide barcodes,
wherein the population of cfDNA molecules is tagged with n different unique identifiers, wherein n is at least 2 and no more than 100,000*z, wherein z is a mean of an expected number of duplicate molecules in the population of cfDNA molecules that map to identical start and stop positions on a reference sequence;
(c) amplifying the population of tagged parent polynucleotides to produce a corresponding population of amplified progeny polynucleotides;
(d) sequencing the population of amplified progeny polynucleotides to produce a set of sequence reads;
(e) mapping sequence reads of the set of sequence reads to the reference sequence;
(f) grouping the sequence reads into a plurality of families using sequence information from the mapping in (e), each of the families comprising sequence reads comprising the same identifier sequence and having the same start and stop positions, whereby each of the families comprises sequence reads amplified from the same tagged parent polynucleotide;
(g) at each genetic locus of a plurality of genetic loci in the reference sequence, collapsing sequence reads in each family to yield a base call for each family at the genetic locus; and
(h) determining a frequency of one or more bases called at the locus from among the families.
2. The method of claim 1 , wherein the bodily sample is blood, plasma, or serum.
3. The method of claim 1 , wherein the bodily sample is obtained or derived from a subject having cancer.
4. The method of claim 1 , wherein the bodily sample comprises between 1 nanogram (ng) and 100 ng of nucleic acid molecules.
5. The method of claim 1 , wherein the bodily sample of cfDNA molecules comprises 100 to 100,000 haploid human genome equivalents.
6. The method of claim 1 , wherein the converting comprises attaching one or more identifier sequences to the cfDNA using blunt-end ligation or sticky-end ligation.
7. The method of claim 1 , wherein n different unique identifiers is no more than 10,000*z.
8. The method of claim 1 , wherein n different unique identifiers is no more than 1,000*z.
9. The method of claim 1 , wherein the population of cfDNA molecules is tagged with from 2 to 100,000 different unique identifiers.
10. The method of claim 7 , wherein the population of cfDNA molecules is tagged with from 2 to 10,000 different unique identifiers.
11. The method of claim 1 , wherein the one or more polynucleotide barcodes have a length of between 5 and 20 nucleotides.
12. The method of claim 1 , further comprising selectively enriching regions from a genome or transcriptome of the subject prior to sequencing.
13. The method of claim 1 , wherein (d) comprises sequencing a panel of actionable cancer-related genes.
14. The method of claim 1 , further comprising filtering out one or more of the sequence reads from the set of sequence reads that fail to meet a quality threshold.
15. The method of claim 1 , wherein the reference sequence comprises a sequence from a human genome assembly.
16. The method of claim 1 , wherein the base call comprises voting, averaging, maximum a posteriori or maximum likelihood detection, dynamic programming, Bayesian methods, hidden Markov methods or support vector machine methods.
17. The method of claim 1 , wherein determining the frequency comprises detecting a rare mutation.
18. The method of claim 1 , further comprising generating a set of consensus sequences from the sequence reads, wherein determining the frequency of one or more bases comprises detecting a presence of a sequence variation in the set of consensus sequences compared with the reference sequence.
19. The method of claim 18 , further comprising determining that the subject has cancer when the presence of the sequence variation is detected.
20. The method of claim 1 , further comprising detecting, at one or more loci, at least one single nucleotide variant (SNV).