UNIQUE MAPPER TOOL FOR EXCLUDING REGIONS WITHOUT ONE-TO-ONE MAPPING BETWEEN A SET OF TWO REFERENCE GENOMES
A first reference genome is segmented into a plurality of bins and high-quality sequenced reads are mapped on a bin-by-bin basis to the plurality of bins in the first reference genome, and a second reference genome is segmented into a plurality of bins and high-quality sequenced reads are mapped on a bin-by-bin basis to the plurality of bins in the second reference genome. A best-mapped bin is identified in the second reference genome based on the greatest degree of match between the best-mapped bin in the second reference genome and a corresponding bin in the first reference genome.
1 . A computer-implemented method of identifying and excluding regions that do not have one-to-one mapping between a first reference genome and a second reference genome, including:
accessing sequenced reads of a sample of a target species;
identifying and removing, from the sequenced reads, low-quality sequenced reads based on applying a mapping quality filter to the sequenced reads, thereby pruning high-quality sequenced reads from the sequenced reads;
segmenting a non-target reference genome of a non-target species into a plurality of bins, and then, on a bin-by-bin basis, mapping the high-quality sequenced reads to the plurality of bins in the non-target reference genome;
segmenting a pseudo-target reference genome of a pseudo-target species into the plurality of bins, and then, on the bin-by-bin basis, mapping the high-quality sequenced reads to the plurality of bins in the pseudo-target reference genome;
identifying a best-mapped bin in the pseudo-target reference genome based on a greatest degree of match between the best-mapped bin in the pseudo-target reference genome and a corresponding bin in the non-target reference genome, wherein a degree of match between corresponding bins in the pseudo-target reference genome and the non-target reference genome is determined by a number of reads mapped between the corresponding bins;
generating a unique-mapper score for the pseudo-target reference genome based on a number of reads mapped between the best-mapped bin in the pseudo-target reference genome and the corresponding bin in the non-target reference genome; and
using the unique-mapper score to identify and exclude low-quality sequenced reads.
2 . The computer-implemented method of claim 1 , wherein low-quality sequenced reads include stop-gained variants.
3 . The computer-implemented method of claim 1 , wherein a plurality of cascading filters the applied to the sequenced reads of the sample of the target species to filter low-quality sequenced reads.
4 . The computer-implemented method of claim 3 , wherein a filter within the plurality of cascading filters is configured to detect and exclude genetic regions possessing incorrect gene annotation in a reference genome.
5 . The computer-implemented method of claim 3 , wherein a filter within the plurality of cascading filters is configured to detect and exclude codons that do not match between the pseudo-target species reference genome and the non-target species reference genome.
6 . The computer-implemented method of claim 3 , wherein a filter within the plurality of cascading filters is configured to detect and exclude genes within a reference genome possessing a skewed distribution of variant classifier scores in comparison to a distribution of variant classifier scores for a complete reference genome.
7 . The computer-implemented method of claim 3 , wherein a filter within the plurality of cascading filters is configured to detect and exclude genes within a reference genome possessing deviations from a Hardy-Weinberg equilibrium.
8 . The computer-implemented method of claim 3 , wherein a filter within the plurality of cascading filters is configured to detect and remove single nucleotide polymorphisms with a random forest score of greater than 0.17.
9 . The computer-implemented method of claim 1 , wherein a one-to-one mapping describes a fraction of a number of reads in one bin within non-target reference genome map to a single corresponding region within a pseudo-target reference genome.
10 . The computer-implemented method of claim 9 , wherein consecutive identical bins are collectively considered as a single bin and do not exclude a possibility of a one-to-one mapping.
11 . The computer-implemented method of claim 9 , wherein more than two nonconsecutive identical bins are considered duplicate regions and exclude a possibility of a one-to-one mapping.
12 . The computer-implemented method of claim 1 , wherein a bin describes a one kilobase (kb) region within a reference genome.
13 . The computer-implemented method of claim 1 , wherein a fraction of reads mapped is detected for a best-mapped region as determined by the unique-mapper score.
14 . The computer-implemented method of claim 1 , wherein the unique-mapper score is determined by an average of a top fraction across samples for each reference genome.
15 . The computer-implemented method of claim 1 , wherein the unique-mapper score is configured as a filter to exclude sequenced reads with a mapper score less than 20.
16 . The computer-implemented method of claim 1 , wherein the pseudo-target species is a human.
17 . The computer-implemented method of claim 1 , wherein the pseudo-target species is a non-human primate.
18 . The computer-implemented method of claim 1 , wherein the non-target species is a human.
19 . The computer-implemented method of claim 1 , wherein the non-target species is a non-human primate.
20 . The computer-implemented method of claim 1 , wherein the target species is a human.
21 . The computer-implemented method of claim 1 , wherein the target species is a non-human primate.
22 . The computer-implemented method of claim 1 , wherein the target species and non-target species are homologous.
23 . The computer-implemented method of claim 1 , wherein the target species and pseudo-target species are homologous.
24 . The computer-implemented method of claim 1 , wherein the quality of a variant identified from sequenced reads of a target species is a proxy for evolutionary constraint on a variant gene.
25 . A system including one or more processors coupled to memory, the memory loaded with computer instructions to identify and exclude regions that do not have one-to-one mapping between a first reference genome and a second reference genome, the instructions, when executed on the processors, implement actions comprising:
accessing sequenced reads of a sample of a target species;
identifying and removing, from the sequenced reads, low-quality sequenced reads based on applying a mapping quality filter to the sequenced reads, thereby pruning high-quality sequenced reads from the sequenced reads;
segmenting a non-target reference genome of a non-target species into a plurality of bins, and then, on a bin-by-bin basis, mapping the high-quality sequenced reads to the plurality of bins in the non-target reference genome;
segmenting a pseudo-target reference genome of a pseudo-target species into the plurality of bins, and then, on the bin-by-bin basis, mapping the high-quality sequenced reads to the plurality of bins in the pseudo-target reference genome;
identifying a best-mapped bin in the pseudo-target reference genome based on a greatest degree of match between the best-mapped bin in the pseudo-target reference genome and a corresponding bin in the non-target reference genome, wherein a degree of match between corresponding bins in the pseudo-target reference genome and the non-target reference genome is determined by a number of reads mapped between the corresponding bins;
generating a unique-mapper score for the pseudo-target reference genome based on a number of reads mapped between the best-mapped bin in the pseudo-target reference genome and the corresponding bin in the non-target reference genome; and
using the unique-mapper score to identify and exclude low-quality sequenced reads.
26 . The system of claim 25 , wherein low-quality sequenced reads include stop-gained variants.
27 . The system of claim 25 , wherein a plurality of cascading filters the applied to the sequenced reads of the sample of the target species to filter low-quality sequenced reads.
28 . The system of claim 27 , wherein a filter within the plurality of cascading filters is configured to detect and exclude genetic regions possessing incorrect gene annotation in a reference genome.
29 . The system of claim 27 , wherein a filter within the plurality of cascading filters is configured to detect and exclude codons that do not match between the pseudo-target species reference genome and the non-target species reference genome.
30 . The system of claim 27 , wherein a filter within the plurality of cascading filters is configured to detect and exclude genes within a reference genome possessing a skewed distribution of variant classifier scores in comparison to a distribution of variant classifier scores for a complete reference genome.