VARIANT CALLING WITHOUT A TARGET REFERENCE GENOME
The technology disclosed relates to determining feasibility of using a reference genome of a non-target species for variant calling a sample of a target species. In particular, the technology disclosed relates to mapping sequenced reads of a sample of a target species to a reference genome of a non-target species to detect a first set of variants in the sequenced reads of the sample of the target species, and mapping the sequenced reads of the sample of the target species to a reference genome of a pseudo-target species to detect a second set of variants in the sequenced reads of the sample of the target species.
1 . A computer-implemented method of determining feasibility of using a reference genome of a non-target species for variant calling a sample of a target species, including:
mapping sequenced reads of a sample of a target species to a reference genome of a non-target species to detect a first set of variants in the sequenced reads of the sample of the target species;
mapping the sequenced reads of the sample of the target species to a reference genome of a pseudo-target species to detect a second set of variants in the sequenced reads of the sample of the target species;
comparing the first set of variants and the second set of variants, and identifying a subset of true positive variants that are common between the first set of variants and the second set of variants;
comparing the first set of variants and the second set of variants, and identifying a subset of false positive variants that are present in the second set of variants but absent from the first set of variants; and
based on a count of the subset of false positive variants determining the feasibility of using the reference genome of the non-target species for variant calling the target species.
2 . The computer-implemented method of claim 1 , wherein the pseudo-target species is the target species.
3 . The computer-implemented method of claim 1 , wherein the pseudo-target species is different from the target species.
4 . The computer-implemented method of claim 3 , wherein the pseudo-target species is homologous with the target species.
5 . The computer-implemented method of claim 1 , wherein the non-target species is a human.
6 . The computer-implemented method of claim 1 , wherein the target species is a non-human primate.
7 . The computer-implemented method of claim 1 , further including detecting the second set of variants by mapping the sequenced reads of the sample of the target species to the reference genome of the target species, and then lifting-over the mapped sequenced reads of the sample of the target species to the reference genome of the non-target species.
8 . The computer-implemented method of claim 1 , further including applying a first filter to filter out low-quality variants from the first set of variants and the second set of variants.
9 . The computer-implemented method of claim 1 , further including applying a second filter to filter out, from the first set of variants and the second set of variants, fixed substitutions shared between the reference genome of the non-target species and the reference genome of the pseudo-target species.
10 . The computer-implemented method of claim 1 , wherein false positive variants in the subset of false positive variants occur because a particular region in the sequenced reads of the sample of the target species map to a first region in the reference genome of the non-target species and a second region in the reference genome of the pseudo-target species, wherein the first region and the second region are different.
11 . The computer-implemented method of claim 10 , wherein the false positive variants occur because the particular region in the sequenced reads of the sample of the target species maps multiple regions in the reference genome of the non-target species.
12 . A system including one or more processors coupled to memory, the memory loaded with computer instructions to determine feasibility of using a reference genome of a non-target species for variant calling a sample of a target species, the instructions, when executed on the processors, implement actions comprising:
mapping sequenced reads of a sample of a target species to a reference genome of a non-target species to detect a first set of variants in the sequenced reads of the sample of the target species;
mapping the sequenced reads of the sample of the target species to a reference genome of a pseudo-target species to detect a second set of variants in the sequenced reads of the sample of the target species;
comparing the first set of variants and the second set of variants, and identifying a subset of true positive variants that are common between the first set of variants and the second set of variants;
comparing the first set of variants and the second set of variants, and identifying a subset of false positive variants that are present in the second set of variants but absent from the first set of variants; and
based on a count of the subset of false positive variants determining the feasibility of using the reference genome of the non-target species for variant calling the target species.
13 . The system of claim 12 , wherein the pseudo-target species is the target species.
14 . The system of claim 12 , wherein the pseudo-target species is different from the target species.
15 . The system of claim 12 , wherein the pseudo-target species is homologous with the target species.
16 . The system of claim 12 , wherein the non-target species is a human.
17 . The system of claim 12 , wherein the target species is a non-human primate.
18 . The system of claim 12 , further including detecting the second set of variants by mapping the sequenced reads of the sample of the target species to the reference genome of the target species, and then lifting-over the mapped sequenced reads of the sample of the target species to the reference genome of the non-target species.
19 . The system of claim 12 , further including applying a first filter to filter out low-quality variants from the first set of variants and the second set of variants.
20 . The system of claim 12 , further including applying a second filter to filter out, from the first set of variants and the second set of variants, fixed substitutions shared between the reference genome of the non-target species and the reference genome of the pseudo-target species.
21 . The system of claim 12 , wherein false positive variants in the subset of false positive variants occur because a particular region in the sequenced reads of the sample of the target species map to a first region in the reference genome of the non-target species and a second region in the reference genome of the pseudo-target species, wherein the first region and the second region are different.
22 . The system of claim 12 , wherein the false positive variants occur because the particular region in the sequenced reads of the sample of the target species maps multiple regions in the reference genome of the non-target species.
23 . A non-transitory computer readable storage medium impressed with computer program instructions to determine feasibility of using a reference genome of a non-target species for variant calling a sample of a target species, the instructions, when executed on a processor, implement a method comprising:
mapping sequenced reads of a sample of a target species to a reference genome of a non-target species to detect a first set of variants in the sequenced reads of the sample of the target species;
mapping the sequenced reads of the sample of the target species to a reference genome of a pseudo-target species to detect a second set of variants in the sequenced reads of the sample of the target species;
comparing the first set of variants and the second set of variants, and identifying a subset of true positive variants that are common between the first set of variants and the second set of variants;
comparing the first set of variants and the second set of variants, and identifying a subset of false positive variants that are present in the second set of variants but absent from the first set of variants; and
based on a count of the subset of false positive variants determining the feasibility of using the reference genome of the non-target species for variant calling the target species.
24 . The non-transitory computer readable storage medium of claim 23 , wherein the pseudo-target species is the target species.
25 . The non-transitory computer readable storage medium of claim 23 , wherein the pseudo-target species is different from the target species.
26 . The non-transitory computer readable storage medium of claim 3 , wherein the pseudo-target species is homologous with the target species.
27 . The non-transitory computer readable storage medium of claim 23 , wherein the non-target species is a human.
28 . The non-transitory computer readable storage medium of claim 23 , wherein the target species is a non-human primate.
29 . The non-transitory computer readable storage medium of claim 23 , further including detecting the second set of variants by mapping the sequenced reads of the sample of the target species to the reference genome of the target species, and then lifting-over the mapped sequenced reads of the sample of the target species to the reference genome of the non-target species.
30 . The non-transitory computer readable storage medium of claim 23 , further including applying a first filter to filter out low-quality variants from the first set of variants and the second set of variants.
31 . The non-transitory computer readable storage medium of claim 23 , further including applying a second filter to filter out, from the first set of variants and the second set of variants, fixed substitutions shared between the reference genome of the non-target species and the reference genome of the pseudo-target species.
32 . The non-transitory computer readable storage medium of claim 23 , wherein false positive variants in the subset of false positive variants occur because a particular region in the sequenced reads of the sample of the target species map to a first region in the reference genome of the non-target species and a second region in the reference genome of the pseudo-target species, wherein the first region and the second region are different.
33 . The non-transitory computer readable storage medium of claim 23 , wherein the false positive variants occur because the particular region in the sequenced reads of the sample of the target species maps multiple regions in the reference genome of the non-target species.