Systems and methods using DNA sequence strings as a common data format for forensic DNA typing applications
The present disclosure relates, generally, to nucleotide sequence data and, more particularly, to computer files and methods supporting forensic DNA analysis. In one illustrative embodiment, a method may comprise identifying a locus corresponding to each item of short tandem repeat (STR) profiling data stored in an existing computer file, wherein the STR profiling data stored in the existing computer file is repeat-based and/or length-based; identifying start and stop coordinates of an STR region of the corresponding locus for each item of STR profiling data stored in the existing computer file; creating an ambiguous text string corresponding to each item of STR profiling data stored in the existing computer file, wherein each ambiguous text string consists of a sequence of ambiguous characters extending from the start coordinate to the stop coordinate identified for the corresponding item of STR profiling data; and storing each ambiguous text string in a sequence-based computer file.
1 . A method comprising:
operating a massively parallel sequencing (MPS) instrument to read nucleotide sequences present in a DNA sample to generate a first computer file comprising a plurality of target text strings representing the nucleotide sequences, wherein each target text string of the plurality of target text strings is associated with a locus of the DNA sample;
comparing the plurality of target text strings to a plurality of ambiguous text strings stored in a second computer file to detect whether each target text string aligns to a corresponding ambiguous text string of the plurality of ambiguous text strings that is associated with the same locus, wherein each ambiguous text string of the plurality of ambiguous text strings consists of a sequence of ambiguous characters used to represent ambiguous base calls by the MPS instrument; and
identifying a human individual associated with the second computer file as a likely contributor to the DNA sample in response to detecting that a threshold number of the plurality of target text strings align to the plurality of ambiguous text strings of the second computer file.
2 . The method of claim 1 , wherein the sequence of ambiguous characters of each ambiguous text string extends from a start coordinate to a stop coordinate of a short tandem repeat (STR) region represented by that ambiguous text string.
3 . The method of claim 2 , wherein the second computer file further comprises (i) locus data associated with each ambiguous text string, wherein the locus data identifies the locus containing the STR region represented by the associated ambiguous text string, and (ii) coordinate data associated with each ambiguous text string, wherein the coordinate data identifies the start and stop coordinates of the STR region represented by the associated ambiguous text string.
4 . The method of claim 1 , wherein the second computer file was created from a legacy computer file that stored repeat-based and/or length-based short tandem repeat (STR) profiling data but no sequence data.
5 . The method of claim 4 , wherein the legacy computer file was obtained from the Combined DNA Index System (CODIS) database.
6 . The method of claim 4 , wherein the legacy computer file was obtained from the United Kingdom National DNA Database (NDNAD).
7 . The method of claim 1 , further comprising:
comparing the plurality of target text strings to another plurality of ambiguous text strings stored in a third computer file to detect whether each target text string aligns to a corresponding ambiguous text string of the another plurality of ambiguous text strings that is associated with the same locus, wherein each ambiguous text string of the another plurality of ambiguous text strings consists of a sequence of ambiguous characters used to represent ambiguous base calls by the MPS instrument; and
determining whether the plurality of target text strings of the first computer file better aligns to the plurality of ambiguous text strings of the second computer file or to the another plurality of ambiguous text strings of the third computer file.
8 . The method of claim 7 , wherein the second computer file and the third computer file are each stored as a record in a database.
9 . The method of claim 7 , further comprising identifying a different human individual associated with the third computer file as a more likely contributor to the DNA sample than the human individual associated with the second computer file in response to determining that the plurality of target text strings of the first computer file better aligns to the another plurality of ambiguous text strings of the third computer file than to the plurality of ambiguous text strings of the second computer file.
10 . The method of claim 7 , wherein the sequence of ambiguous characters of each ambiguous text string extends from a start coordinate to a stop coordinate of a short tandem repeat (STR) region represented by that ambiguous text string.
11 . The method of claim 10 , wherein each of the second and third computer files further comprises (i) locus data associated with each ambiguous text string, wherein the locus data identifies the locus containing the STR region represented by the associated ambiguous text string, and (ii) coordinate data associated with each ambiguous text string, wherein the coordinate data identifies the start and stop coordinates of the STR region represented by the associated ambiguous text string.
12 . The method of claim 7 , wherein the second computer file was created from a legacy computer file that stored repeat-based and/or length-based short tandem repeat (STR) profiling data but no sequence data, and wherein the third computer file was created from a different legacy computer file that stored repeat-based and/or length-based STR profiling data but no sequence data.
13 . The method of claim 12 , wherein the legacy computer file and the different legacy computer file were both obtained from the Combined DNA Index System (CODIS) database.
14 . The method of claim 12 , wherein the legacy computer file and the different legacy computer file were both obtained from the United Kingdom National DNA Database (NDNAD).
15 . A non-transitory computer readable medium storing a plurality of instructions that are configured to cause a processor that executes the plurality of instructions to:
operate a massively parallel sequencing (MPS) instrument to read nucleotide sequences present in a DNA sample to generate a first computer file comprising a plurality of target text strings representing the nucleotide sequences, wherein each target text string of the plurality of target text strings is associated with a locus of the DNA sample;
compare the plurality of target text strings to a plurality of ambiguous text strings stored in a second computer file to detect whether each target text string aligns to a corresponding ambiguous text string of the plurality of ambiguous text strings that is associated with the same locus, wherein each ambiguous text string of the plurality of ambiguous text strings consists of a sequence of ambiguous characters used to represent ambiguous base calls by the MPS instrument; and
identify a human individual associated with the second computer file as a likely contributor to the DNA sample in response to detecting that a threshold number of the plurality of target text strings align to the plurality of ambiguous text strings of the second computer file.
16 . The non-transitory computer readable medium of claim 15 , wherein the plurality of instructions are further configured to cause the processor to:
compare the plurality of target text strings to another plurality of ambiguous text strings stored in a third computer file to detect whether each target text string aligns to a corresponding ambiguous text string of the another plurality of ambiguous text strings that is associated with the same locus, wherein each ambiguous text string of the another plurality of ambiguous text strings consists of a sequence of ambiguous characters used to represent ambiguous base calls by the MPS instrument; and
determine whether the plurality of target text strings of the first computer file better aligns to the plurality of ambiguous text strings of the second computer file or to the another plurality of ambiguous text strings of the third computer file.