SYSTEM, METHOD AND COMPUTER PROGRAM FOR NON-BINARY SEQUENCE COMPARISON
A system and method for performing non-binary comparison of biological sequences includes a new measure ω 0 , which is a non-binary counting measure that is used in a stand alone module called VaSSA-1. This measure obtains substantially more information about sequences and comparisons between them than is gathered by conventional bioinformatics techniques.
1 . A method for sequence analysis, comprising:
reading a sequence file;
selecting a target sequence and a base sequence from said file;
performing a non-binary comparison between each base pair of said target and said base sequences, wherein said non-binary comparison generates a comparison value (ω o ) for each base pair;
determining a similarity between said target and said base sequences based on said comparison values; and
generating at least one of a two-dimensional spectral array plot or a two-dimensional single strand plot, wherein generating said spectral array plot comprises:
calculating ω N ;
performing a radial comparison;
extracting alignment coefficients; and
plotting said alignment coefficients.
2 . The method of claim 1 , further comprising: reversing one of said base or said target; and reversing a calculation.
3 . The method of claim 1 , wherein the similarity is determined by
∑
i
=
0
N
s
i
/
t
i
16
*
N
.
4 . The method of claim 1 , wherein said comparison value (ω o ) is generated using a computer system comprising an analysis module having an align sequences module adapted to align the target sequence to the base sequence and a ω o module adapted to produce a comparison value (ω o ) based on the sequence alignment.
5 . The method of claim 1 , wherein said two-dimensional spectral array plot or two-dimensional single strand plot is generated using a computer system comprising a plots module, wherein said plots module comprises: a spectral array module, adapted to plot aligning coefficients for a base sequence and a target sequence, and a single strand module adapted to plot a single strand for said base sequence and said target sequence; a slopes module adapted to calculate a slope for each nucleotide position in said base sequence and to display a plot of said slopes; and a ω N module adapted to calculate ω N for said base sequence and to display a plot of said ω N .
6 . The method of claim 3 , wherein s i /t i is a non-binary function representing the omega similarity score at base position i of the target sequence and the base sequence, and N is the number of nucleotides in the shorter of the two sequences being compared.