Compiling Co-associating Bioattributes
A bioinformatics method, software, database and system are presented in which attribute profiles of query-attribute-positive individuals and query-attribute-negative individuals are compared, and combinations of pangenetic and non-pangenetic attributes that occur at a higher frequency in the group of query-attribute-positive individuals are identified and stored to generate a compilation of bioattribute combinations that co-associate with the query attribute (i.e., an attribute of interest).
1 . A bioinformatics method for generating a compilation containing combinations of pangenetic and non-pangenetic attributes that co-associate with a query attribute associated with an individual, comprising:
a) receiving a query attribute;
b) accessing attribute profiles of query-attribute-positive individuals and attribute profiles of query-attribute-negative individuals, wherein the attribute profiles comprise pangenetic attributes and non-pangenetic attributes;
c) determining the frequencies of occurrence of attribute combinations in the attribute profiles of the query-attribute-positive individuals and in the attribute profiles of the query-attribute-negative individuals, wherein each of the attribute combinations contains at least one pangenetic attribute and at least one non-pangenetic attribute; and
d) storing the attribute combinations having higher frequencies of occurrence in the attribute profiles of the query-attribute-positive individuals than in the attribute profiles of the query-attribute-negative individuals to generate a compilation of attribute combinations that co-associate with the query attribute.
2 . The bioinformatics method of claim 1 , further comprising:
e) storing, in the compilation, corresponding statistical results that are computed based on the frequencies of occurrence to indicate the strength of association of each of the attribute combinations in the compilation with the query attribute.
3 . The bioinformatics method of claim 2 , further comprising:
f) ranking the attribute combinations based on the corresponding statistical results.
4 . The bioinformatics method of claim 2 , further comprising:
f) eliminating at least one attribute combination from the compilation based on the corresponding statistical results.
5 . The bioinformatics method of claim 2 , further comprising:
f) eliminating at least one attribute combination from the compilation based on at least one predetermined statistical threshold.
6 . The bioinformatics method of claim 2 , further comprising:
f) transmitting the attribute combinations and the corresponding statistical results as output.
7 . The bioinformatics method of claim 1 , further comprising:
e) repeating steps (a)-(d) for a succession of query attributes.
8 . The bioinformatics method of claim 1 , further comprising:
e) ranking the attribute combinations based on the frequencies of occurrence.
9 . The bioinformatics method of claim 1 , further comprising:
e) eliminating at least one attribute combination from the compilation based on the frequencies of occurrence.
10 . The bioinformatics method of claim 1 , wherein the query-attribute-positive individuals and the query-attribute-negative individuals are derived from a single population of individuals preselected based on association or lack of association with one or more user specified attributes.
11 . The bioinformatics method of claim 1 , wherein the query attribute is a combination of two or more attributes.
12 . The bioinformatics method of claim 1 , wherein the identity of one or more individuals is masked or anonymized.
13 . The bioinformatics method of claim 1 , wherein each of the attribute combinations contains at least one pangenetic attribute, at least one physical attribute, at least one behavioral attribute, and at least one situational attribute.
14 . A program storage device readable by a machine and containing a set of instructions which, when read by the machine, causes execution of a bioinformatics method for generating a compilation containing combinations of pangenetic and non-pangenetic attributes that co-associate with a query attribute associated with an individual, comprising:
a) receiving a query attribute;
b) accessing attribute profiles of query-attribute-positive individuals and attribute profiles of query-attribute-negative individuals, wherein the attribute profiles comprise pangenetic attributes and non-pangenetic attributes;
c) determining the frequencies of occurrence of attribute combinations in the attribute profiles of the query-attribute-positive individuals and in the attribute profiles of the query-attribute-negative individuals, wherein each of the attribute combinations contains at least one pangenetic attribute and at least one non-pangenetic attribute; and
d) storing the attribute combinations having higher frequencies of occurrence in the attribute profiles of the query-attribute-positive individuals than in the attribute profiles of the query-attribute-negative individuals to generate a compilation of attribute combinations that co-associate with the query attribute.
15 . The program storage device of claim 14 , wherein each of the attribute combinations contains at least one pangenetic attribute, at least one physical attribute, at least one behavioral attribute, and at least one situational attribute.
16 . A bioinformatics database system for generating a compilation containing combinations of pangenetic and non-pangenetic attributes that co-associate with a query attribute associated with an individual, comprising:
a) a memory containing:
i) a first data structure containing attribute profiles of query-attribute-positive individuals and attribute profiles of query-attribute-negative individuals, wherein the attribute profiles comprise pangenetic attributes and non-pangenetic attributes:
b) a processor for:
i) receiving a query attribute;
ii) accessing the first data structure;
iii) determining the frequencies of occurrence of attribute combinations in the attribute profiles of the query-attribute-positive individuals and in the attribute profiles of the query-attribute-negative individuals, wherein each of the attribute combinations contains at least one pangenetic attribute and at least one non-pangenetic attribute; and
iv) storing, to generate a 2nd data structure in the memory, the attribute combinations having higher frequencies of occurrence in the attribute profiles of the query-attribute-positive individuals than in the attribute profiles of the query-attribute-negative individuals to generate a compilation of attribute combinations that co-associate with the query attribute.
17 . The bioinformatics database system of claim 16 , wherein the processor is also for storing, in the 2nd data structure in the memory, a set of corresponding statistical results computed based on the frequencies of occurrence to indicate the strength of association of each of the attribute combinations in the compilation with the query attribute.
18 . The bioinformatics database system of claim 16 , wherein each of the attribute combinations contains at least one pangenetic attribute, at least one physical attribute, at least one behavioral attribute, and at least one situational attribute.
19 . A bioinformatics computer-based system for generating a compilation containing combinations of pangenetic and non-pangenetic attributes that co-associate with a query attribute associated with an individual, comprising:
a) a communications subsystem for receiving a query attribute;
b) a data accessing subsystem for accessing attribute profiles of query-attribute-positive individuals and attribute profiles of query-attribute-negative individuals, wherein the attribute profiles comprise pangenetic attributes and non-pangenetic attributes;
c) a data processing subsystem for determining the frequencies of occurrence of attribute combinations in the attribute profiles of the query-attribute-positive individuals and in the attribute profiles of the query-attribute-negative individuals, wherein each of the attribute combinations contains at least one pangenetic attribute and at least one non-pangenetic attribute; and
d) a data storage subsystem for storing the attribute combinations having higher frequencies of occurrence in the attribute profiles of the query-attribute-positive individuals than in the attribute profiles of the query-attribute-negative individuals to generate a compilation of attribute combinations that co-associate with the query attribute.
20 . The bioinformatics computer-based system of claim 19 , wherein each of the attribute combinations contains at least one pangenetic attribute, at least one physical attribute, at least one behavioral attribute, and at least one situational attribute.