Creation of Attribute Combination Databases
A method and system are presented in which attributes of query-attribute-positive individuals and query-attribute-negative individuals are compared in order to identify combinations of attributes that have higher frequencies of occurrence for query-attribute-positive individuals than for query-attribute-negative individuals, and to create a compilation of attribute combinations that co-occur with the query attribute.
1 . A computer based method for compiling attribute combinations, comprising:
a) accessing a query attribute;
b) accessing attributes of query-attribute-positive individuals and query-attribute-negative individuals; and
c) storing those combinations of attributes that have higher frequencies of occurrence for the query-attribute-positive individuals than for the query-attribute-negative individuals, to create a compilation of attribute combinations that co-occur with the query attribute.
2 . The computer based method of claim 1 , further comprising:
d) storing statistical results that are based on the frequencies of occurrence and indicate strength of association of each combination with the query attribute.
3 . The computer based method of claim 1 , further comprising:
d) repeating steps (a)-(c) for a succession of query attributes.
4 . The computer based method of claim 1 , further comprising:
d) eliminating attribute combinations from the compilation based on the frequencies of occurrence.
5 . The computer based method of claim 1 , further comprising:
d) ranking the attribute combinations based on the frequencies of occurrence.
6 . The computer based method of claim 2 , further comprising:
e) ranking the attribute combinations based on the statistical results.
7 . The computer based method of claim 2 , further comprising:
e) eliminating attribute combinations from the compilation that do not meet one or more statistical requirements.
8 . The computer based method of claim 1 , further comprising:
d) transmitting one or more attributes, attribute combinations, frequencies of occurrence as output.
9 . The computer based method of claim 1 , further comprising:
d) storing statistical results that are based on the frequencies of occurrence and indicate strength of association of each combination with the query attribute;
e) eliminating attribute combinations from the compilation that do not meet one or more statistical requirements;
f) ranking the attribute combinations based on the statistical results;
g) repeating steps (a)-(f) for a succession of query attributes; and
h) transmitting one or more attributes, attribute combinations, frequencies of occurrence and statistical results as output.
10 . The computer based method of claim 1 , wherein the query attribute can be a query attribute combination consisting of two or more attributes.
11 . The computer based method of claim 1 , wherein the query-attribute-positive individuals and the query-attribute-negative individuals constitute a population or subpopulation that is preselected based on having an association with one or more attributes.
12 . The computer based method of claim 7 , wherein the one or more statistical requirements are selected from the group consisting of: greater than a minimum statistical value, less than a maximum statistical value, and statistical significance.
13 . The computer based method of claim 9 , wherein the one or more statistical requirements are selected from the group consisting of: greater than a minimum statistical value, less than a maximum statistical value, and statistical significance.
14 . The computer based method of claim 1 , wherein the identity of one or more individuals is masked or anonymized.
15 . A computer based system for compiling attribute combinations, comprising:
a) a first data accessing subsystem for accessing a query attribute;
b) a second data accessing subsystem for accessing attributes associated with query-attribute-positive individuals and attributes associated with query-attribute-negative individuals; and
c) a data processing subsystem for identifying attribute combinations that are more frequently associated with query-attribute-positive individuals than with query-attribute-negative individuals.
16 . The computer based system of claim 15 , wherein the data processing subsystem comprises:
i) a data comparison subsystem for identifying attributes associated with the query-attribute-positive individual that are not associated with a portion of the query-attribute-negative individuals; and
ii) a statistical computing subsystem for calculating frequencies of occurrence of attribute combinations and statistical results indicating their strength of association with query attributes.
17 . The system of claim 16 , further comprising:
d) a data storage subsystem for storing one or more attributes, attribute combinations, frequencies of occurrence, statistical results and individual identifiers, to create a compilation of attribute combinations that co-occur with the query attribute.
18 . The computer based system of claim 17 , further comprising:
e) a communications subsystem for:
i) retrieving or receiving at least some attributes from at least one external database; and
ii) transmitting one or more attributes, attribute combinations, frequencies of occurrence and statistical results as output.
19 . The computer based system of claim 18 , wherein the data processing subsystem further comprises:
iii) an attribute elimination subsystem for eliminating attribute combinations from the compilation based on their frequencies of occurrence or the statistical results.
20 . The computer based system of claim 19 , wherein the data processing subsystem further comprises:
iv) a data ranking subsystem for ranking the attribute combinations based on the statistical results.