Missing pattern generation in a data simulation system
Systems, methods, and computer program products for generating data according to patterns of missing data. A method for automatically identifying patterns of missing data and generating data accordingly may comprise reading a plurality of values, each associated a subject and a variable; determining a count of missing values for each variable; selecting one or more variables of the plurality of variables based on their respective counts of missing values; identifying a pattern type characterizing a pattern of missing values for each of the one or more selected variables; and generating simulated data for each of the one or more selected variables based on the identified pattern types.
1 . A computer-implemented method for simulating data, the method comprising:
reading a plurality of values, the plurality of values being associated with a plurality of subjects and a plurality of variables, each value being associated with one subject of the plurality of subjects and one variable of the plurality of variables;
determining a plurality of counts of missing values, the plurality of counts being associated with the plurality of variables, each count being based on one or more subjects not associated with any value for a corresponding variable of the plurality of variables;
selecting one or more variables of the plurality of variables based on the plurality of counts of missing values associated with the plurality of variables, thereby producing one or more selected variables, further comprising:
computing, for each variable of the plurality of variables, an associated missing value percentage based on a corresponding count of missing values of the plurality of counts of missing values and a quantity of subjects of the plurality of subjects; and
selecting the one or more variables of the plurality of variables having associated missing value percentages that satisfy a sparsity threshold;
identifying, for each selected variable of the one or more selected variables, a corresponding pattern type characterizing a pattern of missing values for the selected variable, wherein said identifying comprises:
determining whether a presence of a value for the selected variable is statistically independent of variables of the plurality of variables other than the selected variable,
identifying a set of associated variables, wherein the set of associated variables is a subset of the plurality of variables that are not statistically independent of the presence of a value for the selected variable,
identifying the corresponding pattern type for the selected variable based on the set of associated variables,
determining that the set of associated variables comprises at least one variable,
determining, for each associated variable of the set of associated variables, an associated level of correlation between the associated variable and the presence of a value for the selected variable,
identifying a subset of the set of associated variables, wherein each associated variable of the subset of the set of associated variables has an associated level of correlation of at least a first correlation threshold, and
identifying the corresponding pattern type as structurally missing data responsive to identifying the subset of the set of associated variables; and
generating simulated data for each of the one or more selected variables based on the corresponding pattern type.
2 . The computer-implemented method of claim 1 , wherein the selecting of the one or more variables of the plurality of variables based on the plurality of counts of missing values associated with the plurality of variables further comprises:
ranking one or more selected variables of the plurality of variables based on a corresponding associated missing value percentage.
3 . The computer-implemented method of claim 1 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:
identifying a strongly correlated variable of the subset of associated variables, wherein the strongly correlated variable has an associated level of correlation of at least a second correlation threshold; and
determining the corresponding pattern type as structurally missing data responsive to identifying the strongly correlated variable.
4 . The computer-implemented method of claim 1 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:
determining an associated variable of a subset of the set of associated variables, wherein the associated variable has an associated level of correlation of less than a second correlation threshold; and
identifying the corresponding pattern type as missing at random responsive to determining that the associated variable has the associated level of correlation of less than the second correlation threshold.
5 . The computer-implemented method of claim 1 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:
determining that the set of associated variables is an empty set; and
performing principal component analysis on at least one of the one or more selected variables to identify a first principal component.
6 . The computer-implemented method of claim 5 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:
determining that the first principal component and the presence of a value for the selected variable are not statistically independent; and
identifying the corresponding pattern type as missing not at random.
7 . The computer-implemented method of claim 5 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:
determining that the first principal component and the presence of a value for the selected variable are statistically independent; and
identifying the corresponding pattern type as missing completely at random.
8 . The computer-implemented method of claim 1 , wherein said generating simulated data for each of the one or more selected variables based on the corresponding pattern type comprises:
randomly selecting a random number from a standard uniform distribution for a subject of the plurality of subjects;
determining that the random number is at least a target probability; and generating, as part of the simulated data, a value associated with the subject of the plurality of subjects and a selected variable of the one or more selected variables responsive to determining that the random number is at least the target probability.
9 . The computer-implemented method of claim 8 , wherein said generating simulated data for each selected variable of the one or more selected variables based on the corresponding pattern type further comprises:
determining the target probability based on a count of subjects of the plurality of subjects not associated with any value for the selected variable.
10 . The computer-implemented method of claim 8 , wherein said generating simulated data for each selected variable of the one or more selected variables based on the corresponding pattern type further comprises:
identifying an associated variable of the set of associated variables identified for the selected variable;
determining a pattern count corresponding to a number of subjects of the plurality of subjects that are associated with a value of the associated variable equivalent to a particular value and not associated with any value for the selected variable; and determining the target probability based on the pattern count.
11 . The computer-implemented method of claim 10 , wherein said generating simulated data for each selected variable of the one or more selected variables based on the corresponding pattern type further comprises:
determining that the associated variable is a continuous variable; and
organizing values associated with the associated variable into a plurality of bins, wherein values within each bin are treated as equivalent.
12 . The computer-implemented method of claim 8 , wherein said generating simulated data for each selected variable of the one or more selected variables based on the corresponding pattern type further comprises:
performing principal component analysis on variables of the plurality of variables other than the selected variable to identify a first principal component; and
organizing values associated with the first principal component into a plurality of bins, wherein values within each bin are treated as equivalent.
13 . The computer-implemented method of claim 1 , wherein the corresponding pattern type identified for each selected variable of the one or more selected variables is selected from the group comprising structurally missing data.
14 . A computer program product comprising:
one or more computer-readable storage media; and
program instructions stored on the one or more computer-readable storage media to perform operations comprising:
reading a plurality of values, the plurality of values being associated with a plurality of subjects and a plurality of variables, each value being associated with one subject of the plurality of subjects and one variable of the plurality of variables;
determining a plurality of counts of missing values, the plurality of counts being associated with the plurality of variables, each count being based on one or more subjects not associated with any value for a corresponding variable of the plurality of variables;
selecting one or more variables of the plurality of variables based on the plurality of counts of missing values associated with the plurality of variables, thereby producing one or more selected variables, further comprising:
computing, for each variable of the plurality of variables, an associated missing value percentage based on a corresponding count of missing values of the plurality of counts of missing values and a quantity of subjects of the plurality of subjects; and
selecting the one or more variables of the plurality of variables having associated missing value percentages that satisfy a sparsity threshold;
identifying, for each selected variable of the one or more selected variables, a corresponding pattern type characterizing a pattern of missing values for the selected variable, wherein said identifying comprises:
determining whether a presence of a value for the selected variable is statistically independent of variables of the plurality of variables other than the selected variable,
identifying a set of associated variables, wherein the set of associated variables is a subset of the plurality of variables that are not statistically independent of the presence of a value for the selected variable,
identifying the corresponding pattern type for the selected variable based on the set of associated variables,
determining that the set of associated variables comprises at least one variable,
determining, for each associated variable of the set of associated variables, an associated level of correlation between the associated variable and the presence of a value for the selected variable,
identifying a subset of the set of associated variables, wherein each associated variable of the subset of the set of associated variables has an associated level of correlation of at least a first correlation threshold, and
identifying the corresponding pattern type as structurally missing data responsive to identifying the subset of the set of associated variables; and
generating simulated data for each of the one or more selected variables based on the corresponding pattern type.
15 . The computer program product of claim 14 , wherein the operations further comprise:
ranking the one or more selected variables based on the associated missing value percentages.
16 . The computer program product of claim 14 , wherein the operations further comprise:
determining that a set of associated variables identified for a selected variable of the one or more selected variables is an empty set; and
responsive to determining that the set of associated variables is an empty set,
performing principal component analysis on at least one of the one or more selected variables to identify a first principal component.
17 . The computer program product of claim 14 , wherein the operations further comprise:
randomly selecting a random number from a standard uniform distribution for a subject of the plurality of subjects;
determining that the random number is at least a target probability; and
generating, as part of the simulated data, a value associated with the subject of the plurality of subjects and a selected variable of the one or more selected variables responsive to determining that the random number is at least the target probability.
18 . A computer system comprising:
a processor set;
one or more computer-readable storage media; and
program instructions stored on the one or more computer-readable storage media to perform operations comprising:
reading a plurality of values, the plurality of values being associated with a plurality of subjects and a plurality of variables, each value being associated with one subject of the plurality of subjects and one variable of the plurality of variables;
determining a plurality of counts of missing values associated with the plurality of variables, the plurality of counts being associated with the plurality of variables, each count being based on one or more subjects not associated with any value for a corresponding variable of the plurality of variables;
selecting one or more variables of the plurality of variables based on the plurality of counts of missing values associated with the plurality of variables, thereby producing one or more selected variables, further comprising:
computing, for each variable of the plurality of variables, an associated missing value percentage based on a corresponding count of missing values of the plurality of counts of missing values and a quantity of subjects of the plurality of subjects; and
selecting the one or more variables of the plurality of variables having associated missing value percentages that satisfy a sparsity threshold;
identifying, for each selected variable of the one or more selected variables, a corresponding pattern type characterizing a pattern of missing values for the selected variable, wherein said identifying comprises:
determining whether a presence of a value for the selected variable is statistically independent of variables of the plurality of variables other than the selected variable,
identifying a set of associated variables, wherein the set of associated variables is a subset of the plurality of variables that are not statistically independent of the presence of the value for the selected variable,
identifying the corresponding pattern type for the selected variable based on the set of associated variables,
determining that the set of associated variables comprises at least one variable;
determining, for each associated variable of the set of associated variables, an associated level of correlation between that associated variable and the presence of a value for the selected variable,
identifying a subset of the set of associated variables, wherein each associated variable of the subset of the set of associated variables has an associated level of correlation of at least a first correlation threshold, and
identifying the corresponding pattern type as structurally missing data responsive to identifying the subset of the set of associated variables; and
generating simulated data for each of the one or more selected variables based on the corresponding pattern type.