IP Library › Granted Patent US 12,632,511
Granted Patent B1
US 12,632,511 · App. 19/015,043 · Granted May 19, 2026

Missing pattern generation in a data simulation system

Inventors: Xue Ying Zhang (Xi'an, CN); Jing James Xu (Xi'an, CN); Si Er Han (Xi'an, CN); Xiao Ming Ma (Xi'an, CN); Wen Pei Yu (Xi'an, CN); Jing Xu (Xi'an, CN)
Assignee: International Business Machines Corporation
G06F18/15G06F16/215G06F16/2365G06F16/2462G06F16/258G06F16/283G06F18/2113G06N3/088G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,511
App. No.
19/015,043
Granted
May 19, 2026
Kind
B1
Abstract

Systems, methods, and computer program products for generating data according to patterns of missing data. A method for automatically identifying patterns of missing data and generating data accordingly may comprise reading a plurality of values, each associated a subject and a variable; determining a count of missing values for each variable; selecting one or more variables of the plurality of variables based on their respective counts of missing values; identifying a pattern type characterizing a pattern of missing values for each of the one or more selected variables; and generating simulated data for each of the one or more selected variables based on the identified pattern types.

Claims (92)

1 . A computer-implemented method for simulating data, the method comprising:

reading a plurality of values, the plurality of values being associated with a plurality of subjects and a plurality of variables, each value being associated with one subject of the plurality of subjects and one variable of the plurality of variables;

determining a plurality of counts of missing values, the plurality of counts being associated with the plurality of variables, each count being based on one or more subjects not associated with any value for a corresponding variable of the plurality of variables;

selecting one or more variables of the plurality of variables based on the plurality of counts of missing values associated with the plurality of variables, thereby producing one or more selected variables, further comprising:

computing, for each variable of the plurality of variables, an associated missing value percentage based on a corresponding count of missing values of the plurality of counts of missing values and a quantity of subjects of the plurality of subjects; and

selecting the one or more variables of the plurality of variables having associated missing value percentages that satisfy a sparsity threshold;

identifying, for each selected variable of the one or more selected variables, a corresponding pattern type characterizing a pattern of missing values for the selected variable, wherein said identifying comprises:

determining whether a presence of a value for the selected variable is statistically independent of variables of the plurality of variables other than the selected variable,

identifying a set of associated variables, wherein the set of associated variables is a subset of the plurality of variables that are not statistically independent of the presence of a value for the selected variable,

identifying the corresponding pattern type for the selected variable based on the set of associated variables,

determining that the set of associated variables comprises at least one variable,

determining, for each associated variable of the set of associated variables, an associated level of correlation between the associated variable and the presence of a value for the selected variable,

identifying a subset of the set of associated variables, wherein each associated variable of the subset of the set of associated variables has an associated level of correlation of at least a first correlation threshold, and

identifying the corresponding pattern type as structurally missing data responsive to identifying the subset of the set of associated variables; and

generating simulated data for each of the one or more selected variables based on the corresponding pattern type.

2 . The computer-implemented method of claim 1 , wherein the selecting of the one or more variables of the plurality of variables based on the plurality of counts of missing values associated with the plurality of variables further comprises:

ranking one or more selected variables of the plurality of variables based on a corresponding associated missing value percentage.

3 . The computer-implemented method of claim 1 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:

identifying a strongly correlated variable of the subset of associated variables, wherein the strongly correlated variable has an associated level of correlation of at least a second correlation threshold; and

determining the corresponding pattern type as structurally missing data responsive to identifying the strongly correlated variable.

4 . The computer-implemented method of claim 1 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:

determining an associated variable of a subset of the set of associated variables, wherein the associated variable has an associated level of correlation of less than a second correlation threshold; and

identifying the corresponding pattern type as missing at random responsive to determining that the associated variable has the associated level of correlation of less than the second correlation threshold.

5 . The computer-implemented method of claim 1 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:

determining that the set of associated variables is an empty set; and

performing principal component analysis on at least one of the one or more selected variables to identify a first principal component.

6 . The computer-implemented method of claim 5 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:

determining that the first principal component and the presence of a value for the selected variable are not statistically independent; and

identifying the corresponding pattern type as missing not at random.

7 . The computer-implemented method of claim 5 , wherein said identifying, for each selected variable of the one or more selected variables, the corresponding pattern type characterizing the pattern of missing values for the selected variable further comprises:

determining that the first principal component and the presence of a value for the selected variable are statistically independent; and

identifying the corresponding pattern type as missing completely at random.

8 . The computer-implemented method of claim 1 , wherein said generating simulated data for each of the one or more selected variables based on the corresponding pattern type comprises:

randomly selecting a random number from a standard uniform distribution for a subject of the plurality of subjects;

determining that the random number is at least a target probability; and generating, as part of the simulated data, a value associated with the subject of the plurality of subjects and a selected variable of the one or more selected variables responsive to determining that the random number is at least the target probability.

9 . The computer-implemented method of claim 8 , wherein said generating simulated data for each selected variable of the one or more selected variables based on the corresponding pattern type further comprises:

determining the target probability based on a count of subjects of the plurality of subjects not associated with any value for the selected variable.

10 . The computer-implemented method of claim 8 , wherein said generating simulated data for each selected variable of the one or more selected variables based on the corresponding pattern type further comprises:

identifying an associated variable of the set of associated variables identified for the selected variable;

determining a pattern count corresponding to a number of subjects of the plurality of subjects that are associated with a value of the associated variable equivalent to a particular value and not associated with any value for the selected variable; and determining the target probability based on the pattern count.

11 . The computer-implemented method of claim 10 , wherein said generating simulated data for each selected variable of the one or more selected variables based on the corresponding pattern type further comprises:

determining that the associated variable is a continuous variable; and

organizing values associated with the associated variable into a plurality of bins, wherein values within each bin are treated as equivalent.

12 . The computer-implemented method of claim 8 , wherein said generating simulated data for each selected variable of the one or more selected variables based on the corresponding pattern type further comprises:

performing principal component analysis on variables of the plurality of variables other than the selected variable to identify a first principal component; and

organizing values associated with the first principal component into a plurality of bins, wherein values within each bin are treated as equivalent.

13 . The computer-implemented method of claim 1 , wherein the corresponding pattern type identified for each selected variable of the one or more selected variables is selected from the group comprising structurally missing data.

14 . A computer program product comprising:

one or more computer-readable storage media; and

program instructions stored on the one or more computer-readable storage media to perform operations comprising:

reading a plurality of values, the plurality of values being associated with a plurality of subjects and a plurality of variables, each value being associated with one subject of the plurality of subjects and one variable of the plurality of variables;

determining a plurality of counts of missing values, the plurality of counts being associated with the plurality of variables, each count being based on one or more subjects not associated with any value for a corresponding variable of the plurality of variables;

selecting one or more variables of the plurality of variables based on the plurality of counts of missing values associated with the plurality of variables, thereby producing one or more selected variables, further comprising:

computing, for each variable of the plurality of variables, an associated missing value percentage based on a corresponding count of missing values of the plurality of counts of missing values and a quantity of subjects of the plurality of subjects; and

selecting the one or more variables of the plurality of variables having associated missing value percentages that satisfy a sparsity threshold;

identifying, for each selected variable of the one or more selected variables, a corresponding pattern type characterizing a pattern of missing values for the selected variable, wherein said identifying comprises:

determining whether a presence of a value for the selected variable is statistically independent of variables of the plurality of variables other than the selected variable,

identifying a set of associated variables, wherein the set of associated variables is a subset of the plurality of variables that are not statistically independent of the presence of a value for the selected variable,

identifying the corresponding pattern type for the selected variable based on the set of associated variables,

determining that the set of associated variables comprises at least one variable,

determining, for each associated variable of the set of associated variables, an associated level of correlation between the associated variable and the presence of a value for the selected variable,

identifying a subset of the set of associated variables, wherein each associated variable of the subset of the set of associated variables has an associated level of correlation of at least a first correlation threshold, and

identifying the corresponding pattern type as structurally missing data responsive to identifying the subset of the set of associated variables; and

generating simulated data for each of the one or more selected variables based on the corresponding pattern type.

15 . The computer program product of claim 14 , wherein the operations further comprise:

ranking the one or more selected variables based on the associated missing value percentages.

16 . The computer program product of claim 14 , wherein the operations further comprise:

determining that a set of associated variables identified for a selected variable of the one or more selected variables is an empty set; and

responsive to determining that the set of associated variables is an empty set,

performing principal component analysis on at least one of the one or more selected variables to identify a first principal component.

17 . The computer program product of claim 14 , wherein the operations further comprise:

randomly selecting a random number from a standard uniform distribution for a subject of the plurality of subjects;

determining that the random number is at least a target probability; and

generating, as part of the simulated data, a value associated with the subject of the plurality of subjects and a selected variable of the one or more selected variables responsive to determining that the random number is at least the target probability.

18 . A computer system comprising:

a processor set;

one or more computer-readable storage media; and

program instructions stored on the one or more computer-readable storage media to perform operations comprising:

reading a plurality of values, the plurality of values being associated with a plurality of subjects and a plurality of variables, each value being associated with one subject of the plurality of subjects and one variable of the plurality of variables;

determining a plurality of counts of missing values associated with the plurality of variables, the plurality of counts being associated with the plurality of variables, each count being based on one or more subjects not associated with any value for a corresponding variable of the plurality of variables;

selecting one or more variables of the plurality of variables based on the plurality of counts of missing values associated with the plurality of variables, thereby producing one or more selected variables, further comprising:

computing, for each variable of the plurality of variables, an associated missing value percentage based on a corresponding count of missing values of the plurality of counts of missing values and a quantity of subjects of the plurality of subjects; and

selecting the one or more variables of the plurality of variables having associated missing value percentages that satisfy a sparsity threshold;

identifying, for each selected variable of the one or more selected variables, a corresponding pattern type characterizing a pattern of missing values for the selected variable, wherein said identifying comprises:

determining whether a presence of a value for the selected variable is statistically independent of variables of the plurality of variables other than the selected variable,

identifying a set of associated variables, wherein the set of associated variables is a subset of the plurality of variables that are not statistically independent of the presence of the value for the selected variable,

identifying the corresponding pattern type for the selected variable based on the set of associated variables,

determining that the set of associated variables comprises at least one variable;

determining, for each associated variable of the set of associated variables, an associated level of correlation between that associated variable and the presence of a value for the selected variable,

identifying a subset of the set of associated variables, wherein each associated variable of the subset of the set of associated variables has an associated level of correlation of at least a first correlation threshold, and

identifying the corresponding pattern type as structurally missing data responsive to identifying the subset of the set of associated variables; and

generating simulated data for each of the one or more selected variables based on the corresponding pattern type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2025
From: ZHANG, XUE YING; XU, JING JAMES; HAN, SI ER; MA, XIAO MING; YU, WEN PEI; XU, JING
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 069842/0140 →
References Cited (29)
US 10860651B2 · Convertino · 2020 [cited by examiner]
US 11422995B1 · Marcus · 2022 [cited by examiner]
US 20110307437A1 · Aliferis · 2011 [cited by examiner]
US 20130226842A1 · Chu · 2013 [cited by examiner]
US 20140207493A1 · Sarrafzadeh · 2014 [cited by examiner]
US 20140324752A1 · Statnikov · 2014 [cited by examiner]
US 20160117588A1 · Muraoka · 2016 [cited by examiner]
US 20190258743A1 · Convertino · 2019 [cited by examiner]
US 20210182602A1 · V · 2021 [cited by examiner]
US 20210374164A1 · Ghoula · 2021 [cited by examiner]
US 20220374446A1 · Savir · 2022 [cited by examiner]
US 20230040284A1 · Ali-Tolppa et al. · 2023 [cited by applicant]
US 20240338559A1 · Zhang · 2024 [cited by examiner]
EP 4535186A1 · 2025 [cited by examiner]
WO WO2023003676A1 · 2023 [cited by examiner]
WO WO2024259083A1 · 2024 [cited by examiner]
Elham Kalantari et al., “Evaluating traditional versus ensemble machine learning methods for predicting missing data of daily PM10 concentration”, Atmospheric Pollution Research, vol. 15, Issue 5, May 2024, 102063, pp. … [cited by examiner]
Wouter van Loon et al., “Imputation of missing values in multi-view data”, Information Fusion, vol. 111, Nov. 2024, 102524 , pp. 1-18. [cited by examiner]
Roderick J. Little, “Missing Data Analysis”, Department of Biostatistics, University of Michigan, Ann Arbor, Michigan, USA; Feb. 12, 2024, pp. 149-173. [cited by examiner]
Bechný et al. “Missing Data Patterns: From Theory to an Application in the Steel Industry”, SSDBM '21: Proceedings of the 33rd International Conference on Scientific and Statistical Database Management, Aug. 11, 2021, p… [cited by applicant]
Poudevigne-Durance et al. “MaWGAN: A Generative Adversarial Network to Create Synthetic Data from Datasets with Missing Data”, Electronics, Mar. 8, 2022, 10 pages. [cited by applicant]
Ruddle et al. “Using Set Visualisation to Find and Explain Patterns of Missing Values: A Case Study with NHS Hospital Episode Statistics Data”, BMJ Open, 2022, 9 pages. [cited by applicant]
Wang et al. “Planned Missing Data Design: Through Intended Missing Data Make Research More Effective”, Advances in Psychological Science, Jan. 2014, pp. 1025-1035 (22 pages), vol. 22, Issue No. 6. [cited by applicant]
Wang et al. “Preserving Missing Data Distribution in Synthetic Data”, WWW'23: Proceedings of the ACM Web Conference 2023, Apr. 30, 2023, pp. 2110-2121. [cited by applicant]
Zhang Xijuan. “Tutorial: How to Generate Missing Data for Simulation Studies”, Generating Missing Data, 2023, 50 pages. [cited by applicant]
International Searching Authority, “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or Declaration,” Patent Cooperation Treaty, Feb. 25, 2… [cited by applicant]
P. Vateekul et al, “Tree-Based Approach to Missing Data Imputation,” 2009 IEEE International Conference on Data Mining Workshops, Miami, FL, USA, 2009, pp. 70-75, doi: 10.1109/ICDMW.2009.92. [cited by applicant]
Samad Manar et al., “Missing value estimation using clustering and deep learning within multiple imputation framework”, Knowledge-Based Systems, Aug. 5, 2022, 12 pages, vol. 249, doi: https://doi.org/10.1016/j.knosys.20… [cited by applicant]
Zhou et al., “Review for Handling Missing Data with special missing mechanism”, arXiv, Apr. 7, 2024, 53 pages, doi: https://arxiv.org/abs/2404.04905v1. [cited by applicant]