IP Library › Granted Patent US 11,010,365
Granted Patent B2
US 11,010,365 · App. 15/939,521 · Granted May 18, 2021

Missing value imputation using adaptive ordering and clustering analysis

Inventors: Sunhwan Lee (San Mateo, CA); Lingtao Cao (Hayward, CA); Sarah E. Knoop (San Jose, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/2365G06F16/285G06F16/9535G06T11/206H04L67/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,010,365
App. No.
15/939,521
Granted
May 18, 2021
Kind
B2
Abstract

As received, a data value of an expected input set of received data values is missing from user input. A subset of known data with data values similar to a subset of the received data values is determined. A data sample average for the missing data value is determined from data values within the subset of the known data. An initial estimate of the missing data value is initialized using the data sample average. Boundary data clusters near the initial estimate of the missing data value are identified within the subset of the known data. A data harvesting region encapsulated according to the boundary clusters is defined. Data support clusters within at least one subset of the known data inside the data harvesting region are selected. The initial estimate of the missing data value is updated based upon data of the boundary clusters and the data support clusters.

Claims (50)

1. A computer-implemented method, comprising:

by a data collection interface processor that adaptively imputes missing data values based on data clustering responsive to user input via an operatively-coupled user input device:

receiving, in the user input, data values of an expected input set of data values, where at least one data value of the expected input set of data values is missing from the user input; and

imputing each of the at least one missing data value by, for each missing data value:

determining at least one subset of known data with data values similar to at least a subset of the received data values;

determining, from data values associated with the missing data value within the at least one subset of the known data, a data sample average for the missing data value;

initializing, using the determined data sample average, an initial estimate of the missing data value;

identifying, within the at least one subset of the known data, a plurality of boundary data clusters near the initial estimate of the missing data value;

defining a rectangular data harvesting region encapsulated according to the plurality of boundary data clusters using centroids of coordinates of each of the plurality of boundary data clusters;

selecting multiple data support clusters within the at least one subset of the known data inside the defined data harvesting region; and

updating the initial estimate of the missing data value based upon data of the plurality of boundary data clusters and the selected multiple data support clusters.

2. The computer-implemented method of claim 1 , further comprising the data collection interface processor updating a confidence interval of the updated estimate of the missing data value.

3. The computer-implemented method of claim 1 , further comprising the data collection interface processor determining an order of processing of a plurality of missing data values using one of a random selection and an uncertainty-based selection.

4. The computer-implemented method of claim 1 , where the data collection interface processor identifying the plurality of boundary data clusters near the initial estimate of the missing data value is based upon a confidence interval of the initial estimate of the missing data value.

5. The computer-implemented method of claim 1 , where the data collection interface processor updating the initial estimate of the missing data value based upon the data of the plurality of boundary data clusters and the selected multiple data support clusters is performed using a programmatic calculation technique selected from a set consisting of a majority vote, an average, a comparison with a population statistic of a specified data type of the missing data value, and a user's choice.

6. The computer-implemented method of claim 1 , where the at least one missing data value comprises a plurality of missing data values, and where the data collection interface processor imputing each of the at least one missing data value further comprises the data collection interface processor sequentially ordering imputation of the plurality of missing data values using a next largest uncertainty missing data value selection process.

7. A system, comprising:

an operatively-coupled user input device; and

a data collection interface processor that adaptively imputes missing data values based on data clustering responsive to user input via the user input device, the processor being programmed to:

receive, in the user input, data values of an expected input set of data values, where at least one data value of the expected input set of data values is missing from the user input; and

impute each of the at least one missing data value by being programmed to, for each missing data value:

determine at least one subset of known data with data values similar to at least a subset of the received data values;

determine, from data values associated with the missing data value within the at least one subset of the known data, a data sample average for the missing data value;

initialize, using the determined data sample average, an initial estimate of the missing data value;

identify, within the at least one subset of the known data, a plurality of boundary data clusters near the initial estimate of the missing data value;

define a rectangular data harvesting region encapsulated according to the plurality of boundary data clusters using centroids of coordinates of each of the plurality of boundary data clusters;

select multiple data support clusters within the at least one subset of the known data inside the defined data harvesting region; and

update the initial estimate of the missing data value based upon data of the plurality of boundary data clusters and the selected multiple data support clusters.

8. The system of claim 7 , where the processor is further programmed to one of:

update a confidence interval of the updated estimate of the missing data value; or

determine an order of processing of a plurality of missing data values using one of a random selection and an uncertainty-based selection.

9. The system of claim 7 , where the processor being programmed to identify the plurality of boundary data clusters near the initial estimate of the missing data value is based upon a confidence interval of the initial estimate of the missing data value.

10. The system of claim 7 , where the processor being programmed to update the initial estimate of the missing data value based upon the data of the plurality of boundary data clusters and the selected multiple data support clusters is performed using a programmatic calculation technique selected from a set consisting of a majority vote, an average, a comparison with a population statistic of a specified data type of the missing data value, and a user's choice.

11. The system of claim 7 , where the at least one missing data value comprises a plurality of missing data values, and where in being programmed to impute each of the at least one missing data value, the processor is further programmed to sequentially order imputation of the plurality of missing data values using a next largest uncertainty missing data value selection process.

12. A computer program product, comprising:

a computer readable storage medium having computer readable program code embodied therewith, where the computer readable storage medium is not a transitory signal per se and where the computer readable program code when executed on a computer adaptively imputes missing data values based on data clustering responsive to user input via an operatively-coupled user input device by causing the computer to:

receive, in the user input, data values of an expected input set of data values, where at least one data value of the expected input set of data values is missing from the user input; and

impute each of the at least one missing data value by causing the computer to, for each missing data value:

determine at least one subset of known data with data values similar to at least a subset of the received data values;

determine, from data values associated with the missing data value within the at least one subset of the known data, a data sample average for the missing data value;

initialize, using the determined data sample average, an initial estimate of the missing data value;

identify, within the at least one subset of the known data, a plurality of boundary data clusters the initial estimate of the missing data value;

define a rectangular data harvesting region encapsulated according to the plurality of boundary data clusters using centroids of coordinates of each of the plurality of boundary data clusters;

select multiple data support clusters within the at least one subset of the known data inside the defined data harvesting region; and

update the initial estimate of the missing data value based upon data of the plurality of boundary data clusters and the selected multiple data support clusters.

13. The computer program product of claim 12 , where the computer readable program code when executed on the computer further causes the computer to update a confidence interval of the updated estimate of the missing data value.

14. The computer program product of claim 12 , where the computer readable program code when executed on the computer further causes the computer to determine an order of processing of a plurality of missing data values using one of a random selection and an uncertainty-based selection.

15. The computer program product of claim 12 , where causing the computer to identify the plurality of boundary data clusters near the initial estimate of the missing data value is based upon a confidence interval of the initial estimate of the missing data value.

16. The computer program product of claim 12 , where causing the computer to update the initial estimate of the missing data value based upon the data of the plurality of boundary data clusters and the selected multiple data support clusters is performed using a programmatic calculation technique selected from a set consisting of a majority vote, an average, a comparison with a population statistic of a specified data type of the missing data value, and a user's choice.

17. The computer program product of claim 12 , where the at least one missing data value comprises a plurality of missing data values, and where in causing the computer to impute each of the at least one missing data value, the computer readable program code when executed on the computer further causes the computer to sequentially order imputation of the plurality of missing data values using a next largest uncertainty missing data value selection process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2018
From: LEE, SUNHWAN; CAO, LINGTAO; KNOOP, SARAH E.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 045383/0150 →
Continuity (1)
Related Publication 20190303471A1 · Oct 3, 2019
Cited By (1)
US 12,248,446