IP Library Granted Patent US 10,387,419
Granted Patent B2
US 10,387,419 · App. 14/045,656 · Granted Aug 20, 2019

Method and system for managing databases having records with missing values

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,387,419
App. No.
14/045,656
Granted
Aug 20, 2019
Kind
B2
Abstract

The method includes selecting a target record from a dataset, the target record including a missing value, partitioning records of the dataset into at least two groups including co-related data, the partitioned records including records having a value for a same field as the missing value in the target record, predicting the missing value based on a relationship between fields in each of the at least two groups associated with the partitioned records, and setting the missing value of the target record to the predicted value.

Claims (62)

1. A method, comprising:

in a dataset including a plurality of records, at least one of the plurality of records including one or more field having a missing value, wherein completing the dataset includes:

selecting a target record from the dataset, the target record being one of the plurality of records including a missing value;

partitioning a portion of the plurality of records of the dataset into at least two groups of columns including co-related data, the portion of the plurality of records being selected from the plurality of records and including records having a value for a same field as the missing value in the target record;

predicting the missing value based on a relationship between fields in each of the at least two groups of columns associated with the partitioned records; and

setting the field including the missing value of the target record to the predicted value;

wherein predicting the missing value includes:

generating a first linear function based on a first type of co-related data;

generating a second linear function based on a second type of co-related data;

generate a bi-local linear local model based on the first linear function and the second linear function; and

predicting the missing value using the bi-local linear local model.

2. The method of claim 1 , wherein selecting the target record from the dataset includes selecting the target record as a record with a fewest number of missing values.

3. The method of claim 1 , further comprising:

selecting a target field from the target record, the target field including the missing value.

4. The method of claim 1 , wherein partitioning the records of the dataset into at least two groups includes:

filtering the dataset to exclude from the dataset those records missing values for at least one target field, the at least one target field including the missing value.

5. The method of claim 1 , wherein partitioning the records of the dataset into at least two groups includes:

determining a mean or average value for a target field, and

inserting the mean or average value as a temporary value for each record missing a corresponding value for the target field.

6. The method of claim 1 , wherein partitioning the records of the dataset into at least two groups includes:

selecting k-nearest neighbor (KNN) records of the target record.

7. The method of claim 1 , wherein

the first type of co-related data is based on quantitative data; and

the second type of co-related data is based on qualitative data.

8. The method of claim 1 , wherein the dataset is selected from a personal health record (PHR) database.

9. A non-transitory computer-readable storage medium having stored thereon computer executable program code which, when executed on a computer system, causes the computer system to perform steps comprising:

in a dataset including a plurality of records, at least one of the plurality of records including one or more field having a missing value, wherein completing the dataset includes:

select a target record from the dataset, the target record being one of the plurality of records including a missing value;

partition a portion of the plurality of records of the dataset into at least two groups of columns including co-related data, the portion of the plurality of records being selected from the plurality of records and including records having a value for a same field as the missing value in the target record;

predict the missing value based on a relationship between fields in each of the at least two groups of columns associated with the partitioned records; and

set the field including the missing value of the target record to the predicted value;

wherein predicting the missing value includes:

generating a first linear function based on a first type of co-related data;

generating a second linear function based on a second type of co-related data;

generate a bi-local linear local model based on the first linear function and the second linear function; and

predicting the missing value using the bi-local linear local model.

10. The non-transitory computer-readable storage medium of claim 9 , wherein selecting the target record from the dataset includes selecting the target record as a record with a fewest number of missing values.

11. The non-transitory computer-readable storage medium of claim 9 , wherein the step further comprising:

select a target field from the target record, the target field including the missing value.

12. The non-transitory computer-readable storage medium of claim 9 , wherein partitioning the records of the dataset into at least two groups includes:

filtering the dataset to exclude from the dataset those records missing values for at least one target field, the at least one target field including the missing value.

13. The non-transitory computer-readable storage medium of claim 9 , wherein partitioning the records of the dataset into at least two groups includes:

determining a mean or average value for a target field, and

inserting the mean or average value as a temporary value for each record missing a corresponding value for the target field.

14. The non-transitory computer-readable storage medium of claim 9 , wherein partitioning the records of the dataset into at least two groups includes:

selecting k-nearest neighbor (KNN) records of the target record.

15. The non-transitory computer-readable storage medium of claim 9 , wherein

the first type of co-related data is based on quantitative data; and

the second type of co-related data is based on qualitative data.

16. The non-transitory computer-readable storage medium of claim 9 , wherein the dataset is selected from a personal health record (PHR) database.

17. An apparatus including a processor and a non-transitory computer readable medium, the apparatus comprising:

a value prediction module configured to complete a dataset, the dataset including a plurality of records, at least one of the plurality of records including one or more field having a missing value, wherein the completing of the dataset includes:

select a target record from the dataset, the target record being one of the plurality of records including a missing value; and

set the field including the missing value of the target record to a predicted value; and

a model generation module configured to:

partition a portion of the plurality of records of the dataset into at least two groups of columns including co-related data, the portion of the plurality of records being selected from the plurality of records and including records having a value for a same field as the missing value in the target record; and

predict the missing value based on a relationship between fields in each of the at least two groups of columns associated with the partitioned records;

wherein predicting the missing value includes:

generating a first linear function based on a first type of co-related data;

generating a second linear function based on a second type of co-related data;

generate a bi-local linear local model based on the first linear function and the second linear function; and

predicting the missing value using the bi-local linear local model.

Assignments (2)
CHANGE OF NAME Recorded Aug 26, 2014
From: SAP AG
To: SAP SE
Reel/Frame 033625/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2013
From: LI, WEN-SYAN; CHENG, YU
To: SAP AG
Reel/Frame 031343/0379 →