IP Library › Granted Patent US 10,528,534
Granted Patent B2
US 10,528,534 · App. 15/824,012 · Granted Jan 7, 2020

Method and system for deduplicating data

Inventors: Namit Kabra (Hyderabad, IN); Yannick Saillet (Stuttgart, DE)
Assignee: International Business Machines Corporation
G06F16/215G06F16/2365
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,528,534
App. No.
15/824,012
Granted
Jan 7, 2020
Kind
B2
Abstract

A mechanism is provided for deduplicating a set of records of data. The mechanism identifies a subset of records each having one or more invalid attribute values. For each invalid attribute value of a given attribute the mechanism determines one or more associated valid candidates of attribute values of the given attribute using the set of records. For each record of the subset of records the mechanism replaces the one or more invalid attribute values by one or more combinations of the determined valid candidates of attribute values, resulting in a modified set of records. The mechanism selects a subset of records of the modified set of records that satisfy a consistency condition on the attribute values of each record. The mechanism deduplicates the selected subset of records of the modified set of records responsive to determining the subset of records comprises more than one record.

Claims (29)

1. A computer implemented method for deduplicating a set of records of data, each record of the set of records having a set of attributes, the method comprising:

receiving the set of records by receiving a data table from storage, wherein the data table comprises one or more columns each represent a respective attribute and one or more rows each representing an attribute value;

identifying a subset of records of the set of records each having one or more invalid attribute values by retrieving heuristic rules to determine a likelihood that an attribute value is an invalid value and applying the heuristic rules to each of the attribute values;

determining for each invalid attribute value of a given attribute of the subset of the set of records one or more associated valid candidates of attribute values of the given attribute using the set of records, wherein determining for a given invalid attribute value one or more associated valid candidates of attribute values comprises selecting records of the set of records having a same super key as the record comprising the given invalid attribute value and determining the one or more associated valid candidates of the attribute values using the selected records;

for each record of the subset of records of the set of records:

replacing the one or more invalid attribute values by one or more combinations of the determined valid candidates of attribute values resulting in a modified set of records;

selecting a subset of records of the modified set of records that satisfy a consistency condition on the attribute values of each record, wherein the consistency condition comprises a rule for determining whether the attribute values of each record are consistent;

deduplicating, in the data table in the storage, the subset of records of the modified set of records responsive to determining the subset of records of the modified set of records comprises more than one record.

2. The method of claim 1 , further comprising disregarding the inconsistent records by removing the inconsistent records or skipping the inconsistent records fir performing the deduplicating.

3. The method of claim 1 , further comprising: before identifying the subset of records of the set of records, comparing for each record of the set of records the attribute values of a super key of the record with predefined reference values of the super key and correcting the attribute values responsive to an unsuccessful comparison.

4. The method of claim 1 , the heuristic rules being generated for performing a data quality analysis of the set of records.

5. The method of claim 1 , further comprising:

providing a reference dataset;

determining association relationships among attribute values of the reference dataset, wherein the consistency condition comprises a condition for fulfilling one of the association relationships.

6. The method of claim 5 , the reference dataset comprising the remaining non-identified subset of records of the set of records.

7. The method of claim 1 , the consistency condition being a selected rule of a set of heuristic rules on multiple attributes of the set of attributes.

8. The method of claim 1 , the one or more associated valid candidates attribute values comprising all valid attribute values of the given attribute in the set of records.

9. The method of claim 1 , the one or more associated valid candidates of attribute values comprising a selected portion of valid attribute values of the given attribute in the set of records.

10. The method of claim 1 , the set of records comprising duplicate records only.

11. The method of claim 1 , the invalid attribute value comprising one of:

a blank value;

a null value;

an out of range value;

a wrong type value;

a value violating a predefined domain condition;

a value violating user defined data quality rifles;

an infrequent value;

an outlier value; or

a value different from a predefined value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2017
From: KABRA, NAMIT; SAILLET, YANNICK
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044234/0814 →
Continuity (2)
Continuation 15276245 · Sep 26, 2016
Related Publication 20180089235A1 · Mar 29, 2018