Data quality metric-based record manipulation for master data management
Data quality metric-based record manipulation for master data management includes generating for a pair of data records, attribute pair scores quantifying an extent of match between an attribute of a data record of the pair with a corresponding attribute of another data record in the pair. Data quality metric scores for data quality metrics are computed for each attribute of the pair of data records. A dimension match score is generated that aggregates these scores across multiple data quality metrics. The dimension match score and attribute pair score are combined to create a pair feature vector, which is evaluated against a threshold to determine match outcome information for the pair of data records. Manipulation of the data records is executed in accordance with the match outcome information. The manipulated data is output for further processing, reporting, or analytics.
1 . A computer-implemented method, comprising:
obtaining, by a computer, a pair of input data records, wherein each data record of the pair of input data records comprises at least one corresponding attribute;
applying, by the computer, a machine learning (ML) model on the pair of input data records, wherein
the ML model is trained to match the pair of input data records at an attribute level based on a pair feature vector that includes at least one attribute pair score and at least one dimension match score of a plurality of dimension match scores for the pair of input data records,
each dimension match score of the plurality of dimension match scores corresponds to a different data quality metric of a plurality of data quality metrics, and
the at least one attribute pair score quantifies an extent of match between the at least one corresponding attribute of a first data record in the pair of input data records and the at least one corresponding attribute of a second data record in the pair of input data records;
generating, by the computer, match outcome information for the pair of input data records, based on the applying of the ML model on the pair of input data records, wherein the match outcome information indicates an extent of match between the pair of input data records;
merging, by the computer, based on the generated match outcome information, the first data record into the second data record to generate merged data record;
deleting, by the computer, the first data record from a data storage device based on the merging; and
storing, by the computer, the merged data record in the data storage device.
2 . The computer-implemented method of claim 1 , further comprising computing, by the computer, the at least one attribute pair score corresponding to each attribute of the at least one corresponding attribute of each data record of the pair of input data records.
3 . The computer-implemented method of claim 1 , further comprising computing, by the computer, at least one data quality score for each data record of the pair of input data records, wherein the at least one data quality score corresponds to at least one data quality metric of the plurality of data quality metrics.
4 . The computer-implemented method of claim 3 , wherein
the at least one corresponding attribute of each data record of the pair of input data records comprises at least one data field, and
each data quality metric of the plurality of data quality metrics is defined for a respective data field of the at least one data field of each attribute of the at least one corresponding attribute.
5 . The computer-implemented method of claim 3 , wherein the at least one data quality metric is defined for each attribute of the at least one corresponding attribute of each data record of the pair of input data records.
6 . The computer-implemented method of claim 1 , wherein the computer-implemented method further comprises:
computing, by the computer, a first data metric score for the first data record, wherein the first data metric score corresponds to a data quality metric of the plurality of data quality metrics;
computing, by the computer, a second data metric score for the second data record, wherein the second data metric score corresponds to the data quality metric; and
computing, by the computer, the at least one dimension match score for the pair of input data records, based on the first data metric score and the second data metric score.
7 . The computer-implemented method of claim 1 , further comprising generating, by the computer, the pair feature vector for the pair of input data records based on the at least one attribute pair score and the at least one dimension match score.
8 . A computer system, comprising:
a processor set;
one or more computer-readable storage media; and
program instructions stored on the one or more computer-readable storage media, the program instructions executable by the processor set to cause the processor set to:
obtain a pair of input data records, wherein each data record of the pair of input data records comprises at least one corresponding attribute;
apply a machine learning (ML) model on the pair of input data records, wherein
the ML model is trained to match the pair of input data records at an attribute level based on a pair feature vector that includes at least one attribute pair score and at least one dimension match score of a plurality of dimension match scores for the pair of input data records,
each dimension match score of the plurality of dimension match scores corresponds to a different data quality metric of a plurality of data quality metrics, and
the at least one attribute pair score quantifies an extent of match between the at least one corresponding attribute of a first data record in the pair of input data records and the at least one corresponding attribute of a second data record in the pair of input data records;
generate match outcome information for the pair of input data records, based on the application of the ML model on the pair of input data records, wherein the match outcome information indicates an extent of match between the pair of input data records;
merge, based on the generated match outcome information, the first data record into the second data record to generate merged data record;
delete the first data record from a data storage device based on the merger; and
store the merged data record in the data storage device.
9 . The computer system of claim 8 , wherein the program instructions further cause the processor set to compute the at least one attribute pair score corresponding to each attribute of the at least one corresponding attribute of each data record of the pair of input data records.
10 . The computer system of claim 8 , wherein the program instructions further cause the processor set to compute at least one data quality score for each data record of the pair of input data records, and wherein the at least one data quality score corresponds to at least one data quality metric of the plurality of data quality metrics.
11 . The computer system of claim 10 , wherein
the at least one corresponding attribute of each data record of the pair of input data records comprises at least one data field, and
each data quality metric of the plurality of data quality metrics is defined for a respective data field of the at least one data field of each attribute of the at least one corresponding attribute.
12 . The computer system of claim 10 , wherein the at least one data quality metric is defined for each attribute of the at least one corresponding attribute of each data record of the pair of input data records.
13 . The computer system of claim 8 , wherein the program instructions further cause the processor set to:
compute a first data metric score for the first data record, wherein the first data metric score corresponds to a data quality metric of the plurality of data quality metrics;
compute a second data metric score for the second data record, wherein the second data metric score corresponds to the data quality metric; and
compute the at least one dimension match score for the pair of input data records, based on the first data metric score and the second data metric score.
14 . The computer system of claim 8 , wherein the program instructions further cause the processor set to generate the pair feature vector for the pair of input data records based on the at least one attribute pair score and the at least one dimension match score.
15 . A computer-program product for data record manipulation, the computer-program product comprising:
one or more computer-readable storage media; and
program instructions stored on the one or more computer-readable storage media to perform operations comprising:
obtaining a pair of input data records, wherein each data record of the pair of input data records comprises at least one corresponding attribute;
applying a machine learning (ML) model on the pair of input data records, wherein
the ML model is trained to match the pair of input data records at an attribute level based on a pair feature vector that includes at least one attribute pair score and at least one dimension match score of a plurality of dimension match scores for the pair of input data records,
each dimension match of the plurality of dimension match scores corresponds to a different data quality metric of a plurality of data quality metrics, and
the at least one attribute pair score quantifies an extent of match between the at least one corresponding attribute of a first data record in the pair of input data records and the at least one corresponding attribute of a second data record in the pair of input data records;
generating match outcome information for the pair of input data records, based on the applying of the ML model on the pair of input data records, wherein the match outcome information indicates an extent of match between the pair of input data records;
merging, based on the generated match outcome information, the first data record into the second data record to generate merged data record;
deleting the first data record from a data storage device based on the merging; and
storing the merged data record in the data storage device.