IP Library Granted Patent US 12,608,652
Granted Patent B2
US 12,608,652 · App. 18/120,099 · Granted Apr 21, 2026

Tuning a trained data record matching model using customer data and representation learning

Inventors: Abhishek Seth (Deoband, IN); Devbrat Sharma (Bangalore, IN); Mahendra Singh Kanyal (Banbasa, IN); Soma Shekar Naganna (Bangalore, IN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,652
App. No.
18/120,099
Granted
Apr 21, 2026
Kind
B2
Abstract

A method, system, and computer program product are configured to create a tuned data record matching model by adjusting values of one or more parameters in a data record matching model based on a second training data set labeled at a data record level, wherein the data record matching model is initially trained using a first training data set labeled at an attribute level.

Claims (53)

1 . A method which provides an accurate matching model which is tuned by adjusting values based on customer data, comprising:

training, by a processor set, a classification model to determine adjusted distance measures using distance vectors as inputs and labels of a second training data set as target outputs;

training, by the processor set, at least one regression model to determine an adjusted comparison coefficient vector for each attribute using the determined adjusted distance measures from the trained classification model;

obtaining, by the processor set, a data record matching model trained using a first training data set labeled at an attribute level; and

creating, by the processor set, a tuned data record matching model by adjusting values of one or more parameters in the data record matching model based on the second training data set labeled at a data record level.

2 . The method of claim 1 , wherein the one or more parameters comprise comparison coefficient vectors.

3 . The method of claim 2 , wherein the data record matching model uses a respective one of the comparison coefficient vectors with a respective feature vector to determine a distance measure of a respective attribute of a pair of data records.

4 . The method of claim 2 , wherein the adjusting the values of the one or more parameters is performed using a two-phase learning process.

5 . The method of claim 4 , wherein:

a first phase of the two-phase learning process comprises determining the adjusted distance measures using the labels of the second training data set as the target outputs; and

a second phase of the two-phase learning process comprises determining adjusted values of the comparison coefficient vectors using the adjusted distance measures as the target outputs.

6 . The method of claim 5 , wherein:

the first phase comprises training the classification model that is used to determine the adjusted distance measures; and

the second phase comprises training plural regression models, wherein respective ones of the plural regression models are used to determine the adjusted values of respective ones of the comparison coefficient vectors.

7 . The method of claim 1 , wherein:

the first training data set comprises computer generated and labeled pairs of attribute values; and

the second training data set comprises labeled pairs of customer data records.

8 . The method of claim 1 , further comprising using the tuned data record matching model to classify pairs of data records as matching or unmatching.

9 . A computer program product which provides an accurate matching model which is tuned by adjusting values based on customer data, the computer program product comprising one or more computer readable storage media having program instructions collectively stored on the one or more computer readable storage media, the program instructions executable to:

train a classification model to determine adjusted distance measures using distance vectors as inputs and labels of a second training data set as target outputs;

train at least one regression model to determine an adjusted comparison coefficient vector for each attribute using the determined adjusted distance measures from the trained classification model;

create a tuned data record matching model by adjusting values of one or more parameters in a data record matching model based on the second training data set labeled at a data record level, wherein the data record matching model is initially trained using a first training data set labeled at an attribute level; and

use the tuned data record matching model to classify pairs of data records as matching or unmatching.

10 . The computer program product of claim 9 , wherein the one or more parameters comprise comparison coefficient vectors.

11 . The computer program product of claim 10 , wherein the data record matching model uses a respective one of the comparison coefficient vectors with a respective feature vector to determine a distance measure of a respective attribute of a pair of data records.

12 . The computer program product of claim 10 , wherein:

the adjusting the values of the one or more parameters is performed using a two-phase learning process;

a first phase of the two-phase learning process comprises the determining adjusted distance measures using the labels of the second training data set as the target outputs; and

a second phase of the two-phase learning process comprises determining adjusted values of the comparison coefficient vectors using the adjusted distance measures as the target outputs.

13 . The computer program product of claim 12 , wherein:

the first phase comprises training the classification model that is used to determine the adjusted distance measures; and

the second phase comprises training plural regression models, wherein respective ones of the plural regression models are used to determine the adjusted values of respective ones of the comparison coefficient vectors.

14 . The computer program product of claim 9 , wherein:

the first training data set comprises computer generated and labeled pairs of attribute values; and

the second training data set comprises labeled pairs of customer data records.

15 . A system which provides an accurate matching model which is tuned by adjusting values based on customer data, the system comprising:

a processor set, one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable to:

train a classification model to determine adjusted distance measures using distance vectors as inputs and labels of a second training data set as target outputs;

train at least one regression model to determine an adjusted comparison coefficient vector for each attribute using the determined adjusted distance measures from the trained classification model;

create a tuned data record matching model by adjusting values of one or more parameters in a data record matching model based on the second training data set labeled at a data record level, wherein the data record matching model is initially trained using a first training data set labeled at an attribute level; and

provide the tuned data record matching model to a customer for the customer to use the tuned data record matching model to classify pairs of data records as matching or unmatching.

16 . The system of claim 15 , wherein the one or more parameters comprise comparison coefficient vectors.

17 . The system of claim 16 , wherein the data record matching model uses a respective one of the comparison coefficient vectors with a respective feature vector to determine a distance measure of a respective attribute of a pair of data records.

18 . The system of claim 16 , wherein:

the adjusting the values of the one or more parameters is performed using a two-phase learning process;

a first phase of the two-phase learning process comprises determining the adjusted distance measures using the labels of the second training data set as the target outputs; and

a second phase of the two-phase learning process comprises determining adjusted values of the comparison coefficient vectors using the adjusted distance measures as the target outputs.

19 . The system of claim 18 , wherein:

the first phase comprises training the classification model that is used to determine the adjusted distance measures; and

the second phase comprises training plural regression models, wherein respective ones of the plural regression models are used to determine the adjusted values of respective ones of the comparison coefficient vectors.

20 . The system of claim 15 , wherein:

the first training data set comprises computer generated and labeled pairs of attribute values; and

the second training data set comprises labeled pairs of customer data records.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2023
From: SETH, ABHISHEK; SHARMA, DEVBRAT; KANYAL, MAHENDRA SINGH; NAGANNA, SOMA SHEKAR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 062948/0334 →
Continuity (1)
Related Publication 20240303533A1 · Sep 12, 2024
References Cited (19)
US 11386090B2 · Seth et al. · 2022 [cited by applicant]
US 11514321B1 · Chen · 2022 [cited by examiner]
US 20050278139A1 · Glaenzer et al. · 2005 [cited by applicant]
US 20140358829A1 · Hurwitz · 2014 [cited by examiner]
US 20200250576A1 · Jagota · 2020 [cited by examiner]
US 20200250687A1 · Jagota · 2020 [cited by examiner]
US 20200272845A1 · He · 2020 [cited by examiner]
US 20210279604A1 · Seth et al. · 2021 [cited by applicant]
US 20210319026A1 · Seth · 2021 [cited by examiner]
US 20210342353A1 · Jagota · 2021 [cited by examiner]
US 20220222489A1 · Liu · 2022 [cited by examiner]
US 20220309047A1 · Seth et al. · 2022 [cited by applicant]
US 20220358607A1 · Guo · 2022 [cited by examiner]
CN 113569554 · 2021 [cited by applicant]
Bogatu et al., “Cost-effective Variational Active Entity Resolution”, Nov. 20, 2020, 12 pages. [cited by applicant]
Mehta, “How is Maximum Likelihood Estimation used in machine learning?”, https://analyticsindiamag.com/how-is-maximum-likelihood-estimation-used-in-machine-learning/, Apr. 9, 2022, 13 pages. [cited by applicant]
Anonymous, “IBM® Match 360 on Cloud Pak for Data”, https://www.ibm.com/docs/en/cloud-paks/cp-data/4.0?topic=services-match-360-watson, May 26, 2022, 3 pages. [cited by applicant]
Anonymous, “Machine learning”, Wikipedia, https://en.wikipedia.org/wiki/Machine_learning, archived Feb. 20, 2023, 34 pages. [cited by applicant]
Nguyen et al., “Fine-Tuning Pretrained Language Models With Label Attention for Biomedical Text Classification”, Mar. 7, 2022, 5 pages. [cited by applicant]