IP Library › Granted Patent US 12,380,135
Granted Patent B2
US 12,380,135 · App. 18/337,126 · Granted Aug 5, 2025

Records processing based on record attribute embeddings

Inventors: Devbrat Sharma (Bangalore, IN); Soma Shekar Naganna (Bangalore, IN); Abhishek Seth (Deoband, IN); Neeraj Ramkrishna Singh (Bangalore, IN); Muhammed Abdul Majeed Ameen (Kozhikode, IN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/285G06F16/215
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,135
App. No.
18/337,126
Filed
Jun 19, 2023
Granted
Aug 5, 2025
Kind
B2
Art Unit
2161
USPC
707/737
Abstract

One or more trained embedding generation artificial intelligence models are executed to generate a plurality of record attribute embeddings. The plurality of record attribute embeddings represents a plurality of attributes of data of a plurality of records. Grouping of the plurality of record attribute embeddings is performed. The grouping of a record attribute embedding includes grouping attribute values of the record attribute embedding into one or more groups of attribute values. The performing grouping provides a plurality of groups of attribute values for the plurality of record attribute embeddings. Selected records are compared to provide a set of matched records. The comparing, based on a group of attribute values, includes comparing records that include one or more attribute values grouped in the group of attribute values providing a subset of matched records of the set of matched records. The set of matched records is stored in an accessible computer location.

Claims (53)

1. A computer-implemented method of facilitating processing within a computing environment, the computer-implemented method comprising:

obtaining, by at least one computing device of the computing environment, data of a plurality of records;

executing, by the at least one computing device of the computing environment, one or more trained embedding generation artificial intelligence models to generate a plurality of record attribute embeddings based on the data, the plurality of record attribute embeddings representing a plurality of attributes of the data, wherein a record attribute embedding corresponds to an attribute of the data and is output as a multi-dimensional vector that represents one or more attributes within its context;

performing grouping, using the at least one computing device of the computing environment, of the plurality of record attribute embeddings to optimize candidate selection of records for pairwise comparison of records to find matches among pairs of records, wherein the performing grouping of the record attribute embedding of the plurality of record attribute embeddings includes grouping attribute values of the record attribute embedding into one or more groups of attribute values based on one or more selected criteria, and wherein the performing grouping provides a plurality of groups of attribute values for the plurality of record attribute embeddings;

comparing, using the at least one computing device of the computing environment and based on the plurality of groups of attribute values, selected records of the plurality of records to provide a set of matched records created from a plurality of subsets of matched records generated from the comparing, wherein the comparing, based on a group of attribute values of the plurality of groups of attribute values, generates a subset of matched records of the plurality of subsets of matched records of the set of matched records, the comparing based on the group of attribute values comprising:

performing a pairwise comparison of a pair of records that includes one or more attribute values of the group of attribute values,

determining, for the pair of records, whether the pairwise comparison has a predefined relationship with a threshold; and

selecting the pair of records for the subset of matched records based on the pairwise comparison having the predefined relationship with the threshold; and

storing, using the at least one computing device of the computing environment, the set of matched records in a computer location accessible by one or more users of the set of matched records, and wherein a record of the set of matched records represents one or more other records that are different but found to be similar based on one or more defined criteria.

2. The computer-implemented method of claim 1 , wherein the executing outputs the plurality of record attribute embeddings as a plurality of multi-dimensional vectors, and wherein the performing grouping of the plurality of record attribute embeddings uses an unsupervised clustering technique.

3. The computer-implemented method of claim 1 , wherein the one or more selected criteria includes a cosine similarity.

4. The computer-implemented method of claim 1 , wherein the one or more trained embedding generation artificial intelligence models are trained using natural language models and processing.

5. The computer-implemented method of claim 1 , further comprising performing further training of the one or more trained embedding generation artificial intelligence models using selected data of the plurality of records.

6. The computer-implemented method of claim 1 , further comprising standardizing the data of the plurality of records to provide standardized data, and wherein the data in the executing the one or more trained embedding generation artificial intelligence models to generate the plurality of record attribute embeddings based on the data is the standardized data.

7. The computer-implemented method of claim 1 , wherein the plurality of attribute embeddings are generated in a finite-dimensional inner product space over real numbers.

8. The computer implemented method of claim 1 , wherein the comparing based on the group of attribute values further comprises repeating the performing the pairwise comparison, the determining whether the pairwise comparison has the predefined relationship and the selecting the pair of records for one or more additional pairs of records that includes the one or more attribute values of the group of attribute values.

9. The computer-implemented method of claim 1 , further comprising performing the comparing for each group of attribute values of the plurality of groups of attribute values to generate the plurality of subsets of matched records.

10. The computer-implemented method of claim 9 , further comprising:

consolidating the plurality of subsets of matched records to provide the set of matched records, the set of matched records including fewer records than the plurality of records; and

wherein the storing the set of matched records in the computer location accessible by the one or more users includes storing the set of matched records in storage, wherein an amount of storage used for the set of matched records is less than the amount of storage used for the plurality of records.

11. A computer system for facilitating processing within a computing environment, the computer system comprising:

a memory; and

one or more computing devices in communication with the memory, wherein the computer system is configured to perform a method, said method comprising:

obtaining data of a plurality of records;

executing one or more trained embedding generation artificial intelligence models to generate a plurality of record attribute embeddings based on the data, the plurality of record attribute embeddings representing a plurality of attributes of the data, wherein a record attribute embedding corresponds to an attribute of the data and is output as a multi-dimensional vector that represents one or more attributes within its context;

performing grouping of the plurality of record attribute embeddings to optimize candidate selection of records for pairwise comparison of records to find matches among pairs of records, wherein the performing grouping of the record attribute embedding of the plurality of record attribute embeddings includes grouping attribute values of the record attribute embedding into one or more groups of attribute values based on one or more selected criteria, and wherein the performing grouping provides a plurality of groups of attribute values for the plurality of record attribute embeddings;

comparing, based on the plurality of groups of attribute values, selected records of the plurality of records to provide a set of matched records created from a plurality of subsets of matched records generated from the comparing, wherein the comparing, based on a group of attribute values of the plurality of groups of attribute values, generates a subset of matched records of the plurality of subsets of matched records of the set of matched records, the comparing based on the group of attribute values comprising:

performing a pairwise comparison of a pair of records that includes one or more attribute values of the group of attribute values,

determining, for the pair of records, whether the pairwise comparison has a predefined relationship with a threshold; and

selecting the pair of records for the subset of matched records based on the pairwise comparison having the predefined relationship with the threshold; and

storing the set of matched records in a computer location accessible by one or more users of the set of matched records, and wherein a record of the set of matched records represents one or more other records that are different but found to be similar based on one or more defined criteria.

12. The computer system of claim 11 , wherein the method further comprises standardizing the data of the plurality of records to provide standardized data, and wherein the data in the executing the one or more trained embedding generation artificial intelligence models to generate the plurality of record attribute embeddings based on the data is the standardized data.

13. The computer system of claim 11 , wherein the plurality of attribute embeddings are generated in a finite-dimensional inner product space over real numbers.

14. The computer system of claim 11 , wherein the method further comprises performing the comparing for each group of attribute values of the plurality of groups of attribute values to generate the plurality of subsets of matched records.

15. The computer system of claim 14 , wherein the method further comprises:

consolidating the plurality of subsets of matched records to provide the set of matched records, the set of matched records including fewer records than the plurality of records; and

wherein the storing the set of matched records in the computer location accessible by the one or more users includes storing the set of matched records in storage, wherein an amount of storage used for the set of matched records is less than the amount of storage used for the plurality of records.

16. A computer program product for facilitating processing within a computing environment, said computer program product comprising:

one or more computer readable storage media and program instructions collectively stored on the one or more computer readable storage media readable by at least one processing circuit to:

obtain data of a plurality of records;

execute one or more trained embedding generation artificial intelligence models to generate a plurality of record attribute embeddings based on the data, the plurality of record attribute embeddings representing a plurality of attributes of the data, wherein a record attribute embedding corresponds to an attribute of the data and is output as a multi-dimensional vector that represents one or more attributes within its context;

perform grouping of the plurality of record attribute embeddings to optimize candidate selection of records for pairwise comparison of records to find matches among pairs of records, wherein the performing grouping of the record attribute embedding of the plurality of record attribute embeddings includes grouping attribute values of the record attribute embedding into one or more groups of attribute values based on one or more selected criteria, and wherein the performing grouping provides a plurality of groups of attribute values for the plurality of record attribute embeddings;

compare, based on the plurality of groups of attribute values, selected records of the plurality of records to provide a set of matched records created from a plurality of subsets of matched records generated from the compare, wherein the compare, based on a group of attribute values of the plurality of groups of attribute values, generates a subset of matched records of the plurality of subsets of matched records of the set of matched records, the compare based on the group of attribute values comprising:

performing a pairwise comparison of a pair of records that includes one or more attribute values of the group of attribute values;

determining, for the pair of records, whether the pairwise comparison has a predefined relationship with a threshold; and

selecting the pair of records for the subset of matched records based on the pairwise comparison having the predefined relationship with the threshold; and

store the set of matched records in a computer location accessible by one or more users of the set of matched records, and wherein a record of the set of matched records represents one or more other records that are different but found to be similar based on one or more defined criteria.

17. The computer program product of claim 16 , wherein the program instructions are further readable by the at least one processing circuit to standardize the data of the plurality of records to provide standardized data, and wherein the data in the executing the one or more trained embedding generation artificial intelligence models to generate the plurality of record attribute embeddings based on the data is the standardized data.

18. The computer program product of claim 16 , wherein the plurality of attribute embeddings generated in a finite-dimensional inner product space over real numbers.

19. The computer program product of claim 16 , wherein the program instructions are further readable by the at least one processing circuit to perform the comparing for each group of attribute values of the plurality of groups of attribute values to generate the plurality of subsets of matched records.

20. The computer program product of claim 19 , wherein the program instructions are further readable by the at least one processing circuit to:

consolidate the plurality of subsets of matched records to provide the set of matched records, the set of matched records including fewer records than the plurality of records; and

wherein the storing the set of matched records in the computer location accessible by the one or more users includes storing the set of matched records in storage, wherein an amount of storage used for the set of matched records is less than the amount of storage used for the plurality of records.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2023
From: SHARMA, DEVBRAT; NAGANNA, SOMA SHEKAR; SETH, ABHISHEK; SINGH, NEERAJ RAMKRISHNA; AMEEN, MUHAMMED ABDUL MAJEED
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063984/0280 →
Continuity (1)
Related Publication 20240419690A1 · Dec 19, 2024
References Cited (14)
US 9767127B2 · Feldschuh · 2017 [cited by examiner]
US 10242019B1 · Shan · 2019 [cited by examiner]
US 11481448B2 · Zhong et al. · 2022 [cited by applicant]
US 11755626B1 · Liu · 2023 [cited by examiner]
US 20100312726A1 · Thompson · 2010 [cited by examiner]
US 20130159310A1 · Birdwell · 2013 [cited by examiner]
US 20140279757A1 · Shimanovsky et al. · 2014 [cited by applicant]
US 20190392073A1 · Gutierrez Munoz · 2019 [cited by examiner]
US 20200089693A1 · Fulzele · 2020 [cited by examiner]
US 20230040412A1 · Ramsl · 2023 [cited by examiner]
US 20240153239A1 · Chen · 2024 [cited by examiner]
Azzalini, Fabio et al., “Blocking Techniques for Entity Linkage: A Semantics-Based Approach,” Data Science and Engineering, Jan. 2021, pp. 6-:20-38. [cited by applicant]
Chen, Xiao et al., “The Best of Both Worlds: Combining Hand-Tuned and Word-Embedding-Based Similarity Measures for Entity Resolution,” Aug. 2022, pp. 215-224. [cited by applicant]
Li, Chen et al., “Supporting Efficient Record Linkage for Large Data Sets Using Mapping Techniques,” Nov. 23, 2022, pp. 1-25. [cited by applicant]