IP Library Granted Patent US 11,321,359
Granted Patent B2
US 11,321,359 · App. 16/706,086 · Granted May 3, 2022

Review and curation of record clustering changes at large scale

Inventors: Timothy Kwok Webber (Boston, MA); George Anwar Dany Beskales (Waltham, MA); Dennis Cunningham (North Reading, MA); Alan Benjamin Wagner Rodriguez (Guaynabo, PR); Liam Cleary (Dublin, IE)
Assignee: TAMR, INC.
G06F16/285G06F16/2282G06F16/2456G06F16/2465G06F16/24573G06F16/288G06F16/953
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,321,359
App. No.
16/706,086
Granted
May 3, 2022
Kind
B2
Abstract

Methods are provided to represent proposed changes to clusterings for ease of review, as well as tools to help subject matter experts identify clusters that warrant review versus those that do not. These tools make overall assessment of proposed clustering changes and targeted curation practical at large scale. Use of these tools and method enables efficient data management operations when dealing with extreme scale, such as where entity resolution involves clusterings created from data sources involving millions of entities.

Claims (59)

1. A method of clustering management for large scale data records, comprising, within a computer-implemented software program to manage data:

providing a current published clustering, the current published clustering having a plurality of clusters, each cluster in the current published clustering having a first plurality of data records, wherein the data records in each cluster refer to the same entity, and wherein the plurality of clusters in the current published clustering define published clusters, and wherein the current published clustering is a current version of entity resolution, the current published clustering being a first dataset;

receiving a proposed clustering that is different from the current published clustering, the proposed clustering also having a plurality of clusters, each cluster in the proposed clustering having a second plurality of data records, wherein the plurality of clusters in the proposed clustering define proposed clusters, and wherein the proposed clustering is a subsequent version of entity resolution, the proposed clustering being a second dataset, wherein the second dataset differs from the first dataset in one or more of the following ways:

(i) the second dataset includes one or more data records that are not present in the first dataset, or

(ii) the second dataset does not include one or more data records that are present in the first dataset, or

(iii) the second dataset includes one or more data records that are modified relative to the first dataset, wherein the modifications are data changes to fields of the data records;

matching clusters within the proposed clustering to clusters within the current published clustering;

identifying differences between the current published clustering and the proposed clustering on both cluster and record levels using the matched clusters, thereby identifying differences between the first and second datasets of the respective current published clustering and the proposed clustering, the identified differences including one or more of the following differences:

(i) the second dataset includes one or more data records that are not present in the first dataset, or

(ii) the second dataset does not include one or more data records that are present in the first dataset, or

(iii) the second dataset includes one or more data records that are modified relative to the first dataset, wherein the modifications are data changes to fields of the data records;

approving or rejecting the proposed clustering based upon a review of the identified differences; and

upon approval of the proposed clustering, creating a new published clustering using the proposed clustering, and upon rejection of the proposed clustering, receiving a new proposed clustering for subsequent review.

2. The method of claim 1 , further comprising, within the software program:

storing metadata about the proposed clusters and the published clusters in a table, wherein the metadata includes cluster size, cluster verification, and cluster status, and

wherein the cluster status is one of new, changed, unchanged, or empty.

3. The method of claim 2 , wherein identifying differences further comprises:

computing differences between matched clusters and storing statistics about the differences within the metadata about the published clusters and proposed clusters.

4. The method of claim 3 , wherein the computing of differences is further a record-level computation of differences between matched clusters, and the differences are used to:

update the metadata about the published clusters, including statistics on number of data records added, removed or changed,

update the metadata about the proposed clusters, including statistics on number of data records added, removed, or changed from the matching published cluster, and

update current and proposed cluster membership for each data record in a data table tracking data record cluster membership.

5. The method of claim 4 , further comprising, within the software program:

creating the data table tracking data record cluster membership by performing, within a big data analytics platform, a join of all data fields for each data record along with cluster metadata for the published cluster and the proposed cluster containing the data record, and

storing the join results as a flat table optimized for fast searching and data retrieval within a search engine.

6. The method of claim 1 , further comprising, within the software program:

computing confidence metrics in the clusters of the proposed clustering; and

storing the computed confidence metrics within metadata about the proposed clusters.

7. The method of claim 6 , wherein computing confidence metrics further comprises:

computing intra-cluster confidence for each cluster through pairwise similarity of record fields across all pairs of data records within a cluster, and inter-cluster confidence for each cluster through pairwise similarity of record fields across all pairs of data records selected to pair data records in one cluster with data records from a different cluster.

8. The method of claim 7 , further comprising, within the software program:

applying a threshold setting to the confidence metrics, or degree of change in confidence metrics between a published cluster and matching proposed cluster, to identify clusters warranting manual review, and excluding clusters not identified as warranting manual review from those requiring approval before the proposed clustering is accepted.

9. The method of claim 8 , further comprising, within the software program:

after differences in proposed clusters warranting manual review have been approved, automatically approving the proposed clustering.

10. The method of claim 1 , further comprising, within the software program:

providing a user interface to review and approve or reject the proposed clustering, the user interface including:

filter and search tools to identify clusters for review,

a clustering pane displaying proposed clusters and proposed cluster metadata,

a cluster pane displaying individual data records within a cluster selected from the clustering pane, and

a data record pane displaying data record details selected of a data record selected from the cluster pane, and

providing tools within the user interface to:

undo a change in a proposed clustering,

move a data record to a different cluster,

edit the data record details,

lock a data record,

lock a proposed clustering,

assign proposed clusters for review by a particular individual, and

approve a proposed clustering.

11. The method of claim 10 , further comprising, within the user interface:

displaying visual indication through color or displayed symbols of data record status within the cluster pane, wherein visual indication of status includes:

data record unchanged from the published clustering,

data record new to the proposed cluster,

data record new to the published clustering,

data record moved out of the proposed cluster, and

data record deleted from the published clustering.

12. The method of claim 1 , further comprising, within the software program:

storing source data records for the current published clustering and the proposed clustering within a big data analytics platform.

13. The method of claim 1 , further comprising, within the software program:

assigning a cluster ID of a matching cluster in the published clustering to the matching cluster in the proposed clustering, and assigning previously unused cluster IDs to clusters in the proposed clustering having no match in the published clustering.

Assignments (7)
RELEASE OF SECURITY INTEREST Recorded Feb 21, 2025
From: JPMORGAN CHASE BANK, N.A.
To: TAMR, INC.
Reel/Frame 070284/0101 →
RELEASE OF SECURITY INTEREST Recorded Feb 21, 2025
From: JPMORGAN CHASE BANK, N.A.
To: TAMR, INC.
Reel/Frame 070284/0092 →
AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jan 30, 2023
From: TAMR, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 062540/0438 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Mar 19, 2021
From: TAMR, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 055662/0240 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 17, 2020
From: TAMR, INC.
To: WESTERN ALLIANCE BANK
Reel/Frame 055205/0909 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2020
From: BESKALES, GEORGE ANWAR DANY
To: TAMR, INC.
Reel/Frame 053712/0568 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2020
From: WEBBER, TIMOTHY KWOK; BESCALES, GEORGE ANWAR DANY; CUNNINGHAM, DENNIS; RODRIGUEZ, ALAN BENJAMIN WAGNER; CLEARY, LIAM
To: TAMR, INC.
Reel/Frame 053039/0930 →
Continuity (2)
Provisional Application 62808060 · Feb 20, 2019
Related Publication 20220004565A1 · Jan 6, 2022