IP Library Granted Patent US 10,467,201
Granted Patent B1
US 10,467,201 · App. 15/380,925 · Granted Nov 5, 2019

Systems and methods for integration and analysis of data records

Inventors: Sears Merritt (Groton, MA); Thom Neale (Bedford, MA)
Assignee: Massachusetts Mutual Life Insurance Company
G06F16/212G06F16/215G06F16/258G06F16/285G06F16/9024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,467,201
App. No.
15/380,925
Granted
Nov 5, 2019
Kind
B1
Abstract

Methods and systems for determining relationships between two or more nominally unrelated data sources utilizing a combination of probabilistic modeling and graphical clustering are described. The systems and methods for utilizing probabilistic model functions as a way of determining and judging the likelihood that two records from different systems are related to the same entity.

Claims (34)

1. A system comprising:

a first and second database configured to store a plurality of data points; and

a server configured to:

convert the plurality of data points stored in the first database associated with a first format and the second database associated with a second format, to a common format in a single dataset wherein the plurality of data points are arranged in a plurality of fields;

upon converting the plurality of data points to the common format, assign a unique id to each field of the plurality of data points;

normalize the plurality of data points arranged in the plurality of fields by applying a master schema, wherein the master schema identifies each semantically equivalent field;

group, based on a blocking key, to produce a hashable value for the plurality of data points into one or more groups upon the plurality of data points satisfying a first pre-defined criteria;

select a subset of groups from the one or more groups having a group size that is less than a pre-defined group size;

classify the plurality of data points by matching each pair of data points within each group from the subset of groups with each other based on a relationship corresponding to the hashable value associated with each data point satisfying a second pre-defined criteria; and

generate a graph comprising classification data comprising a set of vertices, wherein each vertex is represented as a unique long integer and a set of edges in an undirected graph, and wherein an edge associated with each vertex corresponds to each matched pair of the data points, and wherein the graph presents each cluster of data points that are determined to be related to each other based on matched pairs of the data points as a single entity having a distinct identification number.

2. The system of claim 1 , wherein the common format is a single text encoding format.

3. The system of claim 1 , wherein the plurality of fields in the first and second databases are mapped to a common set of fields selected from a predetermined set of fields.

4. The system of claim 1 , wherein the matching is based on a characteristic of the data points.

5. The system of claim 1 , wherein the matching is based on a probability of a relationship between the data points.

6. The system of claim 1 , wherein a set of matching instructions causes the server to generate a pair of the data points from the plurality of data points in each group.

7. The system of claim 6 , wherein the set of matching instructions determines a probabilistic relationship between pair of the data points.

8. The system of claim 1 , wherein the second pre-defined criteria is a probability of a relationship between the data points in each group.

9. The system of claim 1 , wherein the normalization of the plurality of fields is mapped to a common set of fields selected from a library of fields stored in a database.

10. A computer-implemented method comprising:

converting, by a server, a plurality of data points stored in a first database associated with a first format and a second database associated with a second format, to a common format in a single dataset wherein the plurality of data points are arranged in a plurality of fields;

upon converting the plurality of data points to the common format, assigning, by the server, a unique id to each field of the plurality of data points;

normalizing, by the server, the plurality of data points arranged in the plurality of fields by applying a master schema, wherein the master schema identifies each semantically equivalent field;

grouping, by the server, based on a blocking key, to produce a hashable value for the plurality of data points into one or more groups upon the plurality of data points satisfying a first pre-defined criteria;

selecting, by the server, a subset of groups from the one or more groups having a group size that is less than a pre-defined group size;

classifying, by the server, the plurality of data points by matching each data point within each group from the subset of groups with each other based on a relationship corresponding to the hashable value associated with each data point satisfying a second pre-defined criteria; and

generating, by the server, a graph comprising classification data comprising a set of vertices, wherein each vertex is represented as a unique long integer and a set of edges in an undirected graph, wherein an edge associated with each vertex corresponds to each matched pair of the data points, and wherein the graph presents each cluster of data points that are determined to be related to each other based on matched pairs of the data points as a single entity having a distinct identification number.

11. The computer-implemented method of claim 10 , wherein the common format is a single text encoding format.

12. The computer-implemented method of claim 10 , wherein the plurality of fields in the first and second databases are mapped to a common set of fields selected from a predetermined set of fields.

13. The computer-implemented method of claim 10 , wherein the matching is based on a characteristic of the data points.

14. The computer-implemented method of claim 10 , wherein the matching is based on a probability of a relationship between the data points.

15. The computer-implemented method of claim 10 , wherein a set of matching instructions causes the server to generate a pair of the data points from the plurality of data points in each group.

16. The computer-implemented method of claim 15 , wherein the set of matching instructions determines a probabilistic relationship between pair of the data points.

17. The computer-implemented method of claim 10 , wherein the second pre-defined criteria is a probability of a relationship between the data points in each group.

18. The computer-implemented method of claim 10 , wherein the normalization of the plurality of fields is mapped to a common set of fields selected from a library of fields stored in a database.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2016
From: MERRITT, SEARS; NEALE, THOM
To: MASSACHUSETTS MUTUAL LIFE INSURANCE COMPANY
Reel/Frame 040638/0306 →
Continuity (1)
Provisional Application 62387164 · Dec 23, 2015
Cited By (7)
US 12,461,711 US 12,566,727 US 12,602,352 US 12,657,178 US 12,670,133 US 12,670,186 US 12,718,914